Cloth material multi-feature fusion AI search optimization method

CN122817531APending Publication Date: 2026-09-25SHANWEI MEIBAO INTELLIGENT DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611032439.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-13
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

视觉特征仅能反映材质的外观信息,文本标签仅能覆盖有限的语义描述,两者均无法全面表达布艺材质的多维度属性

Benefits of technology

1、本发明通过跨模态注意力融合,以视觉特征分量为查询矩阵,动态聚合物理特征和语义特征中的关联信息,使增强融合表征向量能够根据材质的视觉呈现自适应地强化对应物理属性和工艺语义的表达,有助于提升查询匹配的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817531A_ABST
    Figure CN122817531A_ABST
Patent Text Reader

Abstract

The application is a cloth material multi-feature fusion AI search optimization method, which relates to the technical field of information retrieval and artificial intelligence, comprising: the cloth material multi-modal embedding vector is split into visual component, physical component and semantic component, input into the cross-modal attention fusion module, the visual component is used as the query matrix, the physical component and the semantic component after splicing are used as the key matrix and the value matrix, the enhanced fusion representation vector is output after attention calculation and residual connection, and the original embedding vector of the corresponding material in the vector database is replaced. In the application, the cross-modal attention fusion takes the visual feature component as the query matrix, dynamically aggregates the related information in the physical feature and the semantic feature, so that the enhanced fusion representation vector can adaptively strengthen the expression of the corresponding physical attribute and process semantics according to the visual presentation of the material, which helps to improve the accuracy of query matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of information retrieval and artificial intelligence technology, and in particular to an AI search optimization method that integrates multiple features of fabric materials. Background Technology

[0002] Fabric materials possess multiple attributes, including visual appearance, physical properties, and technological semantics, making them an important basis for users to select materials in scenarios such as textile e-commerce and interior design.

[0003] Existing fabric material search solutions mostly rely on single visual features or simple text tags for candidate material matching. Visual features can only reflect the appearance information of the material, and text tags can only cover limited semantic descriptions; neither can fully express the multi-dimensional attributes of fabric materials. In addition, existing solutions typically treat user search queries as static input, lacking the ability to model user historical behavior sequences and long-term preferences, and thus failing to infer the user's deep search intent.

[0004] The aforementioned defects result in semantic discrepancies between search results and users' actual needs, limiting the relevance and diversity of search results and making it difficult to meet users' actual usage needs in the context of fabric material selection. Summary of the Invention

[0005] The purpose of this invention is to propose a multi-feature fusion AI search optimization method for fabric materials in order to solve the above-mentioned problems.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: A multi-feature fusion AI search optimization method for fabric materials includes: Obtain fabric material sample data, extract feature vectors from visual modality, physical modality and semantic modality respectively, map the three types of feature vectors to the same embedding space through projection layer, splice them to generate fabric material multimodal embedding vector, and store them in vector database; The multimodal embedding vector of the fabric material is split into visual components, physical components and semantic components, and input into the cross-modal attention fusion module. The visual components are used as the query matrix, and the physical components and semantic components are concatenated as the key matrix and value matrix. After attention calculation and residual connection, the enhanced fusion representation vector is output to replace the original embedding vector of the corresponding material in the vector database. The system receives user search query input, converts the search query input into a query vector and maps it to the embedding space, calculates the similarity score between the query vector and the enhanced fusion representation vectors of each material in the vector database, and outputs the search ranking results in descending order of similarity score.

[0007] Preferably, the attention calculation process of the cross-modal attention fusion module is as follows: ; in, The query matrix is ​​composed of visual components. The key matrix is ​​formed by concatenating physical and semantic components. For value matrices, The feature dimension of the key matrix; The visual enhancement vector is generated by weighting and aggregating the value matrix based on the attention weight matrix. The visual enhancement vector is then residually connected with the original visual components and processed by layer normalization to output the first layer fusion result. The first layer fusion result is used as a new query matrix, and the enhanced fusion representation vector is output after multi-layer cross-attention iteration.

[0008] Preferably, the step of extracting feature vectors from the visual modality includes: acquiring front, back, and macro images of the fabric sample, inputting the images into a pre-trained visual feature extraction network, and outputting visual feature vectors; the step of extracting feature vectors from the physical modality includes: acquiring the weight, thickness, density, yarn count, softness, drape, and surface friction coefficient of the fabric sample, and encoding the above numerical parameters into physical feature vectors after Z-score normalization; the step of extracting feature vectors from the semantic modality includes: acquiring the material name, composition information, weaving process category, and finishing process category of the fabric sample, inputting semantic label data into a word embedding model, and outputting semantic feature vectors.

[0009] Preferably, the step of converting the search query input into a query vector includes: when the search query input is a text keyword, converting the text keyword into a query semantic vector through a word embedding model; when the search query input is a reference image, converting the reference image into a query visual vector through a visual feature extraction network; and mapping the query semantic vector or query visual vector to the same embedding space as the enhanced fusion representation vector through a projection layer to generate a query vector.

[0010] Preferably, after converting the search query input into a query vector, the method further includes a step of dynamically inferring the user's search intent: Perform word segmentation and named entity recognition on the user's current query text, extract key entities and generate key entity codes; Obtain the user's short-term behavior sequence, including the most recent L search records, browsing history, clicked material icons, and dwell time, and generate a short-term behavior sequence code, where L is the length of the short-term behavior sequence; obtain the user's long-term preference profile, including the statistical distribution of material categories in historical collections and comparison records, and generate a long-term preference profile code. The key entity encoding, short-term behavior sequence encoding, and long-term preference profile encoding are concatenated and input into the Transformer query understanding model to output a user intent vector. The user intent vector is input into a pre-trained text generation model to generate extended query term text. The extended query term text is then encoded by a word embedding model and weighted and summed with the original query vector according to learnable weights to generate an enhanced query vector that replaces the original query vector for similarity score calculation.

[0011] Preferably, the Transformer query understanding model uses explicit user feedback labels as supervision signals, is trained using the cross-entropy loss function and the Adam optimization algorithm, and has a fully connected output layer that outputs a fixed-dimensional user intent vector.

[0012] Preferably, after calculating the similarity score, a multi-objective ranking step is further included: The cosine similarity algorithm is used to calculate the correlation score between the enhanced query vector and the enhanced fusion representation vector of each candidate material; The maximum marginal relevance method is used to apply a penalty term to the similarity between the materials already selected in the result set and the current candidate materials, and output the diversity score of each candidate material. Based on the historical user click-through rate, collection rate and rating data of each candidate material, after mean normalization based on the range, the data is multiplied by three learnable weight coefficients and summed. The sum of the three weight coefficients is one, and the quality score of each candidate material is output. The relevance score, diversity score, and quality score are input into the multi-objective ranking model, which outputs a comprehensive score for each candidate material. The final search ranking results are then output by sorting the comprehensive scores from high to low.

[0013] Preferably, the multi-objective ranking model adopts an MMoE architecture, including multiple expert networks sharing underlying parameters and an upper-level gating network. The gating network dynamically outputs the weight coefficients of each objective task with the search scenario identifier as input. When the search scenario is a new product search, the gating network reduces the weight of the quality score and increases the weight of the relevance score. When the search scenario is an inspiration search, the gating network increases the weight of the diversity score. The multi-objective ranking model uses the labeled tags of each objective task as supervision signals and is trained using a multi-task joint loss function and the Adam optimization algorithm.

[0014] Preferably, during the process of displaying search results, explicit and implicit feedback signals from users are collected, and the feedback signals are associated with the corresponding query identifier, candidate material identifier, and sorting position to generate interaction log records. Based on the interaction log records, the weight parameters of the multi-objective ranking model are incrementally updated using the FTRL-Proximal algorithm. The matching weight is increased for material and query combinations with high positive feedback, and the matching weight is decreased for results that users quickly skip. Based on the material co-occurrence relationship data recorded in the interaction log, the parameters of the cross-modal attention fusion module are incrementally trained periodically to reduce the distance between materials with high-frequency co-collection or comparison relationships in the embedding space.

[0015] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: 1. This invention uses cross-modal attention fusion, with visual feature components as the query matrix, to dynamically combine the correlation information in physical and semantic features. This enables the enhanced fusion representation vector to adaptively strengthen the expression of corresponding physical properties and process semantics according to the visual presentation of the material, which helps to improve the accuracy of query matching.

[0016] 2. This invention dynamically infers user search intent, transforms short-term behavior sequences and long-term preference profiles into user intent vectors and integrates them into query vectors, enabling search results to respond to users' implicit needs rather than relying solely on the user's current query text; by introducing diversity scores and penalty items in multi-objective ranking, it helps avoid homogenization of search results, ensuring that top results cover different material types and styles. Attached Figure Description

[0017] Further details, features, and advantages of this application are disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which: Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0018] Several embodiments of this application will now be described in more detail with reference to the accompanying drawings to enable those skilled in the art to implement this application. This application may be embodied in many different forms and for various purposes and should not be limited to the embodiments set forth herein. These embodiments are provided to make this application thorough and complete, and to fully convey the scope of this application to those skilled in the art. The embodiments described do not limit this application.

[0019] Unless otherwise defined, all terms used herein (including technical and scientific terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It will be further understood that terms such as those defined in commonly used dictionaries shall be interpreted as having a meaning consistent with their meaning in the relevant field and / or the context of this specification, and shall not be interpreted in an idealized or overly formal sense unless expressly defined herein.

[0020] Example 1 Its specific implementation method is combined with the appendix Figure 1 A detailed explanation will be provided.

[0021] In this embodiment, it includes: In the field of fabric material search, existing solutions mostly rely on single visual features or simple text tags for matching. Fabric materials have multi-dimensional attributes such as visual, tactile, and craftsmanship, and a single modality representation cannot fully describe the material characteristics, resulting in semantic discrepancies between search results and users' actual needs. In addition, existing solutions lack the ability to dynamically understand users' search intent and cannot infer their deep preferences based on user behavior sequences, further limiting the relevance and diversity of search results.

[0022] A multimodal intelligent search method for fabric materials according to an embodiment of the present invention includes the following steps: Step 1: Extract multimodal features from fabric material samples and generate a unified vectorized representation. Obtain sample data of the fabric materials to be indexed, and extract feature vectors from the visual, physical, and semantic modalities respectively. Map the three types of feature vectors to the same high-dimensional embedding space through a projection layer, concatenate them to generate a unified multimodal embedding vector for the fabric materials, and store it in a vector database.

[0023] Step 1 may specifically include: Step S101: Obtain the front, back, and macro images of the fabric sample, input the images into a pre-trained visual feature extraction network, and output a visual feature vector. The visual feature extraction network uses a pre-trained ResNet or ViT model.

[0024] Step S102: Obtain physical parameter data of the fabric material sample, including weight, thickness, density, yarn count, softness, drape and surface friction coefficient. The above numerical parameters are standardized using Z-score to eliminate dimensional differences, and the standardized parameters are encoded into physical feature vectors.

[0025] Step S103: Obtain semantic label data for the fabric material sample, including material name, composition information, weaving process category, and finishing process category. Input the semantic label data into the word embedding model and output the semantic feature vector. The word embedding model uses Word2Vec or BERT.

[0026] Step S104: Apply linear projection layers to the visual feature vector, physical feature vector, and semantic feature vector respectively, map the three types of vectors to the same-dimensional embedding space, splice them together to form a multimodal embedding vector of fabric material, and store it in the vector database.

[0027] It should be noted that the above physical parameter data is obtained in two ways: firstly, by collecting dynamic parameters such as softness, drape, and surface friction coefficient in real time using a fabric hand feel tester; secondly, by reading static structured fields such as basis weight, thickness, density, and yarn count from an existing material parameter library. The parameters obtained in both ways are standardized using Z-scores and then uniformly encoded into physical feature vectors.

[0028] It should be noted that the above composition information refers to the fiber composition of the fabric and its blending ratio, including cotton, linen, silk, and synthetic fibers, as well as the percentage of each component. Weaving process categories include woven, knitted, jacquard, and printed fabrics. These category labels are converted into numerical vectors using one-time thermal encoding before being used in feature encoding.

[0029] Step 2: Perform cross-modal attention interaction fusion on the multimodal features to generate an enhanced fusion representation vector. The multimodal embedding vector of fabric material is split into visual, physical and semantic components, input into the cross-modal attention fusion module for interactive calculation, and outputs an enhanced fusion multimodal material representation vector to replace the original embedding vector of the corresponding material in the vector database.

[0030] Step 2 may specifically include: Step S201: Use visual feature components as the query matrix The key matrix is ​​obtained by concatenating the physical feature components and the semantic feature components. Sum matrix The cross-modal attention weight distribution is calculated using the following formula: ; in, For querying the matrix, The key matrix, For value matrices, Key matrix Feature dimensions, This represents the matrix transpose operation. For normalized exponential functions, To query the matrix product of the key matrix and its transpose, the output attention weight matrix reflects the degree of attention paid by visual features to each component of physical and semantic features. It should be noted that the visual, physical, and semantic feature components have all been mapped to the same-dimensional embedding space through a linear projection layer in step S104. Therefore, the feature dimensions of the three types of components are consistent, satisfying the dimension matching requirement of the matrix product.

[0031] Step S202: Based on the attention weight matrix, pair the value matrix Weighted aggregation is performed to generate visual enhancement vectors.

[0032] Step S203: Perform residual connection between the visual enhancement vector and the original visual feature components, and output the first layer fusion result after layer normalization.

[0033] Step S204: Use the first-layer fusion result as the new query matrix. Repeat the calculation process from step S201 to step S203. After multiple layers of cross-attention iteration, output the enhanced fusion multimodal material representation vector.

[0034] It should be noted that when the visual features present a rough texture, the attention weight increases at the corresponding positions of the friction coefficient and yarn thickness in the physical features; when the visual features present a complex pattern topology, the attention weight increases at the corresponding position of the weaving process in the semantic features.

[0035] Step 3: Receive user search queries, perform matching calculations based on enhanced fusion representation vectors, and output the search ranking results for fabric materials. The system receives the user's search query input and converts it into a query vector. It calculates the similarity score between the query vector and the enhanced fusion representation vectors of each material in the vector database, sorts them in descending order of similarity score, and outputs the ranked search results for fabric materials.

[0036] Step 3 may specifically include: Step S301: Receive the user's search query input. When the search query input is text keywords, the text keywords are transformed into a query semantic vector through a word embedding model; when the search query input is a reference image, the reference image is transformed into a query visual vector through a visual feature extraction network. The query semantic vector or query visual vector is mapped to the same embedding space as the enhanced fusion representation vector through a projection layer to generate a query vector.

[0037] Step S302: Using the cosine similarity algorithm, with the query vector and the enhanced fusion representation vectors of each candidate material as input, calculate the similarity score between the query vector and the enhanced fusion representation vectors of each material in the vector database, and output a list of similarity scores for each candidate material.

[0038] Step S303: Sort the candidate materials from high to low according to their similarity scores and output the search ranking results.

[0039] In this embodiment of the application, to improve the matching degree between search results and users' deeper needs, a dynamic inference step of user search intent is further included in step S301. The dynamic inference step of user search intent obtains the user's current query text, short-term behavior sequence, and long-term preference profile, inputs them into the Transformer query understanding model, and outputs a user intent vector. Specifically, it includes: Step S304: Perform word segmentation and named entity recognition on the user's current query text, extract key entities such as material composition, color, and purpose, and generate key entity codes.

[0040] Step S305: Obtain the user's short-term behavior sequence, including recent... Search history, browsing history, clicked material icons, and dwell time are used to generate short-term behavior sequence codes; a long-term user preference profile is obtained, including the statistical distribution of material categories in historical favorites and comparison records, to generate a long-term preference profile code. Among these, Numerical behavioral data such as the length and duration of short-term behavioral sequences are standardized using Z-score before being used in the encoding.

[0041] Step S306: Concatenate the key entity encoding, short-term behavior sequence encoding, and long-term preference profile encoding, input them into the Transformer query understanding model, and output a user intent vector. The output layer of the Transformer query understanding model is a fully connected layer, and the output is a fixed-dimensional user intent vector, which includes dimensional information such as material composition preference, style preference, usage scenario, and price range. The Transformer query understanding model uses explicit user feedback labels as supervision signals, employs the cross-entropy loss function, and is trained using the Adam optimization algorithm.

[0042] Step S307: Input the user intent vector into the pre-trained text generation model, and generate extended query term text related to the user intent based on the material component tendency, style preference and usage scenario corresponding to the user intent vector; after encoding the extended query term text by the word embedding model, perform a weighted sum with the original query vector according to the learnable weights to generate an enhanced query vector, which replaces the original query vector for the matching calculation in step S302.

[0043] In this embodiment of the application, in order to balance relevance and diversity in search results, multi-objective ranking is used instead of single similarity ranking, based on steps S302 and S303. Specifically, this includes: Step S308: Using the cosine similarity algorithm, with the enhanced query vector and the enhanced fusion representation vector of each candidate material as input, output a list of relevance scores for each candidate material.

[0044] Step S309: Employing the maximum marginal relevance method, using the enhanced fusion representation vectors of the materials already selected in the result set and the enhanced fusion representation vectors of the current candidate material as input, a penalty term is applied to the similarity between the materials already selected in the result set and the current candidate material. The diversity score of each candidate material is then output, ensuring the search results appear higher. Each location covers different material types and styles. Among them, The number of top positions displayed in the search results.

[0045] Step S310: Based on the historical user click-through rate, collection rate and rating data of each candidate material, mean normalization based on range is used to eliminate the difference in the three types of indicators before weighted summation. The normalized click-through rate, collection rate and rating data are multiplied by three learnable weight coefficients and then summed. The sum of the three weight coefficients is one, and the weighted summation result is the quality score of each candidate material.

[0046] Step S311: Input the relevance score, diversity score, and quality score into the multi-objective ranking model, and output the comprehensive score of each candidate material. The multi-objective ranking model adopts the MMoE architecture, including multiple expert networks sharing underlying parameters and an upper-layer gating network. The gating network takes the search scenario identifier as input and dynamically outputs the weight coefficients of each objective task: when the search scenario is a new product search, the gating network reduces the weight of the quality score and increases the weight of the relevance score; when the search scenario is an inspiration search, the gating network increases the weight of the diversity score. The output layer of the multi-objective ranking model is a fully connected layer, which outputs the comprehensive score scalar of each candidate material. The multi-objective ranking model uses the labeled tags of each objective task as supervision signals, adopts a multi-task joint loss function, and is trained using the Adam optimization algorithm.

[0047] Step S312: Sort the candidate materials from high to low according to the comprehensive score, and output the final search ranking results.

[0048] In this embodiment of the application, in order to continuously improve the performance of the search system with use, a closed-loop online learning step based on user interaction feedback is also included: Step S313: During the search results display process, collect explicit and implicit user feedback signals. Explicit feedback signals include positive reviews, negative reviews, favorites, adding to comparisons, and sharing; implicit feedback signals include mouse hover duration, click depth, scrolling behavior, and dwell time on the search results page.

[0049] Step S314: Associate the explicit feedback signal and the implicit feedback signal with the corresponding query identifier, candidate material identifier and sorting position to generate an interaction log record.

[0050] Step S315: Based on the interaction log records, the weight parameters of the multi-objective ranking model are incrementally updated using the FTRL-Proximal algorithm. The input of the FTRL-Proximal algorithm is the query and material pairings and corresponding feedback labels in the interaction log records, and the output is the updated weights of the multi-objective ranking model. The matching weight is increased for material and query combinations with high positive feedback, and decreased for results that the user quickly skips.

[0051] Step S316: Periodically perform incremental training on the parameters of the cross-modal attention fusion module based on the material co-occurrence relationship data recorded in the interaction log. Inject the material similarity information implicit in user interactions into the enhanced fusion representation, thereby reducing the distance between materials with high-frequency co-collection or contrast relationships in the embedding space.

[0052] Example 2 A home furnishing fabric e-commerce platform (hereinafter referred to as "Platform A") provides fabric material search services for interior designers and home furnishing buyers. Platform A's material library includes various fabric samples such as woven, knitted, and jacquard fabrics. Users can initiate searches using text keywords or reference images. During the spring purchasing season of 20XX, a long-term active designer user (user ID U-0472) was searching for heavyweight jacquard fabrics suitable for curtains for a hotel interior design project. Their past behavior showed a clear preference for silk-cotton blends and Chinese-style fabrics. The following example, using four candidate samples (material IDs M-101 to M-104) from the material library and the user's complete search process, illustrates the specific operation of each step.

[0053] Steps S101 to S104 extract multimodal features from the fabric samples in the material library. Taking candidate sample M-101 (heavyweight jacquard satin) as an example, the multimodal feature extraction module collects raw data from three dimensions. For the visual modality, three images of M-101—front, back, and macro—are acquired, input into a pre-trained ViT model, and the visual feature vector is output. For the physical modality, weight, thickness, density, and yarn count are read from the material parameter library. Softness, drape, and surface friction coefficient are collected using a fabric feel tester. These seven numerical parameters are Z-score standardized and encoded into physical feature vectors. For the semantic modality, the material name, composition information, weaving process category, and finishing process category are acquired. The weaving process category is converted using one-hot encoding and then input into the BERT model to output a semantic feature vector. The three types of feature vectors are mapped to the same embedding space through linear projection layers and then concatenated to generate the multimodal embedding vector of the fabric material M-101, which is then stored in the vector database. The same procedure was performed on the remaining samples M-102 to M-104.

[0054] Table 1 Original physical parameters and semantic labels of candidate material samples ; Steps S201 to S204 perform cross-modal attention fusion on the multimodal embedding vectors of the four samples. Taking M-101 as an example, the cross-modal fusion module splits its multimodal embedding vector into visual, physical, and semantic components. The visual component is used as the query matrix. The key matrix is ​​obtained by concatenating the physical and semantic components. Sum matrix ,according to ; Calculate the attention weight distribution, where For querying the matrix, The key matrix, For value matrices, Key matrix The feature dimensions are as follows: The frontal image of M-101 exhibits a distinct jacquard three-dimensional texture, and the attention weight in the semantic features corresponding to the weaving process (jacquard) increases. Its macro image displays silky luster and a smooth surface, and the attention weight in the physical features corresponding to the surface friction coefficient and softness increases synchronously. After multi-layer cross-attention iteration and residual connections, an enhanced fusion representation vector of M-101 is output, replacing the corresponding record in the vector database.

[0055] Table 2 Comparison of key feature weights before and after cross-modal attention fusion ; The macro images of M-103 exhibit a rough, burr-like texture, with attention weights concentrated on the positions corresponding to the coefficient of friction. This contrasts with the weight distribution of M-101, demonstrating the adaptive adjustment of attention weights guided by visual features.

[0056] Steps S304 to S307 perform dynamic intent inference on user U-0472's search query. The user's current query text is "heavyweight jacquard curtain fabric, silk cotton, Chinese style". The intent inference module first performs word segmentation and named entity recognition, extracting three key entities: material composition (silk cotton), purpose (curtains), and style (Chinese style), generating key entity codes. Then, it obtains the user's short-term behavior sequence (recent...). This record, Numerical behavioral data such as 8) and long-term preference profiles, dwell time, etc., are used in the encoding after being standardized by Z-score.

[0057] Table 3 Short-term behavior sequence and long-term preference profile of user U-0472 ; After concatenating the three types of codes, the input is given to the Transformer query understanding model, which outputs a user intent vector. Among these, the material composition preference (silk-cotton blend), style preference (Chinese jacquard), and usage scenario (hotel curtains) have higher weights. In step S307, the intent vector is input into the text generation model to generate the extended query term "silk-cotton jacquard heavy-weight satin hotel curtains Chinese style with good drape". After being encoded by the word embedding model, this extended query term is weighted and summed with the original query vector according to learnable weights to generate an enhanced query vector.

[0058] Steps S308 to S312 perform multi-objective ranking on the four candidate materials. The query processing module takes the enhanced query vector and the enhanced fusion representation vector of each material as input, and outputs a relevance score using a cosine similarity algorithm. The multi-objective ranking module simultaneously calculates the diversity score and the quality score. The quality score is based on the historical click-through rate, collection rate, and rating data of each material. After eliminating dimensional differences through range-mean normalization, it is weighted and summed with three learnable weight coefficients. In this search scenario, "procurement search," the gating network correspondingly increases the weight of the relevance score.

[0059] Table 4. Candidate Material Multi-Objective Ranking Score and Final Ranking ; M-101 ranked first in both relevance and quality dimensions, achieving the highest overall score and ranking first. Although M-102 had the lowest relevance score, it had the highest diversity score (significantly different in style from the selected M-101), ranking second in overall score under the balance of the multi-objective ranking model, thus avoiding the homogenization phenomenon where all the top results were concentrated in the heavyweight jacquard category.

[0060] Steps S313 to S316 collect user U-0472's interaction feedback for this search and execute closed-loop online learning. The user added M-101 to favorites and comparisons (explicit positive feedback), and spent a relatively long time on the M-101 details page (implicit positive feedback); the user only clicked briefly on M-104 before leaving (implicit negative feedback).

[0061] Table 5 User U-0472 Interaction Feedback Log Record ; The feedback learning module associates the aforementioned interaction log records with corresponding query identifiers, material identifiers, and sorting positions. It then uses the FTRL-Proximal algorithm to incrementally update the weight parameters of the multi-objective ranking model, increasing the matching weight of M-101 with queries related to "heavyweight jacquard silk cotton curtains" and decreasing the matching weight of M-104 with similar queries. Simultaneously, based on the co-occurrence comparison relationship between M-101 and M-102 in the interaction logs, the cross-modal attention fusion module undergoes incremental training, appropriately reducing the distance between them in the embedding space.

[0062] Throughout the implementation process, the data starts from the original multimodal acquisition data of the material library. After three-modal feature extraction and vectorization in steps S101 to S104, a multimodal embedding vector containing visual, physical, and semantic information is formed. After cross-modal attention fusion in steps S201 to S204, the dynamic association between visual presentation and physical and semantic features is explicitly encoded into an enhanced fusion representation vector. The query text of user U-0472 is expanded into an enhanced query vector through intent inference in steps S304 to S307, introducing the user's implicit preferences into the matching calculation. The multi-objective ranking in steps S308 to S312 integrates three types of signals—relevance, diversity, and quality—into the final ranking, with M-101 ranking first due to its high match with the user's intent. Steps S313 to S316 write the feedback of this interaction back to the system, forming a complete data closed loop from search to feedback to model update, ensuring that the ranking results of subsequent similar queries are continuously optimized with the accumulation of data.

[0063] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

[0064] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0065] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A multi-feature fusion AI search optimization method for fabric materials, characterized in that, include: Obtain fabric material sample data, extract feature vectors from visual modality, physical modality and semantic modality respectively, map the three types of feature vectors to the same embedding space through projection layer, splice them to generate fabric material multimodal embedding vector, and store them in vector database; The multimodal embedding vector of the fabric material is split into visual components, physical components and semantic components, and input into the cross-modal attention fusion module. The visual components are used as the query matrix, and the physical components and semantic components are concatenated as the key matrix and value matrix. After attention calculation and residual connection, the enhanced fusion representation vector is output to replace the original embedding vector of the corresponding material in the vector database. The system receives user search query input, converts the search query input into a query vector and maps it to the embedding space, calculates the similarity score between the query vector and the enhanced fusion representation vectors of each material in the vector database, and outputs the search ranking results in descending order of similarity score.

2. The fabric material multi-feature fusion AI search optimization method according to claim 1, characterized in that, The attention calculation process of the cross-modal attention fusion module is as follows: ; in, The query matrix is ​​composed of visual components. The key matrix is ​​formed by concatenating physical and semantic components. For value matrices, The feature dimension of the key matrix; The visual enhancement vector is generated by weighting and aggregating the value matrix based on the attention weight matrix. The visual enhancement vector is then residually connected with the original visual components and processed by layer normalization to output the first layer fusion result. The first layer fusion result is used as a new query matrix, and the enhanced fusion representation vector is output after multi-layer cross-attention iteration.

3. The fabric material multi-feature fusion AI search optimization method according to claim 1, characterized in that, Feature vector extraction from the visual modality includes: acquiring the front, back, and macro images of the fabric sample, inputting the images into a pre-trained visual feature extraction network, and outputting visual feature vectors; feature vector extraction from the physical modality includes: acquiring the weight, thickness, density, yarn count, softness, drape, and surface friction coefficient of the fabric sample, and encoding the above numerical parameters into physical feature vectors after Z-score normalization; feature vector extraction from the semantic modality includes: acquiring the material name, composition information, weaving process category, and finishing process category of the fabric sample, inputting semantic label data into a word embedding model, and outputting semantic feature vectors.

4. The fabric material multi-feature fusion AI search optimization method according to claim 1, characterized in that, Converting the search query input into a query vector includes: when the search query input is a text keyword, converting the text keyword into a query semantic vector through a word embedding model; when the search query input is a reference image, converting the reference image into a query visual vector through a visual feature extraction network; and mapping the query semantic vector or query visual vector to the same embedding space as the enhanced fusion representation vector through a projection layer to generate a query vector.

5. The multi-feature fusion AI search optimization method for fabric materials according to claim 1, characterized in that, After converting the search query input into a query vector, the process also includes a step of dynamically inferring the user's search intent: Perform word segmentation and named entity recognition on the user's current query text, extract key entities and generate key entity codes; Obtain the user's short-term behavior sequence, including the most recent L search records, browsing history, clicked material icons, and dwell time, and generate a short-term behavior sequence code, where L is the length of the short-term behavior sequence; obtain the user's long-term preference profile, including the statistical distribution of material categories in historical collections and comparison records, and generate a long-term preference profile code. The key entity encoding, short-term behavior sequence encoding, and long-term preference profile encoding are concatenated and input into the Transformer query understanding model to output a user intent vector. The user intent vector is input into a pre-trained text generation model to generate extended query term text. The extended query term text is then encoded by a word embedding model and weighted and summed with the original query vector according to learnable weights to generate an enhanced query vector that replaces the original query vector for similarity score calculation.

6. The fabric material multi-feature fusion AI search optimization method according to claim 5, characterized in that, The Transformer query understanding model uses explicit user feedback labels as supervision signals, is trained using the cross-entropy loss function and the Adam optimization algorithm, and has a fully connected output layer that outputs a fixed-dimensional user intent vector.

7. The multi-feature fusion AI search optimization method for fabric materials according to claim 5, characterized in that, After calculating the similarity score, a multi-objective ranking step is also included: The cosine similarity algorithm is used to calculate the correlation score between the enhanced query vector and the enhanced fusion representation vector of each candidate material; The maximum marginal relevance method is used to apply a penalty term to the similarity between the materials already selected in the result set and the current candidate materials, and output the diversity score of each candidate material. Based on the historical user click-through rate, collection rate and rating data of each candidate material, after mean normalization based on the range, the data is multiplied by three learnable weight coefficients and summed. The sum of the three weight coefficients is one, and the quality score of each candidate material is output. The relevance score, diversity score, and quality score are input into the multi-objective ranking model, which outputs a comprehensive score for each candidate material. The final search ranking results are then output in descending order of the comprehensive score.

8. The fabric material multi-feature fusion AI search optimization method according to claim 7, characterized in that, The multi-objective ranking model adopts the MMoE architecture, which includes multiple expert networks sharing underlying parameters and an upper-level gating network. The gating network dynamically outputs the weight coefficients of each objective task with the search scenario identifier as input. When the search scenario is a new product search, the gating network reduces the weight of the quality score and increases the weight of the relevance score. When the search scenario is inspiration search, the gating network increases the weight of diversity score; the multi-objective ranking model uses the labeled tags of each objective task as supervision signals and is trained using a multi-task joint loss function and Adam optimization algorithm.

9. The multi-feature fusion AI search optimization method for fabric materials according to claim 1, characterized in that, During the search results display process, explicit and implicit feedback signals from users are collected, and these feedback signals are associated with the corresponding query identifier, candidate material identifier, and sorting position to generate interaction log records. Based on the interaction log records, the FTRL-Proximal algorithm is used to incrementally update the weight parameters of the multi-objective ranking model, increase the matching weight of material and query combinations with high positive feedback, and decrease the matching weight of results that users quickly jump out of. Based on the material co-occurrence relationship data recorded in the interaction log, the parameters of the cross-modal attention fusion module are incrementally trained periodically to reduce the distance between materials with high-frequency co-collection or comparison relationships in the embedding space.