Cross-modal multi-scale information fusion method and system supporting design knowledge pushing

Through the cross-modal multi-scale information fusion method, multi-scale features are extracted using text and image encoder, and combined with cross-modal attention fusion module and comparison learning, the problem of insufficient deep semantic correlation of existing models in cross-modal information processing is solved, and efficient push of design knowledge and support for innovative design is realized.

CN120408491APending Publication Date: 2025-08-01SICHUAN UNIV

Patent Information

Application Number
CN202510410219.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

Existing cross-modal information processing models such as CLIP lack a targeted cross-modal fusion mechanism when processing complex cross-modal relationships, making it difficult to deeply explore the deep semantic relationships between images and texts, especially inadequate performance in matching features of long texts and fine-grained image, resulting in limited correlation and accuracy of cross-modal information fusion.

Method used

By acquiring and cleaning patent data, multi-scale features are extracted using text encoder and image encoder, combining the cross-modal attention fusion module and the contrast-learning image-text matching loss function, deep semantic interaction and fusion of text and image features are realized, cross-modal fusion features are generated, and cosine similarity is used to calculate the similarity between innovative design concepts and patent information.

Benefits of technology

It significantly improves the accuracy and efficiency of cross-modal information fusion, supports efficient push of design knowledge, improves the relevance and accuracy of image-to-text and text-to-image retrieval, and helps designers obtain accurate knowledge support in innovative design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408491A_ABST
    Figure CN120408491A_ABST
Patent Text Reader

Abstract

The invention provides a cross-modal multi-scale information fusion method and system supporting design knowledge push, and the method comprises the steps: constructing a patent data set containing patent abstracts and abstract drawings, extracting long text features and short text features through a text encoder, and extracting fine-grained and coarse-grained image features through an image encoder; the text and image features are input into a cross-modal attention fusion module, cross-modal information interaction of the text and the image is realized through a multi-head attention mechanism, and fused text and image features are generated; a loss function based on comparative learning is adopted, and the alignment effect of the image and text features in the semantic space is optimized; and inputting an innovative design concept, extracting feature representation of the innovative design concept, calculating the matching degree of the feature representation and patent data based on cosine similarity, and pushing related patent abstracts and attached drawings to designers according to similarity sorting. According to the method, the limitation in existing image-text matching and cross-modal information fusion can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of cross-modal information push, and in particular to a cross-modal multi-scale information fusion method and system for supporting design knowledge push. Background Art

[0002] Current cross-modal information processing tasks are usually solved by deep learning-based models, among which CLIP (Contrastive Language-Image Pretraining) is a widely used model. CLIP achieves image-text alignment by mapping images and texts to the same embedding space. This model uses large-scale image and text data for joint pre-training, has strong generality, and has achieved good performance in various image-text matching tasks. However, CLIP still has limitations in dealing with complex cross-modal relationships.

[0003] First of all, CLIP lacks a targeted cross-modal fusion mechanism and cannot deeply explore the deep semantic associations between images and texts. When there are context differences or implicit information between images and texts, the matching effect of CLIP may be relatively rough, making it difficult to capture complex semantic relationships. Secondly, the training data of CLIP mainly consists of short texts. Although it can process text descriptions, when faced with long texts, its understanding ability is weak, and the extraction of deep semantic information in texts is limited. In addition, as a general model, CLIP is not optimized for specific fields (such as medical image analysis or innovation design field), and its performance in these scenarios may be insufficient.

[0004] Currently, simply adding a fully connected layer to the CLIP model for fine-tuning can adjust the output to adapt to specific tasks, but it still cannot achieve true cross-modal fusion. This method is usually limited to surface feature mapping and is difficult to capture fine-grained cross-modal semantic information. For complex cross-modal tasks, most existing image-text matching methods ignore the interaction and matching of multi-scale features between different modalities, and this deficiency leads to limited relevance and accuracy of the matching.

[0005] Taking patent data as an example, most existing design knowledge push methods based on patent data fail to fully consider the multi-scale feature matching problem between text and images in cross-modal information fusion. This results in insufficient relevance and accuracy of the push results, making it difficult to effectively support the knowledge acquisition and combination needs of designers in innovation design.

[0006] Therefore, there is an urgent need for a method that can achieve cross-modal multi-scale feature fusion, improve the accuracy and efficiency of patent knowledge push, and meet the application requirements of specific fields. Summary of the Invention

[0007] This application aims to solve at least one of the technical problems in the related technologies to some extent.

[0008] To this end, the first object of the present application is to propose a cross-modal multi-scale information fusion method that supports the push of design knowledge.

[0009] The second object of the present application is to propose a cross-modal multi-scale information fusion system that supports the push of design knowledge.

[0010] The third object of the present application is to propose an electronic device.

[0011] The fourth object of the present application is to propose a computer-readable storage medium.

[0012] The fifth object of the present application is to propose a computer program product.

[0013] To achieve the above object, the first aspect embodiment of the present application proposes a cross-modal multi-scale information fusion method that supports the push of design knowledge, including:

[0014] Obtain the original patent data, clean and extract it to form a patent data set including long text, short text, and abstract drawings;

[0015] Use a text encoder to extract the features of the long text and short text respectively, and use an image encoder to extract the fine-grained image features of the abstract drawings, and generate coarse-grained image features through principal component analysis for dimensionality reduction;

[0016] Input the obtained text features and image features into a cross-modal attention fusion module, and realize the information interaction and fusion of text and image features through a multi-head attention mechanism to generate cross-modal fusion features;

[0017] Use a contrastive learning image-text matching loss function to optimize the cross-modal attention fusion module, maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs;

[0018] Input an innovative design concept, extract its feature representation through a cross-modal multi-scale information fusion module, and calculate the similarity between the innovative design concept features and the patent abstracts and drawings using cosine similarity;

[0019] Rank according to the similarity scores, and push the patent abstracts or drawings with a preset ranking to designers to support the combination of design knowledge and the generation of innovative solutions. [[ID=X]]

[0020] Optionally, the obtaining the original patent data, cleaning and extracting it to form a patent data set including long text, short text, and abstract drawings includes:

[0021] Obtain the original patent documents from the patent database through a crawler tool, and the original patent documents include patent abstracts and abstract drawings;

[0022] Clean the patent abstract, including removing redundant characters, special symbols, and content irrelevant to semantics;

[0023] Use the TF-IDF technique to extract keywords from the patent abstract, and select the sentences that contain multiple keywords and are at the forefront in the abstract as short texts;

[0024] Retain all the content of the patent abstract as the long text;

[0025] Match the long text, short text with the corresponding abstract drawings to form a patent dataset containing the long text, short text, and abstract drawings.

[0026] Optionally, the utilization of the text encoder to extract the features of the long text and short text respectively includes:

[0027] Input the long text and the short text into the text encoder respectively, perform word segmentation on the long text and short text, and convert them into fixed-length token sequences. The text encoder is a Transformer-based CLIP text encoder;

[0028] Input the token sequences into the embedding layer, convert each token into an embedding vector of a fixed dimension, and add positional embeddings to encode the position information of each token in the sequence;

[0029] Input the token sequences processed by the embedding layer into a standard Transformer architecture, capture the context relationship between tokens through the multi-head self-attention mechanism, and further extract the features processed by the multi-head self-attention mechanism through a feed-forward neural network;

[0030] Extract the outputs of the first tokens of the long text and short text processed by the text encoder, and map them to a 512-dimensional feature space through a linear projection layer to generate the long text feature H LT and the short text feature H ST .

[0031] Optionally, the utilization of the image encoder to extract the fine-grained image features of the abstract drawings and generate coarse-grained image features through principal component analysis includes:

[0032] Input the abstract drawings into the image encoder, divide the abstract images into multiple fixed-size image patches, and the size of each image patch is 32×32 pixels;

[0033] Flatten each image patch into a 768 - dimensional vector to form a sequence of vectors containing all image patches, and add a global identity vector at the beginning of the vector sequence to aggregate the global information of the entire image;

[0034] Input the image patch vector sequence into the multi - head self - attention mechanism in the image encoder to capture the context relationships between image patches;

[0035] Input the pre - processed image patch vector sequence into the Transformer block, calculate the correlations between image patches using the multi - head self - attention mechanism, and extract the features of each block;

[0036] Input the feature vector processed by the attention mechanism into the feed - forward neural network to further extract the fine - grained image feature H FI ;

[0037] Use the principal component analysis method to perform dimensionality reduction on the fine - grained image feature, extract the main feature components, and generate the coarse - grained image feature H CI ;

[0038] Output the fine - grained image feature and the coarse - grained image feature for subsequent cross - modal information fusion processing.

[0039] Optionally, the cross - modal attention fusion model is composed of three stacked cross - modal attention fusion text modules and three stacked cross - modal attention fusion image modules, where:

[0040] In each cross - modal attention fusion text module, use the text feature H T as the query Q of the multi - head attention mechanism, use the image feature H I as the key K and value V, calculate the query, key, and value vectors through a learnable weight matrix, and the calculation formula is:

[0041] Q = XW Q , K = XW K , V = XW V

[0042] where, W Q , W K , W V are the learnable weight matrices for query, key, and value respectively;

[0043] Calculate the weight matrix of the query and the key through the dot product, normalize it using the softmax function, and then multiply it with the value V to generate the attention weights. The calculation formula is:

[0044]

[0045] where, d kis the dimension of the key vector;

[0046] Concatenate the outputs of multiple attention heads and pass them through the projection matrix W o Project to the specified dimension to generate the output of multi-head attention. The calculation formula is:

[0047] MultiHead(Q, K, V) = Concat(head1,..., head h )W o

[0048] where head i = Attention(QW i Q , KW i K , VW i V ), and h represents the number of heads;

[0049] Pass the output of the multi-head attention through the MLP module to further extract features. The MLP module consists of two fully connected layers, two Dropout layers, and one ReLU activation function. The calculation formula is:

[0050] MLP(x) = d2(d1(max(W1x + b1, 0))·W2 + b2)

[0051] where W1 is the weight of the fully connected layer, b1 and b2 are bias terms, and d1, d2 are Dropout operations;

[0052] Directly add the outputs of the multi-head attention and the MLP module to the input through residual connection to generate the fused text feature F T , and the calculation formula is:

[0053] F T = H T + MultiHead(Q t , K i , V i ) + MLP(H T + MultiHead(Q t , K i , V i ))

[0054] In each of the cross-modal attention fusion image modules, use the image feature H I as the query Q of the multi-head attention mechanism, use the text feature H T as the key K and value V, and generate the fused image feature through the same calculation process as above. The calculation formula is:

[0055] FI = H I + MultiHead(Q i , K t , V t ) + MLP(H I + MultiHead(Q i , K t , V t ))

[0056] The three cross-modal attention fusion image modules and the three cross-modal attention fusion text modules are alternately stacked, and the feature transfer and fusion are maintained through the residual connection technology, for generating the final cross-modal fusion features.

[0057] Optionally, inputting the obtained text features and image features into the cross-modal attention fusion module, and realizing the information interaction and fusion of the text and image features through the multi-head attention mechanism to generate cross-modal fusion features, including:

[0058] Inputting the long text feature H LT , the short text feature H ST , the fine-grained image feature H FI and the coarse-grained image feature H Cl into the cross-modal attention fusion module to obtain the long text feature F LT , the short text feature F ST , the fine-grained image feature F FI and the coarse-grained image feature F CI after cross-modal information fusion;

[0059] Adding the fused long text feature F LT and the short text feature F ST to obtain the total text feature F T , and the calculation formula is:

[0060] F T = F LT + F ST

[0061] Adding the fused fine-grained image feature F FI and the coarse-grained image feature F CI to obtain the total image feature F I , and the calculation formula is:

[0062] F I = F FI + F CI

[0063] Outputting the total text feature F T and the total image feature F I, for subsequent semantic matching and downstream task processing.

[0064] Optionally, optimizing the cross-modal feature fusion using the image-text matching loss function of contrastive learning, maximizing the similarity between positive sample pairs and minimizing the similarity between negative sample pairs, includes:

[0065] Performing L2 normalization on the total text feature F T and the total image feature F I The normalization calculation formula is:

[0066]

[0067] where represents the normalized vector, ‖x‖ represents the L2 norm of vector x, and its calculation formula is:

[0068]

[0069] Constructing positive and negative sample pairs, taking the matching text feature f i T and the image feature f j T in the same batch as the positive sample pair, and taking the non-matching text feature f i T and the image feature f j T in the same batch as the negative sample pair, j≠i;

[0070] Using cosine similarity to calculate the similarity between the text feature vector and the image feature vector, the calculation formula is:

[0071]

[0072] where, u1 = [a1, a2,..., a N and u2 = [b1, b2,..., b N represent the text feature and the image feature respectively, and N is the dimension of the vector;

[0073] Defining an image-text matching loss function based on contrastive learning, where the text-to-image contrast loss and the image-to-text contrast loss The calculation formulas are respectively:

[0074]

[0075] The final total contrast loss function is:

[0076]

[0077] Among them, τ is the temperature parameter used to adjust the smoothness of the distribution; N is the number of samples in the current batch;

[0078] The cross-modal attention fusion model is optimized through the total contrast loss function to enhance the alignment effect of text features and image features in the semantic space.

[0079] Optionally, the cross-modal multi-scale information fusion module includes a text encoder, an image encoder, and a cross-modal attention fusion module. The input innovative design concept extracts its feature representation through the cross-modal multi-scale information fusion module, and the cosine similarity is used to calculate the similarity between the innovative design concept feature and the patent abstract and drawings, including:

[0080] The collected patent abstracts and drawings are input into the cross-modal multi-scale information fusion module for feature extraction to obtain the feature vectors corresponding to each sentence and each image, and an image vector library and a sentence vector library are established;

[0081] The innovative design concept is also input into the cross-modal multi-scale information fusion module for feature extraction to obtain the corresponding innovative design concept vector representation;

[0082] The cosine similarity is used to calculate the similarity between the innovative design concept vector and all image vectors and sentence vectors.

[0083] To achieve the above object, an embodiment of the second aspect of the present application proposes a cross-modal multi-scale information fusion system for supporting design knowledge push, including:

[0084] A preprocessing unit for obtaining the original patent data, cleaning and extracting it to form a patent data set including long texts, short texts, and abstract drawings;

[0085] A feature extraction unit for respectively extracting the features of long texts and short texts by using a text encoder, and extracting the fine-grained image features of abstract drawings by using an image encoder, and generating coarse-grained image features through principal component analysis for dimensionality reduction;

[0086] A fusion unit for inputting the obtained text features and image features into a cross-modal attention fusion module, and realizing the information interaction and fusion of text and image features through a multi-head attention mechanism to generate cross-modal fusion features;

[0087] An optimization unit for optimizing the cross-modal attention fusion module by using a contrastive learning image-text matching loss function to maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs;

[0088] A similarity calculation unit, configured to input an innovative design concept, extract its feature representation through a cross-modal multi-scale information fusion module, and calculate the similarity between the innovative design concept features and the patent abstract and drawings by using cosine similarity;

[0089] A knowledge push unit, configured to push patent abstracts or drawings with a preset ranking to designers according to the similarity score ranking, for supporting the combination of design knowledge and the generation of innovative solutions.

[0090] To achieve the above object, an embodiment of the third aspect of the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0091] The memory stores computer execution instructions;

[0092] The processor executes the computer execution instructions stored in the memory to implement the method according to any one of the first aspect.

[0093] To achieve the above object, an embodiment of the fourth aspect of the present application provides a computer-readable storage medium, in which computer execution instructions are stored, and when the computer execution instructions are executed by a processor, they are used to implement the method according to any one of the first aspect.

[0094] To achieve the above object, an embodiment of the fifth aspect of the present application provides a computer program product, which implements the method according to any one of the first aspect when executed by a processor.

[0095] The technical solutions provided by the embodiments of the present application at least bring the following beneficial effects:

[0096] By aligning and fusing text and image features, the deficiencies in existing image-text matching and cross-modal information fusion are effectively solved, especially the problem of insufficient complex and fine-grained information processing capabilities; by combining long text and short text, and fine-grained and coarse-grained image features, the accuracy of feature extraction and fusion is optimized, the efficient push of innovative design knowledge in patent data is realized, the relevance and accuracy of image-to-text and text-to-image retrieval are significantly improved, providing accurate and efficient knowledge support for designers, helping them to inspire inspiration and combine concepts in the initial stage of design, and promoting the generation of innovative design solutions and the progress of the design field.

[0097] The additional aspects and advantages of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. Description of the Drawings

[0098] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where:

[0099] Figure 1 Schematic diagram of the simple process of a cross-modal multi-scale information fusion method provided by an embodiment of the present application for supporting the push of design knowledge;

[0100] Figure 2 Detailed schematic diagram of a cross-modal multi-scale information fusion method provided by an embodiment of the present application for supporting the push of design knowledge;

[0101] Figure 3 Schematic diagram of the structure of a cross-modal attention fusion module provided by an embodiment of the present application;

[0102] Figure 4 Schematic diagram of the feature fusion of long and short texts and thick and fine-grained images provided by an embodiment of the present application;

[0103] Figure 5 Schematic diagram of the structure for text and image feature extraction provided by an embodiment of the present application

[0104] Figure 6 Schematic diagram of the calculation of the multi-head attention mechanism provided by an embodiment of the present application;

[0105] Figure 7 Schematic diagram of the calculation principle of the cosine similarity provided by an embodiment of the present application;

[0106] Figure 8 Schematic diagram of the structure of a cross-modal multi-scale information fusion system provided by an embodiment of the present application for supporting the push of design knowledge. Detailed implementation manners

[0107] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions from beginning to end. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application, but should not be construed as limiting the present application.

[0108] In view of the limitations in existing image-text matching and cross-modal information fusion, the embodiments of the present application provide a cross-modal multi-scale information fusion method for supporting the push of design knowledge, Figure 1 and Figure 2 Schematic diagram of the process of a cross-modal multi-scale information fusion method provided by an embodiment of the present application for supporting the push of design knowledge. Referring to Figure 1 and Figure 2 , the method includes the following steps:

[0109] Step S1, obtaining the original patent data, cleaning and extracting it to form a patent data set including long texts, short texts and abstract drawings.

[0110] This step aims to provide high-quality input data for the training and testing of cross-modal multi-scale information fusion models, ensure the integrity and standardization of the data through preprocessing, and lay the data foundation for model training.

[0111] Step S1 specifically includes the following steps:

[0112] First, the original patent documents are obtained from the patent database using a crawler tool. The original patent documents include patent abstracts and abstract drawings.

[0113] In an example of this application, all patent data was crawled from the patent database of the United States Patent and Trademark Office (USPTO) authorized in 2023, and 290,225 valid patent documents were collected. Each patent document contains a patent abstract and at least one corresponding abstract figure.

[0114] The obtained patent abstracts are then cleaned to remove redundant characters (such as extra spaces and line breaks), special symbols (such as non-linguistic characters), and redundant information unrelated to the patent subject. The purpose of cleaning is to ensure the standardization of patent abstract content and provide high-quality data input for subsequent text processing and feature extraction.

[0115] Next, we use TF-IDF technology to extract keywords from the cleaned patent abstracts. We calculate the importance of each word based on its frequency in the document and its inverse document frequency across all documents. After extraction, we select sentences containing multiple keywords at the top of the abstract as short text to highlight the key content of the patent abstract.

[0116] Furthermore, the full content of the patent abstract is retained as long text. This provides comprehensive semantic context for subsequent training, compensating for details that may be missed in short text.

[0117] Finally, the generated short texts, retained long texts and corresponding abstract drawings are matched to form a patent dataset containing long texts, short texts and abstract drawings.

[0118] In one example of this application, the constructed dataset contains 290,225 patent documents, of which 10,000 were randomly selected as the test set, and the remaining data was used to train a cross-modal multi-scale information fusion model. Through these processes, the structural integrity and high quality of the dataset are ensured, providing support for subsequent model training.

[0119] In step S2, the features of the long text and the short text are extracted respectively by using a text encoder, and the fine-grained image features of the abstract drawings are extracted by using an image encoder, and coarse-grained image features are generated by dimensionality reduction through principal component analysis.

[0120] In the embodiments of the present application, by using a text encoder to extract the features of long text and short text respectively, and using an image encoder to extract the fine-grained image features of the abstract drawing, and generating coarse-grained image features through principal component analysis for dimensionality reduction, high-quality multi-scale feature inputs are provided for subsequent cross-modal information fusion.

[0121] Step S2 specifically includes the following two parts:

[0122] (1) Text feature extraction process:

[0123] First, as Figure 4 shown, the long text and short text are respectively input into the CLIP text encoder based on the Transformer architecture, and the structure of the text encoder is as Figure 5 shown. The goal of text feature extraction is to capture multi-scale semantic information in the text, where the long text is used to provide rich global semantics with context, and the short text highlights the key information of the text.

[0124] Then, the input long text and short text are tokenized and transformed into a fixed-length token sequence. The purpose of tokenization is to break down the text into basic language units so that the model can learn semantic relationships word by word.

[0125] Next, the generated token sequence is input into the embedding layer, and the embedding layer maps each token to an embedding vector of a fixed dimension (usually 512 dimensions). At the same time, a position embedding is added to each token to encode its position information in the sequence, thereby helping the model perceive the sequential structure of the text.

[0126] The embedded token sequence enters the Transformer architecture. The Transformer includes the following key modules: multi-head self-attention mechanism: captures the context relationships between tokens and identifies global and local semantic associations in long text and short text; feed-forward neural network: further processes the features obtained through the multi-head attention mechanism to enhance semantic expression ability; residual connection and layer normalization: stabilizes the feature extraction process, avoids the problem of gradient vanishing, and improves training efficiency at the same time.

[0127] Finally, the first token output of the sequence processed by the Transformer architecture is extracted, representing the semantic information of the entire input text. This output is mapped to a 512-dimensional feature space through a linear projection layer to generate the long text feature H LT and the short text feature H ST . Among them, the long text feature provides comprehensive semantic coverage, while the short text feature highlights key information, providing multi-scale semantic support for subsequent multi-modal fusion.

[0128] (2) Image feature extraction process:

[0129] Similarly, first input the abstract drawing into the image encoder. The goal is to extract fine-grained global features and coarse-grained main features from the image to achieve the expression of multi-scale image semantic information.

[0130] Next, preprocess the abstract drawing by dividing the image into multiple image patches of a fixed size. The size of each image patch is 32×32 pixels. The purpose of image patch division is to decompose the entire image into local regions so that the model can learn the spatial features of the image block by block.

[0131] Then, flatten each image patch into a 768-dimensional vector to form a vector sequence containing all image patches. At the same time, add a global identification vector at the head of the vector sequence to aggregate the global information of the entire image and ensure that the model can learn the semantic features of the whole image.

[0132] Then input the generated image patch vector sequence into the Transformer block in the image encoder. The subsequent operations are the same as those of the text encoder, specifically including the following steps: use the multi-head self-attention mechanism to capture the context relationship between image patches and identify the spatial correlation of each image patch; input the feature vector processed by the attention mechanism into the feed-forward neural network to further extract the fine-grained image feature H FI .

[0133] In addition, to obtain a more abstract image representation, the embodiment of this application uses the principal component analysis method (PCA) to reduce the dimension of the fine-grained image features, extract the main feature components, and generate the coarse-grained image feature H CI . The fine-grained features provide detailed image region information, while the coarse-grained features highlight the overall semantic information of the image.

[0134] Finally, the image encoder outputs the fine-grained image features and the coarse-grained image features, providing multi-scale image feature support for subsequent cross-modal information fusion.

[0135] In the embodiment of this application, through the joint processing of the above text encoder and image encoder, multi-scale text and image features are respectively generated. The text features include the long text feature H LT and the short text feature H ST , and the image features include the fine-grained image feature H FI and the coarse-grained image feature H CI . These features can express semantic information at different granularities and provide comprehensive and high-quality input feature support for the subsequent cross-modal information fusion module.

[0136] Step S3: Input the obtained text features and image features into the cross-modal attention fusion module. Through the multi-head attention mechanism, implement the information interaction and fusion between the text and image features to generate cross-modal fusion features.

[0137] In the embodiment of the present application, the extracted text features and image features are input into the cross-modal attention fusion module. Through the multi-head attention mechanism, implement the deep semantic interaction and fusion between the text and image features to generate cross-modal fusion features, providing high-quality semantic expressions for multi-modal tasks. This step can not only perform deep semantic alignment on the text and image features, but also effectively retain the multi-scale information of the input features.

[0138] Next, refer to Figure 3 , and elaborate on the cross-modal attention fusion module in detail:

[0139] The cross-modal attention fusion module is alternately stacked by three cross-modal attention fusion text modules (CMAFT) and three cross-modal attention fusion image modules (CMAFI). Among them: The role of the CMAFT module is to mainly use text features and combine image features for fusion, enabling the model to learn the supplementary information of the image to the text semantics; the role of the CMAFI module is to mainly use image features and combine text features for fusion, enabling the model to capture the supplementary information of the text to the image semantics. Through this cross-modal information fusion, CMAFT and CMAFI can effectively achieve the mutual mapping and understanding between text and image, and are applicable to multi-modal tasks that need to process text and image simultaneously. Moreover, each module is composed of a multi-head attention mechanism and a multi-layer perceptron (MLP) module, with efficient semantic capture and feature interaction capabilities. The two modules are alternately stacked, and the stability of feature transmission and fusion is maintained through residual connections, finally generating high-quality cross-modal fusion features.

[0140] The attention mechanism is used to calculate the correlation between each word in the input sequence and other words, and use these correlations as weights to generate the weighted representation of each word. Figure 6 shows the process of a word vector sequence input into the attention mechanism to obtain the context representation. Specifically, assume the input vector matrix X, where each row represents the vector representation of a word. There are three weight matrices to be learned in the model, namely W Q 、W K 、W V , which are multiplied by X to obtain three vectors in the attention mechanism: the query vector Q, the key vector K, and the value vector V.

[0141] For the cross-modal attention fusion text module CMAFT, it uses the text feature H T as the query Q of the multi-head attention mechanism, and uses the image feature H IAs the key K and value V, the calculation formula is:

[0142] Q = XW Q , K = XW K , V = XW V

[0143] where W Q , W K , W V are learnable weight matrices for queries, keys, and values respectively.

[0144] Next, calculate the weight matrix W of the query and the key through dot product, normalize it using the softmax function, and then multiply it by V to finally obtain the weighted vector sequence. The calculation formula is:

[0145]

[0146] where d k is the dimension of the key vector, used to smooth the distribution of the similarity matrix. In this way, the attention mechanism can effectively capture the dependencies between words in the input sequence, and then generate a weighted representation with context information.

[0147] Multi-head attention projects the input sequence onto different heads respectively. Each head calculates a set of attention weights, then concatenates all the weights, and projects them onto a low-dimensional space with a specified dimension for output. Therefore, concatenate the outputs of multiple attention heads and project them through the projection matrix Q o onto the specified dimension to generate the final output of multi-head attention. The calculation formula is:

[0148] MultiHead(Q, K, V) = Concat(head1,..., head h )W o

[0149] where head i = Attention(QW i Q , KW i K , VW i V ), and h represents the number of heads. Each attention head independently learns different semantic correlations, and finally realizes multi-perspective feature expression through concatenation.

[0150] Next, input the output of multi-head attention into the MLP module to further extract non-linear features. The MLP module consists of two fully connected layers, two Dropout layers, and a ReLU activation function. The calculation formula is as follows:

[0151] MLP(x) = d2(d1(max(W1x + b1, 0))·W2 + b2)

[0152] Among them, Q1 is the weight of the fully connected layer, b1 and b2 are bias terms, and d1 and d2 are Dropout operations used to prevent the model from overfitting.

[0153] In addition, to maintain the transitivity of features and the training stability of the model, the residual connection technology is used in each module in the embodiments of the present application. The residual connection directly adds a part of the input to the corresponding output, which can make the model easier to learn the identity mapping, avoid the problem of gradient disappearance, improve the training stability, enable the model to be stacked to deeper layers, and further enhance its expression ability.

[0154] In the CMAFT module, the calculation formula for the fused text features is:

[0155] F T = H T + MultiHead(Q t , K i , V i ) + MLP(H T + MultiHead(Q t , K i , V i ))

[0156] The calculation process of the CMAFI module is similar to that of the CMAFT module, except that the modalities of the input and query are swapped. Specifically, in each cross-modal attention fusion image module, the image feature H I is used as the query Q of the multi-head attention mechanism, and the text feature H T is used as the key K and value V, and the fused image feature is generated through the same above calculation process. The calculation formula is:

[0157] F I = H I + MultiHead(Q i , K t , V t ) + MLP(H I + MultiHead(Q i , K t , V t ))

[0158] Generally speaking, three cross-modal attention fusion image modules and three cross-modal attention fusion text modules are stacked alternately, and the feature transfer and fusion are maintained through the residual connection technology. The CMAFT module enables the text features to combine with the image semantic information, achieving deep alignment between long texts and images. The CMAFI module enables the image features to combine with the text semantic information, achieving deep fusion between fine-grained image features and short texts. Finally, the stacked fusion features contain multi-scale and multi-modal information, providing a more accurate semantic expression for subsequent tasks.

[0159] During application, the long text feature H LT , short text feature H ST , fine-grained image feature H GI and coarse-grained image feature H CI are input into the cross-modal attention fusion module to obtain the long text feature F LT , short text feature F ST , fine-grained image feature F FI and coarse-grained image feature F CI after cross-modal information fusion. Further, the fused long text feature F LT and short text feature F ST are added together to obtain the total text feature F T . The fused fine-grained image feature F FI and coarse-grained image feature F CI are added together to obtain the total image feature F I . The calculation formula is:

[0160] F T = F LT + F ST

[0161] F I = F FI + F CI

[0162] Finally, the total text feature F T and total image feature F I are output for subsequent semantic matching and downstream task processing.

[0163] Step S4: Optimize the cross-modal attention fusion module using the contrastive learning image-text matching loss function to maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs.

[0164] In the embodiments of the present application, a contrastive learning-based image-text matching loss function is used to optimize the cross-modal attention fusion module to achieve the alignment of text features and image features in the semantic space. This method maximizes the similarity between positive sample pairs while minimizing the similarity between negative sample pairs, thereby realizing the cross-modal semantic consistency modeling and significantly improving the matching performance and robustness of the model.

[0165] First, perform L2 normalization on the total text feature F T and the total image feature F I Normalization is a key step in standardizing feature vectors, which can effectively improve the stability of model training and enhance the efficiency of numerical calculations. The calculation formula for L2 normalization is:

[0166]

[0167] where represents the normalized vector, and ‖x‖ represents the L2 norm of vector x, and its calculation formula is:

[0168]

[0169] Through normalization, the vector lengths of text features and image features are standardized to 1, thus ensuring the comparability of different modal features in similarity calculation.

[0170] Next, construct positive and negative sample pairs for contrastive learning. In the embodiments of the present application, the matching text feature f i T and image feature f j T in the same batch are used as positive sample pairs, and the unmatched text feature f i T and image feature f j T in the same batch are used as negative sample pairs, where j≠i. By distinguishing positive and negative sample pairs, the model can learn the semantic correlation between cross-modal features.

[0171] To measure the semantic correlation between text features and image features, the embodiments of the present application use cosine similarity as the metric. The larger the cosine value, the higher the similarity between the image and the text. Assume that the text feature vector is u1 = [a1, a2,..., a N , the image feature vector is u2 = [b1, b2,..., b N , and the dimension of the vector is N. The calculation formula for the cosine similarity between the two is:

[0172]

[0173] Among them, the larger the value of the cosine similarity, the stronger the semantic correlation between the text features and the image features.

[0174] Figure 7 Taking two-dimensional vectors as an example, a method for measuring the similarity between two vectors by cosine similarity is shown as in the figure: u1 is more similar to u2.

[0175] To further optimize the performance of the model, the embodiments of this application construct an image-text matching loss function based on contrastive learning, and calculate the contrastive loss from text to image and the contrastive loss from image to text respectively. The contrastive loss from text to image and the contrastive loss from image to text are calculated by the following formulas respectively:

[0176]

[0177] The final total contrastive loss function is defined by combining the above two losses as:

[0178]

[0179] Among them, τ is the temperature parameter, which is used to adjust the smoothness of the distribution; N is the number of samples in the current batch.

[0180] The optimization goal of this loss function is to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs at the same time, so as to ensure the deep alignment of text and image features in the semantic space.

[0181] Through the above training process of contrastive learning, the model can capture the cross-modal semantic associations between text and image features more accurately. The introduction of the total contrastive loss function not only significantly improves the matching effect of text and image features, but also enhances the generalization ability of the model, providing a more accurate and efficient semantic basis for the push of innovative design knowledge.

[0182] As an example, the experimental environment uses an NVIDIA GTX3090 GPU to train the cross-modal attention fusion module to ensure the efficiency of large-scale model calculations. The optimizer selects the Adaptive Moment Estimation optimizer (Adam), and its excellent dynamic learning rate adjustment ability can significantly improve the convergence speed of the model. The number of training epochs is set to 50, the batch size is 64, and the initial learning rate is set to 0.001 to gradually optimize the model in a large parameter space.

[0183] To further improve the stability of model training, the embodiments of the present application also set up a learning rate scheduler. The strategy adopted by the scheduler is that when the loss value does not decrease for three consecutive epochs, the learning rate is automatically reduced to 0.2 times the current value until the training ends. This dynamic adjustment mechanism can effectively prevent the model from falling into local optima and maintain control over gradient vanishing.

[0184] The specific hyperparameter settings are shown in Table 1. The specific parameter settings for each module of the MCMAF model are shown in Table 2.

[0185] Table 1

[0186]

[0187] Table 2

[0188]

[0189]

[0190] Step S5: Input the innovative design concept, extract its feature representation through the cross-modal multi-scale information fusion module, and calculate the similarity between the innovative design concept features and the patent abstract and drawings using cosine similarity.

[0191] In the embodiments of the present application, by inputting the innovative design concept and using the cross-modal multi-scale information fusion module to extract its feature representation, the semantic similarity between it and the patent abstract and drawings is further calculated. By constructing a semantically aligned vector library and an accurate similarity calculation mechanism, the matching support for the innovative design concept and the existing patent information is realized. The cross-modal multi-scale information fusion module includes a text encoder, an image encoder, and a cross-modal attention fusion module.

[0192] First, the collected patent abstracts and abstract drawings are input into the cross-modal multi-scale information fusion module to extract the feature representations of each sentence and each image. The patent abstracts are processed by the text encoder to extract long-text features and short-text features, and the abstract drawings are processed by the image encoder to extract fine-grained features and coarse-grained features. These features are interacted and integrated through the cross-modal attention fusion module, and finally, text vectors and image vectors are generated and stored in the sentence vector library and the image vector library respectively. The establishment of the vector library provides a comprehensive semantic representation for subsequent similarity calculations.

[0193] Next, the designer obtains innovative design concepts from the Internet or real life, providing an inspiration source for the system input. The input innovative design concepts are entered into the cross-modal multi-scale information fusion module in text form. Similar to the processing method of patent abstracts, the innovative design concepts are extracted with semantic features by the text encoder and further generate their corresponding vector representations through the cross-modal attention fusion module. The innovative design concept vectors can capture their potential semantic information, providing semantic expressions for the subsequent matching process.

[0194] Finally, the cosine similarity is used to calculate the similarity between the innovative design concept vectors and each feature vector in the sentence vector library and the image vector library. The higher the similarity, the stronger the semantic relevance between the innovative design concept and the patent abstract or the attached drawing.

[0195] Through the above steps, the innovative design concept vectors and the features in the vector library can be sorted based on the cosine similarity, so as to provide the designer with the patent abstracts or attached drawings most relevant to the innovative design concept. This process effectively supports the precise push and combination of design knowledge, helps the designer obtain inspiration from the existing patent knowledge, and quickly generate innovative design solutions.

[0196] Step S6: According to the similarity score ranking, push the patent abstracts or attached drawings with a preset ranking to the designer to support the combination of design knowledge and the generation of innovative solutions.

[0197] In the embodiment of the present application, based on the cosine similarity scores calculated in the previous step, the semantic relevance between the innovative design concept and the patent abstracts and attached drawings is ranked. According to the magnitude of the similarity scores, the top n abstract sentences or abstract attached drawings are selected and pushed to the designer. These pushed patent design knowledge can highly meet the requirements of the innovative design concept, providing multi-dimensional inspiration sources and reference materials for the designer.

[0198] After receiving the pushed patent abstracts and attached drawings, the designer can combine the relevant knowledge according to the actual design requirements. By comprehensively using this highly relevant design knowledge, the designer can quickly conceive a complete innovative design solution. This solution is highly consistent with the existing patent knowledge semantically, and at the same time, combined with the creative thinking of the designer, further optimizes the innovative results, thus realizing the efficient utilization of design knowledge and the rapid generation of innovative solutions.

[0199] The above steps achieve precise knowledge push based on semantic similarity, significantly improving the relevance and practicality of design knowledge recommendation, providing strong support for the work of designers, and accelerating the realization of design innovation.

[0200] In addition, after training the cross-modal multi-scale information fusion module, the present application conducted experimental tests on the 10,000 patent image-text pair test sets selected in step S1 to verify the effectiveness and performance of the module proposed in the present application.

[0201] In the image-text matching task, to comprehensively measure the accuracy and effectiveness of the model in cross-modal alignment and retrieval, the present application selected R@k as the main evaluation metric. R@k represents the proportion of queries in which at least one correct match appears among the top k retrieval results for all queries. Specifically, if the top k results returned by the model for each query contain the correctly matched text or image, the query is considered successful, and the performance of the model is evaluated by calculating the proportion of successful queries.

[0202] This experiment compared several methods on the patent test set, including CLIP, CLIP-linear, CMAF, loss addition, pre-fusion, and MCMAF, and the results are shown in Table 3.

[0203] Table 3

[0204]

[0205] The results show that the cross-modal multi-scale information fusion method (MCMAF) proposed in the present application outperforms other methods in all metrics of R@k for all tasks (text-to-image retrieval and image-to-text retrieval). Specifically, the CLIP method directly uses the text encoder and image encoder for testing without any fusion or optimization; the CLIP-linear method adds a fully connected layer on the basis of the CLIP model and is tested after training, with its performance slightly improved compared to CLIP; the CMAF method introduces a cross-modal information fusion module, which, although not fusing long and short texts with fine and coarse-grained image features, significantly outperforms the original CLIP model in all tasks, verifying the effectiveness of the cross-modal attention fusion module.

[0206] In addition, to further improve the performance of the CMAF model, the embodiments of the present application made two improvement attempts on CMAF. One is to use the loss addition method to combine multiple loss terms to enhance the feature alignment ability, i.e., CMAF-Loss in the table; the other is to fuse long and short texts with fine and coarse-grained image features before the features are input into the cross-modal information fusion module, i.e., CMAF-fusion in the table. These two methods have some improvements in some metrics, but ultimately the MCMAF method proposed in the present application significantly improves the model performance through multi-scale fusion. MCMAF combines long and short texts with fine and coarse-grained image features, optimizing the multi-modal semantic alignment ability of the model. The experimental results show that it performs excellently in all R@k metrics, especially achieving significant performance improvements in R@1 and R@10.

[0207] To verify the universality of the method of this application, this application also conducted training and testing on two classification datasets, CIFAR10 and CIFAR100, as well as COCO2017 (5k data). The results are shown in Table 4.

[0208] Table 4

[0209]

[0210]

[0211] The experimental results show that compared with the original CLIP, the classification accuracy of both CMAF and MCMAF has been significantly improved, but the performance gap between the two is small. MCMAF performs very strongly on CIFAR-10, approaching an accuracy of 92%, and also performs well on CIFAR-100, reaching an accuracy of 71.96%.

[0212] Meanwhile, to verify the performance of the model in the long text and image matching retrieval task, the embodiment of this application also merged all the descriptions of each picture in the COCO dataset into a long sentence as the text description of the corresponding image, and conducted an experimental evaluation of the R@k metric. The results show that MCMAF is significantly better than other methods in most metrics.

[0213] From the above experimental results, it can be seen that the cross-modal multi-scale information fusion method (MCMAF) proposed in this application demonstrated excellent adaptability and effectiveness in the experimental tests on the patent test set and public datasets. This application can not only significantly improve the performance of the image-text matching task, but also be widely applicable to multi-modal tasks in different datasets, providing an innovative technical solution for cross-modal semantic alignment and information fusion.

[0214] To implement the above embodiment, this application also proposes a cross-modal multi-scale information fusion system that supports the push of design knowledge. Figure 8 The structural schematic diagram of a cross-modal multi-scale information fusion system 10 provided for the embodiment of this application. As Figure 8 shown, the system includes:

[0215] A preprocessing unit 100, configured to obtain the original patent data, clean and extract it, and form a patent dataset containing long text, short text, and abstract drawings;

[0216] A feature extraction unit 200, configured to respectively extract the features of the long text and short text using a text encoder, and extract the fine-grained image features of the abstract drawings using an image encoder, and generate coarse-grained image features through principal component analysis for dimensionality reduction;

[0217] The fusion unit 300 is configured to input the obtained text features and image features into a cross-modal attention fusion module, and realize the information interaction and fusion of text and image features through a multi-head attention mechanism to generate cross-modal fusion features;

[0218] The optimization unit 400 is configured to optimize the cross-modal attention fusion module using a contrastive learning image-text matching loss function, maximize the similarity between positive sample pairs, and minimize the similarity between negative sample pairs;

[0219] The similarity calculation unit 500 is configured to input an innovative design concept, extract its feature representation through a cross-modal multi-scale information fusion module, and calculate the similarity between the innovative design concept features and the patent abstracts and drawings using cosine similarity;

[0220] The knowledge push unit 600 is configured to push patent abstracts or drawings with a preset ranking to designers according to the similarity score ranking, for supporting the combination of design knowledge and the generation of innovative solutions.

[0221] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0222] To implement the above embodiments, the present application also provides an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiments.

[0223] To implement the above embodiments, the present application also provides a computer-readable storage medium storing computer execution instructions, and when the computer execution instructions are executed by a processor, they are used to implement the method provided in the foregoing embodiments.

[0224] To implement the above embodiments, the present application also provides a computer program product including a computer program, and when the computer program is executed by a processor, it implements the method provided in the foregoing embodiments.

[0225] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the present application and other processing all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0226] It should be noted that personal information from users should be collected for legal and reasonable purposes and should not be shared or sold outside of these legitimate uses. Additionally, such collection / sharing should be carried out after obtaining the informed consent of the users, including but not limited to notifying the users to read the user agreement / user notice and sign an agreement / authorization covering the authorization of relevant user information before the users use the function. Moreover, any necessary steps should be taken to safeguard and protect access to such personal information data and ensure that others with access to the personal information data comply with their privacy policies and procedures.

[0227] This application is expected to provide embodiments where users can selectively block the use or access of personal information data. That is, this disclosure is expected to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, the risks can be minimized by restricting data collection and deleting the data. Additionally, when applicable, personal identifiers are removed from such personal information to protect the privacy of the users.

[0228] In the description of the foregoing embodiments, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of this application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. Additionally, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0229] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of this application can be achieved, and no limitations are imposed herein.

[0230] The above specific implementation manners do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A cross-modal multi-scale information fusion method for supporting the push of design knowledge, characterized in that Including the following steps: Obtain the original patent data, clean and extract it to form a patent dataset containing long texts, short texts, and abstract drawings; Use a text encoder to extract the features of the long text and short text respectively, and use an image encoder to extract the fine-grained image features of the abstract drawing, and generate coarse-grained image features through principal component analysis for dimensionality reduction; Input the obtained text features and image features into a cross-modal attention fusion module, and realize the information interaction and fusion of text and image features through the multi-head attention mechanism to generate cross-modal fusion features; Use the image-text matching loss function of contrastive learning to optimize the cross-modal attention fusion module, maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs; Input the innovative design concept, extract its feature representation through the cross-modal multi-scale information fusion module, and calculate the similarity between the innovative design concept feature and the patent abstract and drawings using cosine similarity; Rank according to the similarity score, and push the patent abstracts or drawings with a preset ranking to the designer to support the combination of design knowledge and the generation of innovative solutions.

2. The method according to claim 1, wherein The obtaining of the original patent data, cleaning and extraction to form a patent dataset containing long texts, short texts, and abstract drawings includes: Obtain the original patent documents from the patent database through a crawler tool, and the original patent documents include patent abstracts and abstract drawings; Clean the patent abstracts, including removing redundant characters, special symbols, and content irrelevant to semantics; Use the TF-IDF technology to extract keywords from the patent abstracts, and select the sentences containing multiple keywords and ranked in the front of the abstract as short texts; Keep all the content of the patent abstract as the long text; Match the long text, short text with the corresponding abstract drawing to form a patent dataset containing the long text, short text, and abstract drawing.

3. The method according to claim 2, wherein The using of the text encoder to extract the features of the long text and short text respectively includes: Input the long text and the short text into the text encoder respectively, perform word segmentation on the long text and short text, and convert them into fixed-length token sequences. The text encoder is a CLIP text encoder based on Transformer; Input the token sequence into the embedding layer, convert each token into a fixed-dimensional embedding vector, and add position embeddings at the same time to encode the position information of each token in the sequence; Input the token sequence processed by the embedding layer into a standard Transformer architecture, capture the context relationship between tokens through the multi-head self-attention mechanism, and further extract the features processed by the multi-head self-attention mechanism through a feed-forward neural network; Extract the outputs of the first tokens of the processed long text and short text by the text encoder, and map them to a 512-dimensional feature space through a linear projection layer to generate long text feature H LT and short text feature H ST .

4. The method according to claim 3, characterized in that, The using of the image encoder to extract the fine-grained image features of the abstract drawing and generate coarse-grained image features through principal component analysis for dimensionality reduction includes: Input the abstract drawing into the image encoder, divide the abstract image into multiple fixed-size image patches, and the size of each image patch is 32×32 pixels; Flatten each image patch into a 768-dimensional vector to form a sequence of vectors containing all image patches, and add a global identity vector at the head of the vector sequence to aggregate the global information of the entire image; Input the image patch vector sequence into the multi-head self-attention mechanism in the image encoder to capture the context relationship between image patches; Input the preprocessed image patch vector sequence into the Transformer block, use the multi-head self-attention mechanism to calculate the correlation between image patches, and extract the features of each block; Input the feature vector processed by the attention mechanism into the feedforward neural network to further extract the fine-grained image feature H FI ; Dimensionality reduction processing is performed on the fine-grained image features by using the principal component analysis method, the main feature components are extracted, and the coarse-grained image feature H is generated CI ; Output the fine-grained image features and the coarse-grained image features for subsequent cross-modal information fusion processing.

5. The method according to claim 4, wherein The cross-modal attention fusion model is composed of three stacked cross-modal attention fusion text modules and three cross-modal attention fusion image modules, where: In each of the cross-modal attention fusion text modules, the text feature H T is used as the query Q of the multi-head attention mechanism, and the image feature H I is used as the key K and value V. Query, key, and value vectors are calculated through a learnable weight matrix, and the calculation formula is: Q = XW Q , K = XW K , V = XW V Among them, W Q , W K , W V are learnable weight matrices for query, key, and value respectively; Calculate the weight matrix of the query and the key through dot product, and use the softmax function to normalize it and then multiply it with the value V to generate the attention weight. The calculation formula is: where d k is the dimension of the key vector; Concatenate the outputs of multiple attention heads and project them through the projection matrix W o Project them to the specified dimension to generate the output of multi-head attention. The calculation formula is as follows: MultiHead(Q,K,V)=Concat(head1,...,head n )W o Among them, h represents the number of heads; Further extract features from the output of the multi-head attention through the MLP module. The MLP module consists of two fully connected layers, two Dropout layers, and a ReLU activation function. The calculation formula is: MLP(x) = d2(d1(max(W1x + b1, 0)) · W2 + b2) Where, W1 is the weight of the fully connected layer, b1 and b2 are bias terms, and d1, d2 are Dropout operations; The output of the multi-head attention and MLP modules is directly added to the input through residual connections to generate the fused text feature F T , and the calculation formula is: F T = H T + MultiHead(Q t , K i , V i ) + MLP(H T + MultiHead(Q t , K i , V i )) In each of the cross-modal attention fusion image modules, the image feature H I is used as the query Q of the multi-head attention mechanism, and the text feature H T is used as the key K and value V. The fused image feature is generated through the same calculation process as above, and the calculation formula is: F I = H I + MultiHead(Q i , K t , V t ) + MLP(H I + MultiHead(Q i , K t , V t )) The three cross-modal attention fusion image modules and the three cross-modal attention fusion text modules are alternately stacked, and the feature transfer and fusion are maintained through the residual connection technology to generate the final cross-modal fusion features.

6. The method according to claim 5, wherein Input the obtained text features and image features into the cross-modal attention fusion module, and realize the information interaction and fusion of text and image features through the multi-head attention mechanism to generate cross-modal fusion features, including: Input the long text feature H LT , short text feature H ST , fine-grained image feature H FI and coarse-grained image feature H CI into the cross-modal attention fusion module to obtain the long text feature F LT , short text feature F ST , fine-grained image feature F FI and coarse-grained image feature F CI ; Add the fused long text feature F LT and the short text feature F ST to obtain the total text feature F T , and the calculation formula is: F T = F LT + F ST Add the fused fine-grained image feature F FI and the coarse-grained image feature F CI to obtain the total image feature F I , and the calculation formula is: F I = F FI + F CI Output the total text feature F T and the total image feature F I for subsequent semantic matching and downstream task processing.

7. The method according to claim 6, characterized in that, Optimize the cross-modal feature fusion using the image-text matching loss function of contrastive learning, maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs, including: For the total text feature F T and the total image feature F I perform L2 normalization. The normalization calculation formula is as follows: Among them, represents the normalized vector, ‖x‖ represents the L2 norm of vector x, and its calculation formula is: Construct positive and negative sample pairs, and use the matching text features and image features in the same batch as positive sample pairs. Use the unmatched text features and image features in the same batch as negative sample pairs, where j ≠ i; ​ Among them, u1 = [a1, a2,..., a N and u2 = [b1, b2,..., b N respectively represent the text feature and the image feature, and N is the dimension of the vector; Define the contrastive learning-based image-text matching loss function, where the text-to-image contrastive loss l text2img and the image-to-text contrastive loss l img2text are calculated as follows: ​ wherein, τ is the temperature parameter for adjusting the smoothness of the distribution; N is the number of samples in the current batch; ​ 8. The method according to claim 7, characterized in that ​ ​ ​ Calculate the similarity between the innovative design concept vector and all image vectors and statement vectors using cosine similarity.

9. A cross-modal multi-scale information fusion system supporting the push of design knowledge, characterized in that, It includes: A preprocessing unit for obtaining original patent data, cleaning and extracting it to form a patent dataset containing long texts, short texts, and abstract drawings; A feature extraction unit for respectively extracting the features of long texts and short texts using a text encoder, and extracting fine-grained image features of abstract drawings using an image encoder, and generating coarse-grained image features through principal component analysis for dimensionality reduction; A fusion unit for inputting the obtained text features and image features into a cross-modal attention fusion module, and realizing information interaction and fusion of text and image features through a multi-head attention mechanism to generate cross-modal fusion features; An optimization unit for optimizing the cross-modal attention fusion module using a contrastive learning image-text matching loss function to maximize the similarity between positive sample pairs and minimize the similarity between negative sample pairs; A similarity calculation unit for inputting an innovative design concept, extracting its feature representation through a cross-modal multi-scale information fusion module, and calculating the similarity between the innovative design concept features and patent abstracts and drawings using cosine similarity; A knowledge push unit for pushing patent abstracts or drawings with a preset ranking to designers according to the similarity score ranking to support the combination of design knowledge and the generation of innovative solutions.

10. An electronic device, characterized in that, It includes: A processor and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • John f

    US290225A

Cited By

  • Open vocabulary multi-task image classification method based on continuous learning

    CN120599385A