An intelligent product recommendation system based on image and text similarity fusion

Through the intelligent product recommendation system that integrates image and text similarity, the problem of insufficient multimodal data fusion in the existing technology is solved, more accurate and personalized product recommendations are achieved, and the performance and user experience of the recommendation system are improved.

CN120106940BActive Publication Date: 2025-09-05BEIJING JINGNENG TENDERING & COLLECTIVE PROCUREMENT CENT CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510176047.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-09-05
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

The existing recommendation systems have shortcomings in multimodal data fusion, deep model optimization, similarity calculation and recommendation strategy design, resulting in deviations from user needs.

Method used

An intelligent product recommendation system based on the fusion of image and text similarity is adopted, and image feature vectors and text feature vectors are extracted through the feature extraction module, and a high-dimensional embedded vector is generated using the transformer network model, multimodal deep neural network model and cross attention network model. Comprehensive similarity calculation is performed based on other key information of the product, and the optimal recommendation is finally determined in the candidate product collection.

Benefits of technology

It improves the accuracy and diversity of the recommendation system, can express product characteristics more comprehensively, and improves the reliability and user satisfaction of recommendation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120106940B_ABST
    Figure CN120106940B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and in particular to an intelligent product recommendation system based on image and text similarity fusion, comprising: a data acquisition module for acquiring display information of a specified product; a feature extraction module for extracting an image feature vector and a text feature vector based on the display information; a feature fusion module for acquiring a first fusion feature based on the image feature vector and the text feature vector; a model processing module for inputting the first fusion feature into a plurality of pre-trained processing models, each processing model obtaining a corresponding high-dimensional embedding vector; a vector fusion module for acquiring a first comprehensive embedding vector based on the high-dimensional embedding vector; a similarity calculation module for acquiring a final similarity between the specified product and each product based on first information of the specified product, a first comprehensive embedding vector, first information of each product, and the corresponding first fusion feature; and a recommendation module for determining the final recommended product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an intelligent product recommendation system based on the fusion of image and text similarity. Background Art

[0002] In the field of artificial intelligence, with the explosive growth in the variety and quantity of goods, using intelligent technology to recommend products that meet user needs from this vast array of products has become a core issue for e-commerce platforms. Recommender systems, as a key technology for addressing this issue, have been widely researched and applied.

[0003] Existing recommendation systems typically analyze single-modal data (such as product text descriptions or image information). However, single-modal data cannot fully reflect the actual characteristics of a product. For example, recommendation systems based solely on textual features may not accurately capture the visual characteristics of a product, while those based solely on image features may lack an understanding of product descriptions and attribute information. Therefore, these methods still have significant room for improvement in terms of product recommendation accuracy and user satisfaction. In product recommendation tasks, product images and textual information often carry different semantic information. Effectively integrating this multimodal information is key to improving the performance of recommendation systems. Traditional feature fusion methods typically use simple weighted averaging or concatenation. These linear fusion methods struggle to capture the complex interactions between multimodal data, limiting the performance of recommendation systems. With the advancement of deep learning technology, recommendation systems based on deep neural networks have gradually become a research hotspot. These systems leverage the powerful representational capabilities of neural networks to extract high-dimensional features from multimodal product data, significantly improving recommendation accuracy. However, current deep learning recommendation systems still lack efficient model architectures and optimization solutions for multimodal feature fusion and the generation and application of high-dimensional embedding vectors. In existing recommendation systems, similarity calculations are typically based on simple distance or correlation metrics, which fail to fully reflect the multidimensional similarity between items. Furthermore, recommendation strategies often fail to fully consider the weighting of multimodal data in practical applications, leading to discrepancies between recommendation results and user needs.

[0004] In summary, the existing technology has certain shortcomings in multimodal data fusion, deep model optimization, similarity calculation and recommendation strategy design. Therefore, it is necessary to propose an intelligent product recommendation system based on image and text similarity fusion to overcome the above problems and improve the recommendation accuracy and user satisfaction of the recommendation system. Summary of the Invention

[0005] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides an intelligent product recommendation system based on the fusion of image and text similarity, which solves the technical problems in the prior art that the recommendation method based on a single modality (such as image or text) is difficult to fully reflect the actual characteristics of the product and that multimodal information is difficult to effectively integrate.

[0006] In order to achieve the above objectives, the main technical solutions adopted by the present invention include:

[0007] An embodiment of the present invention provides an intelligent product recommendation system based on image and text similarity fusion, characterized in that the system includes:

[0008] A data acquisition module is used to acquire display information of a specified product; the display information includes a product image, title, and first information; the first information includes: product supplier, product selling price, and product code 69;

[0009] A feature extraction module is used to extract the image feature vector and text feature vector of the specified product based on the display information of the specified product;

[0010] a feature fusion module, configured to obtain a first fused feature of the specified product based on the image feature vector and the text feature vector of the specified product;

[0011] A model processing module is used to input the first fused feature of the specified product into multiple pre-trained processing models, and each processing model obtains a corresponding high-dimensional embedding vector;

[0012] Among them, the pre-set multiple processing models include transformer network model, multimodal deep neural network model, and cross-attention network model;

[0013] A vector fusion module, configured to obtain a first comprehensive embedding vector for the specified product based on the high-dimensional embedding vector corresponding to each processing model;

[0014] a similarity calculation module, configured to respectively obtain a final similarity between the designated product and each product in the first product set based on the first information of the designated product, the first comprehensive embedding vector, and the pre-acquired first information and corresponding first fusion features of each product in the first product set; wherein the first product set includes display information of multiple products pre-stored in the database;

[0015] The recommendation module is configured to determine a final recommended product in the first product set based on a final similarity between the designated product and each product in the first product set.

[0016] Preferably, the feature extraction module specifically includes:

[0017] an image feature extraction unit, configured to extract an image feature vector from the product image in the display information of the designated product using a trained convolutional neural network;

[0018] A text feature extraction unit, configured to extract a text feature vector from the product title text in the display information of the designated product using a trained natural language processing model based on deep learning;

[0019] Wherein, the convolutional neural network is a ResNet convolutional neural network;

[0020] The natural language processing model based on deep learning is the BERT model.

[0021] Preferably, the feature fusion module obtains the first fusion feature based on the image feature vector and the text feature vector in the display information of the designated product, specifically including:

[0022] Based on the image feature vector and text feature vector in the display information of the specified product, the first fusion feature is obtained using formula (1);

[0023] Wherein, the formula (1) is:

[0024]

[0025] ⊙ is the Hadamard product operation; is the element-by-element addition operation; f( ) is the nonlinear transformation function; I is the image feature vector; T is the text feature vector; I a is the weighted feature corresponding to the image feature vector calculated by the attention mechanism; T a is the weighted feature corresponding to the text feature vector calculated by the attention mechanism; α is a preset first weighting coefficient; β is a preset second weighting coefficient; where the sum of α and β is equal to 1; F is the first fusion feature.

[0026] Preferably, the feature fusion module obtains the first fusion feature based on the image feature vector and the text feature vector in the display information of the specified product, specifically including:

[0027] Based on the image feature vector and text feature vector in the display information of the specified product, the first fusion feature is obtained using formula (2);

[0028] Wherein, the formula (2) is:

[0029]

[0030] Wherein, W1 is the first learnable weight matrix obtained in advance;

[0031] W2 is the pre-acquired second learnable weight matrix;

[0032] ε is a pre-set weight hyperparameter;

[0033] I is the image feature vector;

[0034] T is the text feature vector;

[0035] cosine(I, T) is the cosine similarity between the image feature vector and the text feature vector;

[0036] F is the first fusion feature.

[0037] Preferably, the pre-set multiple processing models include a transformer network model, a multimodal deep neural network model, and a cross-attention network model;

[0038] Accordingly, the model processing module inputs the first fusion feature of the specified product into multiple pre-trained processing models, and each processing model obtains a corresponding high-dimensional embedding vector, specifically including:

[0039] The first fusion feature of the specified product is input into the transformer network model, the multimodal deep neural network model, and the cross-attention network model respectively, and the first high-dimensional embedding vector corresponding to the transformer network model, the second high-dimensional embedding vector corresponding to the multimodal deep neural network model, and the third high-dimensional embedding vector corresponding to the cross-attention network model are obtained respectively.

[0040] Preferably, the vector fusion module is configured to obtain a first comprehensive embedding vector of the specified product based on the high-dimensional embedding vector corresponding to each processing model, specifically comprising:

[0041] Based on the high-dimensional embedding vector corresponding to each processing model, the first comprehensive embedding vector is obtained using formula (3) or formula (4);

[0042] Wherein, the formula (3) is:

[0043] E c =E a ⊙E b ⊙K(X 12 )⊙K(X 13 )⊙K(X 23 );

[0044] Among them, E c is the first comprehensive embedding vector;

[0045] E a =b1E1+b2E2+b3E2;

[0046] E b =b1E12 +b2X 13 +b3X 23 ;

[0047] X 12 =E1⊙E2,X 13 =E1⊙E3,X 23 =E2⊙E3;

[0048]

[0049] Where E1 is the first high-dimensional embedding vector; E2 is the second high-dimensional embedding vector; E3 is the third high-dimensional embedding vector; exp() is the natural exponential function; C1 is the first preset center; σ1 is the first scale parameter; C2 is the second preset center; σ2 is the second scale parameter; C3 is the third preset center; σ3 is the third scale parameter;

[0050] Wherein, the formula (4) is:

[0051]

[0052] Among them, σ4 is the width hyperparameter.

[0053] Preferably, the similarity calculation module is configured to obtain the final similarity between the specified product and each product in the first product set based on the first information of the specified product, the first comprehensive embedding vector, and the pre-acquired first information of each product in the first product set and the corresponding first fusion feature, in the following manner:

[0054] Calculate the first comprehensive embedding vector of the specified product and the first fusion feature of any product in the first product set using the cosine similarity calculation method to obtain the first similarity between the specified product and the product in the first product set;

[0055] Based on the first information of the designated product and the first information of any product in the first product set, respectively obtain the product supplier similarity, product sales price similarity, and product 69 code similarity between the designated product and the product in the first product set;

[0056] Based on the first similarity, product supplier similarity, product sales price similarity, and product 69 code similarity between the specified product and the product in the first product set, a final similarity between the specified product and the product in the first product set is obtained.

[0057] Preferably, the obtaining of the product supplier similarity, product sales price similarity, and product 69 code similarity between the specified product and the product in the first product set based on the first information of the specified product and the first information of any product in the first product set specifically includes:

[0058] If the product supplier of the specified product is the same as the product supplier of any product in the first product set, the similarity between the specified product and the product supplier of the product in the first product set is determined to be 1; if they are different, the similarity between the specified product and the product supplier of the product in the first product set is determined to be 0;

[0059] Based on the similarity of the sales price of the specified product and the similarity of the sales price of any product in the first product set, the similarity of the sales price of the product is calculated using formula (5);

[0060] The formula (5) is:

[0061]

[0062] Among them, r Ai is the similarity between the sales price of the specified product and the i-th product in the first product set; P A is the sales price of the specified product; P i is the sales price of the i-th product in the first product set;

[0063] If the product supplier of the specified product is the same as the product 69 code of any product in the first product set, the similarity between the product 69 code of the specified product and the product in the first product set is determined to be 1; if they are different, the similarity between the product 69 code of the specified product and the product in the first product set is determined to be 0.

[0064] Preferably, obtaining the final similarity between the specified product and the product in the first product set based on the first similarity between the specified product and the product in the first product set, the product supplier similarity, the product sales price similarity, and the product 69 code similarity specifically includes:

[0065] Based on the first similarity between the specified product and the product in the first product set, the product supplier similarity, the product sales price similarity, and the product 69 code similarity, the final similarity between the specified product and the product in the first product set is obtained using formula (6);

[0066] Wherein, the formula (6) is:

[0067]

[0068] Where R is the final similarity between the specified product and the product in the first product set;

[0069] s1 is the first similarity between the specified product and the product in the first product set;

[0070] s2 is the similarity between the specified product and the product supplier of the product in the first product set;

[0071] s3 is the product 69 code similarity between the specified product and the product in the first product set.

[0072] Preferably, the final recommended products include products in the first product set corresponding to the five largest final similarities.

[0073] The beneficial effects of the present invention are:

[0074] The present invention provides an intelligent product recommendation system based on image and text similarity fusion. A feature extraction module is used to extract the image feature vector and text feature vector of the product, and the image feature vector and text feature vector are fused through a feature fusion module to generate a first fusion feature. At the same time, a transformer network model, a multimodal deep neural network model and a cross-attention network model are used through a model processing module to generate multiple high-dimensional embedding vectors, and finally a first comprehensive embedding vector is generated in the vector fusion module. Compared with the existing technology, it can express the characteristics of the product more comprehensively, thereby achieving the effect of improving the accuracy and diversity of the recommendation system.

[0075] The intelligent product recommendation system based on image and text similarity fusion of the present invention adopts a pre-trained transformer network model, a multimodal deep neural network model and a cross-attention network model. Compared with the existing technology, it can more efficiently capture the complex interactive relationship between multimodal features, thereby achieving the effect of improving the accuracy of multimodal data fusion.

[0076] The intelligent product recommendation system based on the fusion of image and text similarity of the present invention adopts a comprehensive similarity calculation method, which combines multimodal feature embedding with other key information of the product (such as supplier, sales price and product code 69). Compared with the existing technology, it can more comprehensively and accurately evaluate product similarity, thereby achieving the effect of improving the reliability and relevance of recommendation results.

[0077] The intelligent product recommendation system based on the fusion of image and text similarity of the present invention can better meet the personalized needs of users and achieve the effect of improving user satisfaction compared with the existing technology because it screens out the products with the highest final similarity to the specified products from the candidate product set through optimizing the recommendation strategy.

[0078] The intelligent product recommendation system based on the fusion of image and text similarity of the present invention adopts a multimodal deep learning model to generate multi-level, high-dimensional product embedding vectors. Compared with the existing technology, it can more fully express the characteristics of the products and achieve the effect of enhancing the recommendation system's ability to process complex data. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1Schematic diagram of an intelligent product recommendation system based on image and text similarity fusion according to the present invention. DETAILED DESCRIPTION

[0080] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below through specific implementation methods in conjunction with the accompanying drawings.

[0081] To better understand the above technical solutions, exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Instead, these embodiments are provided to enable a clearer and more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.

[0082] Example 1

[0083] See also Figure 1 This embodiment provides an intelligent product recommendation system based on image and text similarity fusion, the system comprising:

[0084] The data acquisition module is used to obtain display information for a specified product. This display information includes the product image, title, and primary information. This primary information includes the product supplier, sales price, and product code 69. In this embodiment, the acquisition of product information such as the product image, title, supplier, and sales price ensures that complete display data is provided for the product, enhancing the recommendation system's understanding of the product. Acquiring not only product image features but also textual information (such as product title and supplier information) allows for recommendations from multiple dimensions, helping to improve recommendation accuracy.

[0085] The 69 code for a product refers to the barcode in the EAN-13 barcode that begins with the number 69. EAN-13 (European Article Number) is an internationally recognized barcode consisting of 13 digits that uniquely identifies a product.

[0086] The feature extraction module extracts the image and text feature vectors of a specific product based on its display information. By extracting these feature vectors separately, the system can better capture the product's visual characteristics and semantic information, ensuring that the recommendation system can identify product characteristics from multiple perspectives. Combining image features with text information enables a more comprehensive understanding of the product, especially in scenarios involving multimodal information, avoiding the biases associated with relying solely on a single type of data.

[0087] The feature fusion module is configured to obtain a first fused feature for a specific product based on its image feature vector and text feature vector. By fusing image and text features, the system can obtain a more accurate product representation, improving the recommendation system's capabilities. This is particularly true when product information is incomplete or when there are significant discrepancies between image and text. The fused features can effectively compensate for any deficiencies in the product. The fused features represent the product's full picture, enabling the system to provide users with more tailored product recommendations and reduce errors and biases.

[0088] A model processing module is used to input the first fused feature of the specified product into multiple pre-trained processing models, and each processing model obtains a corresponding high-dimensional embedding vector;

[0089] Among them, the pre-set multiple processing models include transformer network model, multimodal deep neural network model, and cross-attention network model;

[0090] By utilizing different pre-trained models (such as transformer networks, multimodal deep neural networks, and cross-attention networks), the system can conduct in-depth product analysis from different perspectives, increasing the diversity and accuracy of recommendations. Each model has its own strengths, and through multi-model fusion, the system can fully leverage the strengths of each model, avoid the potential limitations of a single model, and improve recommendation quality.

[0091] It should be noted that the transformer network model, multimodal deep neural network model, and cross-attention network model in this embodiment are all existing models. For example, the transformer network model is a deep learning model based on the self-attention mechanism that performs well in multiple tasks such as text generation, machine translation, and image processing. In multimodal recommendation systems, it can effectively combine text and image information to capture richer semantic information.

[0092] A multimodal deep learning model (MDL) utilizes information from different modalities (such as images, text, and audio) through deep learning methods for joint learning. This model can process different types of input data within a single framework, integrating information from different modalities for prediction or classification. It is widely used in fields such as video analysis, sentiment analysis, and cross-modal search. In particular, multimodal models can simultaneously consider both visual and textual descriptions of products, improving the accuracy and personalization of recommendations.

[0093] The Cross-Attention Network (CAN) is a model that combines attention mechanisms from different modalities and is commonly used in multimodal learning tasks. It enhances the information from each modality and improves synergy between modalities by focusing attention on each other. In product recommendation systems, the CAN can effectively integrate product images and text information to provide more accurate recommendations.

[0094] Transformer networks effectively capture long-range dependencies through a self-attention mechanism, making them suitable for processing complex sequence data and particularly useful in text processing and cross-modal learning. Multimodal deep neural networks fuse information from different modalities, enabling the model to fully leverage the characteristics of each modality and making them suitable for multimodal tasks such as image and text processing. Cross-attention networks, on the other hand, focus on information interaction between modalities, enhancing intermodal synergy through a cross-attention mechanism, demonstrating their advantages in multimodal recommendation. The combined use of these models in intelligent product recommendation systems can effectively improve the accuracy and diversity of product recommendations.

[0095] The vector fusion module is used to obtain the first comprehensive embedding vector for a specific product based on the high-dimensional embedding vector corresponding to each processing model. Based on the high-dimensional embedding vectors of multiple processing models, the system can obtain a unified comprehensive embedding vector. This vector combines the characteristics of different models to help more accurately represent product characteristics. By fusing vectors generated by different models, it can reduce the decline in recommendation accuracy caused by the failure of a single model, and enhance the robustness and adaptability of the system.

[0096] a similarity calculation module, configured to respectively obtain a final similarity between the designated product and each product in the first product set based on the first information of the designated product, the first comprehensive embedding vector, and the pre-acquired first information and corresponding first fusion features of each product in the first product set; wherein the first product set includes display information of multiple products pre-stored in the database;

[0097] By calculating a comprehensive similarity between a given product and the features of other products (including text information and embedding vectors), the system can more accurately assess the similarity between products and avoid incorrect recommendations. Combining both the product's text information and image features yields a more comprehensive similarity calculation, considering not only the semantic similarity of the product content but also its visual similarity, thereby improving the quality of recommendations.

[0098] The recommendation module is configured to determine a final recommended product in the first product set based on a final similarity between the designated product and each product in the first product set.

[0099] By comprehensively evaluating product similarities, the system can filter the most suitable items from a collection, reducing redundancy and irrelevance in recommendations and improving the user experience. Dynamically adjusting recommendations based on similarity allows for continuous optimization of recommended products based on the needs of different users, ensuring the relevance and timeliness of recommendations.

[0100] In this embodiment, the feature extraction module specifically includes:

[0101] an image feature extraction unit, configured to extract an image feature vector from the product image in the display information of the designated product using a trained convolutional neural network;

[0102] A text feature extraction unit is used to extract a text feature vector from the product title text in the display information of the specified product through a trained deep learning-based natural language processing model; wherein the convolutional neural network is a ResNet convolutional neural network; and the deep learning-based natural language processing model is a BERT model.

[0103] In this embodiment, ResNet (residual network) is a very effective convolutional neural network (CNN) architecture. By using ResNet, the ability to extract image features has been significantly improved, especially for the recognition of complex images and details. The ResNet structure can effectively capture the hierarchical information in the product image, thereby providing high-quality image feature vectors, which is crucial for accurate product recognition and subsequent recommendation results. Through precise image feature extraction, the system can better understand the appearance of the product and perform visual similarity calculations. ResNet makes it easier to train deeper neural networks through the design of residual connections, thereby reducing the demand for computing resources during image feature extraction and improving training efficiency. BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language model based on the Transformer architecture that can deeply understand the semantics of text. The BERT model, through the design of a bidirectional encoder, can fully understand the meaning of words from the context and has stronger semantic modeling capabilities than traditional unidirectional language models. The feature extraction module in this embodiment extracts rich and accurate features from product images and title texts by combining the ResNet convolutional neural network and the BERT natural language processing model. This bimodal, multi-level feature extraction can improve the recommendation system's ability to understand products and enhance its comprehensive analysis of product visual and semantic information, thereby improving the accuracy of product recommendations and user experience.

[0104] Specifically, the feature fusion module obtains a first fusion feature based on the image feature vector and the text feature vector in the display information of the specified product, specifically including:

[0105] Based on the image feature vector and text feature vector in the display information of the specified product, the first fusion feature is obtained using formula (1);

[0106] Wherein, the formula (1) is:

[0107]

[0108] ⊙ is the Hadamard product operation; is an element-by-element addition operation; f() is a nonlinear transformation function; I is the image feature vector; T is the text feature vector; I a is the weighted feature corresponding to the image feature vector calculated by the attention mechanism; T a is the weighted feature corresponding to the text feature vector calculated by the attention mechanism; α is a preset first weighting coefficient; β is a preset second weighting coefficient; where the sum of α and β is equal to 1; F is the first fusion feature.

[0109] In this embodiment, formula (1) combines the image feature vector and the text feature vector, as well as their respective weighted features I a and T a , which can capture the multimodal information of products more comprehensively. This comprehensive processing method helps to improve the accuracy and robustness of the recommendation system. a and weighted features T a It indicates that it can automatically learn and highlight important feature parts, thereby enhancing the model's focus on key information. This helps to filter out more valuable information from complex data and improve the quality of recommendation results. The nonlinear transformation function f() can further adjust the fused features to make them more suitable for subsequent classification or regression tasks. This nonlinear transformation helps capture more complex feature relationships and improves the expressiveness of the model. The introduction of the first weighting coefficient and the second weighting coefficient allows the model to flexibly adjust the importance of image and text features according to specific application scenarios. For example, in some cases, image features may be more important than text features, and vice versa. By adjusting these weights, different recommendation scenarios can be better adapted.

[0110] It should be noted that, in this embodiment, the first fusion feature is obtained by using formula (1), and other methods can also be used to obtain the first fusion feature. For example, the image feature vector and the text feature vector are directly spliced ​​into a longer vector, and the spliced ​​vector is used as the first fusion feature.

[0111] Preferably, in another specific embodiment, the feature fusion module obtains the first fusion feature based on the image feature vector and the text feature vector in the display information of the designated product, specifically including:

[0112] Based on the image feature vector and text feature vector in the display information of the specified product, the first fusion feature is obtained using formula (2);

[0113] Wherein, the formula (2) is:

[0114]

[0115] Among them, W1 is the pre-acquired first learnable weight matrix; W2 is the pre-acquired second learnable weight matrix; ε is the pre-set weight hyperparameter; I is the image feature vector; T is the text feature vector; cosin(I, T) is the cosine similarity between the image feature vector and the text feature vector; F is the first fusion feature.

[0116] Formula (2) combines the image feature vector and the text feature vector by weighted summation and introduces cosine similarity (I, T). This comprehensive processing method can more comprehensively capture the multimodal information of the product and improve the accuracy and robustness of the recommendation system. W1 and W2 are pre-acquired learnable weight matrices, which means that they can be automatically adjusted through the training process to better adapt to specific task requirements. This adaptive adjustment capability enables the model to dynamically adjust the weights according to the data distribution, thereby improving the generalization ability and performance of the model. The introduction of cosine similarity can further enhance the model's understanding of the relationship between image and text features. When the image and text features are highly correlated, the cosine similarity will be larger, which will have a greater impact on the final fusion feature. The weight hyperparameter ε is a pre-set weight hyperparameter used to control the degree of influence of cosine similarity on the final fusion feature. By adjusting, the importance of image and text features and the correlation between them can be flexibly controlled according to the specific application scenario.

[0117] It should be noted that, in the specific implementation process, the first fusion feature in this embodiment can be obtained in different ways.

[0118] In the practical application of this embodiment, the pre-set multiple processing models include a transformer network model, a multimodal deep neural network model, and a cross-attention network model;

[0119] Accordingly, the model processing module inputs the first fusion feature of the specified product into multiple pre-trained processing models, and each processing model obtains a corresponding high-dimensional embedding vector, specifically including:

[0120] The first fusion feature of the specified product is input into the transformer network model, the multimodal deep neural network model, and the cross-attention network model respectively, and the first high-dimensional embedding vector corresponding to the transformer network model, the second high-dimensional embedding vector corresponding to the multimodal deep neural network model, and the third high-dimensional embedding vector corresponding to the cross-attention network model are obtained respectively.

[0121] In this embodiment, different processing models have different advantages and focuses. The transformer network model (Transformer) is good at processing sequence data and capturing long-distance dependencies, the multimodal deep neural network model (MultimodalDNN) can effectively fuse information from multiple modalities, and the cross-attention network model (Cross-AttentionNetwork) focuses on enhancing the relationship between different modalities. By passing the same input features (first fusion features) into different models, the system can model the product more comprehensively from multiple perspectives, avoiding the limitations of a single model. By inputting the first fusion features into multiple processing models, the system can model and understand the product from different angles and dimensions, which provides a rich source of information for the recommendation system. The high-dimensional embedding vector of each model provides a guarantee for the diversity and accuracy of product features, and can avoid the deviations or limitations that may exist in a single model. Ultimately, this multi-model synergy improves the accuracy, robustness and personalization of the recommendation system and optimizes the user experience.

[0122] Specifically, the vector fusion module is used to obtain the first comprehensive embedding vector of the specified product based on the high-dimensional embedding vector corresponding to each processing model, specifically including:

[0123] Based on the high-dimensional embedding vector corresponding to each processing model, the first comprehensive embedding vector is obtained using formula (3) or formula (4); wherein, the formula (3) is:

[0124] E c =E a ⊙E b ⊙K(X 12 )⊙K(X 13 )⊙K(X 23 );

[0125] Among them, E c is the first comprehensive embedding vector;

[0126] E a =b1E1+b2E2+b3E2;

[0127] E b =b1X 12 +b2X 13 +b3X 23 ;

[0128] X 12 =E1⊙E2,X 13 =E1⊙E3,X 23 =E2⊙E3;

[0129]

[0130]

[0131] Where E1 is the first high-dimensional embedding vector; E2 is the second high-dimensional embedding vector; E3 is the third high-dimensional embedding vector; exp() is the natural exponential function; C1 is the first preset center; σ1 is the first scale parameter; C2 is the second preset center; σ2 is the second scale parameter; C3 is the third preset center; σ3 is the third scale parameter;

[0132] In this embodiment, each part of formula (3) has a clear physical meaning, for example, E a represents the vector obtained by linear combination, E b represents the vector obtained by a specific operation, and K(X 12 ) represents the similarity between different embedding vectors. This clear definition makes the principle of the formula easy to understand and explain, and facilitates debugging and improvement. By integrating the information of different embedding vectors in a variety of ways, the multimodal information of the product can be captured more comprehensively. The application of the Gaussian kernel function enhances the robustness of the model to noisy data and can adapt to complex data distributions. The combination of linear combination and specific operations enables the model to flexibly adjust the importance of different embedding vectors to adapt to different task requirements. The Gaussian kernel function can enhance the expressive power of the model and capture more complex feature relationships. Each part of formula (3) in this embodiment has a clear physical meaning, which is easy to understand and improve.

[0133] Wherein, the formula (4) is:

[0134]

[0135] Among them, σ4 is the width hyperparameter.

[0136] In this embodiment, formula (4) is used to calculate the similarity between multiple high-dimensional embedding vectors and fuse them through an exponential function to obtain a first comprehensive embedding vector. Formula (4) combines the similarities between different embedding vectors through a weighted summation. This comprehensive processing method can more comprehensively capture the multimodal information of products and improve the accuracy and robustness of the recommendation system. The width hyperparameter σ4 can flexibly adjust the similarity weights between different embedding vectors, allowing the model to better adapt to different task requirements.

[0137] Specifically, the similarity calculation module is configured to obtain the final similarity between the specified product and each product in the first product set based on the first information of the specified product, the first comprehensive embedding vector, and the pre-acquired first information of each product in the first product set and the corresponding first fusion feature, in the following manner:

[0138] Calculate the first comprehensive embedding vector of the specified product and the first fusion feature of any product in the first product set using the cosine similarity calculation method to obtain the first similarity between the specified product and the product in the first product set;

[0139] Based on the first information of the designated product and the first information of any product in the first product set, respectively obtain the product supplier similarity, product sales price similarity, and product 69 code similarity between the designated product and the product in the first product set;

[0140] Based on the first similarity, product supplier similarity, product sales price similarity, and product 69 code similarity between the specified product and the product in the first product set, a final similarity between the specified product and the product in the first product set is obtained.

[0141] In this embodiment, by introducing similarity across multiple dimensions (image, text, and product attributes), the system can more comprehensively measure the similarity between products, avoiding the limitations of single-dimensional similarity calculations. Product similarity is not only reflected in image or text content, but also involves its commercial attributes (such as supplier, sales price, and product code). Images and text can reflect the appearance and description of the product, while the supplier, sales price, and code provide the actual market attributes of the product. By combining these dimensions, the system can better understand the comprehensive characteristics of the product, thereby providing more accurate recommendation results. This embodiment combines deep learning-based image and text features (such as high-dimensional embedding vectors calculated using cosine similarity) with traditional similarity calculations based on product attributes (such as supplier, price, code, etc.), allowing the system to comprehensively leverage the advantages of different technologies. Deep learning models (such as multimodal deep neural networks) can extract complex features from images and text, but these features often cannot fully express the market attributes of a product. Therefore, incorporating traditional attributes (such as supplier, price, etc.) can compensate for the shortcomings of deep learning models in these areas. This combination makes similarity calculation more comprehensive, simultaneously considering various aspects of a product, such as appearance, function, brand, and price. Using cosine similarity to calculate feature vector similarity between products is a common and effective method, particularly suitable for high-dimensional data. Cosine similarity measures the similarity between two vectors in their orientation, rather than their magnitude, thus avoiding bias caused by differences in feature scale. In product recommendations, feature vectors (such as image and text embedding vectors) are often high-dimensional and sparse. Using cosine similarity ensures fair similarity calculations between different products, regardless of the absolute value of their feature vectors. This allows for greater focus on the angular similarity between features, ensuring high stability and consistency in the recommendation system's similarity calculations. By weightedly combining the similarity of image and text features with the similarity of attributes such as supplier, price, and size, the final similarity is calculated, comprehensively considering the impact of various factors on product similarity. A single similarity calculation (e.g., relying solely on images or text) may overlook important product features or attributes. However, information such as supplier, price, and size can play a crucial role in a user's purchasing decision. For example, a user may prefer a certain brand or be more concerned about price differences. By combining the similarities of these different dimensions, we can better capture the user's preferences and needs, thereby improving the accuracy and personalization of the recommendation system.

[0142] In this embodiment, obtaining the product supplier similarity, product sales price similarity, and product 69 code similarity between the specified product and the product in the first product set based on the first information of the specified product and the first information of any product in the first product set, specifically includes:

[0143] If the product supplier of the specified product is the same as the product supplier of any product in the first product set, the similarity between the specified product and the product supplier of the product in the first product set is determined to be 1; if they are different, the similarity between the specified product and the product supplier of the product in the first product set is determined to be 0;

[0144] Based on the similarity of the sales price of the specified product and the similarity of the sales price of any product in the first product set, the similarity of the sales price of the product is calculated using formula (5);

[0145] The formula (5) is:

[0146]

[0147] Among them, r Ai is the similarity between the sales price of the specified product and the i-th product in the first product set; P A is the sales price of the specified product; P i is the sales price of the i-th product in the first product set. This calculation method is intuitive and easy to implement, and has good adaptability to the similarity calculation of sales prices of multiple products.

[0148] If the product supplier of the specified product is the same as the product 69 code of any product in the first product set, the similarity between the product 69 code of the specified product and the product in the first product set is determined to be 1; if they are different, the similarity between the product 69 code of the specified product and the product in the first product set is determined to be 0.

[0149] The method of obtaining the final similarity between the specified product and the product in the first product set based on the first similarity between the specified product and the product in the first product set, the product supplier similarity, the product sales price similarity, and the product 69 code similarity specifically includes:

[0150] Based on the first similarity between the specified product and the product in the first product set, the product supplier similarity, the product sales price similarity, and the product 69 code similarity, the final similarity between the specified product and the product in the first product set is obtained using formula (6);

[0151] Wherein, the formula (6) is:

[0152]

[0153] Where R is the final similarity between the specified product and the product in the first product set;

[0154] s1 is the first similarity between the specified product and the product in the first product set;

[0155] s2 is the similarity between the specified product and the product supplier of the product in the first product set;

[0156] s3 is the product 69 code similarity between the specified product and the product in the first product set.

[0157] In this embodiment, formula (6) comprehensively considers multiple similarity indicators, including first similarity, product supplier similarity, product sales price similarity, and product code 69 similarity. This comprehensive processing method can more comprehensively capture the multimodal information of products and improve the accuracy of the recommendation system. By using different weight coefficients (0.4, 0.3, 0.2, 0.1), the importance of different similarity indicators can be flexibly adjusted, making the model more adaptable to different task requirements.

[0158] The final recommended products include the products in the first product set corresponding to the five largest final similarities.

[0159] In the actual application of this embodiment, after obtaining the final recommended products (products in the first product set corresponding to the five final similarities), they are displayed according to the first ranking, wherein the first ranking of the final recommended products is sorted from high to low according to the final scores of the products, wherein the final scores of the products are calculated based on the final similarities of the products and the user's stay time and historical click rate of the products in the specified historical time period obtained in advance, using formula (7);

[0160] Wherein, the formula (7) is:

[0161]

[0162] Among them, G i is the final score of the i-th item in the final recommended items; DT i represents the time that the user stays on the i-th item in the final recommended items during the historical period; H i R represents the historical click rate of users on the i-th item in the final recommended items during the historical period; i Indicates the final similarity corresponding to the i-th item in the final recommended items.

[0163] Example 2

[0164] An embodiment of the present invention provides an intelligent product recommendation system based on image and text similarity fusion, characterized in that the system includes:

[0165] The data acquisition module is used to obtain product images, titles, and related attributes (such as price, brand, and category) from e-commerce platforms. This data is obtained through API interfaces or batch data crawling and stored in a relational database or distributed file system. The images are then normalized, including resizing, denoising, and normalization. Title text is cleaned, including stop word removal, word segmentation, and stemming.

[0166] The feature extraction module uses pre-trained convolutional neural networks (such as ResNet and EfficientNet) to extract high-level semantic features. Image feature dimensions are reduced using PCA (Principal Component Analysis), preserving key information and reducing computational complexity. The BERT model is used to convert product titles into embedding vectors to capture semantic information. Long titles are truncated or weighted to ensure semantic integrity.

[0167] In this embodiment, a deep learning model, such as a convolutional neural network (CNN), is used to extract features from images. These features typically include color, texture, shape, edges, and more advanced semantic features (such as objects, scenes, etc.).

[0168] a feature fusion module, configured to obtain a first fused feature of the specified product based on the image feature vector and the text feature vector of the specified product;

[0169] A model processing module is used to input the first fused feature of the specified product into multiple pre-trained processing models, and each processing model obtains a corresponding high-dimensional embedding vector;

[0170] Among them, the pre-set multiple processing models include transformer network model, multimodal deep neural network model, and cross-attention network model;

[0171] A vector fusion module, configured to obtain a first comprehensive embedding vector for the specified product based on the high-dimensional embedding vector corresponding to each processing model;

[0172] The similarity calculation module is configured to obtain the final similarity between the specified product and each product in the first product set based on the first information of the specified product, the first comprehensive embedding vector, and the pre-acquired first information and corresponding first fusion features of each product in the first product set. The first product set includes display information of multiple products pre-stored in the database. The similarity calculation module is further configured to calculate the Euclidean distance or cosine similarity between the extracted image feature vector of the specified product and the image feature vectors of other products as an evaluation indicator of image similarity. The cosine similarity between the text embedding vector of the specified product and the text embedding vectors of other products is calculated. For product titles in different languages, a multilingual model (such as MUSE) can be introduced to perform similarity calculation to obtain the comprehensive similarity between the specified product and any other product, where comprehensive similarity = image similarity × 60% + title similarity × 40%.

[0173] The recommendation module is configured to determine a final recommended product in the first product set based on a final similarity and a comprehensive similarity between the designated product and each product in the first product set.

[0174] The final recommended product in this embodiment is the product with the largest sum of the final similarity and the comprehensive similarity.

[0175] In this embodiment, image features and text features are combined, and the accuracy of recommendations is effectively improved through comprehensive analysis of multimodal features (such as cross-attention networks, transformer networks, etc.). Image features capture the appearance characteristics of products (color, texture, shape, etc.), and text features capture semantic information (description of product titles). Multimodal fusion can more comprehensively understand the characteristics of products. Combining image similarity and text similarity, according to weighted fusion (60% image similarity + 40% title similarity), the recommendation results are more reasonable. For systems based only on single-dimensional similarity, this fusion can effectively avoid the failure of a single feature. By introducing multilingual models (such as MUSE), it adapts to different language environments, expands the scope of application of the system, and enhances the recommendation capabilities of cross-border e-commerce. The images are standardized (resize adjustment, denoising, normalization), and the title text is cleaned (stop word removal, word segmentation, stemming, etc.), significantly improving data quality and avoiding noise interference model. Use pre-trained deep learning models (such as ResNet, EfficientNet, BERT) to extract high-level semantic features to ensure the accuracy and efficiency of feature extraction.

[0176] In this embodiment, the process of using an intelligent product recommendation system based on image and text similarity fusion includes:

[0177] Collect product data from various channels, including product pictures, titles, supplier information, prices (market price, negotiated price, sales price), product 69 codes, etc.

[0178] Clean the collected data to ensure its accuracy and completeness, such as removing duplicate data and processing missing values.

[0179] Calculate the similarity between the main product image and title, and use the configured similarity threshold (e.g., 80% or higher) to identify similar products. Calculate the overall similarity based on the similarity between the image and title and the configured weights for each (the sum of the weights is 1).

[0180] Using the intelligent retrieval model, the number of similar products for each product is calculated based on dimensions such as product type, external product code, product code, and product title.

[0181] Retrieve items from the database that are identical to the current item, but may come from different suppliers or have different prices. By default, the first five items are displayed, sorted by price from lowest to highest. Use the left and right arrow keys to view more items (if available). Long item titles are indicated by ellipsis, and the full title is displayed on mouseover.

[0182] Displays products that are highly similar or potentially similar to the current product, with a "Highly Similar / Possibly Similar" label. By default, the top five similar products are displayed, sorted by price from low to high, with a maximum of 15 products displayed. Also supports left and right arrows for switching and omitting or displaying the full title.

[0183] In this embodiment, the user experience is friendly, and the default sorting is from low to high by price, which meets the user's economic choice needs. It supports left and right arrow switching, which is convenient for users to view more recommended products and improves the convenience of operation. Long titles are processed with ellipsis and support mouse hovering to display the full content, which not only optimizes the aesthetics of the interface but also ensures the integrity of the information. It clearly distinguishes between the same product (same product but different suppliers or prices) and similar products (highly or potentially similar products) to provide users with a wealth of choices. The top 5 of the same product and similar products are displayed respectively, sorted by price, which is in line with the decision-making habits of most users. At the same time, it supports up to 15 displays, which increases the diversity of recommendations.

[0184] In the description of the present invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.

[0185] In the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," "connect," "fixed," etc. should be understood broadly. For example, they may refer to fixed connection, detachable connection, or integration; mechanical connection or electrical connection; direct connection or indirect connection through an intermediate medium; and internal communication between two components or interaction between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0186] In the present invention, unless otherwise expressly specified or limited, when a first feature is "above" or "below" a second feature, it may mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. Furthermore, when a first feature is "above," "above," or "above" a second feature, it may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is at a higher level than the second feature. When a first feature is "below," "below," or "below" a second feature, it may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is at a lower level than the second feature.

[0187] In the description of this specification, the terms "one embodiment", "some embodiments", "embodiments", "examples", "specific examples" or "some examples" refer to the specific features, structures, materials or characteristics described in conjunction with the embodiment or example and included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and features of different embodiments or examples, unless they are mutually inconsistent.

[0188] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may alter, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. An intelligent product recommendation system based on image and text similarity fusion, characterized by: The system comprises: A data acquisition module is used to acquire display information of a specified product; the display information includes a product image, title, and first information; the first information includes: product supplier, product selling price, and product code 69; A feature extraction module is used to extract the image feature vector and text feature vector of the specified product based on the display information of the specified product; a feature fusion module, configured to obtain a first fused feature of the specified product based on the image feature vector and the text feature vector of the specified product; A model processing module is used to input the first fused feature of the specified product into multiple pre-trained processing models, and each processing model obtains a corresponding high-dimensional embedding vector; Among them, the pre-set multiple processing models include transformer network model, multimodal deep neural network model, and cross-attention network model; A vector fusion module, configured to obtain a first comprehensive embedding vector for the specified product based on the high-dimensional embedding vector corresponding to each processing model; a similarity calculation module, configured to respectively obtain a final similarity between the designated product and each product in the first product set based on the first information of the designated product, the first comprehensive embedding vector, and the pre-acquired first information and corresponding first fusion features of each product in the first product set; wherein the first product set includes display information of multiple products pre-stored in the database; The recommendation module is configured to determine a final recommended product in the first product set based on a final similarity between the designated product and each product in the first product set.

2. The intelligent product recommendation system based on image and text similarity fusion according to claim 1 is characterized in that: The feature extraction module specifically includes: an image feature extraction unit, configured to extract an image feature vector from the product image in the display information of the designated product using a trained convolutional neural network; A text feature extraction unit, configured to extract a text feature vector from the product title text in the display information of the designated product using a trained natural language processing model based on deep learning; Wherein, the convolutional neural network is a ResNet convolutional neural network; The natural language processing model based on deep learning is the BERT model.

3. The intelligent product recommendation system based on image and text similarity fusion according to claim 2 is characterized in that: The feature fusion module obtains a first fusion feature based on the image feature vector and the text feature vector in the display information of the specified product, specifically including: Based on the image feature vector and text feature vector in the display information of the specified product, the first fusion feature is obtained using formula (1); Wherein, the formula (1) is: ⊙ is the Hadamard product operation; is an element-by-element addition operation; f() is a nonlinear transformation function; I is the image feature vector; T is the text feature vector; I a is the weighted feature corresponding to the image feature vector calculated by the attention mechanism; T a is the weighted feature corresponding to the text feature vector calculated by the attention mechanism; α is a preset first weighting coefficient; β is a preset second weighting coefficient; where the sum of α and β is equal to 1; F is the first fusion feature.

4. The intelligent product recommendation system based on image and text similarity fusion according to claim 2, characterized in that: The feature fusion module obtains a first fusion feature based on the image feature vector and the text feature vector in the display information of the specified product, specifically including: Based on the image feature vector and text feature vector in the display information of the specified product, the first fusion feature is obtained using formula (2); Wherein, the formula (2) is: Wherein, W1 is the first learnable weight matrix obtained in advance; W2 is the pre-acquired second learnable weight matrix; ε is a pre-set weight hyperparameter; I is the image feature vector; T is the text feature vector; cosine(I, T) is the cosine similarity between the image feature vector and the text feature vector; F is the first fusion feature.

5. The intelligent product recommendation system based on image and text similarity fusion according to claim 3 or 4, characterized in that: in, Multiple pre-set processing models include transformer network model, multimodal deep neural network model, and cross-attention network model; Accordingly, the model processing module inputs the first fusion feature of the specified product into multiple pre-trained processing models, and each processing model obtains a corresponding high-dimensional embedding vector, specifically including: The first fusion feature of the specified product is input into the transformer network model, the multimodal deep neural network model, and the cross-attention network model respectively, and the first high-dimensional embedding vector corresponding to the transformer network model, the second high-dimensional embedding vector corresponding to the multimodal deep neural network model, and the third high-dimensional embedding vector corresponding to the cross-attention network model are obtained respectively.

6. The intelligent product recommendation system based on image and text similarity fusion according to claim 5, characterized in that: The vector fusion module is used to obtain the first comprehensive embedding vector of the specified product based on the high-dimensional embedding vector corresponding to each processing model, specifically including: Based on the high-dimensional embedding vector corresponding to each processing model, the first comprehensive embedding vector is obtained using formula (3) or formula (4); Wherein, the formula (3) is: E c =E a ⊙E b ⊙K(X 12 )⊙K(X 13 )⊙K(X 23 ); Among them, E c is the first comprehensive embedding vector; <h2 style=";text-align:left;direction:ltr">E<h2 style=";text-align:left;direction:ltr"> a <h2 style=";text-align:left;direction:ltr"> (b1E1+b2E2+b3E2) E b =b1X 12 +b2X 13 +b3X 23 ; X 12 =E1⊙E2,X 13 =E1⊙E3,X 23 =E2⊙E3; Where E1 is the first high-dimensional embedding vector; E2 is the second high-dimensional embedding vector; E3 is the third high-dimensional embedding vector; exp() is the natural exponential function; C1 is the first preset center; σ1 is the first scale parameter; C2 is the second preset center; σ2 is the second scale parameter; C3 is the third preset center; σ3 is the third scale parameter; Wherein, the formula (4) is: Among them, σ4 is the width hyperparameter.

7. The intelligent product recommendation system based on image and text similarity fusion according to claim 6, characterized in that: The similarity calculation module is used to obtain the final similarity between the specified product and each product in the first product set based on the first information of the specified product, the first comprehensive embedding vector, and the pre-acquired first information of each product in the first product set and the corresponding first fusion feature. The specific method is as follows: Calculate the first comprehensive embedding vector of the specified product and the first fusion feature of any product in the first product set using the cosine similarity calculation method to obtain the first similarity between the specified product and the product in the first product set; Based on the first information of the designated product and the first information of any product in the first product set, respectively obtain the product supplier similarity, product sales price similarity, and product 69 code similarity between the designated product and the product in the first product set; Based on the first similarity, product supplier similarity, product sales price similarity, and product 69 code similarity between the specified product and the product in the first product set, a final similarity between the specified product and the product in the first product set is obtained.

8. The intelligent product recommendation system based on image and text similarity fusion according to claim 7 is characterized in that: The obtaining, based on the first information of the designated product and the first information of any product in the first product set, respectively the product supplier similarity, product sales price similarity, and product 69 code similarity between the designated product and the product in the first product set, specifically includes: If the product supplier of the specified product is the same as the product supplier of any product in the first product set, the similarity between the specified product and the product supplier of the product in the first product set is determined to be 1; if they are different, the similarity between the specified product and the product supplier of the product in the first product set is determined to be 0; Based on the similarity of the sales price of the specified product and the similarity of the sales price of any product in the first product set, the similarity of the sales price of the product is calculated using formula (5); The formula (5) is: Among them, r Ai is the similarity between the sales price of the specified product and the i-th product in the first product set; P A is the sales price of the specified product; P i is the sales price of the i-th product in the first product set; If the product supplier of the specified product is the same as the product 69 code of any product in the first product set, the similarity between the product 69 code of the specified product and the product in the first product set is determined to be 1; if they are different, the similarity between the product 69 code of the specified product and the product in the first product set is determined to be 0.

9. The intelligent product recommendation system based on image and text similarity fusion according to claim 8, characterized in that: Based on the first similarity, product supplier similarity, product sales price similarity, and product code 69 similarity between the specified product and the product in the first product set, obtaining the final similarity between the specified product and the product in the first product set, specifically including: Based on the first similarity between the specified product and the product in the first product set, the product supplier similarity, the product sales price similarity, and the product 69 code similarity, the final similarity between the specified product and the product in the first product set is obtained using formula (6); Wherein, the formula (6) is: Where R is the final similarity between the specified product and the product in the first product set; s1 is the first similarity between the specified product and the product in the first product set; s2 is the similarity between the specified product and the product supplier of the product in the first product set; s3 is the product 69 code similarity between the specified product and the product in the first product set.

10. The intelligent product recommendation system based on image and text similarity fusion according to claim 9, characterized in that: in, The final recommended products include the products in the first product set corresponding to the five largest final similarities.

Citation Information

Patent Citations

  • Method for generating diversity recommendation list of complementary articles in multiple modes

    CN112232929A

  • Commodity recommendation method and device and electronic equipment

    CN118350894A