Multi-modal Industrial Product Recommendation Method and System Based on Hybrid Attention Mechanism
By adopting a hybrid attention mechanism and a multimodal industrial product knowledge graph in the multimodal industrial product recommendation system, the integration of multimodal data and capturing user refinement preferences in the existing technology is solved, and higher recommendation accuracy and timeliness are achieved.
Patent Information
- Application Number
- CN202510096360.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The prior art is difficult to effectively integrate multimodal industrial product data, especially in capturing user refinement preferences and information fusion, resulting in low accuracy of recommendation results.
The multimodal industrial product recommendation method based on a hybrid attention mechanism is adopted, and the multimodal industrial product knowledge graph is constructed, combined with the time-sensitive attention mechanism and the multi-head attention mechanism, and the user's short-term interest characteristics and industrial product characteristics are integrated to obtain the user's comprehensive preference characteristics for recommendation.
It improves the accuracy and timeliness of recommendations, can meet the complex and changing interests of users in all aspects, and improves the compatibility between the recommendation system and the actual needs of users.
Smart Images

Figure CN119537703B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent recommendation, and particularly to a multi-modal industrial product recommendation method and system based on a hybrid attention mechanism. Background Art
[0002] With the development of industrial Internet and intelligent manufacturing, the data related to industrial products presents multi-modal characteristics. On the one hand, this data covers various forms, including not only the text description of products but also product images. On the other hand, the data scale shows an exponential growth, the user behavior shows an increasingly diverse trend, and the user's choice is comprehensively affected by multiple factors.
[0003] In practical applications, using recommendation algorithms to recommend information of interest to users has become an increasingly important technical means. However, due to the diversity of user behavior patterns, directly using multi-modal information for recommendation faces problems such as information heterogeneity. Existing technologies cannot effectively integrate this heterogeneous information to construct a unified representation. Although some advanced deep learning models can handle the relationships between modalities to a certain extent, there are still deficiencies in maximizing the utilization of potential associations and complementarities between different modalities while maintaining information integrity.
[0004] At the same time, in the field of industrial products, users' requirements for recommendations are not only to simply find relevant products, but also to deeply understand whether the products truly meet their complex needs, including production process requirements, quality standards, cost-effectiveness, and other aspects. In the face of complex and changing user interests, the current recommendation methods still have limitations, especially in the ability to capture and understand users' refined preferences and information fusion. They fail to fully consider users' short-term preferences and long-term overall preferences, and the recommendation results are overly dependent on historical behavior, ignoring the changes in user interests, resulting in low accuracy of recommendation results. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a multi-modal industrial product recommendation method and system based on a hybrid attention mechanism, which uses a hybrid attention mechanism that combines a time-sensitive attention mechanism and a multi-head attention mechanism, and fully utilizes multi-modal information to fully model user preferences and improve recommendation performance.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] In the first aspect, the present invention provides a multi-modal industrial product recommendation method based on a hybrid attention mechanism, including:
[0008] Obtain multi-modal product data for preprocessing, where the multi-modal product data includes text data and image data;
[0009] Construct a multi-modal industrial product knowledge graph based on the preprocessed multi-modal product data. The multi-modal industrial product knowledge graph includes product information triples, text triples, image triples, and user-product interaction quadruples;
[0010] Use one-hot encoding to extract the user behavior sequence features in the user-product interaction quadruples; based on the CLIP model and the ConvE model, extract the entity node feature embeddings of the multi-modal industrial product knowledge graph;
[0011] Adopt a time-sensitive attention mechanism based on the user behavior sequence features to obtain the user's short-term interest features; adopt a multi-head attention mechanism based on the entity node feature embeddings to extract the industrial product features;
[0012] Fuse the user's short-term interest features and the industrial product features to obtain the user's comprehensive preference features; make recommendations based on the user's comprehensive preference features and score the user's comprehensive preference features.
[0013] Preferably, the construction of the multi-modal industrial product knowledge graph based on the preprocessed multi-modal product data specifically includes:
[0014] Obtain the preprocessed text data and image data, extract the product ID, relationship, and attributes to construct the product knowledge graph triples, extract the product ID, description, and text information to construct the text triples, extract the product ID, image, and image information to construct the image triples, and extract the user ID, product ID, score, and timestamp to construct the user-product interaction quadruples.
[0015] Preferably, the extraction of the user behavior sequence features in the user-product interaction quadruples using one-hot encoding specifically includes:
[0016] Extract the user ID, the user's interaction behavior with the product, and its corresponding timestamp in the user-product interaction quadruples, and obtain the preliminary user features through one-hot encoding;
[0017] Input the preliminary user features into the linear graph convolution module and the attention mechanism layer to obtain the user's product interaction sequence;
[0018] Obtain the user behavior sequence features based on the user's product interaction sequence, interaction timestamp, and product ID.
[0019] Preferably, the extraction of the entity node feature embeddings of the multi-modal industrial product knowledge graph based on the CLIP model and the ConvE model specifically includes:
[0020] Extract the text feature embeddings of the text triples and the image feature embeddings of the image triples based on the CLIP model;
[0021] The ConvE model is used to extract user nodes, product nodes, and attribute nodes from the user-product interaction quadruples and product knowledge graph triples, obtaining initial entity features, and based on the initial entity features, initial product entity feature embeddings are obtained;
[0022] Entity node feature embeddings are obtained based on text feature embeddings, image feature embeddings, and initial product entity feature embeddings.
[0023] Preferably, the time-sensitive attention mechanism is adopted based on the user behavior sequence features to obtain user short-term interest features, specifically including:
[0024] Extract the interaction time information in the user-product interaction quadruples and calculate the time interval features of the user;
[0025] Construct an intermediate vector of the time-sensitive attention mechanism based on the time interval features and user behavior sequence features;
[0026] Merge the preliminary user features with the intermediate vector of the time-sensitive attention mechanism and perform normalization to obtain normalized attention weights;
[0027] Perform weighted summation on the normalized attention weights and user behavior sequence features to obtain user short-term interest features.
[0028] Preferably, the multi-head attention mechanism is adopted based on the entity node feature embeddings to extract industrial product features, specifically including:
[0029] Calculate attention scores based on entity node feature embeddings;
[0030] Based on the attention scores, sum the adjacent embeddings of the head entity to obtain first-order structure information;
[0031] Use the double-interaction aggregation function to splice the structure information of all heads, and aggregate the head entity and its first-order structure information, and capture high-order structure information by stacking multiple attention layers to obtain industrial product features.
[0032] Preferably, the scoring of the user's comprehensive preference features specifically includes:
[0033]
[0034] Among them, represents the fused user comprehensive preference features, represents the final representation of the industrial product features after splicing different layers, represents transpose.
[0035] Second, the present invention provides a multi-modal industrial product recommendation system based on a hybrid attention mechanism, including:
[0036] A data processing layer for obtaining multi-modal product data for preprocessing, where the multi-modal product data includes text data and image data;
[0037] A knowledge graph layer for constructing a multi-modal industrial product knowledge graph based on the preprocessed multi-modal product data, where the multi-modal industrial product knowledge graph includes product information triples, text triples, image triples, and user-product interaction quadruples;
[0038] A knowledge feature extraction layer for using one-hot encoding to extract user behavior sequence features in the user-product interaction quadruple; based on the CLIP model and the ConvE model, extracting entity node feature embeddings of the multi-modal industrial product knowledge graph;
[0039] An attention mechanism layer for using a time-sensitive attention mechanism based on the user behavior sequence features to obtain user short-term interest features; using a multi-head attention mechanism based on the entity node feature embeddings to extract industrial product features;
[0040] An intelligent recommendation and application evaluation layer for fusing user short-term interest features and industrial product features to obtain comprehensive user preference features; making recommendations based on the comprehensive user preference features and scoring the comprehensive user preference features.
[0041] In a third aspect, the present invention provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the steps in a multi-modal industrial product recommendation method based on a hybrid attention mechanism described in the first aspect are implemented.
[0042] In a fourth aspect, the present invention provides a computer device including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the steps in a multi-modal industrial product recommendation method based on a hybrid attention mechanism described in the first aspect are implemented.
[0043] Compared with the prior art, the beneficial effects of the present invention are:
[0044] (1) By introducing multi-modal information as auxiliary information, the present invention constructs a multi-modal industrial product knowledge graph to integrate different types of data, thereby considering more dimensions of information in user behavior interest modeling and better capturing user preferences. The multi-modal industrial product knowledge graph not only contains triple knowledge around product information, enabling rich and complete project product information to be obtained; on this basis, it also includes user-product interaction quadruples considering user and product interaction information, where the "rating" element serves as a quantitative representation of the interaction behavior, achieving the effect of accurately measuring the degree of user preference for products and demand tendencies. At the same time, the timestamp records the time when the user queries the product, introducing the time characteristics between the user and the product, thus realizing the tracking and analysis of the dynamic changes in user interests. The construction of the multi-modal industrial product knowledge graph effectively improves the accuracy and timeliness of recommendations and can fully meet the complex and changing interest needs of users.
[0045] (2) The present invention proposes a hybrid attention mechanism that combines a time-sensitive attention mechanism and a multi-head attention mechanism. The time-sensitive attention mechanism is used to effectively capture the short-term interest preferences of users containing time information by weighting recent interaction behaviors. At the same time, the multi-head attention mechanism is used to divide the input information into multiple heads to mine user preferences from multiple dimensions, thereby comprehensively capturing the overall interests of users. Through the synergistic effect of these two mechanisms, the comprehensive preferences of users can be obtained, thereby deeply analyzing the complex interaction relationship between users and industrial product projects and providing rich representations for the calculation of recommendation results. This helps to improve the fit between the recommendation system and the actual needs of users, making the recommendation results more accurate and effective. It can not only enable users to obtain a more satisfactory experience during the product selection process but also help industrial products achieve more efficient supply-demand docking in the market circulation link and promote the development of industrial products in the market.
[0046] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The specification drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute a limitation to the present invention.
[0048] Figure 1 It is the main flowchart of a multi-modal industrial product recommendation method based on a hybrid attention mechanism provided by an embodiment of the present invention;
[0049] Figure 2 It is the detailed flowchart of a multi-modal industrial product recommendation method based on a hybrid attention mechanism provided by an embodiment of the present invention;
[0050] Figure 3 Flow chart of text and image feature extraction provided by an embodiment of the present invention;
[0051] Figure 4 Flow chart of time-sensitive attention mechanism provided by an embodiment of the present invention;
[0052] Figure 5 Flow chart of multi-head attention mechanism provided by an embodiment of the present invention. Detailed implementation manners
[0053] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0054] Embodiment 1
[0055] As Figure 1 shown, this embodiment discloses a multi-modal industrial product recommendation method based on a hybrid attention mechanism, including the following steps:
[0056] S1: Obtain multi-modal product data for preprocessing, and the multi-modal product data includes text data and image data;
[0057] S2: Construct a multi-modal industrial product knowledge graph based on the preprocessed multi-modal product data, and the multi-modal industrial product knowledge graph includes product information triples, text triples, image triples, and user-product interaction quadruples;
[0058] S3: Use one-hot encoding to extract the user behavior sequence features in the user-product interaction quadruple; based on the CLIP model and the ConvE model, extract the entity node feature embeddings of the multi-modal industrial product knowledge graph;
[0059] S4: Adopt a time-sensitive attention mechanism based on the user behavior sequence features to obtain the user's short-term interest features; adopt a multi-head attention mechanism based on the entity node feature embeddings to extract the industrial product features;
[0060] S5: Fuse the user's short-term interest features and the industrial product features to obtain the user's comprehensive preference features; make recommendations based on the user's comprehensive preference features and score the user's comprehensive preference features.
[0061] Next, in combination with Figure 2 , a multi-modal industrial product recommendation method based on a hybrid attention mechanism disclosed in this embodiment will be described in detail.
[0062] In S1, the original data is preprocessed to make the data more standardized, concise and easy to analyze.
[0063] (1) Data collection
[0064] Obtain a multi-modal industrial product recommendation dataset containing text and image information from open recommendation data sources, and download the selected industrial dataset through a data platform or an open data warehouse. Then, perform a decompression operation on the data to extract the original data.
[0065] (2)Data preprocessing
[0066] 1) Data cleaning
[0067] After obtaining the original dataset, clean the data to ensure data quality and consistency. First, filter the data in the industrial product dataset, removing records with missing text information or image information and duplicate records, as these incomplete samples cannot provide effective support for subsequent multi-modal analysis. Second, perform further screening by filtering based on the interaction frequency between users and industrial products, and select data information with more than 20 text interactions to facilitate dataset division.
[0068] 2) Text processing
[0069] During the text data processing, entries containing garbled characters, unrecognizable characters, or incorrect formats must be excluded to ensure data quality and consistency. For example, text containing garbled characters, unrecognizable characters, or incorrect formats.
[0070] 3) Image processing
[0071] During the image processing, to ensure the quality and effectiveness of the images used for products, select images with higher resolution and clarity from all available images as the final image resources.
[0072] In S2, based on the processed data, mine the data information. By identifying entities, relationships, and their attributes in the data, extract knowledge multi-tuples to construct a multi-modal industrial product knowledge graph. At the same time, to simplify data processing in the knowledge graph and improve computational efficiency, map the knowledge graph data and use digital encoding to replace string representation, which can occupy less storage space.
[0073] (1)Construction of multi-modal industrial product knowledge graph
[0074] When faced with the processed dataset, to achieve efficient data utilization and in-depth mining, according to different data types, specifically extract, classify, and structurally represent knowledge multi-tuples.
[0075] Specifically, extract "product ID, relationship, and attribute" to construct product knowledge graph triples to accurately represent various relationships and related attribute information between industrial products.
[0076] To enrich the representation of industrial product projects, in addition to the attribute information related to industrial products, other product descriptions and product image information are used to construct text triples and image triples.
[0077] Specifically, "product ID, description, and text information" are extracted to construct text triples for representing industrial products and their related descriptive text information, enabling better structured and semantic processing of text data. "Product ID, image, and image information" are extracted to construct image triples for extracting visual content related to industrial products and converting it into processable image data.
[0078] To accurately grasp users' complex needs for industrial products, deeply understand users' interest preferences at different times, and comprehensively reflect the interaction details between users and products, "user ID, product ID, rating, and timestamp" are extracted to construct user-product interaction quadruples to capture users' evaluations of industrial products and their interaction times, thereby providing support for applications such as recommendation systems.
[0079] The quadruples proposed in this embodiment quantify users' degree of preference and demand tendency for products with "ratings", accurately grasping users' complex needs. By recording users' query times with timestamps and introducing time characteristics, the dynamic changes in users' interests can be tracked and analyzed, and their interest preferences at different times can be deeply understood. It can more comprehensively reflect the interaction details between users and products, providing strong support for recommendation systems.
[0080] Through the structured representation of multi-tuples in the above form, multi-modal data can be effectively integrated, enhancing the system's knowledge expression and reasoning capabilities.
[0081] (2)Multi-tuple data mapping
[0082] In the process of processing knowledge multi-tuples, by analyzing various types of information in the multi-tuples, each independent element is mapped to a unique integer identifier, thereby converting the original symbolic or text information into a digital representation. Using digital representation can simplify data processing and calculation, occupy less storage space, be more easily vectorized, and improve calculation efficiency.
[0083] A variety of feature extraction algorithms are adopted to comprehensively extract various types of feature information in the knowledge graph.
[0084] (1)User behavior sequence features
[0085] To extract the interaction features between users and products in the user-product interaction quadruples, the user ID, the interaction behavior between the user and the product, and its corresponding timestamp in the user-product interaction quadruples are extracted, and one-hot encoding is input to obtain the feature information of the interaction between the user and the industrial product, that is, the preliminary user features 。
[0086] It should be understood that the interaction behavior between the user and the product can be indirectly obtained by analyzing and processing the user-product interaction quadruple. For example, the score can be regarded as a quantitative result of an interaction behavior, and different score values can represent different degrees of interaction behavior (such as a high score indicating that the user likes the product and actively interacts, and a low score indicating dissatisfaction, etc.); the timestamp can reflect information such as the time sequence and frequency of the occurrence of the interaction behavior. By analyzing the distribution and interval of the timestamps, behavioral characteristics such as the activity level and periodicity of the user's interaction with the product can be inferred. Those skilled in the art can define the interaction behavior between the user and the product according to the actual situation.
[0087] These preliminary user characteristics are passed as input to the linear graph convolutional module. In this module, operations such as linear aggregation and index lookup are used to optimize the embeddings of the user and the industrial product, generating preliminary embedded representations of the user and the industrial product items. Subsequently, through the attention mechanism layer, the feature representation of the embedding is further enriched and strengthened, capturing more detailed and complex user-industrial product interaction relationships, and then the user behavior sequence features are obtained according to the user's product interaction sequence, interaction timestamp, and product ID .
[0088] (2) Embedding of entity node features
[0089] To enhance the feature mining of multimodal data, the CLIP model is used to extract features from the text triples and image triples in the multimodal industrial product knowledge graph. As Figure 3 shown, through joint training, CLIP can effectively align text and images in a shared feature space, thereby extracting the semantic features of the text and the visual features of the image .
[0090] In the process of feature extraction of the product knowledge graph triples, the ConvE method is used to deeply analyze the user-product interaction quadruple and the product knowledge graph triples to extract the semantic features of entities and relationships. In this part, the user node, product node, and attribute node in the multi-tuple are fused into an entity node containing all nodes. When using the ConvE method to extract entity node features, the relevant entity features and the relationship features need to be initialized first, and then the text features, image features, and the product entity features in the initialized original entity features are fused. The fusion method is as follows:
[0091]
[0092] where, represents the original entity feature embedding, Represents the product entity feature embedding in the original entity embedding, represents the image feature embedding, represents the text feature embedding, and concat() represents the concatenation operation, represents the fused entity embedding.
[0093] After obtaining the fused entity embedding, convolution operations are performed through two convolutional layers to extract local features. The feature map after the convolution operation is flattened into a one-dimensional vector through the view operation, and is ready to be input into the fully connected layer for further processing. According to different aggregation methods, the adjacent node information of the graph and the features extracted by convolution are combined and fused through addition. Finally, the fused features are input into the fully connected layer for output, and all node features are output , including the head entity and the tail entity .
[0094] The attention mechanism layer mainly contains two attention mechanisms. One is the time-sensitive attention mechanism to extract the short-term interests of users, and the other is the multi-head attention mechanism to comprehensively focus on the overall interests of users from multiple aspects.
[0095] In S3, the time-sensitive attention mechanism is used to weight the user-product interactions containing time information to dynamically adjust the weights of user interests, so as to extract its short-term interest representation.
[0096] (1) Time-sensitive attention mechanism
[0097] As Figure 4 shown, the time-sensitive attention mechanism analyzes the time interval and the interaction history between the user and industrial products by considering the user's historical behavior and time factors, and effectively identifies the user's preference for product items within a specific time period.
[0098] 1) Using the interaction time information in the user-product interaction quadruple, calculate the time interval feature of the user. The calculation formula is:
[0099]
[0100] where, and represent the weight and bias parameters respectively, represents the tanh activation function, and represent the timestamps of the user at different time steps, and the time interval feature represents the encoded time step and the time step the absolute time distance between them.
[0101] 2) Construct the intermediate vector of the time-sensitive attention mechanism using the time interval feature and the user behavior sequence feature. The calculation formula is as follows:
[0102]
[0103] Among them, is the user behavior sequence feature, representing the vector representation related to the product at time step in the user u's behavior sequence, represents the time interval feature, and represent learnable weight matrices, represents the bias vector.
[0104] 3) Associate the user's embedding representation with the intermediate vector of the time-sensitive attention mechanism and normalize it using the normalization function to obtain the normalized attention weights. The calculation formula is as follows:
[0105]
[0106] Among them, is the preliminary user feature, representing the user's embedding representation; represents the intermediate vector of the time-sensitive attention mechanism: represents the transpose; represents the intermediate vector obtained by user u at .
[0107] 4) After obtaining the attention weights, perform weighted summation using the attention weights and the user behavior sequence features to obtain the feature representation of the user's short-term interest. The calculation formula is as follows:
[0108]
[0109] In this embodiment, by calculating the time interval feature using the interaction time information, it can accurately analyze the user's behavior changes at different time points and effectively identify the preferences within a specific time period. Further, by combining the time interval feature and the user behavior sequence feature to construct an intermediate vector, and then obtaining the attention weights through association and normalization, the importance of recent interactions can be highlighted. Finally, using the attention weights and the user behavior sequence features for weighted summation to obtain the short-term interest feature representation can timely reflect the user's current interest dynamics, avoid the recommendation results being overly dependent on historical behaviors, make the recommendation more in line with the user's current needs, and improve the timeliness and accuracy of the recommendation.
[0110] In S4, through the multi-head attention mechanism, it aims to comprehensively focus on the overall interest representation of the user and the multi-modal feature representation of industrial products from multiple perspectives using multiple attention heads.
[0111] Such as Figure 5As shown, the multi-head attention mechanism aims to capture multi-level features and complex relationships in the input data from different perspectives by operating multiple attention heads in parallel. Each attention head independently calculates the attention weights and generates corresponding feature representations according to its own focus. Finally, by integrating the outputs of each attention head, the model obtains a more rich and diverse representation ability.
[0112] 1) First, ConvE is used to extract features from multi-modal features to obtain the comprehensive features of nodes and relationship features .
[0113] 2) Using the entity and relationship embeddings extracted from the multi-modal industrial product knowledge graph and user product interaction knowledge graph through relevant operations, the calculation formula for the attention score is:
[0114]
[0115] Among them, is the weight matrix of the k-th head, is the relationship transformation matrix, , and respectively represent the embedded vectors after the fusion of the head entity, relationship, and tail entity in the triple composed of user ID, rating, and product ID in the user product interaction quadruple, the triple in the product knowledge graph, the text triple, and the image triple.
[0116] 3) Collect the first-order structure information by summing the embeddings of the adjacent entities of the head entity h (i.e., the tail entity corresponding to the head entity), and the calculation formula is:
[0117]
[0118] Among them, is the neighbor set of the head entity h.
[0119] 4) Use the double-interaction aggregation function to splice the structure information of all heads and perform an aggregation operation on the head entity and its first-order neighbor information. The calculation formula is:
[0120]
[0121] Among them, and represent learnable transformation matrices, is the element-wise product, and concat represents the splicing operation.
[0122] 5) Capture high-order structure information by stacking multiple attention layers, obtain more supervision signals from distant nodes, and the entity representation can be obtained through recursive calculation. The calculation formula is:
[0123]
[0124] Among them, are all entity feature representations obtained by the multi-head attention mechanism, including user entity feature representations and product entity feature representations .
[0125] In this embodiment, by operating multiple attention heads in parallel, multi-level features and complex relationships of the input data can be captured from different perspectives, enabling the model to have richer and more diverse representation capabilities. Secondly, each attention head independently calculates the attention weights and generates corresponding feature representations, and then integrates and outputs them, which can comprehensively focus on the overall interests of users and the multi-modal features of products. In addition, by summing the embeddings of adjacent entities of the head entity to collect first-order structure information and using a double-interaction aggregation function for aggregation operations, higher-order structure information can also be captured, obtaining more supervision signals from distant nodes, thereby improving the model's understanding and processing capabilities of user-product interaction information and optimizing the recommendation effect.
[0126] In S5, by comprehensively utilizing the comprehensive interest representation of users and the multi-modal feature representation of industrial products, the inner product between the user interest vector and the industrial product item feature vector is calculated to evaluate the similarity between the two, so as to be able to provide accurate product recommendations according to the user's interest preferences.
[0127] (1) Multi-layer representation fusion
[0128] By recursively obtaining representations of different layers, node features and connection information at different levels can be captured, and then a layer aggregation mechanism is used to splice the representations of users and industrial products at each layer, integrating the information of different layers into a unified embedding to more comprehensively represent users and industrial product items. The calculation formula is:
[0129]
[0130]
[0131] Among them, L is the number of aggregation layers, and are the feature representations of users and products at each layer, and concat represents the splicing operation.
[0132] (2) User interest feature fusion
[0133] By fusing the user's short-term interest preference features in the time-sensitive attention mechanism and the user's overall interest preference features obtained by the multi-head attention mechanism, the comprehensive preference feature representation of the user is obtained. The calculation formula is:
[0134]
[0135] Among them, represents the overall user interest representation after splicing different layers, represents the user's short-term interest feature representation, represents the fused comprehensive user preference representation.
[0136] This solution adopts a hybrid attention mechanism. The time-sensitive attention mechanism can capture the user's short-term interest, use the interaction time information to accurately analyze the behavior changes, and avoid over-relying on historical behaviors; the multi-head attention mechanism focuses on the overall interest and multi-modal features from multiple perspectives, enriches the representation ability, and captures high-order structural information. By fusing the user interest features, the features obtained by the two attention mechanisms are combined, which can not only reflect the current interest in a timely manner but also comprehensively understand the user's needs, providing strong support for providing accurate and timely recommendations, and improving the recommendation effect and user experience.
[0137] (3)Calculate the prediction score
[0138] The matching score between the user and the industrial product is predicted by calculating the inner product of their final representations, and the calculation formula is:
[0139]
[0140] Among them, represents the fused comprehensive user preference representation, represents the final representation of the industrial product item after splicing different layers.
[0141] Furthermore, the performance of the method is comprehensively evaluated through a series of evaluation metrics, and through the comparative analysis with traditional recommendation algorithms, the user preference modeling and recommendation performance of the method are evaluated.
[0142] In the development and application process of the recommendation system, the evaluation metrics are the core tools for measuring the performance of the model method. Through reasonable evaluation metrics, we can comprehensively understand the performance of the model method in different dimensions and then objectively evaluate its effect. The following are the evaluation metrics used in the method:
[0143] (1)Precision
[0144] Precision refers to the proportion of the number of products accurately recommended (that is, the products that the user is really interested in or has actually interacted with) in the recommendation list to the total number of products in the recommendation list, and the calculation formula is:
[0145]
[0146] Among them, denotes the number of users in the user set U, and k denotes the top k recommended results considered when calculating precision. is a two-dimensional array, indicating whether the i-th product in the recommendation list of user u is a positive sample.
[0147] (2)Recall
[0148] Recall is a metric that measures the proportion of products that the recommendation system can successfully recall (i.e., recommend) that the user is truly interested in or has actually interacted with among all relevant products. The formula is:
[0149]
[0150] where, denotes the number of users in the user set U, and k denotes the top k recommended results considered when calculating precision. is a two-dimensional array, indicating whether the i-th product in the recommendation list of user u is a positive sample, denotes the number of true relevant products of user u.
[0151] (3)Normalized Discounted Cumulative Gain (NDCG)
[0152] Normalized Discounted Cumulative Gain is an evaluation metric used to measure the performance of a recommendation system or an information retrieval system, and its main purpose is to comprehensively evaluate the relevance and ranking quality of the recommended results.
[0153]
[0154] where, Discounted Cumulative Gain is calculated as:
[0155]
[0156] where, denotes the number of users in the user set U, and k denotes the top k recommended results considered when calculating precision. is an indicator function that is 1 when the i-th product in the recommendation list is in the set of true relevant products R(u) of user u, and 0 otherwise.
[0157] Ideal Discounted Cumulative Gain is calculated as:
[0158]
[0159] where, denotes taking the smaller of the two values of
[0160] After preprocessing the multi-modal product data in this specific embodiment, multi-tuples are extracted, and one-hot encoding and two attention mechanisms are used to obtain user short-term interest features and industrial product features respectively, and then they are fused to obtain user comprehensive preference features and score. In this process, by processing heterogeneous information, the potential associations and complementarities between different modalities can be fully explored while maintaining information integrity. Moreover, in industrial product recommendation, not only product relevance is considered, but also the complex requirements of users in aspects such as production processes, quality standards, cost-effectiveness, etc. are taken into account. At the same time, by fusing user short-term interest features and long-term overall preferences, over-reliance on historical behaviors in the recommendation results is avoided, user interest changes can be effectively captured, the deficiencies of existing recommendation methods in capturing user refined preferences and information fusion can be overcome, and thus the accuracy and practicality of industrial product recommendation can be greatly improved.
[0161] Embodiment 2
[0162] This embodiment provides a multi-modal industrial product recommendation system based on a hybrid attention mechanism, including:
[0163] A data processing layer for obtaining and preprocessing multi-modal product data, where the multi-modal product data includes text data and image data;
[0164] A knowledge graph layer for constructing a multi-modal industrial product knowledge graph based on the preprocessed multi-modal product data, where the multi-modal industrial product knowledge graph includes product information triples, text triples, image triples, and user-product interaction quadruples;
[0165] A knowledge feature extraction layer for using one-hot encoding to extract user behavior sequence features in the user-product interaction quadruple; based on the CLIP model and the ConvE model, extracting entity node feature embeddings of the multi-modal industrial product knowledge graph;
[0166] An attention mechanism layer for using a time-sensitive attention mechanism based on the user behavior sequence features to obtain user short-term interest features; using a multi-head attention mechanism based on the entity node feature embeddings to extract industrial product features;
[0167] An intelligent recommendation and application evaluation layer for fusing user short-term interest features and industrial product features to obtain user comprehensive preference features; making recommendations based on the user comprehensive preference features and scoring the user comprehensive preference features.
[0168] Embodiment 3
[0169] This embodiment provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in a multi-modal industrial product recommendation method based on a hybrid attention mechanism as described in Embodiment 1 above.
[0170] Embodiment 4
[0171] This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a multi-modal industrial product recommendation method based on a hybrid attention mechanism as described in Embodiment 1 above.
[0172] The steps or modules involved in Embodiments 2 to 4 above correspond to those in Embodiment 1. For specific implementation manners, reference may be made to the relevant description part of Embodiment 1. The term "computer-readable storage medium" should be understood to include a single medium or multiple media containing one or more instruction sets; it should also be understood to include any medium that can store, encode, or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.
[0173] The foregoing is only the preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multimodal industrial product recommendation method based on a hybrid attention mechanism, characterized in that: include: Acquiring multimodal product data for preprocessing, wherein the multimodal product data includes text data and image data; Constructing a multimodal industrial product knowledge graph based on the preprocessed multimodal product data, wherein the multimodal industrial product knowledge graph includes a product information triple, a text triple, an image triple, and a user-product interaction quadruple; The user behavior sequence features in the user-product interaction quadruple are extracted using one-hot encoding. Specifically, the user ID, the user-product interaction behavior and its corresponding timestamp are extracted from the user-product interaction quadruple, and preliminary user features are obtained through one-hot encoding. The preliminary user features are input into the linear graph convolution module and the attention mechanism layer to obtain the user's product interaction sequence. The user behavior sequence features are obtained based on the user's product interaction sequence, interaction timestamp and product ID. Based on the CLIP model and ConvE model, the entity node feature embedding of the multimodal industrial product knowledge graph is extracted; Based on the user behavior sequence characteristics, the time-sensitive attention mechanism is used to obtain the user's short-term interest characteristics. Specifically, the interaction time information in the user-product interaction quadruple is extracted to calculate the user's time interval characteristics. Based on the time interval features and user behavior sequence features, the intermediate vector of the time-sensitive attention mechanism is constructed; the preliminary user features and the intermediate vector of the time-sensitive attention mechanism are combined and normalized to obtain the normalized attention weight; the normalized attention weight and the user behavior sequence features are weighted summed to obtain the user's short-term interest features; among which, the user's time interval features are expressed as: ; in, and denote weight and bias parameters respectively, represents the tanh activation function, and Represents the timestamp and time interval features of users at different time steps Represents the encoding time step and time step The absolute time distance between The normalized attention weight is expressed as: ; in, is the preliminary user feature, which represents the embedded representation of the user; The intermediate vector representing the time-sensitive attention mechanism: represents transpose; Indicates that user u is The intermediate vector obtained at Based on entity node feature embedding and relationship embedding, a multi-head attention mechanism is used to extract user overall interest preference features and industrial product features; The user's short-term interest characteristics and the user's overall interest preference characteristics are integrated to obtain the user's comprehensive preference characteristics; recommendations are made based on the user's comprehensive preference characteristics, and the user's comprehensive preference characteristics are scored.
2. A multimodal industrial product recommendation method based on a hybrid attention mechanism as claimed in claim 1, characterized in that: The multimodal industrial product knowledge graph is constructed based on the preprocessed multimodal product data, specifically including: Obtain preprocessed text data and image data, extract product ID, relationship and attribute to construct product knowledge graph triples, extract product ID, description and text information to construct text triples, extract product ID, image and image information to construct image triples, extract user ID, product ID, rating and timestamp to construct user-product interaction quadruples.
3. The multimodal industrial product recommendation method based on a hybrid attention mechanism as claimed in claim 1, characterized in that: The method extracts entity node feature embedding of the multimodal industrial product knowledge graph based on the CLIP model and the ConvE model; specifically includes: Extract text feature embedding of text triples and image feature embedding of image triples based on CLIP model; The ConvE model is used to extract user nodes, product nodes and attribute nodes from the user-product interaction quadruple and the product knowledge graph triple to obtain the initial entity features. Based on the initial entity features, the initial product entity feature embedding is obtained. The entity node feature embedding is obtained based on the text feature embedding, image feature embedding and initialized product entity feature embedding.
4. The multimodal industrial product recommendation method based on a hybrid attention mechanism as claimed in claim 1, characterized in that: The multi-head attention mechanism is adopted based on entity node feature embedding and relationship embedding to extract user overall interest preference features and industrial product features, specifically including: Calculate the attention score based on entity node feature embedding and relationship embedding; Based on the attention score, the adjacent embeddings of the head entity are summed to obtain the first-order structural information; The double interactive aggregation function is used to concatenate the structural information of all heads, and the head entities and their first-order structural information are aggregated. By stacking multiple attention layers, high-order structural information is captured to obtain the user's overall interest preference characteristics and industrial product characteristics.
5. The multimodal industrial product recommendation method based on a hybrid attention mechanism as claimed in claim 1, characterized in that: The scoring of the user's comprehensive preference characteristics specifically includes: in, represents the comprehensive preference characteristics of users after fusion, The final representation of industrial product features after splicing different layers is shown. Indicates transpose.
6. A multimodal industrial product recommendation system based on a hybrid attention mechanism, used to execute a multimodal industrial product recommendation method based on a hybrid attention mechanism as claimed in claim 1, characterized in that: include: A data processing layer, used to obtain multimodal product data for preprocessing, wherein the multimodal product data includes text data and image data; A knowledge graph layer, used to construct a multimodal industrial product knowledge graph based on the preprocessed multimodal product data, wherein the multimodal industrial product knowledge graph includes a product information triple, a text triple, an image triple, and a user-product interaction quadruple; The knowledge feature extraction layer is used to extract the user behavior sequence features in the user-product interaction quadruple using one-hot encoding; based on the CLIP model and the ConvE model, it extracts the entity node feature embedding of the multimodal industrial product knowledge graph; The attention mechanism layer is used to obtain the user's short-term interest features by using the time-sensitive attention mechanism based on the user's behavior sequence characteristics; Based on entity node feature embedding and relationship embedding, a multi-head attention mechanism is used to extract user overall interest preference features and industrial product features; Intelligent recommendation and application evaluation layer, used to integrate user short-term interest characteristics with user overall interest preference characteristics to obtain user comprehensive preference characteristics; Recommendations are made based on users' overall preference characteristics and scores are given to users' overall preference characteristics.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in a multimodal industrial product recommendation method based on a hybrid attention mechanism as described in any one of claims 1 to 5 are implemented.
8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the multimodal industrial product recommendation method based on a hybrid attention mechanism as described in any one of claims 1-5 are implemented.
Citation Information
Patent Citations
Recommendation algorithm based on knowledge graph and neural network
CN116955647A
Serialization recommendation method based on knowledge graph
CN117390275A
Group video recommendation system and method fusing multi-modal information and interest similarity
CN118690037A