AI search recommendation method, system and equipment combined with commodity semantic understanding and medium

By combining AI search recommendation methods with product semantic understanding and using convolutional neural networks and recurrent neural networks to generate comprehensive semantic vectors for products and users, the problems of insufficient semantic association and interest modeling in traditional recommendation methods are solved, and efficient personalized recommendations are achieved in cold start scenarios.

CN120612154APending Publication Date: 2025-09-09河北燕鸣科技有限公司
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510759646.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Traditional recommendation methods are unable to fully explore the deep semantic associations between product text descriptions and image information, and are difficult to model the dynamic evolution of user interests and personalized preferences, resulting in decreased accuracy and poor generalization ability of recommendation systems in cold start scenarios.

Method used

A method combining convolutional neural networks and recurrent neural networks is used to extract semantic vectors through multi-layer BERT encoding and multi-head self-attention mechanism of product titles and descriptions. The residual connection and spatial attention of product images are combined to enhance feature extraction to generate a comprehensive semantic vector. The interest vector is updated using user behavior data for joint modeling and recommendation.

Benefits of technology

It improves the performance of the recommendation system in personalized and cold start scenarios, has good diversity and adaptability, can accurately match the deep semantics of users and products, and enhances the expressiveness and discriminative ability of the recommendation model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612154A_ABST
    Figure CN120612154A_ABST
Patent Text Reader

Abstract

The invention discloses an AI search recommendation method, system and device combined with commodity semantic understanding and a medium, and belongs to the technical field of commodity recommendation. The method comprises the steps that semantic vectors of commodity titles and descriptions are generated; generating a feature vector of the commodity image; fusing the semantic vector of the commodity title and description with the feature vector of the commodity image to generate a comprehensive semantic vector of the commodity; generating a user interest vector; updating the user interest vector by using a recurrent neural network based on the user interest vector; and carrying out joint modeling on the comprehensive semantic vector of the commodity and the updated user interest vector to generate a recommendation result. According to the method, a content semantic driving modeling mode is adopted, recommendation judgment can still be conducted through the deep matching relation between commodity image-text semantics and user behavior preferences even under the condition that user historical behaviors are limited or new commodities are online, good cold start adaptability is achieved, and the method is suitable for popularization and application. And meanwhile, high recommendation difference and content diversity are shown for different user groups.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of product recommendation, and in particular to an AI search recommendation method, system, device and medium combined with product semantic understanding. Background Art

[0002] In scenarios like e-commerce, content recommendations, and short video push, with a large user base and diverse product categories, recommendation systems are becoming a key means of improving user conversion rates and satisfaction. Traditional recommendation methods primarily rely on collaborative filtering, matrix decomposition, or user click statistics for modeling, which have the following significant limitations:

[0003] On the one hand, traditional models have limited ability to understand product content and are unable to fully explore the deep semantic associations between product text descriptions and image information. This leads to a decrease in recommendation accuracy when dealing with long-tail products and cold-start products.

[0004] On the other hand, user interests are typically multi-layered and dynamically evolving, with different user behaviors (such as browsing, clicking, and purchasing) expressing different preferences. Existing technologies often use static interest vectors or simple time series models, which cannot effectively separate and integrate short-term user interests and long-term preferences, and also struggle to model the impact of different behavior types on interest evolution.

[0005] In addition, in the process of modeling the matching of user interests and products, traditional dot product or splicing methods are difficult to capture high-order semantic dependencies, and lack personalized modeling mechanisms for user behavior preferences, resulting in poor generalization ability of recommendation results and inability to fully explore the deep connection between users and products.

[0006] Therefore, there is an urgent need for a recommendation modeling method that combines multimodal product semantic understanding and user behavior perception capabilities. Based on the accurate extraction of product text and image information, a dynamic and structured user interest expression is constructed. By introducing a behavior category-sensitive mechanism, a deep semantic matching between user interests and products is achieved, thereby improving the overall performance of the recommendation system in personalization, diversity and cold start scenarios. Summary of the Invention

[0007] To solve the above technical problems, an AI search and recommendation method combined with product semantic understanding is proposed, including obtaining the title, description and image information of the product; performing text processing on the title and description to generate semantic vectors of the product title and description; using a convolutional neural network to extract features of the image information to generate a feature vector of the product image; fusing the semantic vectors of the title and description with the feature vector of the image information to generate a comprehensive semantic vector of the product; extracting the user's browsing, clicking and purchasing records from user behavior data to generate a user interest vector; based on the user interest vector, using a recursive neural network to update the user interest vector; and jointly modeling the comprehensive semantic vector of the product with the updated user interest vector to generate a recommendation result.

[0008] As a preferred solution of the AI ​​search recommendation method combined with product semantic understanding described in the present invention, the generating semantic vectors of product titles and descriptions includes: segmenting the product titles and descriptions, removing stop words, and performing part-of-speech tagging; adding field position codes to the processed product titles and descriptions respectively to represent the field category information of the titles and descriptions; inputting the text after adding the position codes into a pre-trained BERT model, performing multi-layer bidirectional encoding, and obtaining context-related vectors of each word; in the encoding process, using a multi-head self-attention mechanism to model the context dependency of the words in each field, and assigning independent training parameters to the attention heads of different fields; based on the user clicks and purchase behaviors recorded in the training phase, normalizing the average attention weight of the words in the field with the response frequency of the field in the user behavior, and determining the global weight coefficient of each field; weighted averaging the global weight coefficient of each field with the encoding vectors of all words in the field to generate a product title semantic vector and a product description semantic vector.

[0009] As a preferred solution of the AI ​​search recommendation method combined with product semantic understanding described in the present invention, the method further comprises: generating a feature vector of a product image, adjusting the product image to 224 pixels by 224 pixels, and performing histogram equalization processing on each pixel channel; inputting the processed product image into a convolutional neural network with a residual connection structure and a spatial attention module; sequentially performing convolution, activation, normalization, and downsampling operations in the convolutional neural network to extract feature maps of the second, third, and fourth residual modules; adjusting the extracted feature maps to a uniform spatial size using bilinear interpolation; weighting the channel response values ​​of each spatial position on the feature map of the uniform spatial size using the spatial weight map generated by the spatial attention module; splicing the weighted feature maps according to the channel dimension; and performing global average pooling on the spliced ​​feature maps to generate a fixed-dimensional product image feature vector.

[0010] As a preferred solution of the AI ​​search recommendation method combined with product semantic understanding described in the present invention, wherein: the generation of the comprehensive semantic vector of the product includes splicing the product title semantic vector and the product description semantic vector to form a text comprehensive semantic vector; the text comprehensive semantic vector and the product image feature vector are respectively input into the mapping module, and the mapping module includes a linear transformation layer with shared parameters, which is used to map the two input vectors to a fusion space of the same dimension; based on the user's historical click behavior, the frequency normalized ratio of the contribution of text and image modalities to user recommendation clicks in the training set is calculated as the initial fusion modality weight vector; according to the mapped text and image The method comprises the following steps: calculating the cosine similarity between image vectors, calculating the cross-modal attention weight matrix, and performing weighted control according to the initial modal weight vector; applying the controlled cross-modal attention weight matrix to the two modal vectors to obtain a fusion vector; introducing a modal residual gating structure into the fusion vector, which automatically controls whether to retain text modal information, image modal information, or output both at the same time according to context-related parameters; performing comparative training on the fused vector and the semantic vector of the positive example product of the user interaction in the recommendation system, and constructing a multimodal semantic alignment constraint by minimizing the contrast loss; performing normalization processing on the fusion vector after the alignment training to generate a comprehensive semantic vector of the product.

[0011] As a preferred solution of the AI ​​search recommendation method combined with product semantic understanding described in the present invention, the generating of the user interest vector includes extracting the interaction records between the user and the product from the user behavior log, including three types of behaviors: browsing, clicking and purchasing, and each behavior record contains the product identifier, behavior type and timestamp; using the comprehensive semantic vector of the product corresponding to the product identifier in each behavior record as the behavior embedding vector, and assigning fixed behavior weights to different types of behaviors; multiplying the behavior embedding vector of each behavior record by the corresponding behavior weight, and constructing a user behavior sequence in timestamp order; introducing a time decay factor into the behavior sequence, assigning exponential decay weights to behaviors at earlier times, and obtaining a weighted behavior sequence tensor; inputting the weighted behavior sequence into a multi-layer one-dimensional convolution structure for feature extraction to extract user interest trajectory features; and normalizing the extracted feature vector to generate a user interest vector.

[0012] As a preferred solution of the AI ​​search recommendation method combined with product semantic understanding described in the present invention, the method includes: using a recursive neural network to update the user interest vector includes splitting the user interest vector into a short-term interest vector and a long-term interest vector, the short-term interest vector is generated based on the most recent interactive behavior, and the long-term interest vector is generated based on the historical accumulated behavior; the comprehensive semantic vector of the product corresponding to the user's latest behavior record is used as the current behavior vector of the input sequence; a dual-track recursive neural network containing a gated recurrent unit is constructed, the short-term interest track and the long-term interest track respectively contain an update gate and a reset gate, and the parameters of the update gate and the reset gate are dynamically adjusted according to the behavior type, and the behavior type includes browsing, clicking and purchasing; the input sequence is recursively updated in the short-term interest track and the long-term interest track respectively, the short-term track is highly sensitive to the time sequence, and the long-term track introduces a time attenuation coefficient to suppress the influence of outdated behavior; the updated short-term interest vector and the long-term interest vector are fused through a weighting coefficient sensitive to the behavior category to generate an updated user interest vector.

[0013] As a preferred solution of the AI ​​search recommendation method combined with product semantic understanding described in the present invention, the generating of recommendation results includes normalizing the comprehensive semantic vector of the product and the updated user interest vector respectively; inputting the normalized comprehensive semantic vector of the product and the user interest vector into a behavior category-sensitive bidirectional attention matching module to calculate the attention weight matrix, and the weight matrix is ​​adjusted according to the correlation between the dimensions of the two vectors and the category weight of the user's historical behavior type; according to the attention weight matrix, the dimensional response values ​​of the comprehensive semantic vector of the product and the user interest vector are weighted to generate a matching feature vector; the matching feature vector is input into the recommendation scoring module, and the recommendation scoring module linearly transforms the matching feature vector according to the trainable parameters generated by the historical feedback of the user behavior category, and outputs the recommendation score; in the recommendation scoring training process, a contrastive learning loss function sensitive to the user's historical behavior category is introduced to minimize the difference between the scores of positive and negative products, and at the same time, the update rule of the attention matrix is ​​adjusted according to the preference consistency of the user's historical interest behavior to optimize the distinguishing ability of the recommendation score.

[0014] As a preferred solution of the AI ​​search recommendation system combined with product semantic understanding described in the present invention, it is characterized by including: a data acquisition module for obtaining the title, description and image information of the product; a text processing module for performing text processing on the product title and description to generate semantic vectors of the product title and description; an image processing module for using a convolutional neural network to extract features of the product image to generate a feature vector of the product image; a feature fusion module for fusing the semantic vector of the product title and description with the feature vector of the product image to generate a comprehensive semantic vector of the product; a behavior analysis module for extracting the user's browsing, clicking and purchasing records from the user behavior data to generate the user's interest orientation. interest update module, which is used to update the user interest vector based on the user interest vector using a recursive neural network; a recommendation generation module, which is used to jointly model the comprehensive semantic vector of the product with the updated user interest vector to generate recommendation results; a query understanding module, which is used to perform semantic understanding of the query submitted by the user, generate a query vector, and calculate the similarity between the query vector and the comprehensive semantic vector of the product; a ranking weighting module, which is used to sort according to the similarity between the query vector and the comprehensive semantic vector of the product, and weight it according to the price and evaluation information of the product; a feedback update module, which is used to update the user interest vector in real time according to the interactive feedback between the user and the product, and adjust the recommendation results according to the updated interest vector.

[0015] A computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, it implements the steps of the AI ​​search recommendation method combined with product semantic understanding.

[0016] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of an AI search recommendation method combined with product semantic understanding.

[0017] The beneficial effects of the present invention are as follows: by subjecting product titles and description texts to multi-layer BERT encoding, field position embedding, and multi-head self-attention mechanism processing, and combining them with convolutional feature extraction of product images through residual connection and spatial attention enhancement, the fusion modeling of product multimodal semantic information is effectively realized, and the model's ability to understand product semantic features is enhanced, especially in scenarios where image and text information are inconsistent or text descriptions are incomplete.

[0018] By assigning behavioral category weights to user behaviors according to browsing, clicking, and purchasing types, and introducing a time decay factor in combination with timestamps, interest trajectories are modeled in a one-dimensional convolutional network. Further, with the help of a dual-track gated recurrent network, short-term interests and long-term preferences are captured respectively, achieving behavioral sensitivity and temporal dynamics in interest modeling, and avoiding interest drift or information forgetting problems.

[0019] A behavioral category-sensitive bidirectional attention mechanism is introduced. By constructing a correlation matrix between user interest vectors and comprehensive product semantic vectors and integrating historical behavioral preferences to generate weighted matching features, the recommendation modeling is endowed with behavioral adaptability and dimensional response control capabilities, effectively improving the expressiveness and discriminative capabilities of the recommendation scoring model.

[0020] The present invention adopts a content semantics-driven modeling approach. Even when the user's historical behavior is limited or new products are launched, recommendation judgments can still be made through the deep matching relationship between the product's image and text semantics and user behavior preferences. It has good cold start adaptability and shows a high degree of recommendation differentiation and content diversity for different user groups. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0022] Figure 1 This is a general flow chart of an AI search and recommendation method combined with product semantic understanding, provided in one embodiment of the present invention. DETAILED DESCRIPTION

[0023] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.

[0024] Example 1, reference Figure 1 , which is the first embodiment of the present invention, provides an AI search recommendation method combined with product semantic understanding, including:

[0025] Step 1: Get the product's title, description, and image information.

[0026] In step 1, the basic information of the products to be recommended on the platform is first collected and structured. The product information includes two types of fields: text information fields and image information fields, which are used for subsequent semantic modeling and image feature extraction processing respectively.

[0027] Product text information includes the product title and description. The product title is a short text entered by the merchant when listing a product, typically including keywords such as the product name, brand, model, and category. The product description is natural language text entered by the merchant to supplement the product details, typically including information such as the product material, specifications, functions, and usage. Product image information is an image uploaded by the merchant to showcase the product's appearance, typically a main front view or a typical angled image. The image comes from the platform's image and text publishing system or the merchant's content management interface.

[0028] To ensure the consistency and structure of the collected information, the product information collection and processing process includes the following:

[0029] The product title and description are checked for legality and integrity. If there are missing fields or abnormal information, they are handled according to the platform's content integrity policy, including backtracking historical version content or pushing supplementary reminders to merchants to ensure that each product data has three fields: title, description and main image; the character content in the product title and description fields is standardized, including removing invalid symbols, redundant spaces, repeated punctuation, format control characters, etc., unifying the text encoding format, and establishing field identifiers for product text fields to distinguish the source of fields in subsequent semantic modeling; the format and resolution of product images are processed, including unified conversion to three-channel color image format, unified adjustment of image size to a preset size, and completion of image color space conversion and resolution normalization to meet the input requirements of the subsequent image feature extraction module; a mapping relationship between product information and product identification is established, and the product identification is bound to the title text, description text and image data one by one to build a unified product data structure to provide original input support for subsequent semantic vector generation, image feature extraction and multimodal fusion operations.

[0030] The product titles, descriptions, and images are derived from data uploaded by merchants and reviewed by the platform. They do not contain any personal images, identification information, or other potentially personally identifiable content. This information is used solely as input for product semantic understanding and content feature analysis and does not affect user privacy or personal information.

[0031] Step 2: Perform text processing on the product title and description to generate semantic vectors for the product title and description.

[0032] In step 2, generating semantic vectors for product titles and descriptions includes segmenting the product titles and descriptions, removing stop words, and performing part-of-speech tagging; adding field position codes to the processed product titles and descriptions to indicate the field category information of the titles and descriptions; inputting the text after adding the position codes into a pre-trained BERT model, performing multi-layer bidirectional encoding, and obtaining context-related vectors for each word; during the encoding process, using a multi-head self-attention mechanism to model the contextual dependencies of the words in each field, and assigning independent training parameters to the attention heads of different fields; based on the user clicks and purchase behaviors recorded during the training phase, normalizing the average attention weight of the words in the field with the response frequency of the field in the user behavior to determine the global weight coefficient of each field; and taking the weighted average of the global weight coefficient of each field and the encoding vectors of all words in the field to generate the product title semantic vector and the product description semantic vector.

[0033] Furthermore, to enhance the expressive power of product text in recommendation tasks, this invention employs a structurally improved multi-layer bidirectional semantic encoder to generate semantic representations for product titles and descriptions. This module not only builds on the existing pre-trained language model framework but also incorporates field structure information and a dynamic attention mechanism guided by user behavior, thereby ensuring that the semantic modeling process is adaptable and responsive to recommendation scenarios.

[0034] The implementation process of this step includes the following key parts:

[0035] 1. Field text preprocessing and structure encoding:

[0036] The product title text and product description text obtained in step 1 are respectively subjected to word-level segmentation, stop word removal and part-of-speech tagging, retaining the semantic core terms, and embedding field identifier codes in the text, that is, setting identifier vectors for the title and description respectively as part of the input embedding to ensure that the model can distinguish different text sources during training and provide a basis for the subsequent field weighting mechanism.

[0037] 2. Embedding layer and field position control module:

[0038] Before entering the semantic encoding model, product titles and descriptions undergo a word embedding layer and a field position encoding layer. The word embedding layer converts terms into fixed-length vector representations, while the field position encoding indicates the field in which the term resides and controls the attention distribution structure in the encoder. This embedding approach ensures that the model learns distinct representations when processing the same term from different field sources, providing a structural foundation for subsequent field semantic comparison.

[0039] 3. Multi-layer encoder model with improved structure:

[0040] This paper adopts a multi-layer bidirectional semantic encoder based on the Transformer structure. Each layer of the encoder contains a multi-head self-attention substructure and a feedforward substructure, and introduces the following innovative mechanisms in the attention module:

[0041] Field Grouping Attention Mechanism:

[0042] Separate attention channels are constructed for the title field and the description field. That is, within the same encoder layer, two sets of parameters are used to model the internal dependencies of the fields to avoid cross-field interference. Behavior-sensitive weight initialization strategy:

[0043] During the initial parameter setting of the training phase, the initial attention distribution of each field channel is dynamically set based on the historical user click preference ratio for titles or descriptions. This allows the model to have a recommendation scenario bias at the beginning of training, improving convergence speed. The attention training mechanism with enhanced behavioral response:

[0044] During the back-propagation training phase, a field behavior response bias term is introduced to calculate the difference between the user's actual click data and the field attention distribution generated by the current model. This is used as the attention error feedback term and is jointly optimized with the traditional loss to achieve the behavioral preference coupled update of the attention mechanism.

[0045] 4. Field-level semantic aggregation and weighted output:

[0046] After each field is encoded in multiple layers, the context-related vector of each term is obtained, and then aggregation processing is performed on a field-by-field basis:

[0047] The context vectors of the terms in each field are averaged to obtain the basic field representation. A field-level weight coefficient is introduced, which is generated by weighting the field click-through rate and purchase conversion rate in historical behaviors. The value is continuously updated during the training process. The field vector is weighted and fused with the field weight to finally output the product title semantic vector and product description semantic vector respectively.

[0048] 5. Output structure and downstream connection relationship:

[0049] Finally, the two semantic vectors output from step 2 are used for multimodal fusion processing in step 4. Their dimensions, structure, and embedding space remain consistent, facilitating subsequent joint image and text modeling. Furthermore, the training process of this module shares some feedback metrics with the user behavior modeling module, forming a cross-module behavioral response closed loop and improving the global optimization capabilities of the overall recommendation system.

[0050] Step 3: Use a convolutional neural network to extract features from product images and generate feature vectors of product images.

[0051] In step 3, generating a feature vector of the product image includes adjusting the product image to 224 pixels by 224 pixels and performing histogram equalization on each pixel channel; inputting the processed product image into a convolutional neural network with a residual connection structure and a spatial attention module; sequentially performing convolution, activation, normalization, and downsampling operations in the convolutional neural network to extract feature maps of the second, third, and fourth residual modules; adjusting the extracted feature maps to a uniform spatial size using bilinear interpolation; weighting the channel response values ​​of each spatial position on the feature map of uniform spatial size using a spatial weight map generated by the spatial attention module; splicing the weighted feature maps according to the channel dimension; and performing global average pooling on the spliced ​​feature maps to generate a fixed-dimensional product image feature vector.

[0052] It should be noted that in order to achieve the discriminability and user behavior responsiveness of product images in the recommendation system, the present invention constructs a structurally optimized convolutional neural network model, combines the spatial attention mechanism guided by user behavior heat, the channel gating strategy sensitive to product categories, and the feature comparison learning process of image interest alignment, performs multi-level semantic modeling on the main product image, and finally generates a product image feature vector with semantic consistency and recommendation guidance capabilities.

[0053] The process consists of the following interrelated sub-steps:

[0054] Sub-step 3.1: Product image preprocessing:

[0055] First, the input product images are uniformly preprocessed to ensure that the image data format is compatible with the subsequent network structure:

[0056] The original product image is scaled to 224×224 pixels and uniformly converted to RGB three-channel format. Histogram equalization is performed on the pixel values ​​of each channel to enhance image contrast and edge clarity. The image pixel values ​​are normalized so that their value range is distributed in the range of [0,1]. The processed image data is constructed into a tensor format for input into the convolutional neural network.

[0057] The above preprocessing steps ensure the uniformity of image input in resolution, brightness and structural format, and provide a basis for the training stability and feature expression consistency of deep models.

[0058] Sub-step 3.2: Multi-layer residual convolution feature extraction network construction:

[0059] After the image is input, multi-scale semantic features are extracted through the following structural layers in sequence:

[0060] The input image passes through a set of basic convolution modules to complete preliminary feature extraction and signal compression; multiple residual connection modules are set up, including the second, third and fourth layer residual modules, each module contains a dual convolution substructure and an identity mapping path; at the end of each residual module, the corresponding feature map is extracted to retain low-level texture information, mid-level local structure information and high-level semantic information respectively.

[0061] The introduction of the residual structure effectively prevents the gradient vanishing problem when the network is stacked in multiple layers, ensuring the high-level network's ability to abstractly model complex semantic relationships.

[0062] Sub-step 3.3: Spatial size alignment and multi-scale feature fusion preparation:

[0063] Since the spatial sizes of the feature maps output by each residual module are inconsistent, they need to be processed uniformly:

[0064] Bilinear interpolation operations are performed on the feature maps output by the second, third, and fourth residual modules respectively, and adjusted to a unified spatial dimension (such as 28×28); this ensures that in the subsequent processing of the spatial attention module, each feature map can be spliced ​​and fused according to the channel dimension.

[0065] Spatial scale unification is a prerequisite for achieving cross-scale feature fusion and unified attention modeling.

[0066] Sub-step 3.4: Behavioral response-guided spatial attention modeling:

[0067] To guide the model to focus on the user's actual preferred areas, a behavior-responsive thermal guidance mechanism is introduced to construct spatial attention:

[0068] During the training phase, the position distribution of clicked areas in images of similar products is statistically analyzed from historical user click records to generate behavioral thermal templates corresponding to the categories. In the early stages of model training, the behavioral thermal templates are injected into the attention module as the initial spatial attention weights. During the training process, the attention bias loss term is introduced, i.e., the squared error of the deviation between the actual generated attention map and the behavioral thermal template, and is jointly optimized with the main loss. The attention weight is finally used as a channel weighting factor to weight the feature response value of each spatial position to generate a feature map with enhanced behavioral attention.

[0069] This mechanism ensures that the model not only establishes an attention structure based on the image content, but also couples it with the distribution of users' actual click behavior, thereby enhancing the model's interest adaptation ability.

[0070] Sub-step 3.5: Category-sensitive channel gating regulation:

[0071] Considering the differences in image structural characteristics of different product categories, a category-guided channel gating mechanism is introduced:

[0072] The first-level category to which the product belongs is embedded and encoded to form a category vector; the category vector is mapped to a channel control vector through an activation function, and a weighted mask is generated for the output feature channel of the residual module; the response value of each channel is multiplied according to the mask value to achieve gated adjustment of "enhancement / inhibition / ignorance" of the channel response; the gate structure parameters are jointly trained with the main network parameters to enable the model to automatically adapt to the image modeling path according to the category.

[0073] This mechanism improves the modeling flexibility and expression resolution capabilities of the network structure for multi-category product images.

[0074] It should be noted that a spatial attention guidance method based on user behavior heat map is proposed. The click density of users on the focus area of ​​product images in historical interactions is mapped into a thermal template, and the template is used as the initial attention weight in the early stage of training. At the same time, a thermal deviation loss term is introduced during the training process to constrain the deviation between the attention distribution learned by the model and the user's actual focus area, thereby realizing the coupled learning of spatial attention and user behavior.

[0075] This mechanism breaks through the limitation of existing general image modeling methods that "attention is only adaptively generated by image content and lacks user response supervision", avoids the model learning to focus on invalid areas (such as background or non-identifiable textures), and improves the relevance of attention distribution to recommendation targets.

[0076] Compared with the attention mechanism in existing image classification or recognition tasks, this invention introduces the supervision signal of user behavior response feedback to construct a mapping path between behavior thermal distribution and image area response. It is the first time to explicitly introduce user interests into image attention training in recommendation tasks, substantially enhancing the ability of image semantic expression to respond to user preferences.

[0077] Sub-step 3.6: Image feature comparison training and interest alignment mechanism:

[0078] To achieve consistent expression of image features and user preferences, a contrastive learning mechanism for image interest alignment is introduced:

[0079] In the training samples, the product images that the user has clicked on in the past are extracted as positive samples, forming a "positive pair" with the current product image; the product images that have not been interacted with by the user are randomly sampled as negative samples, forming a "negative pair" with the current image; a contrast loss function is constructed between image feature vectors to minimize the distance between positive pairs and maximize the distance between negative pairs; the contrast loss function is combined with the recommendation scoring loss and attention bias loss to form the total loss, achieving multi-objective optimization.

[0080] This mechanism realizes the adaptive adjustment of the user's interest direction in the image vector space, so that the image features have interest-guiding semantics in the vector space.

[0081] It should be noted that a channel activation control structure driven by product categories is constructed. The first-level category to which the product belongs is input into the embedding coding module to generate a category vector, and then the category vector is mapped to a channel-level weight mask to adjust the activation intensity of each channel in the output feature map of the residual convolution module, thereby realizing dynamic activation, inhibition or shielding at the channel level.

[0082] This mechanism addresses the problem of "fixed channel structure and no differentiated response" in existing image modeling methods when processing images of multi-category products. It can automatically adapt image encoding strategies based on the visual structure differences between product types. For example, for home appliances, the model can strengthen the structural outline and background consistency channels; for clothing products, it can enhance the color and texture channels, improving the model's scene adaptability and feature sparsity.

[0083] Compared with the existing general CNN structure, this mechanism does not rely on model reconstruction or multi-model switching. It only dynamically adjusts the channel through lightweight gating vectors, substantially optimizing the model's ability to express multiple categories of product images, improving parameter utilization efficiency and feature difference representation capabilities, and constituting an innovative extension of the convolutional structure with domain perception capabilities.

[0084] Sub-step 3.7: Feature fusion and semantic vector generation:

[0085] The three sets of feature maps after attention adjustment are spliced ​​in the channel dimension to form a high-dimensional fused feature map; a global average pooling operation is performed on the fused feature map to remove the spatial dimension and retain the channel semantics; a fixed-dimensional product image feature vector is output as the semantic expression of the product image; the image semantic vector will be modally fused with the text semantic vector in subsequent steps to generate a comprehensive semantic representation of the product.

[0086] It should be noted that this innovation designs an image feature comparison learning mechanism based on the user's historical click behavior, taking the product images that the user has clicked as positive examples of interest and the non-interacted images as negative examples. By constructing a contrast loss function that maximizes the similarity and minimizes the difference between image pairs, the image feature encoding network is guided to aggregate images with consistent interest and distinguish images with irrelevant interest, thereby improving the coupling consistency between the image semantic vector and the user's preference.

[0087] This mechanism effectively solves the problem of existing image modeling in recommendation systems that "features only express image content but not user interest semantics". It enables image features to not only have structural representation capabilities but also user interest alignment capabilities, and builds an organic semantic closed loop between user behavior → image features → recommendation response.

[0088] Unlike traditional image coding models that only rely on supervised labels for optimization, this mechanism constructs the image semantic space structure through the user interest dimension, expanding the image feature training goal from "correct image classification" to "enhanced interest semantic consistency", and technically achieving a leap from visual representation optimization to recommendation response modeling.

[0089] In the specific design of the spatial attention mechanism, the present invention adopts a lightweight channel-space joint attention structure, whose core process includes three parts: spatial compression, channel fusion, and position weighting. Specifically: first, channel dimension average pooling and maximum pooling are performed on the feature map of uniform size to obtain two spatial response maps respectively; then the two response maps are jointly input into a convolution layer containing a convolution kernel size of 7×7 through a splicing operation, and the initial spatial weight map is output; finally, the weight map is processed by the activation function and bit-by-bit multiplication operation is performed with each spatial position element of the original feature map to achieve response enhancement of key visual areas. While maintaining high efficiency, this structure can effectively capture the salient areas in the product image and adapt to the recommendation system's modeling requirements for key areas of product display.

[0090] The product image feature vector generated in the present invention is a real number vector of fixed dimension, preferably 1024. The dimension value is determined based on the performance evaluation and recommendation accuracy performance during the model training process. During the experiment, the vector dimension was evaluated under multiple sets of values ​​such as 256, 512, 768, and 1024. Finally, 1024 dimensions were selected as the output vector length, which has the best balance between training stability, fusion compatibility and recommendation accuracy. This dimension setting is compatible with the text semantic vector structure in the vector space, facilitates image-text fusion modeling, and ensures sufficient semantic expression capabilities in the recommendation scoring module without excessively increasing the complexity of the model.

[0091] In the image interest alignment training, the present invention adopts an interest neighbor sampling strategy based on user behavior trajectories to construct positive and negative sample pairs. Specifically, the positive image samples are derived from product images that have real click or purchase behaviors with the target user to ensure that they have interest consistency in the semantic space; the negative image samples are randomly sampled from product images of the same category but that have not interacted with the user, and a behavioral relevance threshold is set to eliminate potential candidate products of interest to avoid misjudgment affecting the training effect. In addition, in order to improve training efficiency and generalization ability, at least 1 positive pair and 3 negative pairs are constructed in each batch of training samples to ensure that the contrast loss has sufficient information tension in each round of training, thereby guiding the model to accurately distinguish the user's preference direction.

[0092] In order to achieve the optimal adaptation of the image modeling process to the recommendation target, the present invention designs a joint loss function consisting of three parts, namely recommendation score loss, behavior thermal deviation loss and image contrast loss. The recommendation score loss adopts a comparative training strategy based on the difference in scores between positive and negative product samples; the behavior deviation loss calculates the L2 norm between the current spatial attention map and the historical click heat map to guide the attention optimization direction; the image contrast loss calculates the maximization of the cosine similarity between the positive image features and the target image features, and minimizes the similarity of the negative image features. The three loss functions are weighted summed to form a total loss. The weight parameters are determined by cross-validation before training and remain constant during the training process, so that the model takes into account both interest direction and structural expression ability in the image feature extraction process.

[0093] Step 4: Fuse the semantic vectors of the product title and description with the feature vector of the product image to generate a comprehensive semantic vector for the product.

[0094] In step 4, the generation of the comprehensive semantic vector of the product includes: concatenating the product title semantic vector and the product description semantic vector to form a text comprehensive semantic vector; inputting the text comprehensive semantic vector and the product image feature vector generated in step 3 into the mapping module respectively, and the mapping module includes a linear transformation layer with shared parameters, which is used to map the two input vectors to a fusion space of the same dimension; based on the user's historical click behavior, calculating the frequency normalized ratio of the contribution of the text and image modalities to the user recommendation click in the training set as the initial fusion modality weight vector; calculating the cross-modality weight vector based on the cosine similarity between the mapped text and image vectors. The method comprises the following steps: a modal attention weight matrix is ​​obtained, and weighted control is performed according to the initial modal weight vector; the regulated cross-modal attention weight matrix is ​​applied to the two modal vectors to obtain a fusion vector; a modal residual gating structure is introduced into the fusion vector, and the structure automatically controls whether to retain text modal information, image modal information or output both at the same time according to context-related parameters; the fused vector is compared with the semantic vector of the positive example product of the user interaction in the recommendation system for training, and a multimodal semantic alignment constraint is constructed by minimizing the contrast loss; the fused vector after alignment training is normalized to generate a comprehensive semantic vector of the product.

[0095] In the present invention, in order to achieve a unified expression of the multimodal semantics of the product in terms of text and images, the text semantic vector and the image feature vector are aligned and modeled through a structured modal fusion path to generate a comprehensive semantic vector of the product with multimodal semantic integrity and recommendation response consistency.

[0096] The implementation process of this step includes the following sub-steps:

[0097] Sub-step 4.1: Text semantic vector concatenation and mapping:

[0098] The product title semantic vector and product description semantic vector generated in step 2 are concatenated according to their vector dimensions to form a comprehensive text semantic vector. The concatenation operation maintains the original semantic structure without compression or weight fusion, preserving the independent feature expression capabilities of the title and description fields in subsequent fusion.

[0099] Next, the text's comprehensive semantic vector and the product image feature vector generated in step 3 are fed into a unified mapping module. This module, comprised of parameter-sharing linear transformation layers, maps the text and image vectors to a fusion space of the same dimension, allowing subsequent attention calculation and semantic alignment in the fusion process to be performed within the same spatial dimension.

[0100] Sub-step 4.2: Initial modal weight vector generation:

[0101] During the model training phase, statistics are collected from the user's historical click data to determine the click response frequency when the image and text information are used as the basis for recommendation. That is, the following calculation is performed:

[0102] The proportion of responses that use titles or descriptions for recommendations; the proportion of responses that use images as the main source of information for recommendations.

[0103] After normalizing these frequencies, an initial modality weight vector is formed, which serves as the prior weight for attention regulation between the text and image modalities. This weight is input into the attention module as a guidance vector at the beginning of training and can participate in adaptive updates during training.

[0104] Sub-step 4.3: Cross-modal attention matrix calculation and regulation:

[0105] The mapped text and image vectors are fed into the cross-modal attention mechanism to calculate their mutual semantic response in the fusion space. Cosine similarity is used as the basic metric for attention similarity, resulting in attention matrices for both the text-to-image and image-to-text directions.

[0106] Subsequently, based on the aforementioned initial modal weight vector, the output of the attention matrix is ​​weighted and regulated, that is, the attention result of each modality is scaled by the modal guidance parameters, thereby achieving control of the relative importance between modalities and forming a regulated bidirectional semantic fusion vector.

[0107] Sub-step 4.4: Introduce the modal residual gating structure for dynamic fusion control:

[0108] In order to enhance the contextual adaptability of the fusion vector, a modal residual gating structure is introduced into the fusion result. Specifically:

[0109] A gating layer is constructed to receive the contextual semantic environment as input and generate two scalar gating factors. The gating factors are used to determine whether to retain text modality information, image modality information, or a weighted output of both in the final fusion result. A residual connection path is introduced to prevent the fusion information from disappearing during multiple rounds of training.

[0110] This gating mechanism enables the fusion strategy to no longer rely on fixed weights, but dynamically adjusts the modality-preserving structure according to semantic context and behavioral feedback, thereby enhancing the adaptability of the fusion vector to the recommendation context.

[0111] Sub-step 4.5: Multimodal semantic alignment contrast training:

[0112] To further improve the semantic consistency between the fused semantic vector and user behavior interests, this paper introduces a comparative training mechanism for multimodal semantic alignment:

[0113] During the training phase, the fusion vectors of the positive item clicked by the user and the current item are combined to form a "semantic pair"; non-clicked items or other items that the user has not interacted with are used as negative examples; a contrast loss function is constructed to minimize the distance between positive semantic vectors and maximize the distance between negative examples; and joint training is performed with the recommendation target loss to ensure that the fused semantic space has discriminative and responsive capabilities in the recommendation task.

[0114] Sub-step 4.6: Output the fused semantic vector and normalize it:

[0115] After alignment training, the fused semantic vector is L2 normalized to obtain a fixed-dimensional product comprehensive semantic vector. This vector will be used in subsequent steps to match the user interest vector and serve as one of the input features for recommendation scoring.

[0116] In the process of product semantic fusion modeling, the present invention realizes dynamic fusion modeling between product titles, description text semantics and image features by introducing behavior-guided modal weight initialization, cross-modal attention matrix regulation, and modal residual gating mechanism. Unlike traditional image-text splicing or weighted averaging methods, this fusion strategy not only takes into account the complementarity of the semantic information of each modality itself, but also combines the behavioral feedback of users on the degree of attention to different modalities in actual recommendation interactions. It guides the initial distribution of the attention mechanism by normalizing the weight of the behavior frequency, and introduces a bidirectional attention matrix to jointly model the semantic correlation between image and text modalities. In the fusion process, the residual gating mechanism is further used to realize dynamic retention or suppression of modalities, forming a semantic expression structure that adapts to the context. This fusion path significantly improves the responsiveness of product semantic representation to user preferences and recommendation goals, realizes the combination of nonlinear collaboration between modalities, behavioral response perception and controllability of fusion expression, and realizes a leap from "static modal fusion" to "dynamic user behavior-driven semantic fusion" at the level of semantic modeling technology.

[0117] Step 5: Extract the user's browsing, clicking, and purchasing records from the user behavior data to generate the user interest vector.

[0118] In step 5, generating the user interest vector includes extracting the interaction records between the user and the product from the user behavior log, including three types of behaviors: browsing, clicking, and purchasing, and each behavior record contains a product identifier, a behavior type, and a timestamp; using the comprehensive semantic vector of the product corresponding to the product identifier in each behavior record as a behavior embedding vector, and assigning fixed behavior weights to different types of behaviors, and setting the behavior weight values ​​to 1.0 for purchase, 0.6 for click, and 0.2 for browsing based on the empirical parameters of the behavior conversion rate; multiplying the behavior embedding vector of each behavior record by the corresponding behavior weight, and constructing a user behavior sequence in timestamp order; introducing a time decay factor in the behavior sequence, assigning exponential decay weights to behaviors at earlier times, and obtaining a weighted behavior sequence tensor; inputting the weighted behavior sequence into a multi-layer one-dimensional convolution structure for feature extraction to extract user interest trajectory features; and normalizing the extracted feature vector to generate a user interest vector.

[0119] To accurately model users' personalized preferences in recommendation scenarios and respond to their changing interest trends, this paper proposes a user interest modeling method that combines multi-behavior modeling, time decay control, interest drift structure, and intent abstraction mechanism to construct a user interest vector based on the user's historical interaction behavior with products. Specifically, it includes the following steps:

[0120] Sub-step 5.1: Behavioral data extraction and formatting:

[0121] The target user's interaction records are extracted from the user behavior logs recorded by the system. The behavior types include browsing, clicking, and purchasing. Each behavior record includes the user ID, product ID, behavior type, and the timestamp of the behavior. The behavior logs are reordered in chronological order to form a basic sequence of user behavior.

[0122] Sub-step 5.2: Product semantic vector mapping and behavior type weight assignment

[0123] For each behavior record, the comprehensive semantic vector of the product generated in step 4 is used as the embedded semantic representation of the behavior record through the product identification index; fixed weight parameters are set according to the behavior type: the weight of purchase behavior is 1.0; the weight of click behavior is 0.6; and the weight of browsing behavior is 0.2.

[0124] Perform item-by-item multiplication of the behavior embedding vector and its behavior weight to generate the behavior weighted semantic representation.

[0125] Sub-step 5.3: Behavior type decoupling and convolutional modeling:

[0126] To improve the ability of user interest vectors to distinguish behavioral intentions, the present invention constructs browsing, clicking, and purchasing behavior records into three independent subsequences, and inputs three sets of behavior-type-specific one-dimensional convolution channels with independent convolution kernel parameters:

[0127] Each channel consists of four parts: convolution, activation, normalization, and pooling, and is specifically used to model the semantic changes of this type of behavior in the temporal dimension; the output of the convolution channel is weighted and fused through a behavior gating structure, and the weight is trained based on the contribution ratio of the behavior category to the click results in historical recommendations.

[0128] This structure avoids the problem of information mixing among different types of behaviors in sequence learning, and improves the user preference model's ability to distinguish and learn behavior-driven intentions.

[0129] Sub-step 5.4: Time decay factor generation and personalized control:

[0130] In the process of constructing the behavior sequence, in order to express the time-effect impact of the behavior on the current interest state, the present invention introduces a time decay factor based on the weighted behavior representation. The specific process is:

[0131] The time difference between each behavior and the current time is calculated to generate an exponential decay factor. A dynamic adjustment coefficient is generated based on the user's historical activity periodicity (such as average interaction frequency and recent activity interval). This coefficient scales the time decay function to form a user-personalized decay control factor.

[0132] The decay factor ultimately acts on the weighted representation of behavior to form a behavioral input sequence with time-sensitive weights, enhancing the model's ability to discriminate recent behavioral responses.

[0133] Sub-step 5.5: Interest drift structure construction and stage modeling:

[0134] To express the changing trend of user interests in long-term usage, the above time-weighted series is divided into multiple time windows (such as the past 7 days, the past 30 days, and historical cumulative behavior):

[0135] Each time period is fed into a convolutional module with the same structure. The model outputs sub-interest vectors corresponding to different time periods. Based on the semantic relevance between the product domain and each sub-vector, the attention coefficient is calculated as a weighting factor. Finally, the interest vectors of the three stages are fused to form a stage-fused interest representation.

[0136] This mechanism enhances the model's ability to identify the "interest drift" phenomenon and prevents behavior in a single time period from dominating interest expression.

[0137] Sub-step 5.6: User intent abstraction layer and interest vector output:

[0138] The fused behavioral features are input into the intent abstraction module. The module includes:

[0139] The multi-head attention mechanism is used to identify the core response signals in the behavioral semantic dimension; the sparse activation mechanism is used to compress the behavioral dimension and retain the significant interest subspace; the fully connected mapping and normalization output a fixed-length user interest vector.

[0140] The final output user interest vector will be used to match the comprehensive semantic vector of the product and is the core input of the subsequent recommendation scoring model.

[0141] In the process of constructing user interest vectors, this invention introduces a multi-channel modeling structure that decouples behavior types and combines it with a dynamic time decay mechanism based on user activity cycles. This allows for refined processing and weight control of user historical behaviors, enhancing the accuracy of behavioral semantic modeling and the responsiveness of behavioral timeliness. Compared to existing approaches that uniformly model user behaviors or employ fixed decay strategies, this invention can model and regulate different types of behaviors based on their semantic characteristics and time decay properties, thereby improving the coverage and explanatory power of interest expressions for users' true preferences.

[0142] Furthermore, the present invention achieves the combined expression of users' long-term interests and short-term concerns through a multi-stage interest drift modeling structure and a sparse intent abstraction layer. This significantly reduces semantic redundancy in interest expressions, making interest vectors more focused and discriminable. These improvements not only enhance the recommendation system's adaptability to user preference shifts, but also significantly strengthen the system's ability to maintain consistent and novel recommendations in high-frequency behavior streams, providing a more expressive and responsive technical foundation for personalized recommendations under complex user behavior patterns.

[0143] Step 6: Based on the user interest vector, use the recurrent neural network to update the user interest vector.

[0144] In step 6, the use of a recursive neural network to update the user interest vector includes splitting the user interest vector into a short-term interest vector and a long-term interest vector, the short-term interest vector is generated based on the most recent interactive behavior, and the long-term interest vector is generated based on the historical accumulated behavior; the comprehensive semantic vector of the product corresponding to the user's latest behavior record is used as the current behavior vector of the input sequence; a dual-track recursive neural network containing a gated recurrent unit is constructed, the short-term interest track and the long-term interest track respectively contain an update gate and a reset gate, and the parameters of the update gate and the reset gate are dynamically adjusted according to the behavior type, which includes browsing, clicking and purchasing; the input sequence is recursively updated in the short-term interest track and the long-term interest track respectively, the short-term track is highly sensitive to the time sequence, and the long-term track introduces a time decay coefficient to suppress the influence of outdated behavior; the updated short-term interest vector and the long-term interest vector are fused through a weighting coefficient sensitive to the behavior category to generate an updated user interest vector.

[0145] To improve the recommender system's responsiveness to changes in user behavior and the dynamic adaptability of preference expression, a dual-track gated recurrent neural network (GRN) continuously updates user interest vectors during recursive updates based on user interest vectors. The system splits user interest vectors into short-term and long-term interest vectors, processing recent active interests and historical accumulated preferences separately. This constructs a user interest update path that decouples temporal granularity from behavioral semantics.

[0146] The short-term interest vector is generated based on the sequence of recent interaction behaviors, including the semantic vector of the product in the most recent behavior and its corresponding behavior category label. The long-term interest vector is derived from the user's stable behavior over a period of time and represents the long-term preference profile. These two vectors are input into two GRU tracks with identical structures but independent parameters, and their states are updated separately. The GRU structure of each track consists of an update gate, a reset gate, and a state candidate layer. The update gate controls the fusion ratio between the previous interest state and the current input, while the reset gate controls the influence of the current input on the state reconstruction.

[0147] Define the current behavior semantic vector as x t , the previous state vector is h t-1 , update gate to z t , reset the gate to r t , the candidate state is The final updated state is h t The update gate is calculated as follows:

[0148] z t =σ(W z x t +U z h t-1 +b z )

[0149] Among them, W z is the weight matrix between the action input and the update gate, U z is the weight matrix between the previous state and the update gate, b z is the bias term, and σ represents the sigmoid activation function.

[0150] The reset gate is calculated as follows:

[0151] r t =σ(W r x t +U r h t-1 +b r )

[0152] Among them, W r ,U r ,b r are the weight items corresponding to behavior input, state input and bias respectively.

[0153] Calculate state candidate values ​​based on the reset gate:

[0154]

[0155] Among them, W h ,U h ,b h are the weight terms of input, state, and bias respectively, and ⊙ represents the bit-by-bit multiplication operation. The final updated state vector is:

[0156]

[0157] In the long-term interest track, in order to prevent historical interests from being gradually forgotten in multiple rounds of training, a time decay coefficient γ∈[0,1] is introduced to the state h t-1 For scaling, the revised update formula is as follows:

[0158]

[0159] The value of γ is calculated based on the interval between the behavior timestamp and the current time. The smaller the value, the older the behavior and the weaker its impact on the current interest expression.

[0160] The system takes the comprehensive semantic vector of the product of each user's latest behavior as the current input x t , and adjust the influence parameters of each behavior in the gating weight matrix based on the behavior type (browsing, clicking, purchasing). Specific behavior types are controlled by a scalar weight factor α in the gating parameters: α = 1.0 for purchasing, α = 0.6 for clicking, and α = 0.2 for browsing. These factors act on the weight terms in the gating calculation process, ensuring that different behavior types have different responses to the strong or weak control of interest state updates.

[0161] Finally, the state vector of the short-term interest track output is h s , the state vector of the long-term interest track output is h l The system sets the fusion weight factor β based on the current behavior scenario or channel type. s ,β l , both satisfy β s +β l = 1. User's updated interest vector h f Generated by the following formula:

[0162] h f =β s ·h s +β l ·h l

[0163] Among them, β s Control the weight contribution of short-term interest, preferably taking a higher value in channels with high timeliness requirements, such as new product recommendations, flash sales, etc.; β l Control the influence of long-term interests and give priority to increasing the weight of channels with high demand for stable preference expression, such as Guess What You Like and personalized lists.

[0164] The final output h f This represents the vector representation of the current user's interest expression, incorporating the influence of multiple time scales and behavior types. It serves as the core input for subsequent product matching calculations and recommendation scoring models, improving the system's personalized response accuracy and dynamic behavioral adaptability. The model's structural parameters, weight matrix, and behavioral adjustment factors are all iteratively optimized through supervised learning during the training phase, enabling the system to continuously learn and dynamically adjust interest updates in real-world scenarios.

[0165] The present invention constructs a dual-track recursive interest update structure to model the user's short-term behavioral reactions and long-term behavioral preferences in a decoupled manner, and introduces a behavior type-driven gating mechanism to achieve fine control of the state update path by behavioral semantics. At the same time, a time decay control factor is embedded in the long-term interest modeling to achieve the time-weighted decay of the influence of historical behavior, effectively addressing the "interest expiration" problem. The final interest output combines short-term and long-term expressions in an adjustable fusion manner, dynamically adapting to multi-scenario recommendation needs, and achieving a comprehensive improvement in interest expression ability, response speed, and semantic accuracy.

[0166] In the state update process of the long-term interest trajectory, in order to control the residual influence of historical behavior on the current state of interest, the time decay coefficient γ is introduced to the previous state vector h t-1Scaling is performed. γ is a real scalar related to time, and its value range is [0,1]. In order to achieve a dynamic correspondence between time distance and attenuation degree, the present invention uses an exponential decay function to construct the calculation formula of γ as follows:

[0167] γ=exp(-λ·Δt)

[0168] Among them, λ is the preset decay rate hyperparameter, usually set to 10 -4 ~10 -2 Δt represents the time interval between the current behavior and the previous state of interest, expressed in hours or seconds. Through this calculation, γ dynamically adjusts the state retention ratio based on the timeliness of the behavior, automatically attenuating outdated behaviors in long-term interest modeling and improving the accuracy of expressing interest timeliness.

[0169] Step 7: Jointly model the comprehensive semantic vector of the product and the updated user interest vector to generate recommendation results.

[0170] In step 7, generating the recommendation result includes normalizing the comprehensive semantic vector of the product and the updated user interest vector respectively; inputting the normalized comprehensive semantic vector of the product and the user interest vector into a behavior category-sensitive bidirectional attention matching module to calculate the attention weight matrix, and the weight matrix is ​​adjusted according to the correlation between the dimensions of the two vectors and the category weight of the user's historical behavior type; according to the attention weight matrix, weighting the dimensional response values ​​of the comprehensive semantic vector of the product and the user interest vector to generate a matching feature vector; inputting the matching feature vector into the recommendation scoring module, and the recommendation scoring module linearly transforms the matching feature vector according to the trainable parameters generated by the historical feedback of the user behavior category, and outputs the recommendation score; in the recommendation scoring training process, introducing a contrastive learning loss function that is sensitive to the user's historical behavior category to minimize the difference between the positive and negative product scores, and at the same time adjusting the update rule of the attention matrix according to the preference consistency of the user's historical interest behavior to optimize the distinguishing ability of the recommendation score.

[0171] To jointly model the multimodal, integrated semantic vectors of products with updated user interest vectors and generate highly accurate recommendation outputs, this paper constructs a behaviorally sensitive, bidirectional attention matching structure and a parameter-learnable recommendation scoring module, enabling responsive modeling of the semantic relevance between products and users. The matching module captures the micro-correlations between products and user interests across various semantic dimensions, assigning higher weight to highly focused semantic information. The scoring module further generates a learnable preference scoring mechanism based on user behavioral feedback.

[0172] The comprehensive semantic vector of the product output in the previous step is recorded as The user interest vector is recorded as Where d is the uniform embedding dimension. First, normalize the two vectors using the L2 norm to obtain:

[0173]

[0174] The normalized vector ensures that it is distributed on the unit sphere, which is beneficial to the stability of subsequent calculations based on inner product or similarity measurement operations.

[0175] The two normalized vectors are input into the bidirectional attention matching module. The core of the module consists of the attention weight matrix Calculate the composition, each element A in the matrix ij express The i-th dimension and The response degree between the jth dimension. The attention calculation formula is:

[0176]

[0177] Among them, cos(·,·) represents the cosine similarity, α∈[0,1] is the behavior response weight adjustment factor, and w b (i, j) is the behavior category-sensitive weight matrix, which is derived from the trainable prior weights generated by historical click frequency statistics of users in different categories of behavior. This mechanism ensures that attention weights not only consider the numerical similarity between products and interests in the semantic dimension, but also adjust them based on behavioral feedback preferences.

[0178] The attention matrix A is normalized along the row and column directions to obtain two attention mask vectors Act on the product semantic vector and user interest vector respectively. After weighted processing, the matching feature vector is generated

[0179]

[0180] Among them, ⊙ represents the weighted summation by dimension, and the obtained f m The vector integrates the most responsive semantic dimensions of products and user interests, and is a semantic feature representation that reflects the matching degree of the current recommendation target.

[0181] f m The input recommendation scoring module is a set of learnable fully connected structures that establish independent channels for corresponding historical behavior types. With purchase behavior as the main monitoring target, the following scoring function is constructed:

[0182] s=W T f m +b

[0183] in, is the scoring parameter vector, is the bias term, and both participate in gradient optimization as training parameters.

[0184] To improve recommendation accuracy and model generalization, a contrastive learning loss function is introduced during the model training phase. The user's actual clicked items are used as positive examples, and randomly selected unclicked items are used as negative examples. The loss function is defined as:

[0185]

[0186] Among them, s + With s - are the recommendation scores of positive and negative examples respectively, δ is the score interval threshold, λ is the regularization coefficient, and A b is the attention template matrix generated based on historical preferences. This loss function simultaneously optimizes the difference in recommendation ranking and regularizes the behavior preference preservation of the attention matrix, ensuring that the attention results have semantic continuity with the user's behavior response.

[0187] The final output recommendation score s will serve as the basis for ranking products in front of users, and can be used in the fine ranking stage or click-through rate estimation module to generate the final recommendation list based on the user's current behavior context.

[0188] This architecture achieves a complete closed-loop modeling process, encompassing bidirectional product-user attention perception, weighted fusion of semantic dimensions, behavioral feedback response regulation, matching feature aggregation, and score prediction. This enhances the system's ability to characterize multidimensional semantic relationships and learn consistent user preferences, enabling end-to-end structural modeling from embedding representation to recommendation ranking. The system boasts online deployment capabilities, with vector normalization, attention calculation, and score output all completed within 10ms, making it suitable for real-time personalized output in large-scale recommendation scenarios.

[0189] In the attention weight matrix During the calculation process, A ij Indicates the response strength between the i-th dimension feature of the product comprehensive semantic vector and the j-th dimension feature of the user interest vector. To avoid ambiguity, the variables in the attention calculation formula should be written as:

[0190]

[0191] in, Represents the value of the normalized product semantic vector in the i-th dimension, represents the value of the normalized user interest vector in the jth dimension, cos(·,·) represents the cosine similarity between the two, and α∈[0,1] is the behavior feedback adjustment coefficient. Behavior preference weight matrix w bThe generation of (i, j) is based on historical user interaction responses across various semantic dimensions. The system records a large number of historical user click and purchase logs and normalizes the average weight of each dimension in their behavior. This matrix forms a static preference template, which can also be updated as a learnable parameter during model training to reflect the importance of specific semantic dimensions in behavioral feedback. This preference matrix enhances the attention mechanism's ability to capture users' true preference patterns, further improving the accuracy of recommendation matching.

[0192] Example 2 is the second embodiment of the present invention, which is different from the previous embodiment in that:

[0193] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0194] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0195] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a computer memory.

[0196] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, it can be implemented using a combination of any of the following technologies known in the art: a discrete logic circuit having logic gate circuits for implementing logic functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0197] Example 3, the third embodiment of the present invention, provides an AI search recommendation system combined with product semantic understanding, including:

[0198] Data collection module, used to obtain product title, description and image information;

[0199] A text processing module is used to process product titles and descriptions to generate semantic vectors for product titles and descriptions;

[0200] The image processing module is used to extract features from product images using a convolutional neural network and generate feature vectors of product images;

[0201] The feature fusion module is used to fuse the semantic vectors of product titles and descriptions with the feature vectors of product images to generate a comprehensive semantic vector for the product;

[0202] Behavior analysis module, used to extract user browsing, clicking, and purchasing records from user behavior data and generate user interest vectors;

[0203] An interest updating module is used to update the user interest vector based on the user interest vector using a recurrent neural network;

[0204] The recommendation generation module is used to jointly model the comprehensive semantic vector of the product and the updated user interest vector to generate recommendation results;

[0205] The query understanding module is used to understand the semantics of user-submitted queries, generate query vectors, and calculate the similarity between the query vectors and the comprehensive semantic vectors of products;

[0206] The ranking weighting module is used to sort based on the similarity between the query vector and the comprehensive semantic vector of the product, and to weight the product according to information such as price and evaluation;

[0207] The feedback update module is used to update the user interest vector in real time based on the interaction feedback between the user and the product, and adjust the recommendation results based on the updated interest vector.

[0208] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. An AI search and recommendation method that combines product semantic understanding, characterized by: include, Get the product's title, description, and image information; Perform text processing on titles and descriptions to generate semantic vectors for product titles and descriptions; Use convolutional neural networks to extract features from image information and generate feature vectors of product images; The semantic vectors of the title and description are fused with the feature vectors of the image information to generate a comprehensive semantic vector of the product; Extract users' browsing, clicking, and purchasing records from user behavior data to generate user interest vectors; Based on the user interest vector, a recurrent neural network is used to update the user interest vector; The comprehensive semantic vector of the product and the updated user interest vector are jointly modeled to generate recommendation results.

2. The AI ​​search and recommendation method combined with product semantic understanding according to claim 1, characterized in that: Generating semantic vectors for product titles and descriptions includes: Segment product titles and descriptions, remove stop words, and perform part-of-speech tagging; Add field position codes to the processed product titles and descriptions to indicate the field category information of the titles and descriptions; The text input with positional encoding is fed into the pre-trained BERT model, and multi-layer bidirectional encoding is performed to obtain the context-dependent vectors of each word. During encoding, a multi-head self-attention mechanism is used to model the contextual dependencies of words within each field, and independent training parameters are assigned to attention heads of different fields. Based on the user click and purchase behaviors recorded during the training phase, the average attention weight of the vocabulary within a field is normalized with the response frequency of the field in user behavior to determine the global weight coefficient of each field. The global weight coefficient of each field is weighted averaged with the encoding vectors of all words in the field to generate the product title semantic vector and product description semantic vector.

3. The AI ​​search and recommendation method combined with product semantic understanding according to claim 2, characterized in that: The feature vector for generating the product image includes: Resize the product image to 224 pixels by 224 pixels and perform histogram equalization on each pixel channel; The processed product images are fed into a convolutional neural network with a residual connection structure and a spatial attention module. Convolution, activation, normalization, and downsampling operations are performed sequentially in the convolutional neural network to extract feature maps of the second, third, and fourth layer residual modules; The extracted feature maps are adjusted to a uniform spatial size using bilinear interpolation; On the feature map of uniform spatial size, the channel response value of each spatial position is weighted by the spatial weight map generated by the spatial attention module; Concatenate the weighted feature maps according to the channel dimension; Perform global average pooling on the concatenated feature maps to generate a fixed-dimensional product image feature vector.

4. The AI ​​search and recommendation method combined with product semantic understanding according to claim 3, characterized in that: The generation of a comprehensive semantic vector of a product includes: Concatenate the product title semantic vector and the product description semantic vector to form a comprehensive text semantic vector; Inputting the text comprehensive semantic vector and the product image feature vector into a mapping module respectively, wherein the mapping module includes a linear transformation layer with shared parameters, which is used to map the two input vectors into a fusion space of the same dimension; Based on the user's historical click behavior, the normalized ratio of the frequency of text and image modal contributions to user recommendation clicks in the training set is calculated as the initial modality weight vector for fusion; Calculating a cross-modal attention weight matrix based on the cosine similarity between the mapped text and image vectors, and performing weighted control based on the initial modality weight vector; Apply the regulated cross-modal attention weight matrix to the two modal vectors to obtain a fusion vector; A modality residual gating structure is introduced into the fusion vector, which automatically controls whether to retain text modality information, image modality information, or output both modalities simultaneously based on context-related parameters. Compare the fused vector with the semantic vectors of positive user interaction items in the recommendation system, and construct a multimodal semantic alignment constraint by minimizing the contrast loss. Normalize the fused vector after alignment training to generate a comprehensive semantic vector for the product.

5. The AI ​​search and recommendation method combined with product semantic understanding according to claim 4, characterized in that: Generating the user interest vector includes: Extract user-product interaction records from user behavior logs, including browsing, clicking, and purchasing behaviors. Each behavior record contains the product ID, behavior type, and timestamp. The comprehensive semantic vector of the product corresponding to the product identifier in each behavior record is used as the behavior embedding vector, and fixed behavior weights are assigned to different types of behaviors; Multiply the behavior embedding vector of each behavior record by the corresponding behavior weight, and construct the user behavior sequence in timestamp order; A time decay factor is introduced into the behavior sequence, and an exponential decay weight is assigned to the behavior at earlier times to obtain a weighted behavior sequence tensor. The weighted behavior sequence is input into a multi-layer one-dimensional convolutional structure for feature extraction to extract the user's interest trajectory features; The extracted feature vector is normalized to generate the user interest vector.

6. The AI ​​search and recommendation method combined with product semantic understanding according to claim 5, characterized in that: The updating of the user interest vector using a recurrent neural network includes: Split the user interest vector into short-term interest vector and long-term interest vector. The short-term interest vector is generated based on the recent interaction behavior, and the long-term interest vector is generated based on the historical accumulated behavior. The corresponding comprehensive semantic vector of the product in the user's latest behavior record is used as the current behavior vector of the input sequence; A dual-track recurrent neural network with gated recurrent units was constructed. The short-term interest track and the long-term interest track each contained an update gate and a reset gate. The parameters of the update and reset gates were dynamically adjusted based on the behavior type, which included browsing, clicking, and purchasing. The input sequence is recursively updated in the short-term and long-term interest tracks. The short-term track is more sensitive to the time sequence, and the long-term track introduces a time decay coefficient to suppress the influence of outdated behavior. The updated short-term interest vector and the long-term interest vector are fused through a behavior category-sensitive weighting coefficient to generate an updated user interest vector.

7. The AI ​​search and recommendation method combined with product semantic understanding according to claim 6, characterized in that: Generating the recommendation result includes: Normalize the product comprehensive semantic vector and the updated user interest vector respectively; The normalized product comprehensive semantic vector and user interest vector are input into a behavior category-sensitive bidirectional attention matching module to calculate the attention weight matrix. The weight matrix is ​​adjusted based on the correlation between the dimensions of the two vectors and the category weight of the user's historical behavior type. According to the attention weight matrix, the dimensional response values ​​of the product comprehensive semantic vector and the user interest vector are weighted to generate a matching feature vector; The matching feature vector is input into the recommendation scoring module. The recommendation scoring module performs a linear transformation on the matching feature vector based on the trainable parameters generated by historical feedback of user behavior categories and outputs a recommendation score. During the recommendation scoring training process, a contrastive learning loss function that is sensitive to user historical behavior categories is introduced to minimize the difference between positive and negative product scores. At the same time, the update rule of the attention matrix is ​​adjusted according to the preference consistency of users' historical interest behaviors to optimize the discriminative ability of recommendation scores.

8. An AI search and recommendation system combined with product semantic understanding, applying the AI ​​search and recommendation method combined with product semantic understanding as described in any one of claims 1 to 7, characterized in that: include: Data collection module, used to obtain product title, description and image information; A text processing module is used to process product titles and descriptions to generate semantic vectors for product titles and descriptions; The image processing module is used to extract features from product images using a convolutional neural network and generate feature vectors of product images; The feature fusion module is used to fuse the semantic vectors of product titles and descriptions with the feature vectors of product images to generate a comprehensive semantic vector for the product; Behavior analysis module, used to extract user browsing, clicking, and purchasing records from user behavior data and generate user interest vectors; An interest updating module is used to update the user interest vector based on the user interest vector using a recurrent neural network; The recommendation generation module is used to jointly model the comprehensive semantic vector of the product and the updated user interest vector to generate recommendation results; The query understanding module is used to understand the semantics of user-submitted queries, generate query vectors, and calculate the similarity between the query vectors and the comprehensive semantic vectors of products; The ranking weighting module is used to sort based on the similarity between the query vector and the comprehensive semantic vector of the product, and to weight the product according to its price and evaluation information; The feedback update module is used to update the user interest vector in real time based on the interaction feedback between the user and the product, and adjust the recommendation results based on the updated interest vector.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the processor implements the steps of an AI search recommendation method combined with product semantic understanding as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of an AI search recommendation method combined with product semantic understanding as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Commodity recommendation method and system based on multi-feature fusion

    CN120894107A

  • A commodity recommendation method and system based on multi-feature fusion

    CN120894107B

  • Semantic analysis-based user portrait recognition method and website recommendation method

    CN120950779A

  • AI intelligent marketing content publishing subject matching recommendation method

    CN121071234A

  • Large language model smart search grouping method, system and device and medium

    CN122019569A