A product recommendation method based on multimodal knowledge fusion based on bilinear pooling
By adopting the method of bilinear pooling and multimodal features on e-commerce platforms, the problem of insufficient judgment of consumer preferences in the existing technology is solved, and a more accurate and personalized product recommendation effect is achieved.
Patent Information
- Application Number
- CN202211671672.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-26
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-12-26
AI Technical Summary
The existing technology cannot deeply utilize massive multimodal user data, resulting in insufficient accuracy in judging consumer preferences and inability to provide personalized product recommendations.
The multimodal knowledge fusion method based on bilinear pooling is adopted to extract multimodal features through deep learning, and multi-level feature fusion is used to accurately judge consumers' preferences and push intended products.
It has achieved in-depth utilization of massive multimodal user data, improved the accuracy and personalization of product recommendations, and enabled e-commerce platforms to push products in line with consumer preferences more accurately.
Smart Images

Figure CN115860875B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a commodity recommendation method based on multimodal knowledge fusion of bilinear pooling. Background Art
[0002] At present, shopping platforms can be divided into two categories, general shopping platforms and online shopping platforms. General shopping platforms are shopping malls, stores, etc. in real life. Sometimes sales channels are also collectively referred to as shopping platforms. Online shopping platforms are platforms for shopping activities in the virtual network world. They mostly use digital information transmission to achieve the purpose of physical transactions. Driven by the development of the Internet, online shopping has become a fashion and an important form of shopping at present. With the development of communication technology and the surge in Internet users, the content of e-commerce platforms has become more diversified and vibrant, and competition has become fierce. In this context, if we can make better use of data and use data mining technology to analyze and predict data, empower e-commerce platforms, make platforms more accurate and personalized, inject more vitality into e-commerce platforms, and allow platforms to provide consumers with better products and services. At the same time, let technology create greater profit margins for e-commerce platforms.
[0003] Currently, most existing technologies only recommend products through single-modal information. The problem with existing technologies is that they are unable to make full use of massive multi-modal user data and their judgments on consumer preferences are not accurate enough. Summary of the invention
[0004] 1. Technical issues to be resolved
[0005] In view of the shortcomings of the prior art, the present invention provides a product recommendation method based on multimodal knowledge fusion of bilinear pooling, which makes full use of massive multimodal user data and fuses multimodal features by hierarchical bilinear pooling to accurately judge consumers' preferences and accurately push the advantages of intended products. It solves the problem that most of the prior art only recommends products through single-modal information. The problem of the prior art is that it cannot make full use of massive multimodal user data and the judgment of consumers' preferences is not accurate enough.
[0006] (II) Technical solution
[0007] To achieve the above-mentioned purpose, the present invention provides the following technical solution: a product recommendation method based on multimodal knowledge fusion of bilinear pooling, comprising the following steps:
[0008] S1. Extract user behavior data on the e-commerce platform, pre-process the collected user behavior data, and analyze and evaluate it;
[0009] S2, extracting features from the collected data using a deep learning image feature collector, a text feature collector, a time series feature collector, an audio feature collector, and a video feature collector;
[0010] S3, extracting deep features of modal interaction by using bilinear pooling multimodal knowledge fusion according to the features of various single-modal feature collectors in the above step S2;
[0011] S4, by using the deep feature information of multimodal fusion extracted in the above step S3, the bilinear pooling is used again in the intelligent recommendation module to classify and fuse the features, and the fused features are put into the intelligent recommendation module to obtain the recommended area index and the recommendation index;
[0012] S5. Determine the recommendation form, recommended area and recommended order of the product according to the recommended area index and the recommended index in the above step S4;
[0013] S6. Based on the recommended content in step S5, the feedback module tracks the user's interactive behavior with respect to the recommended content and feeds it back to the e-commerce recommendation model.
[0014] Preferably, the steps S3 to S5 are as follows:
[0015] (1) Initialize the feature collector, feature fusion module, and intelligent recommendation module, and train the model using the preprocessed data;
[0016] (2) During the model training phase, the multimodal feature extraction module needs to minimize the loss function ξ of the product recommendability module in order to improve its ability to correctly judge whether a product can be recommended. h (θ e ,θ h ), and at the same time, the multimodal feature extraction module attempts to maximize the loss function ξ of the recommended product scoring module g (θ e ,θ g ), to capture the common features of all recommended products, and the recommended product scoring module needs to minimize the loss function of the module to score the recommended products, indicating the degree to which the products are recommended. Since the multimodal feature extraction module and the recommended product scoring module are adversarial, the overall loss of the three loss functions is defined as:
[0017] ξ final (θ e ,θ h ,θ g )=ξ h (θ e ,θ h )-λξ g (θ e ,θ g) is the trade-off parameter, where θ e is the parameter set of the multimodal feature extraction module, θ h represents the parameter set of the product recommendation module, θ g is the parameter set of the recommended product scoring module, and λ controls the weights of the two objective functions. For this adversarial strategy, the final target parameters are shown in the following formula:
[0018]
[0019] in, and is the optimal parameter sought by the objective function;
[0020] (3) During the training process, various fusion modules will use different types of large-scale multimodal data for unlabeled training, and then use a small amount of labeled data for fine-tuning testing to obtain the fusion modules of different parts when used.
[0021] Preferably, a gradient reversal layer GRL is added to solve the parameter optimization problem when searching for the final target parameters. The optimization process of the model parameters is updated according to the following formula:
[0022]
[0023] Among them, η is the learning rate. In order to maintain the stability of the model training process, the formula x is used to adjust the learning rate η. The formula is as follows:
[0024]
[0025] Where α=10, β=0.75.
[0026] Preferably, in step S3, when multimodal knowledge fusion using bilinear pooling is used to extract deep features of modal interaction;
[0027] First, define a basic model of the multimodal decomposition bilinear pooling method, the formula is as follows:
[0028]
[0029] Take text images as an example. is the product text feature extracted in step S2, is the product image feature extracted in step S2, is the parameter matrix, Z is the model output;
[0030] Since the model needs to capture the relationship between different features, it needs to learn a large number of parameters, which will lead to high computational costs and the risk of overfitting. i Decomposed into two low-rank matrices U iand V i :
[0031]
[0032] where k is the dimension of the two low-rank matrices, The symbol o is the Hadamard product, Ι is a matrix of all 1s, and the two third-order tensors obtained are and As the weight of the o-dimensional output, on the basis of consistency, use the two-dimensional matrix and Denote U and V respectively, so the above formula can be rewritten as:
[0033]
[0034] The above formula indicates that a one-dimensional non-overlapping window of size k is used to pool the product features.
[0035] Preferably, in order to solve the problem that the output dimension size will change greatly due to the introduction of element-level multiplication and the convergence effect of the model is not good, the following two formulas are used to solve it:
[0036] Z=sign(Z)|Z| 0.5
[0037] Z=Z T / ||Z||
[0038] Finally, the multimodal feature extraction module is represented as E(P; θ e ), we get the formula:
[0039] Z=E(P;θ e
[0040] Where E represents the overall mapping function, P represents the input, and θ e Represents the parameter set of the multimodal feature extraction module.
[0041] Preferably, the steps of the intelligent recommendation module in step S4 include:
[0042] (1) After obtaining the fusion features, they are input into the regional recommendation index calculation module, and then the module output and different fusion features are input into different recommendation index calculation modules;
[0043] (2) The different fusion features in step (1) are combined into image-text fusion, time series image-text fusion, video-audio time series image-text fusion, and video time series image-text fusion. The mixed features are further fused using a bilinear pooling fusion layer to obtain fusion features.
[0044] (3) The regional recommendation index calculation module calculates the regional recommendation vector by integrating features. The value on the regional recommendation vector represents the recommendation index of its region. The recommendation area includes four categories: product page, community, video recommendation area and live e-commerce sequence. The product page includes the homepage, sidebar and bottom bar. The community includes the recommendation below the blogger on the homepage. The video recommendation includes the homepage, the recommendation below the video and the recommendation after the video is played. The live e-commerce sequence only includes the live e-commerce sequence. This module is defined as H(Z; θ h ), where H represents the mapping function of the module, θ h Represents the parameter set of the product recommendation module, a multi-product ρ i The regional recommendation index is but:
[0045]
[0046] Let Y represent a set of labels, where The sum is 1, which represents the probability that the product should appear in each area, and the classification loss function is calculated using cross entropy:
[0047]
[0048] To find the optimal parameters and Need to minimize the loss function ξ h (θ e ,θ h ):
[0049]
[0050] (4) The recommendation index calculation module obtains the recommendation index of the product in the region through the recommendation vector and fusion features of the recommendation region index calculation module, and defines the recommendation product scoring module as G(Z; θ g ), where G represents the mapping function of the module, θ g Represents the parameter set of the recommended product scoring module, using Y g Represents the scoring label, and the loss function of the cross entropy calculation module:
[0051]
[0052] To find the optimal parameters Need to minimize the loss function ξ g (θ e ,θ g ):
[0053]
[0054] The above loss function ξ g (θ e ,θ g ) is used to estimate the difference in the distribution of products with different recommendation levels. A larger loss means that all recommended products have similar characteristics. Therefore, in order to eliminate the uniqueness of the characteristics of a recommended product, it is necessary to find the optimal parameter θ e To maximize the loss function ξ g (θ e ,θ g ).
[0055] Preferably, step S5 will construct multiple recommendation tracks for users based on the collected user personal information and product categories, and place the products into the recommendation channel based on the recommendation points and product category information and price information.
[0056] Preferably, step S6 continuously tracks the feedback of platform users on the recommendation information through distributed crawler technology, and feeds it back into the recommendation model by forming new training data.
[0057] Preferably, a bilinear pooling fusion layer is used in S3 for multi-level fusion, such as image-text fusion, time series image-text fusion, video-audio time series image-text fusion, and video time series image-text fusion. This type of bilinear pooling fusion layer obtains features from a single-modal feature collector for fusion, and then uses the bilinear pooling fusion layer to fuse different types of fusion features to obtain more interactive information.
[0058] Preferably, the user's behavior data on the e-commerce platform in step S1 includes platform multimodal data, community multimodal data, video live e-commerce multimodal data, analysis and evaluation data and user data, wherein the platform multimodal data includes product images, product text descriptions and product reviews, the community multimodal data includes blogger browsing order, browsing blogger text, browsing blogger pictures and blog comments, the video live e-commerce multimodal data includes browsing time period videos, live broadcast comments, live broadcast e-commerce browsing order and live broadcast audio, the analysis and evaluation data includes the number of product searches, the number of add-to-carts, the number of purchases, the number of blog post likes and comments views, the number of video likes and comments views, the popularity of live broadcast sales, the number of active people, sales volume and the number of fans, and the user data includes the user's height, weight, age, occupation, usage time, and partition browsing status.
[0059] (III) Beneficial effects
[0060] Compared with the prior art, the present invention provides a product recommendation method based on multimodal knowledge fusion of bilinear pooling, which has the following beneficial effects:
[0061] By obtaining platform user data and product information, using the unimodal feature collector to collect unimodal data, different unimodal features are fused using bilinear pooling according to different data sources, and then the fusion features of the unimodal features are fused using bilinear pooling. The fused features are put into the recommended area index calculation module to obtain the regional recommendation index matrix, and finally the fused features and the regional recommendation index matrix are put into multiple recommendation index calculation modules, the recommended index is calculated and finally put into a comprehensive recommendation index calculation module to calculate the final recommendation index, so that the platform can accurately push products according to consumers' preferences. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 A block diagram of a product recommendation method based on bilinear pooling and multimodal knowledge fusion proposed by the present invention;
[0063] Figure 2 A flowchart of model training in a product recommendation method based on multimodal knowledge fusion based on bilinear pooling proposed by the present invention;
[0064] Figure 3 This is a flowchart of the model usage in the product recommendation method based on multimodal knowledge fusion based on bilinear pooling proposed by the present invention. DETAILED DESCRIPTION
[0065] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0066] Example:
[0067] See attached Figure 1-3 ,A product recommendation method based on multimodal knowledge fusion of bilinear pooling, including the following steps:
[0068] S1. Extract the user's behavior data on the e-commerce platform, pre-process the collected user behavior data, and analyze and evaluate it. The user's behavior data on the e-commerce platform includes platform multimodal data, community multimodal data, video live e-commerce multimodal data, analysis and evaluation data, and user data. Among them, platform multimodal data includes product images, product text descriptions, and product reviews. Community multimodal data includes blogger browsing order, browsing blogger text, browsing blogger pictures, and blog comments. Video live e-commerce multimodal data includes browsing time period videos, live comments, live e-commerce browsing order, and live audio. The analysis and evaluation data includes the number of product searches, the number of add-to-carts, the number of purchases, the number of blog post likes and comments, the number of video likes and comments, the popularity of live streaming, the number of active users, sales, and the number of fans. User data includes the user's height, weight, age, occupation, usage time, and partition browsing status;
[0069] S2, extracting features from the collected data using a deep learning image feature collector, a text feature collector, a time series feature collector, an audio feature collector, and a video feature collector;
[0070] S3, according to the features of various single-modal feature collectors in the above step S2, the multi-modal knowledge fusion of bilinear pooling is used to extract the deep features of modal interaction, and the bilinear pooling fusion layer is used for multi-level fusion, such as image-text fusion, time series image-text fusion, video-audio time series image-text fusion, video time series image-text fusion, etc., which are obtained from single-modal feature collectors. The bilinear pooling fusion layer is then used to fuse different types of fusion features to obtain more interactive information;
[0071] When using bilinear pooling to fuse multimodal knowledge to extract deep features of modal interaction;
[0072] First, define a basic model of the multimodal decomposition bilinear pooling method, the formula is as follows:
[0073]
[0074] Take text images as an example. is the product text feature extracted in step S2, is the product image feature extracted in step S2, is the parameter matrix, Z is the model output;
[0075] Since the model needs to capture the relationship between different features, it needs to learn a large number of parameters, which will lead to high computational costs and the risk of overfitting. i Decomposed into two low-rank matrices U i and V i :
[0076]
[0077] where k is the dimension of the two low-rank matrices, The symbol o is the Hadamard product, Ι is a matrix of all 1s, and the two third-order tensors obtained are and As the weight of the o-dimensional output, on the basis of consistency, use the two-dimensional matrix and Denote U and V respectively, so the above formula can be rewritten as:
[0078]
[0079] The above formula indicates that a one-dimensional non-overlapping window of size k is used to pool product features;
[0080] In order to solve the problem that the output dimension size will change greatly due to the introduction of element-level multiplication and the convergence effect of the model is not good, the following two formulas are used to solve it:
[0081] Z=sign(Z)|Z| 0.5
[0082] Z=Z T / ||Z||
[0083] Finally, the multimodal feature extraction module is represented as E(P; θ e ), we get the formula:
[0084] Z=E(P;θ e
[0085] Where E represents the overall mapping function, P represents the input, and θ e Represents the parameter set of the multimodal feature extraction module;
[0086] S4. By using the deep feature information of multimodal fusion extracted in the above step S3, the bilinear pooling is used again in the intelligent recommendation module to classify and fuse the features, and the fused features are put into the intelligent recommendation module to obtain the recommended area index and the recommendation index. The steps of the intelligent recommendation module include:
[0087] (1) After obtaining the fusion features, they are input into the regional recommendation index calculation module, and then the module output and different fusion features are input into different recommendation index calculation modules;
[0088] (2) The different fusion features in step (1) are combined into image-text fusion, time series image-text fusion, video-audio time series image-text fusion, and video time series image-text fusion. The mixed features are further fused using a bilinear pooling fusion layer to obtain fusion features.
[0089] (3) The regional recommendation index calculation module calculates the regional recommendation vector by integrating features. The value on the regional recommendation vector represents the recommendation index of its region. The recommendation area includes four categories: product page, community, video recommendation area and live e-commerce sequence. The product page includes the homepage, sidebar and bottom bar. The community includes the recommendation below the blogger on the homepage. The video recommendation includes the homepage, the recommendation below the video and the recommendation after the video is played. The live e-commerce sequence only includes the live e-commerce sequence. This module is defined as H(Z; θ h ), where H represents the mapping function of the module, θ h Represents the parameter set of the product recommendation module, a multi-product ρ i The regional recommendation index is but:
[0090]
[0091] Let Y represent a set of labels, where The sum is 1, which represents the probability that the product should appear in each area, and the classification loss function is calculated using cross entropy:
[0092]
[0093] To find the optimal parameters and Need to minimize the loss function ξ h (θ e ,θ h ):
[0094]
[0095] (4) The recommendation index calculation module obtains the recommendation index of the product in the region through the recommendation vector and fusion features of the recommendation region index calculation module, and defines the recommendation product scoring module as G(Z; θ g ), where G represents the mapping function of the module, θ g Represents the parameter set of the recommended product scoring module, using Y g Represents the scoring label and the loss function of the cross entropy calculation module:
[0096]
[0097] To find the optimal parameters Need to minimize the loss function ξ g(θ e ,θ g ):
[0098]
[0099] The above loss function ξ g (θ e ,θ g ) is used to estimate the difference in the distribution of products with different recommendation levels. A larger loss means that all recommended products have similar characteristics. Therefore, in order to eliminate the uniqueness of the characteristics of a recommended product, it is necessary to find the optimal parameter θ e To maximize the loss function ξ g (θ e ,θ g );
[0100] S5. According to the recommended area index and recommendation index in the above step S4, the recommendation form, recommended area and recommended order of the product are determined. Multiple recommendation tracks are constructed for the user based on the collected user personal information and product categories, and the product is placed in the recommendation channel based on the recommendation score, product category information and price information;
[0101] The steps of steps S3 to S5 are as follows:
[0102] (1) Initialize the feature collector, feature fusion module, and intelligent recommendation module, and train the model using the preprocessed data;
[0103] (2) During the model training phase, the multimodal feature extraction module needs to minimize the loss function ξ of the product recommendability module in order to improve its ability to correctly judge whether a product can be recommended. h (θ e ,θ h ), and at the same time, the multimodal feature extraction module attempts to maximize the loss function ξ of the recommended product scoring module g (θ e ,θ g ), to capture the common features of all recommended products, and the recommended product scoring module needs to minimize the loss function of the module to score the recommended products, indicating the degree to which the products are recommended. Since the multimodal feature extraction module and the recommended product scoring module are adversarial, the overall loss of the three loss functions is defined as:
[0104] ξ final (θ e ,θ h ,θ g )=ξ h (θ e ,θ h )-λξ g (θ e ,θg ) is a trade-off parameter. Among them, θ e is the parameter set of the multimodal feature extraction module, θ h represents the parameter set of the product recommendation module, θ g is the parameter set of the recommended product scoring module, and λ controls the weights of the two objective functions. For this adversarial strategy, the final target parameters are shown in the following formula:
[0105]
[0106] in, and is the optimal parameter sought by the objective function;
[0107] (3) During the training process, various fusion modules will use different types of large batches of multimodal data for unlabeled training, and then use a small amount of labeled data for fine-tuning testing to obtain different parts of the fusion modules when used;
[0108] When searching for the final target parameters, a gradient reversal layer GRL is added to solve the parameter optimization problem. The optimization process of the model parameters is updated according to the following formula:
[0109]
[0110]
[0111] Among them, η is the learning rate. In order to maintain the stability of the model training process, the formula x is used to adjust the learning rate η, the formula is as follows:
[0112]
[0113] Where α = 10, β = 0.75;
[0114] S6. Based on the recommended content in step S5 above, the feedback module will track the user's interactive behavior on the recommended content and feed it back to the e-commerce recommendation model. The platform users' feedback on the recommended information will be continuously tracked through distributed crawler technology, and fed back into the recommendation model by forming new training data.
[0115] The principle of the present invention is as follows:
[0116] S1. Extract user behavior data on the e-commerce platform, pre-process the collected user behavior data, and analyze and evaluate it;
[0117] S2. Use the collected data to train the recommendation model, pre-train the model using a large amount of unlabeled data, and then fine-tune it using a small amount of pre-processed labeled data;
[0118] S3, storing the trained recommendation model for later use;
[0119] S4. Usage process: First, all the data related to the product is collected using the deep learning image feature collector, text feature collector, time series feature collector, audio feature collector, and video feature collector to extract features;
[0120] S5, extracting deep features of modal interaction by using bilinear pooling multimodal knowledge fusion according to the features of various single-modal feature collectors in the above step S4;
[0121] S6. By using the deep feature information of multimodal fusion extracted in the above step S5, the bilinear pooling is used again in the intelligent recommendation module to classify and fuse the features, and the fused features are put into the intelligent recommendation module to obtain the regional recommendation index and the recommendation index;
[0122] S7, according to the regional recommendation index and recommendation index in the above step S6, determine the recommendation form, recommended area and recommended order of the product;
[0123] S8. Based on the recommended content in step S7, the feedback module tracks the user's interactive behavior with respect to the recommended content and feeds it back to the e-commerce recommendation model.
[0124] The model training process is as follows Figure 2 As shown, the principle is as follows:
[0125] S1. Obtain platform user data, multi-modal product blog and video live streaming data in each region, and pre-process them. Perform manual comprehensive evaluation of regional recommendation index and recommendation index based on the collected evaluation data, and put the data into the database.
[0126] S2, using unlabeled data to train a unimodal feature collector and a bilinear pooling layer;
[0127] S3, using labeled data to train the added area recommendation index module and the recommendation index module;
[0128] S4. Storage model.
[0129] The model usage process is as follows Figure 3 As shown, the principle is as follows:
[0130] S1. Obtain platform user data and product information;
[0131] S2. Collecting unimodal data using a unimodal feature collector, fusing different unimodal features using bilinear pooling according to different data sources, and then fusing the fusion features of the unimodal features using bilinear pooling;
[0132] S3, using the fusion features in S2 to enter the recommended area index calculation module to obtain the regional recommendation index matrix;
[0133] S4, using the features and regional recommendation index matrix in S2 and S3, put them into multiple recommendation index calculation modules, calculate the recommendation index, and finally put them into a comprehensive recommendation index calculation module to calculate the final recommendation index.
[0134] It should be noted that the term "comprises" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device that includes a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such process, method, article, or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the presence of other identical elements in the process, method, article, or device that includes the element.
[0135] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A product recommendation method based on multimodal knowledge fusion of bilinear pooling, characterized in that: The steps include: S1. Extract user behavior data on the e-commerce platform, pre-process the collected user behavior data, and analyze and evaluate it; S2, extracting features from the collected data using a deep learning image feature collector, a text feature collector, a time series feature collector, an audio feature collector, and a video feature collector; S3, extracting deep features of modal interaction by using bilinear pooling multimodal knowledge fusion according to the features of various single-modal feature collectors in the above step S2; S4, by using the deep feature information of multimodal fusion extracted in the above step S3, the deep features are classified and fused again by using bilinear pooling in the intelligent recommendation module, and the fused features are put into the intelligent recommendation module to obtain the recommended area index and the recommendation index; S5. Determine the recommendation form, recommended area and recommended order of the product according to the recommended area index and the recommended index in the above step S4; S6. Based on the recommended content in step S5, the feedback module tracks the user's interactive behavior with respect to the recommended content and feeds it back to the e-commerce recommendation model.
2. The product recommendation method based on multimodal knowledge fusion based on bilinear pooling according to claim 1, characterized in that: The steps from step S3 to step S5 are as follows: (1) Initialize the feature collector, feature fusion module, and intelligent recommendation module, and train the model using the preprocessed data; (2) During the model training phase, the multimodal feature extraction module needs to minimize the loss function ξ of the product recommendability module in order to improve its ability to correctly judge whether a product can be recommended. h (θ e ,θ h ), and at the same time, the multimodal feature extraction module attempts to maximize the loss function ξ of the recommended product scoring module g (θ e ,θ g ), to capture the common features of all recommended products, and the recommended product scoring module needs to minimize the loss function of the module to score the recommended products, indicating the degree to which the products are recommended. Since the multimodal feature extraction module and the recommended product scoring module are adversarial, the overall loss of the three loss functions is defined as: ξ final (θ e ,θ h ,θ g )=ξ h (θ e ,θ h )-λξ g (θ e ,θ g ) adjusts the trade-off parameter, where θ e is the parameter set of the multimodal feature extraction module, θ h represents the parameter set of the product recommendation module, θ g is the parameter set of the recommended product scoring module, and λ controls the weights of the two objective functions. For this adversarial strategy, the final target parameters are shown in the following formula: in, and is the optimal parameter sought by the objective function; (3) During the training process, various fusion modules will use different types of large-scale multimodal data for unlabeled training, and then use a small amount of labeled data for fine-tuning testing to obtain the fusion modules of different parts when used.
3. The product recommendation method based on multimodal knowledge fusion based on bilinear pooling according to claim 2, characterized in that: When searching for the final target parameters, a gradient reversal layer GRL is added to solve the parameter optimization problem. The optimization process of the model parameters is updated according to the following formula: Among them, η is the learning rate. In order to maintain the stability of the model training process, the formula x is used to adjust the learning rate η. The formula is as follows: Where α=10, β=0.
75.
4. The product recommendation method based on multimodal knowledge fusion based on bilinear pooling according to claim 1, characterized in that: When the step S3 utilizes bilinear pooling multimodal knowledge fusion to extract deep features of modal interaction; First, define a basic model of the multimodal decomposition bilinear pooling method, the formula is as follows: Take text images as an example. is the product text feature extracted in step S2, is the product image feature extracted in step S2, is the parameter matrix, Z is the model output; Since the model needs to capture the relationship between different features, it needs to learn a large number of parameters, which will lead to high computational costs and the risk of overfitting. i Decomposed into two low-rank matrices U i and V i : where k is the dimension of the two low-rank matrices, The symbol o is the Hadamard product, Ι is a matrix of all 1s, and the two third-order tensors obtained are and As the weight of the o-dimensional output, on the basis of consistency, use the two-dimensional matrix and Denote U and V respectively, so the above formula can be rewritten as: The above formula indicates that a one-dimensional non-overlapping window of size k is used to pool the product features.
5. The product recommendation method based on multimodal knowledge fusion based on bilinear pooling according to claim 4, characterized in that: In order to solve the problem that the output dimension size will change greatly due to the introduction of element-level multiplication and the convergence effect of the model is not good, the following two formulas are used to solve it: Z=sign(Z)|Z| 0.5 Z=Z T / ||From|| Finally, the multimodal feature extraction module is represented as E(P; θ e ), we get the formula: Z=E(P;θ e ) Where E represents the overall mapping function, P represents the input, and θ e Represents the parameter set of the multimodal feature extraction module.
6. The product recommendation method based on multimodal knowledge fusion based on bilinear pooling according to claim 1, characterized in that: The steps of the intelligent recommendation module in step S4 include: (1) After obtaining the fusion features, they are input into the regional recommendation index calculation module, and then the module output and different fusion features are input into different recommendation index calculation modules; (2) The different fusion features in the above step (1) are combined into image-text fusion, time series image-text fusion, video-audio time series image-text fusion, and video time series image-text fusion. The mixed features are further fused using a bilinear pooling fusion layer to obtain fusion features. (3) The regional recommendation index calculation module calculates the regional recommendation vector by integrating features. The value on the regional recommendation vector represents the recommendation index of its region. The recommendation area includes four categories: product page, community, video recommendation area and live e-commerce sequence. The product page includes the homepage, sidebar and bottom bar. The community includes the recommendation below the blogger on the homepage. The video recommendation includes the homepage, the recommendation below the video and the recommendation after the video is played. The live e-commerce sequence only includes the live e-commerce sequence. This module is defined as H(Z; θ h ), where H represents the mapping function of the module, θ h Represents the parameter set of the product recommendation module, a multi-product ρ i The regional recommendation index is but: Let Y represent a set of labels, where The sum is 1, which represents the probability that the product should appear in each area, and the classification loss function is calculated using cross entropy: To find the optimal parameters and Need to minimize the loss function ξ h (θ e ,θ h ): (4) The recommendation index calculation module obtains the recommendation index of the product in the region through the recommendation vector and fusion features of the recommendation region index calculation module, and defines the recommendation product scoring module as G(Z; θ g ), where G represents the mapping function of the module, θ g Represents the parameter set of the recommended product scoring module, using Y g Represents the scoring label, and the loss function of the cross entropy calculation module: To find the optimal parameters Need to minimize the loss function ξ g (θ e ,θ g ): The above loss function ξ g (θ e ,θ g ) is used to estimate the difference in the distribution of products with different recommendation levels. A larger loss means that all recommended products have similar characteristics. Therefore, in order to eliminate the uniqueness of the characteristics of a recommended product, it is necessary to find the optimal parameter θ e To maximize the loss function ξ g (θ e ,θ g ).
7. The product recommendation method based on multimodal knowledge fusion based on bilinear pooling according to claim 1, characterized in that: Step S5 will build multiple recommendation tracks for users based on the collected user personal information and product categories, and place products into the recommendation channel based on the recommendation points, product category information, and price information.
8. The product recommendation method based on multimodal knowledge fusion based on bilinear pooling according to claim 1, characterized in that: Step S6 continuously tracks the feedback of platform users on the recommendation information through distributed crawler technology, and feeds it back into the recommendation model by forming new training data.
9. The product recommendation method based on multimodal knowledge fusion based on bilinear pooling according to claim 1, characterized in that: In step S3, a bilinear pooling fusion layer is used for multi-level fusion. First, the bilinear pooling fusion layer is used for image-text fusion, time series image-text fusion, video-audio time series image-text fusion, and video time series image-text fusion. Then, the bilinear pooling fusion layer is used to fuse different types of fusion features to obtain more interactive information.
10. The product recommendation method based on multimodal knowledge fusion based on bilinear pooling according to claim 1, characterized in that: In step S1, the user's behavior data on the e-commerce platform includes platform multimodal data, community multimodal data, video live e-commerce multimodal data, analysis and evaluation data and user data, among which the platform multimodal data includes product images, product text descriptions and product reviews, the community multimodal data includes blogger browsing order, browsing blogger text, browsing blogger pictures and blog comments, the video live e-commerce multimodal data includes browsing time period videos, live broadcast comments, live broadcast e-commerce browsing order and live broadcast audio, the analysis and evaluation data includes the number of product searches, the number of add-to-carts, the number of purchases, the number of blog post likes and comments, the number of video likes and comments, the popularity of live broadcast sales, the number of active people, sales and the number of fans, and the user data includes the user's height, weight, age, occupation, usage time, and partition browsing status.
Citation Information
Patent Citations
Image content question and answer method based on multi-modality low-rank dual-linear pooling
CN107480206A
Combined commodity retrieval method and system based on multi-modal pre-training model
CN114840705A