An item recommendation method, device, equipment and medium based on multimodal representation

By aligning and encoding the user's historical interactive item text and images, a semantic encoding matrix is ​​generated and combined with the item representation matrix is ​​predicted to predict user preferences, and the problem of insufficient alignment processing of multimodal information in the prior art is solved, and the accuracy of item recommendation is improved.

CN119807547BActive Publication Date: 2025-06-17SHENZHEN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510288376.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-17
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The prior art fails to effectively align multimodal historical interaction information in item recommendation, resulting in the inability to distinguish different modal information of the same item, reducing the accuracy of recommendation.

Method used

By aligning and encoding the user's historical interactive item text and images, a text semantic coding matrix and image semantic coding matrix are generated, and combined with the item representation matrix, the user's preference for each item is predicted, and the recommended items are finally screened out.

Benefits of technology

Through alignment encoding processing, the text and image semantics of the same item are accurately matched, improving the accuracy of item recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807547B_ABST
    Figure CN119807547B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of item recommendation, and specifically relates to an item recommendation method, device, equipment and medium based on multi-modal representation. The present invention first performs alignment encoding processing on text and images to obtain a text semantic encoding matrix and an image semantic encoding matrix, and then predicts the preference degree of a user for each item according to the text semantic encoding matrix, the image semantic encoding matrix and the item representation matrix. Finally, according to the preference degree of the user, the target recommended item is obtained. The alignment in the present invention is to make the corresponding same row in the two matrices of the text semantic encoding matrix and the image semantic encoding matrix point to the same item, so as to effectively represent the same item. Since the alignment can accurately match the text semantics and image semantics corresponding to the same item, subsequent calculation of the preference degree of the user for each item can be performed based on the text semantics and image semantics of the item, thereby improving the recommendation accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of item recommendation, and particularly to an item recommendation method, device, equipment and medium based on multi-modal representation. Background Art

[0002] According to the historical interaction information of users, the preference degree of users for each item is predicted, and then corresponding items are recommended to users according to the preference degree. The historical interaction information is the behavior of users to obtain information related to items. For example, a user watching a teaching video of using an item is the historical interaction information of the user. The historical interaction information includes various modal information such as item identification information, item text information and item image information. Among them, the item identification information is the item number, the item text information describes the item in the form of text, and the item image information shows the item in the form of an image. Although the prior art also recommends items to users according to the multi-modal historical interaction information of users, it does not perform alignment processing on the multi-modal historical interaction information, resulting in the inability to distinguish which modal information belongs to the modal information of the same item, thereby reducing the accuracy of item recommendation.

[0003] Therefore, the prior art still needs to be improved. Summary of the Invention

[0004] To solve the above technical problems, the present invention provides an item recommendation method, device, equipment and medium based on multi-modal representation, which solves the problem of reducing the accuracy of item recommendation in the prior art.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] In the first aspect, the present invention provides an item recommendation method based on multi-modal representation, which includes:

[0007] Obtain the historical interaction item text and historical interaction item image of the user, and perform alignment encoding processing on the historical interaction item text and historical interaction item image to obtain a text semantic encoding matrix and an image semantic encoding matrix. The same row and column elements of the text semantic encoding matrix and the image semantic encoding matrix correspond to the same item;

[0008] Generate an item representation matrix for representing the interaction time and identification information of each item, and predict the preference degree of the user for each item according to the text semantic encoding matrix, the image semantic encoding matrix and the item representation matrix;

[0009] Select recommended items from each item according to the preference degree of the user for each item.

[0010] In one implementation, the historical interactive item text and the historical interactive item image are subjected to alignment encoding processing to obtain a text semantic encoding matrix and an image semantic encoding matrix, including:

[0011] Apply a semantic discriminator to the historical interactive item text and the historical interactive item image to obtain the item semantics of the historical interactive item text and the item semantics of the historical interactive item image;

[0012] Apply a text encoder to the historical interactive item text marked with the item semantics to obtain a text semantic encoding matrix;

[0013] Apply an image encoder to the historical interactive item image marked with the item semantics to obtain an image semantic encoding matrix.

[0014] In one implementation, the text encoder and the image encoder are the trained text encoder and the trained image encoder, and the training method of the text encoder and the image encoder is joint training. A total loss function for joint training is constructed, including:

[0015] During the training process, obtain the text semantic encoding sample matrix output by the text encoder and the image semantic encoding sample matrix output by the image encoder;

[0016] Divide each sample text semantics included in the text semantic encoding sample matrix into text sample pair groups. Each group of the text sample pair groups includes a first positive sample pair and a first negative sample pair. Only one same item is included in the items corresponding to the first positive sample pair and the first negative sample pair;

[0017] Obtain a first loss function based on the first positive sample pair and the first negative sample pair;

[0018] Divide each image sample semantics included in the image semantic encoding sample matrix into image sample pair groups. Each group of the image sample pair groups includes a second positive sample pair and a second negative sample pair. Only one same item is included in the items corresponding to the second positive sample pair and the second negative sample pair;

[0019] Obtain a second loss function based on the second positive sample pair and the second negative sample pair;

[0020] Obtain a text-image sample group based on the text semantic encoding sample matrix and the image semantic encoding sample matrix. Each group of the text-image sample groups includes a text-image positive sample pair and a text-image negative sample pair. The sample text semantics and the image sample semantics included in each text-image positive sample pair correspond to the same item, and the sample text semantics and the image sample semantics included in the text-image negative sample pair correspond to two different items;

[0021] Obtain a third loss function according to the positive text image pairs and the negative text image pairs.

[0022] Construct a total loss function according to the first loss function, the second loss function, and the third loss function.

[0023] In one implementation, obtaining the first loss function according to the first positive sample pair and the first negative sample pair includes:

[0024] Determine the semantic similarity of the two sample texts included in the first positive sample pair.

[0025] Determine the semantic similarity of the two sample texts included in the first negative sample pair.

[0026] Obtain the first loss function according to the similarity corresponding to the first positive sample pair and the similarity corresponding to the first negative sample pair.

[0027] In one implementation, predicting the preference degree of the user for each of the items according to the text semantic encoding matrix, the image semantic encoding matrix, and the item representation matrix includes:

[0028] Obtain a behavior preference sequence according to the item representation matrix, where the behavior preference sequence is used to characterize the preference order of the user for each item.

[0029] Perform an aggregation process on the text semantic encoding matrix and the image semantic encoding matrix to obtain a content preference sequence.

[0030] Perform an aggregation process on the behavior preference sequence and the content preference sequence to obtain a mixed preference sequence.

[0031] Obtain a user's final preference representation vector according to the behavior preference sequence, the content preference sequence, and the mixed preference sequence.

[0032] Predict the preference degree of the user for each of the items according to the user's final preference representation vector.

[0033] In one implementation, apply an item number modality sequence encoder to the item representation matrix to obtain a behavior preference sequence; apply a content modality sequence encoder to the text semantic encoding matrix and the image semantic encoding matrix for aggregation processing to obtain a content preference sequence; apply a mixed modality sequence decoder to the behavior preference sequence and the content preference sequence for aggregation processing to obtain a mixed preference sequence.

[0034] After completing the training of the item number modality sequence encoder, based on the trained item number modality sequence encoder, the content modality sequence encoder and the hybrid modality sequence decoder are jointly trained.

[0035] In one implementation, the final loss function used for the joint training of the content modality sequence encoder and the hybrid modality sequence decoder consists of a contrastive loss function, a cross-entropy loss function, and sample information, and the sample information is generated from a sample identifier, sample text, and sample image.

[0036] In a second aspect, an embodiment of the present invention further provides an item recommendation device based on multi-modal representation, where the device includes the following components:

[0037] An alignment module, configured to obtain the historical interaction item text and historical interaction item image of a user, and perform alignment encoding processing on the historical interaction item text and historical interaction item image to obtain a text semantic encoding matrix and an image semantic encoding matrix, where the same row and column elements of the text semantic encoding matrix and the image semantic encoding matrix correspond to the same item;

[0038] A prediction module, configured to generate an item representation matrix for characterizing the interaction time and identification information of each item, and predict the preference degree of the user for each item according to the text semantic encoding matrix, the image semantic encoding matrix, and the item representation matrix;

[0039] A recommendation module, configured to obtain a target recommended item according to the preference degree of the user for each item.

[0040] In a third aspect, an embodiment of the present invention further provides a terminal device, where the terminal device includes a memory, a processor, and an item recommendation program based on multi-modal representation stored in the memory and executable on the processor. When the processor executes the item recommendation program based on multi-modal representation, the steps of the above-mentioned item recommendation method based on multi-modal representation are implemented.

[0041] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which an item recommendation program based on multi-modal representation is stored. When the item recommendation program based on multi-modal representation is executed by a processor, the steps of the above-mentioned item recommendation method based on multi-modal representation are implemented.

[0042] Beneficial effects: First, the present invention performs alignment encoding processing on the user's historical interaction item texts and historical interaction item images to obtain a text semantic encoding matrix and an image semantic encoding matrix. Then, based on the text semantic encoding matrix, the image semantic encoding matrix, and the item representation matrix, the preference degree of the user for each item is predicted. Finally, according to the preference degree of the user, recommended items are selected from each item and presented to the user. The alignment in the present invention is to make the corresponding same row in the two matrices of the text semantic encoding matrix and the image semantic encoding matrix point to the same item, that is, the purpose of alignment is to make the text semantic vector and the image semantic vector of the same item as similar as possible, so as to narrow the spatial distance represented by these two vectors, thereby being able to effectively represent the same item. Since alignment can accurately match the text semantics and image semantics corresponding to the same item, subsequent calculations can be made based on the text semantics and image semantics of the item to obtain the preference degree of the user for each item, thereby improving the recommendation accuracy. Description of the Drawings

[0043] Figure 1 is the overall flowchart of the present invention;

[0044] Figure 2 is the construction flowchart of the content modality sample pair in the embodiment of the present invention;

[0045] Figure 3 is the alignment schematic diagram in the embodiment of the present invention;

[0046] Figure 4 is the preference prediction schematic diagram in the embodiment of the present invention;

[0047] Figure 5 is the structure diagram of the item recommendation device based on multi-modal representation provided by the present invention;

[0048] Figure 6 is the internal structure principle block diagram of the terminal device provided in the embodiment of the present invention. Detailed Embodiments

[0049] The following combines embodiments and the accompanying drawings of the specification to clearly and completely describe the technical solutions in the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0050] It has been found through research that, based on the user's historical interaction information, the preference degree of the user for each item is predicted, and then corresponding items are recommended to the user according to the preference degree. The historical interaction information is the user's behavior of obtaining information related to the item. For example, when the user watches a teaching video of using an item, it is the user's historical interaction information. The historical interaction information includes various modal information such as item identification information, item text information, and item image information. Among them, the item identification information is the item number, the item text information describes the item in the form of text, and the item image information displays the item in the form of an image. Although the existing technology also recommends items to the user based on the multi-modal historical interaction information of the user, it does not perform alignment processing on the multi-modal historical interaction information, resulting in the inability to distinguish which modal information belongs to the modal information of the same item, thereby reducing the accuracy of item recommendation.

[0051] To solve the above technical problems, the present invention provides an item recommendation method, device, equipment, and medium based on multi-modal representation, which solves the problem of reducing the accuracy of item recommendation in the existing technology.

[0052] The item recommendation method based on multi-modal representation in this embodiment can be applied to a terminal device. The terminal device can be a terminal product with image processing and text processing functions, such as a computer, etc. In this embodiment, as Figure 1 shown, the item recommendation method based on multi-modal representation specifically includes the following steps:

[0053] S100, obtain the historical interaction item text and historical interaction item image of the user, and perform alignment encoding processing on the historical interaction item text and historical interaction item image to obtain a text semantic encoding matrix and an image semantic encoding matrix. The same row and column elements of the text semantic encoding matrix and the image semantic encoding matrix correspond to the same item;

[0054] S200, generate an item representation matrix for representing the interaction time and identification information of each item, and predict the preference degree of the user for each item according to the text semantic encoding matrix, the image semantic encoding matrix, and the item representation matrix;

[0055] S300, obtain the target recommended item according to the preference degree of the user for each item.

[0056] The application scenario of the item recommendation method based on multi-modal representation according to steps S100, S200, and S300 is as follows:

[0057] When a user views texts and images on a mobile phone, and these images and texts describe the functions and appearances of items, then these texts and images are the historical interactive item texts and historical interactive item images generated during the interaction between the user and the mobile phone. For example, during the interaction between the user and the mobile phone, the user browses three groups of historical interactive item texts and three groups of historical interactive item images, which respectively involve three items, namely item A, item B, and item C. The three groups of historical interactive item texts and three groups of historical interactive item images are subjected to alignment encoding processing, and a text semantic encoding matrix (this matrix is a matrix with one row and three columns) and an image semantic encoding matrix (this matrix is a matrix with one row and three columns) are obtained respectively. Alignment means that if the text semantics in the first row and first column of the text semantic encoding matrix correspond to item A, then the image semantics in the first row and first column of the image semantic encoding matrix also correspond to item A; similarly, the first row and second column of the text semantic encoding matrix and the first row and second column of the image semantic encoding matrix also correspond to the same item C, and similarly, the first row and third column of the text semantic encoding matrix and the first row and third column of the image semantic encoding matrix also correspond to the same item B.

[0058] The first row and first column of the item representation matrix are used to represent the ID (the ID is the identification information) of item A and the interaction time of item A (that is, the time when the user views the texts and images related to A on the mobile phone); the first row and second column of the item representation matrix are used to represent the ID of item C and the interaction time of item C; the first row and third column of the item representation matrix are used to represent the ID of item B and the interaction time of item B.

[0059] During the interaction between the user and the mobile phone, the text semantic encoding matrix, the image semantic encoding matrix, and the item representation matrix are statistically obtained. Based on these three matrices, the preference degrees of the user for the above three items are calculated. Finally, the shopping software on the mobile phone recommends corresponding items to the user according to the preference degrees of the user.

[0060] In the first embodiment, the historical interactive item texts and historical interactive item images pass through a semantic discriminator to mark the item semantics on the texts and the item semantics on the images. Then, the historical interactive item texts marked with item semantics are input into a text encoder to obtain a text semantic encoding matrix, and the historical interactive item images marked with item semantics are input into an image encoder to obtain an image semantic encoding matrix. Both the text encoder and the image encoder are encoders after training. In this embodiment, the text encoder and the image encoder are trained in a joint training manner. The training process is as follows:

[0061] The user item interaction sequence is used as a training sample and input into the semantic discriminator. The item information in this sequence includes the item ID and text information, and the most similar item pairs (the similar item pairs are the item pairs in Figure 3 ) are selected based on text similarity. The similar item pairs include two item images or two item texts.

[0062] The semantic discriminator labels the item semantics of the user-item interaction sequence, and then inputs the text pairs containing similar item semantics into the text encoder. The text encoder outputs a text semantic encoding sample matrix. Similarly, the image pairs containing similar item semantics are input into the image encoder, and the image encoder outputs an image semantic encoding sample matrix. Then, a total loss function for joint training is constructed based on these two matrices. (This total loss function is the supervised fine-tuning loss function SFT, and SFT stands for Supervised Fine-Tuning), and finally, the image encoder and the text encoder are fine-tuned according to the total loss function to complete the training of the image encoder and the text encoder.

[0063] The above similar items are item pairs, and the construction process of the item pairs is as follows Figure 2 shown. Call the large model as the semantic discriminator. Construct a prompt based on the user-item interaction behavior sequence and input it into the large model. In the prompt, use the ID and text information of the last item interacted by the user as the target, and the ID and text information of each previous interacted item as candidates. The large model selects the candidate item with the most similar text semantics to the target item from the candidate items and outputs the candidate item ID. According to the candidate item ID selected by the large model, an object item ID-candidate item ID pair can be obtained, and the database is retrieved to obtain (object item text - candidate item text) and (object item image - candidate item image).

[0064] Among them, constructing the total loss function includes the following specific steps S011 to S0110:

[0065] S011, obtain the text semantic encoding sample matrix output by the text encoder, and obtain the image semantic encoding sample matrix output by the image encoder.

[0066] S012, divide each sample text semantics contained in the text semantic encoding sample matrix into text sample pair groups. Each group of text sample pair groups includes a first positive sample pair and a first negative sample pair. Only one same item is included in the items corresponding to the first positive sample pair and the first negative sample pair.

[0067] For example, the text semantic encoding sample matrix contains five groups of sample text semantics: sample text semantics A, sample text semantics B, sample text semantics C, sample text semantics D, and sample text semantics E. Suppose A describes item A, B describes item B, C describes item C, D describes item D, and E also describes item A. If A that describes A and B that describes B are used as the first positive sample pair AB, and C that describes C and D that describes D are used as the first positive sample pair CD, then E that describes A and C that describes C are the first negative sample pair CE, E that describes A and D that describes D are the first negative sample pair ED, and A that describes A and D that describes D are the first negative sample pair AD. Among them, the first positive sample pair AB and the first negative sample pair CE form a group of text sample pairs (the first positive sample pair AB and the first negative sample pair CE only include one same item A), and the first positive sample pair AB and the first negative sample pair ED form another group of text sample pairs.

[0068] The purpose of training the text encoder is to make the text encoder use one text semantic to describe the same item as much as possible.

[0069] S013. Determine the similarity between the two sample text semantics included in the first positive sample pair :

[0070] ;

[0071] In the formula, t2t represents two texts, is one of the sample text semantics (represented by which is one of the first positive sample pairs corresponding to ), represents the text, is the th another sample text semantics in the first positive sample pair (represented by which is the other of the first positive sample pair), is the and vector inner product, represents the vector inner product operation, is a parameter used to control the similarity score, .

[0072] S014. Determine the similarity between the two sample text semantics included in the first negative sample pair :

[0073] ;

[0074] is one of the sample text semantics, in Corresponding to the first negative sample pair, For the The semantics of another sample in the first negative sample pair ( Indicates the other one of the first negative sample pairs), for and The vector inner product of .

[0075] S015: based on the similarity corresponding to the first positive sample pair The similarity with the first negative sample pair , and get the first loss function :

[0076] ;

[0077] is the total number of first positive sample pairs. In this embodiment, given a batch of item-text pairs, each text pair is considered as a positive sample pair, and the item texts from other text pairs in the same batch are considered as negative samples. This allows the encoder to learn the representation of positive sample pairs in the semantic space, shorten the distance between positive sample vectors to make them more similar, and push the distance between negative sample vectors to make them dissimilar. The first loss function is used Jointly training the text encoder and image encoder allows the encoder to learn the alignment task between text and image.

[0078] S016, semantically dividing each image sample contained in the image semantic coding sample matrix into image sample pair groups, each group of the image sample pair groups includes a second positive sample pair and a second negative sample pair, and the objects corresponding to the second positive sample pair and the second negative sample pair include only one identical object.

[0079] The second positive sample pair and the second negative sample pair are obtained in the same manner as step S012, that is, if an image in the second positive sample pair describes an object, then other images describing the object are in the second negative sample pair.

[0080] S017, obtaining a second loss function according to the second positive sample pair and the second negative sample pair ,in Represents an image.

[0081] The first loss function will be calculated The similarity in the formula can be calculated using the semantics of the image samples to obtain the second loss function .

[0082] S018. According to the text semantic encoding sample matrix and the image semantic encoding sample matrix, a text-image sample group is obtained. Each text-image sample group includes a text-image positive sample pair and a text-image negative sample pair. The sample text semantics and image sample semantics included in each text-image positive sample pair correspond to the same item, and the sample text semantics and image sample semantics included in the text-image negative sample pair correspond to two different items.

[0083] S019. According to the text-image positive sample pair and the text-image negative sample pair, a third loss function is obtained :

[0084] ;

[0085] ;

[0086] ;

[0087] is the sample text in the th text-image positive sample pair, is the sample image in the th text-image positive sample pair, represents an image, is the and vector inner product of, is the similarity of the th text-image positive sample pair, is the sample image in the th text-image negative sample pair. is the and similarity of, N is the total number of text-image positive sample pairs, is the base of the natural logarithm.

[0088] S0110. According to the first loss function , the second loss function , and the third loss function , a total loss function is constructed:

[0089] ;

[0090] This embodiment uses to train a text encoder and an image encoder for each downstream recommendation field, and uses the fine-tuned encoder as a feature encoder to encode the text semantic vector and image semantic vector of each item. These semantic vectors are integrated into the semantic encoding matrix and participate in the modeling of the sequence preference learning module.

[0091] Embodiment 2, based on Embodiment 1, step S200 uses the item number modal sequence encoder as shown in Figure 4 (this encoder is a Transformer model, and the Transformer model is a neural network model for processing language data), the content modal sequence encoder (this encoder is used to mine the user content preference from the text sequence and picture sequence of user interaction), and the hybrid modal sequence decoder to predict the preference degree of the user for each item. The item number modal sequence encoder, the content modal sequence encoder, and the hybrid modal sequence decoder are all prior arts.

[0092] Adopt a two-step training strategy optimized by low-rank representation to train the item number modal sequence encoder, the content modal sequence encoder, and the hybrid modal sequence decoder. That is, first train the item number modal sequence encoder and fix the weights of this encoder. Then, based on the output of this encoder, the user's behavior preference sequence (the behavior preference sequence is the preference order of the user for the items) is used to train the content modal sequence encoder and the hybrid modal sequence decoder.

[0093] Among them, the pre-trained item number modal sequence encoder is given fixed weights , and the update process satisfies the following relational expression:

[0094] ;

[0095] In the formula, is the updated weight, and are both learnable lightweight weights, d, r, and k are vector dimensions, A can be a normal distribution matrix with a mean of 0, and C is a matrix of all zeros.

[0096] Adopt a joint training method to train the content modal sequence encoder and the hybrid modal sequence decoder. The final loss function used for training is :

[0097] ;

[0098] In the formula, is a hyperparameter for controlling the contrastive learning task, is the contrastive loss function, is 's regularization parameter.

[0099] ;

[0100] In the formula, is a parameter for controlling the similarity score, , M refers to the batch size of the input data within the same time, and m refers to the interaction sequence of the m-th user. a is the first letter of the English word "aggregation", and aggregation represents the final user preference representation vector , which is aggregated from the outputs of the item number modality sequence encoder, the content modality sequence encoder, and the hybrid modality sequence decoder. B refers to the number of negative sample pairs in the contrast loss used to label the -th negative sample. refers to the content modality representation vector of the next item as the label in the interaction sequence of the m-th user. refers to the content modality representation vector of the item corresponding to the -th negative sample.

[0101] is the cross-entropy loss function , where is the Figure 4 training prediction value of the preference aggregation layer in for the -th item, is the true label, representing the user's preference degree for the item, and I is the total number of items to be predicted.

[0102] In the formula, is the sample information set of the three modalities , is the information set composed of sample identifiers is the item identifier of the number is the information set composed of sample texts is the information set composed of images.

[0103] Example 3. Based on the encoder-decoder trained in Example 1 and Example 2, this example predicts the prediction scores of each item according to the user's historical interaction item text and historical interaction item image, and this score is used to characterize the user's preference degree for the item.

[0104] In this example, the step of performing alignment encoding processing on the historical interaction item text and historical interaction item image in step S100 to obtain the text semantic encoding matrix and the image semantic encoding matrix includes the following specific steps S101, S102, and S103:

[0105] S101. Apply a semantic discriminator to the historical interaction item text and historical interaction item image to obtain the item semantics of the historical interaction item text and the item semantics of the historical interaction item image.

[0106] That is, the item text (i.e., the historical interaction item text) that the user has viewed on the mobile phone is input into the semantic discriminator, and the semantic discriminator analyzes the item semantics on the item text, and the item semantics is the name of the item. Similarly, the image (i.e., the historical interaction item image) that the user has viewed on the mobile phone is input into the semantic discriminator, and the semantic discriminator identifies the item on the image and analyzes the attribute information of the item.

[0107] S102. Apply a text encoder to the historical interaction item text marked with the item semantics to obtain a text semantic coding matrix.

[0108] The text semantic coding matrix is composed of several vectors, and each vector includes the coding information of multiple historical interaction item texts corresponding to the same item semantics. That is, the text encoder divides the historical interaction item text according to the item semantics of the historical interaction item text, so that the information of multiple historical interaction item texts belonging to the same item semantics is encoded in the same vector.

[0109] S103. Apply an image encoder to the historical interaction item image marked with the item semantics to obtain an image semantic coding matrix.

[0110] Similarly, the image semantic coding matrix is composed of several vectors, and each vector includes the coding information of multiple historical interaction item images corresponding to the same item semantics. That is, the image encoder divides the historical interaction item image according to the item semantics of the historical interaction item image, so that the information of multiple historical interaction item images belonging to the same item semantics is encoded in the same vector.

[0111] In this embodiment, the interaction time and identification information of each item are input into the item number modality sequence encoder, and the item number modality sequence encoder outputs an item representation matrix , that is, the elements in the item representation matrix represent the interaction time and identification information of the item.

[0112] ;

[0113] In the formula, is a vector composed of the ID (ID is the item number) of the first item, is a vector composed of the position of the first item, and the position coding is used to represent the position of each item in (this position is related to the item interaction time, for example, the interaction item farther from the current moment is in the matrix the position is more forward); is a vector composed of the ID of the first item, is a vector composed of the position of the second item; A vector composed of the IDs of the nth item, is a vector composed of the positions of the nth item. The potential dimension (the preset vector dimension) of all the above vectors is , .

[0114] In this embodiment, step S200 includes the following specific steps S201 to S205:

[0115] S201, according to the item representation matrix ( ), obtain a behavior preference sequence, which is used to characterize the preference order of the user for each item.

[0116] That is, input the item representation matrix into the item number modality sequence encoder, and the item number modality sequence encoder outputs the behavior preference sequence .

[0117] The item number modality sequence encoder outputs by the following steps:

[0118] The item number modality sequence encoder uses the self-attention mechanism inside it to map the item representation matrix into a query vector Q, a key vector K, and a value vector V:

[0119] , , , where , , are linear mapping matrices in the self-attention mechanism.

[0120] The self-attention mechanism outputs the matrix according to Q, K, and V:

[0121] ;

[0122] Wherein, represents the self-attention mechanism, represents normalization, and T represents the transpose of the matrix.

[0123] Then input into the pointwise feed-forward neural network FNN to output the behavior preference sequence :

[0124] ;

[0125] Wherein, ReLU represents the activation function, is a learnable weight matrix.

[0126] is a learnable bias vector, , is the first behavioral preference, is the second behavioral preference, is the nth behavioral preference.

[0127] S202, perform an aggregation process on the text semantic encoding matrix and the image semantic encoding matrix to obtain a content preference sequence .

[0128] That is, first convert the dimension of the text vector corresponding to each item in the text semantic encoding matrix to the same dimension as the item number to obtain an item text input matrix ; convert the dimension of the image vector corresponding to each item in the image semantic encoding matrix to the same dimension as the item number to obtain an item image input matrix .

[0129] ;

[0130] ;

[0131] represents an adaptation layer, and the adaptation layer is used for dimension size conversion, represents the first item text vector, represents the second item text vector, represents the nth item text vector, represents the first item image vector, represents the second item image vector, represents the nth item image vector, and the vectors involved in the formulas and are all of size , containing n vectors, and the size of each vector is 1xd. Emb refers to the Embedding technology, which is a common encoding technology in deep learning. It maps high-dimensional data (such as text, pictures) to a low-dimensional continuous vector space, mainly used to capture semantic or structural information in the data.

[0132] As Figure 4 shown, apply a content modality sequence encoder to the item image input matrix and the item text input matrix to obtain a content preference sequence .

[0133] That is, Each item text vector in and each item image vector in are aggregated to obtain an aggregated vector , where represents the serial number.

[0134] ;

[0135] ; In the formula, represents the content, represents the second norm. represents the aggregation matrix. This embodiment uses which helps to unify the semantic representations between various modalities and reduces the computational complexity.

[0136] Use the causal learning model Transformer to learn the user's content preference to obtain the content preference sequence :

[0137] ;

[0138] Attention refers to the attention mechanism, and FNN refers to the feed-forward neural network. Transformer is composed of a stack of the attention mechanism and the feed-forward neural network. , is the first content preference, is the second content preference, is the nth content preference.

[0139] S203. Aggregate the behavior preference sequence and the content preference sequence to obtain a mixed preference sequence :

[0140] ;

[0141] where , is the first mixed preference, is the second mixed preference, is the th mixed preference. This aggregation mechanism can capture the relationship between the behavior preference sequence and the content preference sequence .

[0142] S204. Based on the behavior preference sequence , the content preference sequence , and the mixed preference sequence , obtain the user's final preference 。

[0143] That is, for the behavior preference sequence the th behavior preference in the content preference sequence the th content preference in the mixed preference sequence the th mixed preference are aggregated to obtain :

[0144] ;

[0145] In the formula, Projection represents the linear mapping layer, represents the fully connected layer.

[0146] S205. According to the user's final preference representation vector predict the user's preference degree for each of the items :

[0147] ;

[0148] In the formula, represents the th item, is the ID coding representation vector of item .

[0149] In this embodiment, step S300 screens out the target recommended items from each item according to the user's preference degree for each item , or the text and image of the item corresponding to the largest can be analyzed to obtain the characteristics of the item, and then other items can be recommended to the user according to the characteristics.

[0150] The following comparative experiments illustrate that the recommendation method of the present invention is superior to the existing recommendation methods:

[0151] Four real videos collected from the website are used as the dataset. An item in this dataset contains an item number, a picture of the item, and a text title for describing the item. The leave-one-out method is adopted to divide the dataset, that is, the last item of the user interaction sequence is used for testing, the penultimate item is used as the validation set, and the previous items are used as the training set. As shown in Table 1, the proposed recommendation method and the recommendation methods in the prior art are evaluated using four metrics: Metric 1 is the hit rate metric with a parameter of 5, Metric 2 is the hit rate metric with a parameter of 10, Metric 3 is the normalized discounted cumulative gain with a parameter of 5, and Metric 4 is the normalized discounted cumulative gain with a parameter of 10.

[0152] In Table 1, GRU4Rec, SASRec, BERT4Rec that only use IDs, UniSRec, VQ-rec, MISSRec that only use text, and UniSRec, MISSRec that use text plus IDs are all prior art. MISSRec that uses image plus ID plus text is also prior art. Among them, GRU4Rec is a sequence recommendation method based on recurrent neural networks, which uses a gated neural network to model user preferences. SASRec is a sequence recommendation method based on causal transformers. BERT4Rec is a sequence recommendation method based on bidirectional attention. UniSRec that uses text is a sequence recommendation method based on the text modality, which learns a general semantic representation from text information as the item representation. VQ-Rec that uses text is a semantic representation transfer recommendation method based on vector quantization, which maps the text representation into multiple learnable codebook representations. MISSRec that uses text is a hybrid enhanced sequence recommendation method based on the text modality, which mines user preferences from text sequences. UniSRec that uses text plus IDs is an extended version of UniSRec, and the extended version is fine-tuned by combining the item ID representation matrix in the downstream field to improve the recommendation effect. MISSRec that uses text plus IDs is an extended version of MISSRec, and this extended version is fine-tuned by combining the item ID representation matrix and the text representation matrix in the downstream field. MISSRec that uses image plus ID plus text is a method that models user preferences by jointly using the item ID representation matrix, the text representation matrix, and the image representation matrix.

[0153] Table 1

[0154]

[0155] The optimal metric values of the recommendation methods in the prior art are the underlined metric values. As can be seen from Table 1, compared with the optimal metric values corresponding to the prior art, the metric value of Metric 3 of the proposed recommendation method of the present invention has increased by 6.87%, and the metric value of Metric 4 has increased by 5.87%.

[0156] An ablation experiment was conducted on four datasets to verify the effectiveness of the modality semantic alignment module of the present invention. As shown in Table 2, the four datasets are the comic dataset, the dance dataset, the food dataset, and the movie dataset. Six ablation experiments were set up, namely w / o (removing the contrastive learning of the content modality perception of the present invention), w / o LoRA layer (removing the low-rank representation layer of the present invention), w / o Emb L2 norm (removing the L2 regularization of the present invention, where L2 is the second norm), w / o MixDec (removing the mixed modality decoder of the present invention), w / o IDEnc (removing the ID modality sequence encoder of the present invention), and w / o ContentEnc (removing the content modality sequence encoder of the present invention).

[0157] Table 2

[0158]

[0159] Index four was used as the evaluation index to evaluate the above six ablation experiments and the recommendation method of the present invention.

[0160] It can be seen from Table 2 that the supervised fine-tuning of the joint text encoder and image encoder of the present invention effectively improves the effectiveness of content modality representation learning; removing any one of the alignment tasks of the present invention (the alignment tasks include the alignment task of text and text, the alignment task of text and image, and the alignment task of image and image) will lead to a decrease in model performance.

[0161] By comparing different training strategies, the effectiveness of the two-step training strategy based on low-rank representation optimization proposed by the present invention was verified. The comparison results are shown in Table 3:

[0162] Table 3

[0163]

[0164] In Table 3, #epo represents the number of iterations. Fixed IDEmb means fixing the weight of the encoding matrix for item IDs and updating other parts of the model; Fixed IDEnc means fixing the weight of the ID modality sequence encoder and updating other parts of the model; Fixed IDEmb&Enc means fixing the weights of the encoding matrix for item IDs and the ID modality sequence encoder and updating other parts of the model; Fixed ConEmb means fixing the weights of the item content modality matrix (including the image semantic encoding matrix and the text semantic encoding matrix) and updating other parts of the model; Fixed ConEnc means fixing the weight of the item content modality encoder and updating other parts of the model; Fixed ConEmb&ConEnc means fixing the weights of the item content modality matrix and the item content modality encoder and updating other parts of the model; End2end (not fixed) means using end-to-end training without fixing the weights of any part of the model.

[0165] As can be seen from Table 3, end-to-end training is very inefficient and can lead to unstable performance. This is because the joint training of the content modality and the ID modality makes it difficult for the model to converge; training only the sequence encoder of the pure content modality usually requires more training times and brings sub-optimal recommendation performance; pre-training the ID modality encoder and then training the content modality encoder and other components (i.e., the two-step training strategy proposed in the present invention) brings stable performance improvement and requires fewer iterations to achieve model convergence.

[0166] This embodiment also provides an item recommendation device based on multi-modal representation, as Figure 5 shown. The device includes the following components:

[0167] Alignment module 01, configured to obtain the historical interaction item text and historical interaction item image of the user, and perform alignment encoding processing on the historical interaction item text and historical interaction item image to obtain a text semantic encoding matrix and an image semantic encoding matrix, where the same row and column elements of the text semantic encoding matrix and the image semantic encoding matrix correspond to the same item;

[0168] Prediction module 02, configured to generate an item encoding matrix for representing the interaction time and identification information of each item, and predict the preference degree of the user for each item according to the text semantic encoding matrix, the image semantic encoding matrix, and the item encoding matrix;

[0169] Recommendation module 03, configured to obtain the target recommended item according to the preference degree of the user for each item.

[0170] Based on the above embodiments, the present invention also provides a terminal device, and its principle block diagram can be asFigure 6 As shown in the figure. The terminal device includes a processor, a memory, a network interface, and a display screen connected through a system bus. Among them, the processor of the terminal device is used to provide computing and control capabilities. The memory of the terminal device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the terminal device is used to communicate with external terminals through a network connection. When the computer program is executed by the processor, it realizes a method for recommending items based on multimodal representations. The display screen of the terminal device can be a liquid crystal display screen or an electronic ink display screen.

[0171] Those skilled in the art can understand that Figure 6 the block diagram of the principle shown in the figure is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal device to which the solution of the present invention is applied. The specific terminal device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0172] In one embodiment, a terminal device is provided. The terminal device includes a memory, a processor, and a program for recommending items based on multimodal representations stored in the memory and executable on the processor. When the processor executes the program for recommending items based on multimodal representations, the following operation instructions are realized:

[0173] Obtain the historical interaction item text and historical interaction item image of the user, and perform alignment encoding processing on the historical interaction item text and historical interaction item image to obtain a text semantic encoding matrix and an image semantic encoding matrix. The same row and column elements of the text semantic encoding matrix and the image semantic encoding matrix correspond to the same item;

[0174] Generate an item encoding matrix for representing the interaction time and identification information of each item, and predict the preference degree of the user for each item according to the text semantic encoding matrix, the image semantic encoding matrix, and the item encoding matrix;

[0175] Obtain the target recommended item according to the preference degree of the user for each item.

[0176] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0177] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An item recommendation method based on multimodal representation, characterized in that: include: Acquire the user's historical interactive item text and historical interactive item image, and perform alignment coding processing on the historical interactive item text and the historical interactive item image to obtain a text semantic coding matrix and an image semantic coding matrix, wherein the same row and column elements of the text semantic coding matrix and the image semantic coding matrix correspond to the same item; Generate an item representation matrix for representing the interaction time and identification information of each item, and predict the user's preference for each item based on the text semantic encoding matrix, the image semantic encoding matrix and the item representation matrix; According to the user's preference for each of the items, a target recommended item is obtained; Perform alignment coding processing on the historical interactive item text and the historical interactive item image to obtain the text semantic coding matrix and the image semantic coding matrix, including: Applying a semantic discriminator to the historical interactive item text and the historical interactive item image to obtain the item semantics of the historical interactive item text and the item semantics of the historical interactive item image; Applying a text encoder to the historical interaction item texts marked with the item semantics to obtain a text semantic encoding matrix; Applying an image encoder to the historical interaction item images marked with the item semantics to obtain an image semantic encoding matrix; The text encoder and the image encoder are trained text encoders and trained image encoders, and the training method of the text encoder and the image encoder is joint training, wherein constructing the total loss function of the joint training includes: During the training process, a text semantic coding sample matrix output by the text encoder is obtained, and an image semantic coding sample matrix output by the image encoder is obtained; Dividing the sample text semantics contained in the text semantic encoding sample matrix into text sample pair groups, each of the text sample pair groups includes a first positive sample pair and a first negative sample pair, and the items corresponding to the first positive sample pair and the first negative sample pair include only one identical item; Obtaining a first loss function according to the first positive sample pair and the first negative sample pair; Semantically dividing each image sample contained in the image semantic coding sample matrix into image sample pair groups, each of the image sample pair groups includes a second positive sample pair and a second negative sample pair, and the objects corresponding to the second positive sample pair and the second negative sample pair include only one identical object; Obtaining a second loss function according to the second positive sample pair and the second negative sample pair; According to the text semantic coding sample matrix and the image semantic coding sample matrix, a text image sample group is obtained, each of the text image sample groups includes a text image positive sample pair and a text image negative sample pair, the sample text semantics and image sample semantics included in the text image positive sample pair of each group correspond to the same object, and the sample text semantics and image sample semantics included in the text image negative sample pair correspond to two different objects; Obtaining a third loss function according to the text image positive sample pair and the text image negative sample pair; A total loss function is constructed based on the first loss function, the second loss function, and the third loss function.

2. The method for recommending items based on multimodal representation according to claim 1, wherein: According to the first positive sample pair and the first negative sample pair, a first loss function is obtained, including: Determining the semantic similarity of two sample texts included in the first positive sample pair; Determining the semantic similarity between two sample texts included in the first negative sample pair; A first loss function is obtained according to the similarity corresponding to the first positive sample pair and the similarity corresponding to the first negative sample pair.

3. The method for recommending items based on multimodal representation according to claim 1, characterized in that: Predicting the user's preference for each of the items based on the text semantic coding matrix, the image semantic coding matrix, and the item representation matrix includes: According to the item representation matrix, a behavior preference sequence is obtained, where the behavior preference sequence is used to represent the user's preference order for each item; Aggregating the text semantic coding matrix and the image semantic coding matrix to obtain a content preference sequence; Aggregating the behavior preference sequence and the content preference sequence to obtain a mixed preference sequence; Obtaining a user final preference representation vector according to the behavior preference sequence, the content preference sequence, and the mixed preference sequence; The user's preference for each of the items is predicted based on the user's final preference representation vector.

4. The method for recommending items based on multimodal representation according to claim 3, characterized in that: Applying an item number modal sequence encoder to the item representation matrix to obtain a behavior preference sequence; applying a content modal sequence encoder to the text semantic encoding matrix and the image semantic encoding matrix to perform aggregation processing to obtain a content preference sequence; applying a mixed modal sequence decoder to the behavior preference sequence and the content preference sequence to perform aggregation processing to obtain a mixed preference sequence; After completing the training of the article number modal sequence encoder, the content modal sequence encoder and the mixed modal sequence decoder are jointly trained based on the trained article number modal sequence encoder.

5. The method for recommending items based on multimodal representation according to claim 4, characterized in that: The final loss function used for the joint training of the content modality sequence encoder and the mixed modality sequence decoder is composed of a contrast loss function, a cross entropy loss function and sample information, and the sample information is generated by a sample identifier, a sample text and a sample image.

6. An item recommendation device based on multimodal representation, characterized in that: The device comprises the following components: An alignment module is used to obtain the user's historical interactive item text and historical interactive item image, and perform alignment coding processing on the historical interactive item text and historical interactive item image to obtain a text semantic coding matrix and an image semantic coding matrix, wherein the same row and column elements of the text semantic coding matrix and the image semantic coding matrix correspond to the same item; A prediction module, used to generate an item representation matrix for representing the interaction time and identification information of each item, and predict the user's preference for each item based on the text semantic encoding matrix, the image semantic encoding matrix and the item representation matrix; A recommendation module, used to obtain target recommended items according to the user's preference for each of the items; Perform alignment coding processing on the historical interactive item text and the historical interactive item image to obtain the text semantic coding matrix and the image semantic coding matrix, including: Applying a semantic discriminator to the historical interactive item text and the historical interactive item image to obtain the item semantics of the historical interactive item text and the item semantics of the historical interactive item image; Applying a text encoder to the historical interaction item texts marked with the item semantics to obtain a text semantic encoding matrix; Applying an image encoder to the historical interaction item images marked with the item semantics to obtain an image semantic encoding matrix; The text encoder and the image encoder are trained text encoders and trained image encoders, and the training method of the text encoder and the image encoder is joint training, wherein constructing the total loss function of the joint training includes: During the training process, a text semantic coding sample matrix output by the text encoder is obtained, and an image semantic coding sample matrix output by the image encoder is obtained; Dividing the sample text semantics contained in the text semantic encoding sample matrix into text sample pair groups, each of the text sample pair groups includes a first positive sample pair and a first negative sample pair, and the items corresponding to the first positive sample pair and the first negative sample pair include only one identical item; Obtaining a first loss function according to the first positive sample pair and the first negative sample pair; Semantically dividing each image sample contained in the image semantic coding sample matrix into image sample pair groups, each of the image sample pair groups includes a second positive sample pair and a second negative sample pair, and the objects corresponding to the second positive sample pair and the second negative sample pair include only one identical object; Obtaining a second loss function according to the second positive sample pair and the second negative sample pair; According to the text semantic coding sample matrix and the image semantic coding sample matrix, a text image sample group is obtained, each of the text image sample groups includes a text image positive sample pair and a text image negative sample pair, the sample text semantics and image sample semantics included in the text image positive sample pair of each group correspond to the same object, and the sample text semantics and image sample semantics included in the text image negative sample pair correspond to two different objects; Obtaining a third loss function according to the text image positive sample pair and the text image negative sample pair; A total loss function is constructed based on the first loss function, the second loss function, and the third loss function.

7. A terminal device, characterized in that: The terminal device includes a memory, a processor, and an item recommendation program based on multimodal representation stored in the memory and executable on the processor. When the processor executes the item recommendation program based on multimodal representation, the steps of the item recommendation method based on multimodal representation as described in any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores an item recommendation program based on multimodal representation. When the item recommendation program based on multimodal representation is executed by a processor, the steps of the item recommendation method based on multimodal representation as described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Commodity feature processing method, electronic equipment and computer storage medium

    CN115952313A