Recommendation model training method and device, article recommendation method and device, equipment and storage medium
By mining the space-time constraint information in the interaction records between users and items and training the recommendation model, the problem of insufficient semantic capture of space-time description in the prior art is solved, and a more efficient and adaptable recommendation effect is achieved.
Patent Information
- Application Number
- CN202510408221.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
Existing recommendation techniques are unable to effectively capture the rich semantics in the original spatial and temporal description, resulting in poor recommendation results and inefficiency.
By obtaining the interaction records between sample users and items, mining the spatiotemporal constraint information for direct and indirect interaction, the recommended model is trained to integrate the spatiotemporal information for direct and indirect interaction.
It significantly improves the accuracy, timeliness and scenario adaptability of recommendations, and improves recommendation efficiency.
Smart Images

Figure CN120336850A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and particularly to a method for training a recommendation model, an item recommendation method, an apparatus, a device, and a storage medium. Background Art
[0002] With the development of computer technologies, users can browse various kinds of information through computer devices. Correspondingly, various merchants or service platforms can recommend items to users to meet their needs. Therefore, how to make the recommended items meet the needs of users is a research direction.
[0003] Currently, it is usually adopted to analyze users' historical behaviors (such as clicks, purchases, comments, etc.) and spatio-temporal context information (such as geographical location, timestamp, etc.), and simplify the spatio-temporal information into predefined entities (such as "holidays", "working days", "area ID", etc.) to provide personalized recommendations for users' surrounding services.
[0004] However, the predefined entities cannot capture the rich semantics in the original spatio-temporal description, such as the cultural activity preferences implied by "holidays", the social scene requirements associated with "coffee shops", etc., resulting in poor recommendation effects and low recommendation efficiency. Summary of the Invention
[0005] The present disclosure provides a method for training a recommendation model, an item recommendation method, an apparatus, a device, and a storage medium. This solution enables the spatio-temporal constraint-based recommendation model to fuse the directly interactive spatio-temporal information and the indirectly interactive spatio-temporal information, significantly improving the accuracy, timeliness, and scenario adaptability of recommendations and enhancing the recommendation efficiency.
[0006] According to one aspect of the embodiments of the present disclosure, there is provided a method for training a recommendation model, the method including:
[0007] Determining first spatio-temporal constraint information based on sample interaction information, where the sample interaction information includes interaction records between a sample user and a sample item, and the first spatio-temporal constraint information is used to indicate the spatio-temporal connection when the sample user directly interacts with the sample item;
[0008] Determining second spatio-temporal constraint information based on the sample interaction information and the first spatio-temporal constraint information, where the second spatio-temporal constraint information is used to indicate the spatio-temporal connection when the sample user indirectly interacts with the sample item;
[0009] Training a recommendation model based on the first spatio-temporal constraint information and the second spatio-temporal constraint information.
[0010] According to another aspect of the embodiments of the present disclosure, there is provided an item recommendation method, the method including:
[0011] Based on the historical behavior information of the target user, first spatio-temporal constraint information is determined, and the first spatio-temporal constraint information is used to indicate the spatio-temporal connection when the target user directly interacts with each item;
[0012] Based on the historical behavior information and the first spatio-temporal constraint information, second spatio-temporal constraint information is determined, and the second spatio-temporal constraint information is used to indicate the spatio-temporal connection when the target user indirectly interacts with each item;
[0013] Based on a recommendation model, the first spatio-temporal constraint information and the second spatio-temporal constraint information are processed to obtain at least one target item. The recommendation model is trained based on the above training method of the recommendation model, and the target item is an item to be recommended to the target user.
[0014] According to another aspect of the embodiments of the present disclosure, a training device for a recommendation model is provided. The device includes:
[0015] A first determination unit configured to determine first spatio-temporal constraint information based on sample interaction information. The sample interaction information includes interaction records between a sample user and a sample item, and the first spatio-temporal constraint information is used to indicate the spatio-temporal connection when the sample user directly interacts with the sample item;
[0016] A second determination unit configured to determine second spatio-temporal constraint information based on the sample interaction information and the first spatio-temporal constraint information. The second spatio-temporal constraint information is used to indicate the spatio-temporal connection when the sample user indirectly interacts with the sample item;
[0017] A training unit configured to train a recommendation model based on the first spatio-temporal constraint information and the second spatio-temporal constraint information.
[0018] In some embodiments, the first determination unit is configured to determine user direct spatio-temporal information and item direct spatio-temporal information based on the sample interaction information; extract the spatio-temporal constraints of the user from the user direct spatio-temporal information by using a first large language model to obtain first direct spatio-temporal constraint information; extract the spatio-temporal constraints of the item from the item direct spatio-temporal information by using the first large language model to obtain second direct spatio-temporal constraint information; and determine the first spatio-temporal constraint information based on the first direct spatio-temporal constraint information and the second direct spatio-temporal constraint information.
[0019] In some embodiments, the first determination unit is configured to encode the first direct spatio-temporal constraint information through an alignment network to obtain a first user spatio-temporal embedding; encode the second direct spatio-temporal constraint information through the alignment network to obtain a first item spatio-temporal embedding; splice the first user spatio-temporal embedding and the first item spatio-temporal embedding to obtain the first spatio-temporal constraint information.
[0020] In some embodiments, the training unit is further configured to, for any sample user, determine a plurality of positive sample users and a plurality of negative sample users corresponding to the sample user based on the sample interaction information, where the similarity between the positive sample user and the sample user is greater than a similarity threshold; for any sample item, determine a plurality of positive sample items and a plurality of negative sample items corresponding to the sample item based on the sample interaction information, where the similarity between the positive sample item and the sample item is greater than a similarity threshold; determine a contrast loss based on the plurality of positive sample users, the plurality of negative sample users, the plurality of positive sample items, and the plurality of negative sample items; and train the alignment network based on the contrast loss.
[0021] In some embodiments, the training unit is further configured to decode the first user spatio-temporal embedding through the alignment network to obtain a second user spatio-temporal embedding; decode the first item spatio-temporal embedding through the alignment network to obtain a second item spatio-temporal embedding; determine a reconstruction loss based on the user direct spatio-temporal information, the item direct spatio-temporal information, the second user spatio-temporal embedding, and the second item spatio-temporal embedding; and train the alignment network based on the reconstruction loss.
[0022] In some embodiments, the first determination unit is further configured to structurally process the user direct spatio-temporal information based on a first prompt template; and structurally process the item direct spatio-temporal information based on a second prompt template.
[0023] In some embodiments, the second determination unit is configured to determine user indirect spatio-temporal information and item indirect spatio-temporal information based on the sample interaction information; extract the spatio-temporal constraints of the user from the user indirect spatio-temporal information through a second large language model to obtain first indirect spatio-temporal constraint information; extract the spatio-temporal constraints of the item from the item indirect spatio-temporal information through the second large language model to obtain second indirect spatio-temporal constraint information; and determine the second spatio-temporal constraint information based on the first indirect spatio-temporal constraint information, the second indirect spatio-temporal constraint information, and the first spatio-temporal constraint information.
[0024] In some embodiments, the second determination unit is further configured to perform structured processing on the user's indirect spatio-temporal information based on a third prompt template; and perform structured processing on the item's indirect spatio-temporal information based on a fourth prompt template.
[0025] In some embodiments, the training unit is configured to: based on the sample interaction information, user embedding, and item embedding; concatenate the user embedding, the item embedding, the first spatio-temporal constraint information, and the second spatio-temporal constraint information to obtain an input embedding; input the input embedding into the recommendation model to obtain recommendation result information; and train the recommendation model based on the recommendation result information and the sample interaction information.
[0026] According to another aspect of the embodiments of the present disclosure, there is provided an item recommendation device, which includes:
[0027] A first determination unit, configured to determine first spatio-temporal constraint information based on the historical behavior information of a target user, where the first spatio-temporal constraint information is used to indicate the spatio-temporal connection when the target user directly interacts with each item;
[0028] A second determination unit, configured to determine second spatio-temporal constraint information based on the historical behavior information and the first spatio-temporal constraint information, where the second spatio-temporal constraint information is used to indicate the spatio-temporal connection when the target user indirectly interacts with each item;
[0029] A recommendation unit, configured to process the first spatio-temporal constraint information and the second spatio-temporal constraint information based on a recommendation model to obtain at least one target item, where the recommendation model is trained based on the above-mentioned training method of the recommendation model, and the target item is an item to be recommended to the target user.
[0030] According to another aspect of the embodiments of the present disclosure, there is provided an electronic device, which includes:
[0031] One or more processors;
[0032] A memory for storing program code executable by the processor;
[0033] Wherein, the processor is configured to execute the program code to implement the above-mentioned training method of the recommendation model, or implement the above-mentioned item recommendation method.
[0034] According to another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by the processor of the electronic device, enabling the electronic device to execute the above-mentioned training method of the recommendation model, or implement the above-mentioned item recommendation method.
[0035] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product, which includes a computer program that, when executed by a processor, implements the above-mentioned training method of the recommendation model or implements the above-mentioned item recommendation method.
[0036] The embodiments of the present disclosure provide a training solution for a recommendation model. By obtaining the interaction records between sample users and sample items, the spatio-temporal constraint information of direct and indirect interactions is mined to train the recommendation model, so that the recommendation model based on spatio-temporal constraints can fuse the spatio-temporal information of direct interactions and the spatio-temporal information of indirect interactions, significantly improving the accuracy, timeliness and scenario adaptability of recommendations and improving the recommendation efficiency.
[0037] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation to the present disclosure.
[0039] Figure 1 is a schematic diagram of an implementation environment of a method for training a recommendation model shown according to an exemplary embodiment.
[0040] Figure 2 is a flowchart of a method for training a recommendation model shown according to an exemplary embodiment.
[0041] Figure 3 is a flowchart of another method for training a recommendation model shown according to an exemplary embodiment.
[0042] Figure 4 is a schematic diagram of another method for training a recommendation model provided according to an exemplary embodiment.
[0043] Figure 5 is a flowchart of an item recommendation method shown according to an exemplary embodiment.
[0044] Figure 6 is a block diagram of a training device for a recommendation model shown according to an exemplary embodiment.
[0045] Figure 7 is a block diagram of an item recommendation device shown according to an exemplary embodiment.
[0046] Figure 8 is a block diagram of an electronic device shown according to an exemplary embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] To enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0048] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0049] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present disclosure are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions. For example, the sample interaction information involved in the present disclosure is obtained under full authorization.
[0050] Figure 1 It is a schematic diagram of the implementation environment of a training method for a recommendation model shown according to an exemplary embodiment. Refer to Figure 1 , and the implementation environment specifically includes: an electronic device 101 and a server 102. The electronic device 101 can be connected to the server 102 through a wireless network or a wired network.
[0051] The electronic device 101 can be at least one of devices such as a smart phone, a smart watch, a desktop computer, a laptop computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, and a laptop portable computer.
[0052] The electronic device 101 can generally refer to one of multiple electronic devices, and the electronic device 101 is used as an example in this embodiment. Those skilled in the art can know that the number of the above-mentioned electronic devices can be more or less. For example, the above-mentioned electronic devices can be several, or the above-mentioned electronic devices can be dozens or hundreds, or more, and the present disclosure embodiments do not limit the number and type of the electronic devices.
[0053] The server 102 is at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Optionally, the number of the above-mentioned servers may be more or less, and the embodiments of the present disclosure do not limit this. Of course, the server 102 may further include other functional servers to provide more comprehensive and diverse services. In some embodiments, the server 102 undertakes the main computing work, and the electronic device 101 undertakes the secondary computing work; or, the server 102 undertakes the secondary computing work, and the electronic device 101 undertakes the main computing work; or, the server 102 and the electronic device 101 adopt a distributed computing architecture for collaborative computing. The server 102 can be connected to the electronic device 101 and other electronic devices through a wireless network or a wired network. Optionally, the number of the above-mentioned servers may be more or less, and the embodiments of the present disclosure do not limit this.
[0054] Figure 2 It is a flowchart of a method for training a recommendation model shown according to an exemplary embodiment. As Figure 2 shown, this method is executed by an electronic device and includes the following steps:
[0055] In step S201, based on the sample interaction information, the first spatio-temporal constraint information is determined.
[0056] In the embodiments of the present disclosure, the sample interaction information is a data set, and this sample interaction information records the interaction situation between sample users and sample items. Correspondingly, the sample interaction information includes the interaction records between sample users and sample items. Among them, the sample users are a part of the user group selected for research and analysis on the premise of compliance, and the embodiments of the present disclosure do not limit the selection method of the sample users. The sample items are the items involved by the above-mentioned sample users. The interaction records may include behaviors such as clicks, purchases, views, collections, and evaluations of sample items by sample users.
[0057] For example, in an e-commerce scenario, the sample users can be consumers, and the sample items are commodities. Or, in a video consumption scenario, the sample users can be viewers, and the sample items are various videos.
[0058] By analyzing and processing the interaction records between sample users and sample items, relevant information about time and controls can be mined. Correspondingly, the first spatio-temporal constraint information is used to indicate the spatio-temporal connection when sample users and sample items have direct interactions. Optionally, the spatio-temporal connection can also be referred to as spatio-temporal constraint.
[0059] Among them, direct interaction refers to the direct interaction behavior between the sample user and the sample item without passing through other intermediate links. Direct interaction corresponds to indirect interaction. For example, when user A purchases item a, this purchase behavior is a direct interaction between user A and item a. And when user B purchases item a and also purchases item b, through the relationship chain of user A → item a → user B → item b, it can be seen that there is an indirect interaction between user A and item b. Or, the user sees a product link or recommendation information shared by a friend on a social media platform and then clicks on the link to enter an e-commerce platform to purchase the product. Here, the social media platform and the friend's recommendation constitute intermediate links, and the corresponding user and product belong to indirect interaction.
[0060] Spatio-temporal connection contains associated information in two dimensions of time and space. Among them, in terms of time, the spatio-temporal connection indicates at what specific time point or time period the user interacted with the item; in terms of space, the spatio-temporal connection indicates the specific location where the interaction behavior occurred (such as a shopping mall, a café, etc.).
[0061] In step S202, based on the sample interaction information and the first spatio-temporal constraint information, the second spatio-temporal constraint information is determined.
[0062] In the embodiments of the present disclosure, determining the second spatio-temporal constraint information based on the first spatio-temporal constraint information is a process of deeply analyzing and mining the interaction behavior between the sample user and the sample item. That is, from the behavioral data of indirect interaction, the spatio-temporal connection between the sample user and the sample item during the indirect interaction process is analyzed. Correspondingly, the second spatio-temporal constraint information is used to indicate the spatio-temporal connection when the sample user and the sample item conduct indirect interaction.
[0063] Optionally, indirect interaction can also be referred to as high-order interaction.
[0064] For example, the interaction record between the sample user and the sample item is marked as a user-item bipartite graph. The interaction between the first-order adjacent sample user and the sample item is direct interaction, and the interaction between the second-order adjacent and multi-order adjacent sample user and the sample item is high-order interaction.
[0065] In step S203, based on the first spatio-temporal constraint information and the second spatio-temporal constraint information, a recommendation model is trained.
[0066] In the embodiments of the present disclosure, the first spatio-temporal constraint information models the spatio-temporal constraints in the direct interaction data, and the second spatio-temporal constraint information models the spatio-temporal constraints in the high-order interaction data. Training the recommendation model based on the first spatio-temporal constraint information and the second spatio-temporal constraint information can incorporate comprehensive spatio-temporal dependence relationships into the recommendation model and enhance the ability of the recommendation model to capture the dynamic preferences of users.
[0067] The embodiments of the present disclosure provide a training solution for a recommendation model. By obtaining the interaction records between sample users and sample items, the spatio-temporal constraint information of direct and indirect interactions is mined to train the recommendation model, enabling the spatio-temporal constraint-based recommendation model to fuse the spatio-temporal information of direct interactions and the spatio-temporal information of indirect interactions, significantly improving the quality, timeliness, and scenario adaptability of recommendations, and enhancing the recommendation efficiency.
[0068] In some embodiments, determining the first spatio-temporal constraint information based on the sample interaction information includes:
[0069] Determining the user's direct spatio-temporal information and the item's direct spatio-temporal information based on the sample interaction information;
[0070] Extracting the spatio-temporal constraints of the user from the user's direct spatio-temporal information based on the first large language model to obtain the first direct spatio-temporal constraint information;
[0071] Extracting the spatio-temporal constraints of the item from the item's direct spatio-temporal information based on the first large language model to obtain the second direct spatio-temporal constraint information;
[0072] Determining the first spatio-temporal constraint information based on the first direct spatio-temporal constraint information and the second direct spatio-temporal constraint information.
[0073] In the embodiments of the present disclosure, by determining the user's direct spatio-temporal information and the item's direct spatio-temporal information based on the sample interaction information, and then using the first large language model to accurately extract the spatio-temporal constraints of the user and the item from both, obtaining the first and second direct spatio-temporal constraint information, and further determining the first spatio-temporal constraint information, this method can fully mine the latent spatio-temporal correlations in the sample interaction information, accurately define the spatio-temporal constraints of the user and the item with the powerful extraction ability of the large language model, and improve the accuracy and reliability of the task.
[0074] In some embodiments, determining the first spatio-temporal constraint information based on the first direct spatio-temporal constraint information and the second direct spatio-temporal constraint information includes:
[0075] Encoding the first direct spatio-temporal constraint information through an alignment network to obtain the first user spatio-temporal embedding;
[0076] Encoding the second direct spatio-temporal constraint information through an alignment network to obtain the first item spatio-temporal embedding;
[0077] Concatenating the first user spatio-temporal embedding and the first item spatio-temporal embedding to obtain the first spatio-temporal constraint information.
[0078] In the embodiments of the present disclosure, by using an alignment network to encode the spatio-temporal constraint information extracted from the direct spatio-temporal information of users and items respectively, spatio-temporal embeddings of users and items are obtained, and then the first spatio-temporal constraint information is determined by concatenation. This method uses the encoding ability of the alignment network to transform spatio-temporal constraint information from different sources into a fusible embedding representation, realizing the effective integration of the spatio-temporal information of users and items, and improving the processing accuracy and efficiency of related tasks in the spatio-temporal dimension.
[0079] In some embodiments, the training steps of the alignment network include:
[0080] For any sample user, based on the sample interaction information, a plurality of positive sample users and a plurality of negative sample users corresponding to the sample user are determined, and the similarity between the positive sample users and the sample user is greater than the similarity threshold;
[0081] For any sample item, based on the sample interaction information, a plurality of positive sample items and a plurality of negative sample items corresponding to the sample item are determined, and the similarity between the positive sample items and the sample item is greater than the similarity threshold;
[0082] Based on the plurality of positive sample users, the plurality of negative sample users, the plurality of positive sample items, and the plurality of negative sample items, a contrastive loss is determined;
[0083] Based on the contrastive loss, the alignment network is trained.
[0084] In the embodiments of the present disclosure, by respectively determining a plurality of positive and negative sample users corresponding to the sample user (the similarity between the positive sample users and the sample user exceeds the threshold) and a plurality of positive and negative sample items corresponding to the sample item (the similarity between the positive sample items and the sample item exceeds the threshold) based on the sample interaction information, and then using these samples to determine the contrastive loss and training the alignment network based on this loss, the alignment network can learn the similarity and difference relationships between users and between items, optimize the network parameters, and improve the model's ability to represent the features of users and items and the alignment effect, so as to better apply to tasks such as recommendation and improve the accuracy and effectiveness of recommendation.
[0085] In some embodiments, the method further includes:
[0086] The alignment network decodes the first user spatio-temporal embedding to obtain a second user spatio-temporal embedding;
[0087] The alignment network decodes the first item spatio-temporal embedding to obtain a second item spatio-temporal embedding;
[0088] Based on the user direct spatio-temporal information, the item direct spatio-temporal information, the second user spatio-temporal embedding, and the second item spatio-temporal embedding, a reconstruction loss is determined;
[0089] Based on the reconstruction loss, the alignment network is trained.
[0090] In the embodiments of the present disclosure, the alignment network decodes the spatio-temporal embeddings of the first user and the first item to obtain the spatio-temporal embeddings of the second user and the second item, and determines the reconstruction loss by combining the direct spatio-temporal information of the user and the item, and trains the alignment network, enabling the network to learn effective decoding and reconstruction patterns of the spatio-temporal embeddings, optimizing the network parameters to enhance the representation and reconstruction capabilities of the spatio-temporal features of the user and the item, and enabling the alignment network to retain more semantic information in the input embeddings.
[0091] In some embodiments, the method further includes:
[0092] Structuring the direct spatio-temporal information of the user based on the first prompt template;
[0093] Structuring the direct spatio-temporal information of the item based on the second prompt template.
[0094] In the embodiments of the present disclosure, by respectively structuring the direct spatio-temporal information of the user and the item based on different first and second prompt templates, the two types of spatio-temporal information are effectively converted into a format compatible with the large language model (LLM), which is beneficial to improving the processing efficiency and adaptability of the information in the large language model, facilitating subsequent operations such as relevant analysis, application, or model training based on this information, and improving the data processing efficiency.
[0095] In some embodiments, determining the second spatio-temporal constraint information based on the sample interaction information and the first spatio-temporal constraint information includes:
[0096] Determining the indirect spatio-temporal information of the user and the indirect spatio-temporal information of the item based on the sample interaction information;
[0097] Extracting the spatio-temporal constraints of the user from the indirect spatio-temporal information of the user based on the second large language model to obtain the first indirect spatio-temporal constraint information;
[0098] Extracting the spatio-temporal constraints of the item from the indirect spatio-temporal information of the item based on the second large language model to obtain the second indirect spatio-temporal constraint information;
[0099] Determining the second spatio-temporal constraint information based on the first indirect spatio-temporal constraint information, the second indirect spatio-temporal constraint information, and the first spatio-temporal constraint information.
[0100] In the embodiments of the present disclosure, by first determining the indirect spatio-temporal information of the user and the item based on the sample interaction information, then using the second large language model to extract spatio-temporal constraints from both respectively to obtain the first and second indirect spatio-temporal constraint information, and finally combining this information with the first spatio-temporal constraint information to determine the second spatio-temporal constraint information, it is possible to comprehensively and deeply explore the indirect spatio-temporal correlation between the user and the item in the sample interaction information, accurately extract spatio-temporal constraints using the large language model, integrate spatio-temporal information from multiple aspects, and improve the comprehensiveness, accuracy, and reliability of task processing.
[0101] In some embodiments, the method further includes:
[0102] Structurally process the user's indirect spatio-temporal information based on the third prompt template;
[0103] Structurally process the item's indirect spatio-temporal information based on the fourth prompt template.
[0104] In the embodiments of the present disclosure, by structurally processing the user's indirect spatio-temporal information and the item's indirect spatio-temporal information based on the third prompt template and the fourth prompt template respectively, it is possible to transform the originally scattered and complex indirect spatio-temporal information into a regular, orderly, and easy-to-understand and analyze structural form, improving the accuracy and efficiency of the entire process in processing indirect spatio-temporal information.
[0105] In some embodiments, training a recommendation model based on the first spatio-temporal constraint information and the second spatio-temporal constraint information includes:
[0106] Based on the sample interaction information, user embedding and item embedding;
[0107] Concatenate the user embedding, item embedding, first spatio-temporal constraint information, and second spatio-temporal constraint information to obtain an input embedding;
[0108] Input the input embedding into the recommendation model to obtain recommendation result information;
[0109] Train the recommendation model based on the recommendation result information and the sample interaction information.
[0110] In the embodiments of the present disclosure, by first obtaining the user embedding and item embedding based on the sample interaction information, then concatenating them with the first and second spatio-temporal constraint information to form an input embedding, inputting it into the recommendation model to obtain the recommendation result information, and then training the recommendation model based on the recommendation result information and the sample interaction information, it is possible to fully integrate the basic embedding information of the user and the item as well as rich spatio-temporal constraint information, enabling the recommendation model to learn more comprehensive features during training, thereby improving the performance of the recommendation model, making its recommendation results more in line with the actual needs of users, and improving the accuracy and quality of the recommendation.
[0111] The above Figure 2The following is a flowchart of a method for training a recommendation model according to the present disclosure. The training solution of the recommendation model provided by the present disclosure will be further elaborated below. Figure 3 is a flowchart of another method for training a recommendation model shown according to an exemplary embodiment. Refer to Figure 3 This method is executed by an electronic device and includes the following steps:
[0112] In step S301, based on the sample interaction information, the user's direct spatio-temporal information and the item's direct spatio-temporal information are determined.
[0113] In the embodiments of the present disclosure, the sample interaction information includes the interaction records between the sample user and the sample item. Optionally, there is a user set U and an item set I. The user set U includes the sample user, and the item set I includes the sample item. The implicit interaction matrix between the sample user and the sample item is represented by R. For the interaction matrix R, there is corresponding spatio-temporal information, denoted as ST. In other words, each interaction behavior between the sample object and the sample item has a corresponding spatio-temporal description information, and the content of the spatio-temporal description information is the spatio-temporal text description. That is, for each interaction behavior r ui ∈R, there is a corresponding spatio-temporal text description st ui ∈ST.
[0114] For example, the spatio-temporal description information of the interaction behavior r ui is "20xx-10-01 08:01:00" and "a coffee shop located 3 kilometers away". According to the interaction behavior of the sample user and the spatio-temporal text description, it is possible to predict the sample item that the sample user is most likely to interact with in a specific spatio-temporal environment.
[0115] It should be noted that different from the existing solution that uniformly converts the spatio-temporal description information into predefined labels with fixed interpretations, the embodiments of the present disclosure dynamically generate spatio-temporal constraints through the semantic understanding and reasoning capabilities of the LLM (Large Language Model). These spatio-temporal constraints are generated based on the interaction records between different sample users and sample items, and can adapt and evolve with the changes in the behavior environment. These spatio-temporal constraints can then be used to simulate the influence of the spatio-temporal environment on users and items. In the field of recommendation, spatio-temporal constraints refer to the restrictions and influences on the recommendation results considering time and space factors.
[0116] Optionally, the spatio-temporal constraint is represented as P(ST|R). Correspondingly, the formula for spatio-temporal constraint modeling is:
[0117] P(ST|R) = ∑ u∈U,i∈I P(ST|R u ) + P(ST|R i )
[0118] Among them, R u represents the interaction data of the sample user u, while R i is the interaction data of the sample item i. The interaction data of the sample user u includes the interaction records of the sample user u with multiple sample items.
[0119] It should also be noted that the above is the process of processing the interaction records of the direct interaction between the sample user and the sample item. In order to capture richer interaction relationships between the sample user and the sample item and further improve the accuracy of spatio-temporal constraint modeling, the embodiment of the present disclosure represents the interaction R as a user-item bipartite graph G, and incorporates higher-order adjacent interaction behaviors into the bipartite graph G.
[0120] Correspondingly, the formula for the above spatio-temporal constraint modeling can be updated as:
[0121] P(ST|R) = ∑ u∈U,i∈I P(ST|R u ) + P(ST|R i ) + P(ST|G u ) + P(ST|G i )
[0122] Among them, G u represents the subgraph centered on the user u, sampled from the bipartite graph G. G i represents the subgraph centered on the item i, also sampled from the graph G.
[0123] The steps for determining the spatio-temporal constraints will be described in detail below.
[0124] First, the extraction of spatio-temporal constraints from the direct interaction data will be described. The direct interaction data is the interaction record of the user or the item, such as R u or R i . For each interaction behavior included in the direct interaction data, the corresponding spatio-temporal description information can be obtained, denoted as or In this way, two types of spatio-temporal data based on the direct interaction data can be constructed: and For the convenience of description, is called the user direct spatio-temporal information, and is called the item direct spatio-temporal information. Among them, Among them, |R i | represents the total number of sample users, |ST u | represents the total number of spatio-temporal description information related to the sample users, |R i | represents the total number of sample items, |ST i|Represents the total number of spatio-temporal description information related to the sample item. Embodiments of the present disclosure can utilize this spatio-temporal description information to capture the spatio-temporal preferences of the sample user u and the spatio-temporal constraints of the sample item i. Among them, the spatio-temporal preferences of the sample user u reflect the behavioral changes of the sample user under different spatio-temporal conditions, while the spatio-temporal constraints of the sample item i indicate the spatio-temporal environment in which the sample item is more likely to be interacted with.
[0125] Embodiments of the present disclosure leverage the advantages of LLM in text understanding and reasoning, and use LLM to extract spatio-temporal constraints from spatio-temporal description information. Referring to the following steps S302 to S306, the steps of extracting spatio-temporal constraints from direct interaction data are introduced in detail.
[0126] It should be noted that due to spatio-temporal description information, such as user direct spatio-temporal information is essentially unstructured data, while LLM requires structured text sequences as input. Therefore, embodiments of the present disclosure design two prompt templates to convert the above user direct spatio-temporal information and item direct spatio-temporal information into a format compatible with LLM. Correspondingly, based on the first prompt template, the user direct spatio-temporal information is structured; based on the second prompt template, the item direct spatio-temporal information is structured. By respectively structuring the user direct spatio-temporal information and the item direct spatio-temporal information based on different first and second prompt templates, these two types of spatio-temporal information can be effectively converted into a format compatible with the large language model (LLM), which is beneficial to improving the processing efficiency and adaptability of information in the large language model, facilitating subsequent related analysis, applications, or model training operations based on this information, and improving the data processing efficiency.
[0127] Optionally, for ease of description, the prompt template for the user direct spatio-temporal information is called the first prompt template and is denoted as Then, the structuring process of the user direct spatio-temporal information can be expressed as Optionally, the first prompt template can be: "In the spatio-temporal recommendation scenario, please summarize the preferences of the user in terms of time, space, spatio-temporal, and overall preferences. Among them, the time preference reflects the user's behavioral pattern over time. The space preference reflects the fine-grained spatial changes captured based on the user's behavior. The spatio-temporal preference combines the time and space aspects, revealing the user's behavioral preference under the combined pattern. Finally, the overall preference combines the above three factors, providing a comprehensive representation of the user's recommendation tendency. The above time preference should be inferred from the observed user behavior, and the space preference should consider the fine-grained spatial differences before incorporating the overall preference."
[0128] Optionally, the user direct spatio-temporal information can be expressed as: "The time and space information of user u is The user and item i 0Interaction\n; …\n; …\n; The time and space information of user u is User and item Interaction.” Among them, the time and space information is the spatio-temporal description information, and \n represents the line break character.
[0129] Similarly, for the sake of description, the direct spatio-temporal information of the item The hint template is called the second hint template, denoted as Then, the structuring process of the direct spatio-temporal information of the item can be expressed as Optionally, the second hint template can be: "In the spatio-temporal recommendation scenario, please summarize the preferences of the item in terms of time, space, spatio-temporal, and overall preferences. Among them, the time preference reflects the interaction pattern of the item over time. The space preference reflects the capture of fine-grained spatial changes based on the item interaction data. The spatio-temporal preference combines the time and space aspects to reveal the behavioral preferences of the item under the joint pattern. Finally, the overall preference combines the above three factors to provide a comprehensive representation of the item recommendation tendency. The above time preference should be inferred from the observed item interaction data, and the space preference should consider the fine-grained spatial differences before incorporating the overall preference."
[0130] Similarly, the direct spatio-temporal information of the item can be expressed as: "The time and space information of item i is Item and user u 0 ∈R u Interaction\n; …\n; …\n; The time and space information of item i is Item and user Interaction.” Among them, the time and space information is the spatio-temporal description information, and \n represents the line break character.
[0131] In step S302, based on the first large language model, extract the spatio-temporal constraints of the user from the user's direct spatio-temporal information to obtain the first direct spatio-temporal constraint information.
[0132] In the embodiments of the present disclosure, input the user's direct spatio-temporal information into the LLM through the hint template, and the LLM outputs the preferences of the user regarding spatio-temporal factors. For the sake of description, it is called the first direct spatio-temporal constraint information, denoted as S u . In other words, given the user's direct spatio-temporal information Use the first hint template Generate the corresponding text description Then input it into the LLM to obtain the first direct spatio-temporal constraint information This first direct spatio-temporal constraint information S u Represents the spatio-temporal constraints under which user u is more likely to perform interaction behaviors.
[0133] In step S303, based on the first large language model, the spatio-temporal constraints of the item are extracted from the direct spatio-temporal information of the item, and the second direct spatio-temporal constraint information is obtained.
[0134] In an embodiment of the present disclosure, similarly to the above step S302, the direct spatio-temporal information of the item is input into the LLM through a prompt template, and the LLM outputs the spatio-temporal factors specific to the item, which is called the second direct spatio-temporal constraint information for ease of description, denoted as S. i . In other words, given the direct spatio-temporal information of the item Using the second prompt template Generate the corresponding text description Then input it into the LLM to obtain the second direct spatio-temporal constraint information This second direct spatio-temporal constraint information S i Represents the spatio-temporal constraints under which item i is more likely to be interacted with.
[0135] It should be noted that through the above steps S302 and S303, the spatio-temporal constraints of all users U and items I can be obtained through the LLM, denoted as S. Among them, the representation S of this spatio-temporal constraint directly models the spatio-temporal description information in the direct interaction data. In particular, through the above ∑ u∈U,i∈I P(ST|R u ) + P(ST|R i ) process, it focuses on modeling the spatio-temporal constraints of all users and items. That is, the spatio-temporal constraints from the direct interaction data.
[0136] It should be noted that after aligning the embedding representation of S u , the embedding representation of S i and the embedding representation of the collaborative signal, these aligned embedding representations will be incorporated into the recommendation model to consider the spatio-temporal constraints. However, the dimension of the spatio-temporal constraint embedding (i.e., the embedding representation of the above S u , the embedding representation of S i and the embedding of the collaborative signal) (for example, 4096) is much larger than the embedding dimension in a typical recommendation model (for example, both the user and item dimensions are 64). To address this difference, an embodiment of the present disclosure introduces a dimensionality reduction process during the alignment process so that the spatio-temporal constraint embedding after dimensionality reduction can adapt to the input of a general recommendation model, thereby ensuring dimension compatibility. See the following steps S304 and step S305.
[0137] Among them, in the field of recommendation, collaborative signals refer to various types of information that can reflect the relationships between users and items, between users and users, and between items and items. Collaborative signals can include users' historical behaviors, such as a series of data generated by users during the interaction with various items in the past, such as browsing records, click preferences, purchase behaviors, evaluation feedback, etc. These behavioral data can reveal characteristics such as users' interests, consumption habits, and demand tendencies.
[0138] In step S304, the first direct spatio-temporal constraint information is encoded through an alignment network to obtain the first user spatio-temporal embedding.
[0139] In the embodiments of the present disclosure, the alignment network includes an encoder network. The input of this encoder network is an embedding vector of dimension D1. The encoder network undergoes step-by-step transformation through multiple linear layers and finally outputs an embedding vector of dimension D2, where D2 << D1. The encoder network extracts features and compresses the dimension of the input through continuous linear transformations and activation functions, and can obtain a low-dimensional embedding vector suitable for general recommendation models.
[0140] Optionally, encoding the above first direct spatio-temporal constraint information S u to obtain the first user spatio-temporal embedding S' u can be expressed as: S' u = Encoder(S u ).
[0141] In step S305, the second direct spatio-temporal constraint information is encoded through the alignment network to obtain the first item spatio-temporal embedding.
[0142] In the embodiments of the present disclosure, similar to the above step S304, encoding the above second direct spatio-temporal constraint information S i to obtain the first item spatio-temporal embedding S' i can be expressed as S' i = Encoder(S i ).
[0143] In step S306, the first user spatio-temporal embedding and the first item spatio-temporal embedding are concatenated to obtain the first spatio-temporal constraint information.
[0144] In the embodiments of the present disclosure, by concatenating the above first user spatio-temporal embedding S' u and the first item spatio-temporal embedding S' i , and aligning with the collaborative signal, the first spatio-temporal constraint information S' can be obtained.
[0145] The first spatio-temporal constraint information S' can be expressed as:
[0146] It should be noted that the above - mentioned dimension - compression process may lead to the loss of the discriminability and robustness of the original embeddings (i.e., the first direct spatio - temporal constraint information and the second direct spatio - temporal constraint information). To alleviate this problem, the information in the original embeddings is retained by reconstructing the encoded embeddings obtained by dimensionality reduction, that is, the spatio - temporal constraint information of S y and S i is retained, rather than simply aligning them with the collaborative signals. Therefore, the design of the alignment network focuses on two main objectives: 1) Adapt the spatio - temporal constraint embeddings to the general recommendation model through dimensionality reduction and alignment with the collaborative signals; secondly, reconstruct the encoded embeddings obtained by dimensionality reduction to retain their original semantic content after dimensionality reduction.
[0147] The training steps of the alignment network are introduced below.
[0148] In the embodiments of the present disclosure, the purpose of training the alignment network is to optimize the alignment network. The optimization objectives are: 1) Align the first user spatio - temporal embedding and the first item spatio - temporal embedding obtained by encoding with the embedding of the collaborative signal; 2) Retain more semantic information from the input embedding vectors. To achieve the above two objectives, the embodiments of the present disclosure design two loss functions to guide the optimization process. First, the contrastive loss is used to align the first user spatio - temporal embedding and the first item spatio - temporal embedding obtained by encoding with the embedding of the collaborative signal, and at the same time, the alignment strength is used to ensure effective integration. Secondly, the reconstruction loss is used to minimize the difference between the reconstructed embedding and the original embedding. The first user spatio - temporal embedding and the first item spatio - temporal embedding can be collectively referred to as the encoded embedding. The above method reconstructs the original embedding through the encoded embedding obtained by dimensionality reduction, enabling the alignment network to retain more semantic information in the input embedding.
[0149] Correspondingly, the training steps of the alignment network include: for any sample user, based on the sample interaction information, determine multiple positive sample users and multiple negative sample users corresponding to the sample user, where the similarity between the positive sample user and the sample user is greater than the similarity threshold. For any sample item, based on the sample interaction information, determine multiple positive sample items and multiple negative sample items corresponding to the sample item, where the similarity between the positive sample item and the sample item is greater than the similarity threshold. Based on the multiple positive sample users, multiple negative sample users, multiple positive sample items, and multiple negative sample items, determine the contrastive loss. Based on the contrastive loss, train the alignment network. By respectively determining multiple positive and negative sample users corresponding to the sample user (the similarity between the positive sample user and the sample user exceeds the threshold) and multiple positive and negative sample items corresponding to the sample item (the similarity between the positive sample item and the sample item exceeds the threshold) based on the sample interaction information, and then using these samples to determine the contrastive loss and training the alignment network based on this loss, the alignment network can learn the similarity and difference relationships between users and between items, optimize the network parameters, improve the model's representation ability and alignment effect for user and item features, so as to be better applied to tasks such as recommendation and improve the accuracy and effectiveness of recommendation.
[0150] Among them, the contrastive loss is used to align the collaborative signal with the encoded embedding, and can ensure that the embedding representations of two users (or two items) with similar interaction histories are also similar. Optionally, the encoded embeddings with strong collaborative signals show higher similarity, while those with weak collaborative signals show lower similarity. Based on this, the encoded embedding can be effectively aligned with the collaborative signal. To measure the collaborative signal, the similarity of interaction behaviors can be used. For example, if two users purchase many of the same items, the collaborative signal between them is strong, and their corresponding encoded embeddings become more similar, forming a pair of positive sample users. Similarly, if two items are often purchased together by users, the collaborative signal between them is strong, and their encoded embeddings are also more similar, forming a pair of positive sample items.
[0151] The training steps of the alignment network also include:
[0152] Decode the first user spatio-temporal embedding through the alignment network to obtain the second user spatio-temporal embedding. Decode the first item spatio-temporal embedding through the alignment network to obtain the second item spatio-temporal embedding. Based on the user direct spatio-temporal information, item direct spatio-temporal information, second user spatio-temporal embedding, and second item spatio-temporal embedding, determine the reconstruction loss. Based on the reconstruction loss, train the alignment network. Decoding the first user and first item spatio-temporal embeddings through the alignment network to obtain the second user and second item spatio-temporal embeddings, combining the user and item direct spatio-temporal information to determine the reconstruction loss and train the alignment network can enable the network to learn effective decoding and reconstruction patterns of spatio-temporal embeddings, optimize the network parameters to enhance the representation and reconstruction capabilities of user and item spatio-temporal features, and enable the alignment network to retain more semantic information in the input embeddings.
[0153] Optionally, after obtaining the low-dimensional spatio-temporal constrained embeddings (i.e., the above-mentioned first user spatio-temporal embedding and first item spatio-temporal embedding), the spatio-temporal constrained embedding can be decoded through a decoder network. The input of this decoder network is a D2-dimensional embedding vector. The decoder network undergoes gradual changes through multiple linear layers and finally outputs a D1-dimensional embedding vector to achieve the reconstruction of the original embedding.
[0154] Optionally, decode the above-mentioned first user spatio-temporal embedding S′ u to obtain the second user spatio-temporal embedding which can be expressed as Decode the above-mentioned first item spatio-temporal embedding S′ i to obtain the second item spatio-temporal embedding which can be expressed as
[0155] It should be noted that popular users and items will introduce noise in the above process. That is, popular users often have significant overlap with many other users in the items they purchase, but this overlap does not necessarily indicate a strong collaborative signal. Similarly, popular items are purchased by a large number of users, but their embeddings may not be able to capture meaningful relationships. To solve this problem, positive samples are measured using the following similarity scores:
[0156]
[0157] where u and v are used to represent users, and i and j are used to represent items. α1 and α2 are constant parameters used to adjust the size of the denominator, playing a role in smoothing or normalizing to prevent the denominator from being too small, resulting in unstable or abnormal calculation results.
[0158] Optionally, based on the similarity between users or the similarity between items, samples with high similarity can be selected as positive samples. Correspondingly, the positive sample user is represented as f pos (u) = TopK ({sim(u, u0), …, sim(u, u |U| )}), f pos (u) identifies the top K users most similar to user u as positive sample users, and these positive sample users are from the user set U = u0, …, u |U| . Similarly, the positive sample items are represented as f pos (i) = Top K ({sim(i, i0), …, sim(u, i I )}), f pos (i) identifies the top K items most similar to item i as positive sample items, and these positive sample items are from the item set I = i0, …, i |I| . It should be noted that the similarity between the top K users and user u is greater than the similarity threshold, and the similarity between the top K items and item i is greater than the similarity threshold.
[0159] Since the number of negative samples usually exceeds that of positive samples, this may increase the complexity of the model. Therefore, to solve this problem, the embodiments of the present disclosure adopt a strategy of randomly selecting negative samples to reduce the computational overhead. Accordingly, the negative sample users are represented as f neg (u) = Rand({u0, …, u |U|}), f neg (u) represents the user samples randomly selected from the user set U = u0, …, u |U| . The negative sample items are represented as f neg (i) = Rand({i0, …, i |I|}), f neg (i) represents the item samples randomly selected from the item set I = i0, …, i |I| .
[0160] Optionally, the contrastive loss adopts the InfoNCE loss as the training objective for contrastive learning, and the objective formula is as follows:
[0161]
[0162] where the cos function is used to calculate the cosine similarity between the spatio-temporal embeddings. f pos (u) and f neg (u) respectively represent the positive and negative sample sets of user u, and f pos (i) and f neg (i) are the positive and negative sample sets of item i.
[0163] This contrastive loss can ensure that the embeddings of positive samples (with strong collaborative signals) are more similar to each other, while the embeddings of negative samples have lower similarity, effectively aligning the encoded embeddings with the collaborative signals.
[0164] The purpose of the reconstruction loss is to minimize the distance between the reconstructed embedding and the input embedding, ensuring that the reconstructed embedding retains semantic information from the input embedding as much as possible.
[0165] Optionally, the mean squared error (MSE) is used to quantify the above distance, and the formula is as follows:
[0166]
[0167] where S represents the set of input embeddings, s ∈ S is an input embedding, such as S u , and is the reconstructed embedding corresponding to S u This reconstruction loss can ensure that the reconstruction process effectively maintains the semantic integrity of the original input embedding.
[0168] Second, an explanation is given for extracting spatio-temporal constraints from indirect interaction data.
[0169] In step S307, based on the sample interaction information, the user's indirect spatio-temporal information and the item's indirect spatio-temporal information are determined.
[0170] In the embodiments of the present disclosure, the indirect interaction data is the interaction data extended in the user-item bipartite graph G, also known as the high-order interaction data. A bipartite graph is a special graph structure whose vertex set can be divided into two non-overlapping subsets. For example, it can be divided into a user set U and an item set I, and all the edges in the graph connect the vertices in U and the vertices in I, and there are no edges connecting the vertices within U or I. Select a part of the users from the user set U, denoted as U', and then consider the edges associated with these users and the item vertices connected by these edges. The graph formed by these vertices (the user vertices in U' and the item vertices connected to the users in U' by edges) and the edges between them is the subgraph based on the user. For example, in the bipartite graph of a movie recommendation system, the user set U contains all users, and the item set I contains all movies. If we select the users who like science fiction movies as U', then the science fiction movies associated with these users and the edges between them form a subgraph based on the user. This subgraph can help us focus on the behaviors and preferences of specific user groups for more targeted analysis and recommendation.
[0171] For example, taking the subgraph G u based on the user and the subgraph G i based on the item as an example, by performing a depth-first search on the user u or the item i first, all relevant reachable paths are explored in the graph to retrieve the subgraph G u or G i , which can capture deeper and more complex high-order spatio-temporal relationship patterns, playing a crucial role in modeling spatio-temporal constraints.
[0172] For each high-order interaction data, obtain the corresponding spatio-temporal description information, expressed as or In this way, high-order spatio-temporal data based on direct interaction data can be constructed: and For ease of description, is called user indirect spatio-temporal information, and is called item indirect spatio-temporal information. Among them,
[0173] Similar to the above steps S302 to S304, utilize the text understanding and graph reasoning capabilities of the LLM to filter and reconstruct high-order spatio-temporal data. Refer to the following steps S308 to S310 for a detailed introduction to the steps of extracting spatio-temporal constraints from indirect interaction data.
[0174] Since high-order spatio-temporal data is unstructured data, therefore, this disclosure embodiment designs two prompt templates to convert the above high-order spatio-temporal data into a text sequence compatible with the LLM. Correspondingly, based on the third prompt template, perform structured processing on user indirect spatio-temporal information. Based on the fourth prompt template, perform structured processing on item indirect spatio-temporal information. By performing structured processing on user and item indirect spatio-temporal information based on the third and fourth prompt templates, the compatibility and processing efficiency of high-order spatio-temporal data and indirect spatio-temporal information in the large language model can be effectively improved, facilitating the large language model to understand and utilize this information, and improving the data processing efficiency.
[0175] Optionally, for ease of description, the prompt template for user indirect spatio-temporal information is called the third prompt template, expressed as Then, the structured processing of user indirect spatio-temporal information can be expressed as Optionally, the third prompt template can be: "This is a spatio-temporal recommendation task that needs to model the spatio-temporal preferences of entities. Please filter out confident behavior data from the high-order interaction behaviors of entities."
[0176] Optionally, user indirect spatio-temporal information can be expressed as: "The time and space information of the root node entity is The root node entity interacts with the potential interaction entity ; …; …; The time and space information of the root node entity is The root node entity interacts with the potential interaction entity ." Among them, the time and space information is the spatio-temporal description information, and \n represents the line break character.
[0177] Similarly, for ease of description, the indirect spatio-temporal information of the item prompt template is called the fourth prompt template and is denoted as Then, the structural processing of the indirect spatio-temporal information of the item can be expressed as Optionally, the fourth prompt template can be: "This is a spatio-temporal recommendation task that requires modeling the spatio-temporal preferences of entities. Please filter out confident behavioral data from the high-order interaction behaviors of the entities."
[0178] Similarly, the indirect spatio-temporal information of the item can be expressed as: "The time and space information of the root node entity is the root node entity and the potential interaction entity interact\n; …\n; …\n; the time and space information of the root node entity is the root node entity and the potential interaction entity interact."
[0179] Among them, in the subgraph of the user base, the root node entity generally refers to those specific user nodes that are the basis for constructing the subgraph. These root node users are the core of the subgraph, and the structure and analysis of the entire subgraph are carried out around them. The potential interaction entity refers to the item nodes in the subgraph of the user base that have not been directly interacted with by the root node users but may have the possibility of interaction inferred based on the structure of the graph and the existing information. By analyzing the connection relationship between the root node users and other item nodes in the subgraph (such as through path, neighbor node, etc. information), these potential interaction entities can be mined. For example, if the root node user A and user B have a common interaction item, and user B also interacts with item C, while user A has not interacted with item C, then item C may be a potential interaction entity of user A. The same applies to the subgraph of the item base, which will not be elaborated here.
[0180] In step S308, based on the second large language model, the spatio-temporal constraints of the user are extracted from the user's indirect spatio-temporal information to obtain the first indirect spatio-temporal constraint information.
[0181] In the embodiments of the present disclosure, the user's indirect spatio-temporal information is input into the LLM through a prompt template to obtain a filtered subgraph that retains the high-order relationships related to the spatio-temporal information, which is called the first indirect spatio-temporal constraint information for ease of description and is denoted as In other words, given the user's indirect spatio-temporal information using the third prompt template generate the corresponding text description and then input it into the LLM to obtain the first indirect spatio-temporal constraint information represents the filtered subgraph.
[0182] In step S309, based on the second large language model, the spatio-temporal constraints of the item are extracted from the indirect spatio-temporal information of the item to obtain the second indirect spatio-temporal constraint information.
[0183] In the embodiment of the present disclosure, similar to step S308 above, the indirect spatio-temporal information of the item is input into the LLM through a prompt template to obtain a filtered subgraph that retains the high-order relationships related to the spatio-temporal information, which is called the second indirect spatio-temporal constraint information for convenience of description, denoted as In other words, given the indirect spatio-temporal information of the item Use the fourth prompt template Generate the corresponding text description Then input it into the LLM to obtain the second indirect spatio-temporal constraint information Represents the filtered subgraph.
[0184] In step S310, based on the first indirect spatio-temporal constraint information, the second indirect spatio-temporal constraint information, and the first spatio-temporal constraint information, the second spatio-temporal constraint information is determined.
[0185] In the embodiment of the present disclosure, since the GNN (Graph Neural Network) is a neural network based on the graph data structure, it takes the nodes and edges in the graph as inputs and learns the structure and features of the graph through a series of calculations. Different from traditional neural networks, the graph neural network can directly process graph data without converting it into vector or matrix form. Correspondingly, in the filtered subgraph, the interaction behaviors between users and items are included. Using the spatio-temporal constraint embeddings of users and the spatio-temporal constraint embeddings of items as the initial node representations, and then using GNN for modeling, the depth of users / items can be obtained. Since the input of GNN includes both the subgraph of users and the subgraph of items, the output includes both the spatio-temporal constraints of users and the spatio-temporal constraints of items:
[0186]
[0187] Among them, S′ represents the spatio-temporal constraints from direct interaction data, which is aligned with the collaborative information, that is, the first spatio-temporal constraint information in the above text.
[0188] Correspondingly, the spatio-temporal constraints of user U and item I can be expressed as:
[0189]
[0190] Among them, the spatio-temporal constraint embedding S″ directly models the spatio-temporal description information in the high-order interaction data and is implicitly aligned with the collaborative information. S″ focuses on implementing the above-mentioned ∑ u∈U,i∈I P(ST|G u )+P(ST|G i) to model the spatio-temporal constraints of high-order proximity interaction data.
[0191] It should be noted that after modeling the above spatio-temporal constraints S′ and S″, the next step is to integrate S′ and S″ into the recommendation model, calculate the recommendation score, and train and optimize the recommendation model.
[0192] Currently, most existing ID-based recommendation models are embedding-based. Among them, the embeddings of users and items are usually obtained by random initialization. Based on these initialized embeddings to construct a recommendation network, the above spatio-temporal constraints can be seamlessly integrated with existing ID-based recommendation models.
[0193] The embodiments of the present disclosure provide a combination method, which directly concatenates the user embeddings and item embeddings initialized by the recommendation model with the above spatio-temporal constraints. This method can play a good role when the quality of spatio-temporal constraint modeling is good, so there is no need to invest additional effort to use a more complex combination method.
[0194] In step S311, based on the sample interaction information, user embeddings and item embeddings.
[0195] In the embodiments of the present disclosure, in the initial embedding layer of the recommendation model, first perform random initialization embeddings for each user in the user set U and each item in the item set I, denoted as E = {E U ,E I}. Among them, user embeddings item embeddings Correspondingly, input the above user embeddings and item embeddings into the ID-based recommendation model Rec to calculate the loss function This loss function can measure the quality of the model's recommendation performance. By minimizing this loss function, the recommendation model is optimized to make the recommendation model converge.
[0196] In step S312, concatenate the user embeddings, item embeddings, first spatio-temporal constraint information, and second spatio-temporal constraint information to obtain the input embedding.
[0197] In the embodiments of the present disclosure, by means of concatenation, the above spatio-temporal constraints S′ and S″ are concatenated with the initialized embedding E to obtain the input embedding. This input embedding can be denoted as E all = cat[E, S′, S″].
[0198] In step S313, input the input embedding into the recommendation model to obtain the recommendation result information.
[0199] In the embodiments of the present disclosure, input the above input embedding E allInput into the recommendation model Rec while keeping other parts of the model unchanged to obtain recommendation result information.
[0200] In step S314, train the recommendation model based on the recommendation result information and the sample interaction information.
[0201] In the embodiments of the present disclosure, based on R in the above-mentioned recommendation result information and sample interaction information, the loss of the recommendation model can be obtained, denoted as L rec = Rec(E all , R). Train the recommendation model through this loss.
[0202] It should be noted that during backpropagation, the gradient is not only backpropagated through the recommendation model but also passed to the above GNN network to optimize the GNN parameters that contribute to the spatio-temporal constraint S″.
[0203] It should be noted that the above spatio-temporal constraint representation S′ is obtained by directly optimizing the loss function in the alignment network and This step extracts and aligns the spatio-temporal constraints, making the above spatio-temporal constraints easier to integrate with existing ID-based recommendation models. In some embodiments, since the alignment network is an offline operation, the spatio-temporal constraint S′ can be pre-computed and stored, making the integration into the recommendation system seamless and efficient.
[0204] For example, as shown in Figure 4 shown, Figure 4 is a schematic diagram of another method for training a recommendation model provided according to an exemplary embodiment. As Figure 4 shown, the sample interaction information includes spatio-temporal information and interaction data, and the interaction data includes direct interaction data and indirect interaction data. Structurally process the direct interaction data through the first prompt template and the second prompt template to obtain and Input and into the LLM to obtain the first direct spatio-temporal constraint information S u and the second direct spatio-temporal constraint information S i . Then, encode the above first direct spatio-temporal constraint information S u and the second direct spatio-temporal constraint information S i through the alignment network to obtain the first user spatio-temporal embedding S′ u and the first item spatio-temporal embedding S′ i , that is, the first spatio-temporal constraint information S′. Structurally process the indirect interaction data through the third prompt template and the fourth prompt template to obtain and Input and into the LLM to obtain the first indirect spatio-temporal constraint information and the second indirect spatio-temporal constraint information Process the first spatio-temporal constraint information S′, the first indirect spatio-temporal constraint information and the second indirect spatio-temporal constraint information through the GNN to obtain the second spatio-temporal constraint information S″. Finally, the above spatio-temporal constraints S′ and S″ are concatenated with the initialized embedding E and input into the recommendation model.
[0205] The embodiment of the present disclosure provides a training scheme for a recommendation model. By obtaining the interaction records between sample users and sample items, the spatio-temporal constraint information of direct and indirect interactions is mined to train the recommendation model, so that the recommendation model based on spatio-temporal constraints can fuse the spatio-temporal information of direct interactions and the spatio-temporal information of indirect interactions, significantly improving the accuracy, timeliness and scenario adaptability of recommendations and enhancing the recommendation efficiency.
[0206] Figure 5 is a flowchart of an item recommendation method shown according to an exemplary embodiment. Refer to Figure 5 , which is executed by an electronic device and includes the following steps:
[0207] In step S501, based on the historical behavior information of the target user, the first spatio-temporal constraint information is determined.
[0208] In the embodiment of the present disclosure, the first spatio-temporal constraint information is used to indicate the spatio-temporal connection when the target user directly interacts with each item. This step is the same as steps S301 - S306 above and will not be elaborated here.
[0209] In step S502, based on the historical behavior information and the first spatio-temporal constraint information, the second spatio-temporal constraint information is determined.
[0210] In the embodiment of the present disclosure, the second spatio-temporal constraint information is used to indicate the spatio-temporal connection when the target user indirectly interacts with each item. This step is the same as steps S307 - S310 above and will not be elaborated here.
[0211] In step S503, based on the recommendation model, the first spatio-temporal constraint information and the second spatio-temporal constraint information are processed to obtain at least one target item.
[0212] In the embodiment of the present disclosure, the recommendation model is trained based on the above training method of the recommendation model, and the target item is the item to be recommended to the target user.
[0213] An embodiment of the present disclosure provides an item recommendation solution. By determining the first spatio-temporal constraint information based on the historical behavior information of the target user to clarify the spatio-temporal connection of direct interaction, and then combining the historical behavior information and the first spatio-temporal constraint information to determine the second spatio-temporal constraint information to indicate the spatio-temporal connection of indirect interaction, and finally using the trained recommendation model to process the two types of spatio-temporal constraint information to obtain the target item to be recommended. This way can more comprehensively and deeply explore the relationship between the target user and the item in the spatio-temporal dimension, so as to accurately screen out suitable recommended items for the target user, improve the accuracy and pertinence of the recommendation, and improve the recommendation efficiency.
[0214] The following reflects the effect of the training method of the recommendation model provided by the present disclosure through qualitative experimental results.
[0215] 1) Collect datasets
[0216] Due to the limited publicly available recommendation datasets, the present disclosure collected two large datasets from Platform A and Platform B under the premise of authorization, and applied a 10-core setting during the data processing to ensure high data quality. The first dataset comes from Platform A, including 101,700 users, 81,689 items, and 3,574,071 interaction data. Each interaction behavior includes corresponding spatio-temporal data, where the time information may be "October 1, 20xx, 18:00", and the spatial information may be "a coffee shop located 3 kilometers away". The second dataset comes from Platform B, including 73,417 users, 195,726 items, and 15,211,044 interaction data. Like the dataset of Platform A, each interaction is accompanied by relevant spatio-temporal data. The statistical information of the two datasets is summarized in Table 1 below. For each dataset, the data is divided into a training set, a validation set, and a test set according to the ratio of 8:1:1.
[0217] Table 1 Statistical information of the datasets
[0218] Dataset Platform A Platform B Number of user 101,700 73,417 Number of item 81,689 195,726 Interest interactions 3,574,071 15,211,044 Rating matrix density 0.043% 0.106%
[0219] 2) Set evaluation metrics
[0220] HR (Hit Rate) and NDCG (Normalized Discounted Cumulative Gain) are used to evaluate the performance of the recommendation model. Among them, HR measures the proportion of relevant items in the top K recommendation lists, while NDCG evaluates the ranking quality of these items. The present disclosure reports the results of various top K values, such as HR@10 and NDCG@10, to provide a more comprehensive comparison of recommendation performance. The reason for choosing these metrics is that they are widely used in the evaluation of recommendation systems and can comprehensively evaluate the relevance and ranking quality.
[0221] 3) Comparison method
[0222] Since the solution provided by this disclosure (abbreviated as SIRce) can be widely applied to existing ID-based recommendation models, several representative ID-based models are selected for comparison to demonstrate the effectiveness of using the solution provided by this disclosure to model spatio-temporal insights. In addition, since SIRec uses LLMs to model spatio-temporal environmental information to enhance recommendations, SIRce is also compared with other LLM-based recommendation models. Finally, since spatio-temporal information is often used in knowledge graphs to improve recommendation performance, SIRec is also compared with KG-based recommendation models. These baseline models include a variety of methods to ensure a comprehensive evaluation.
[0223] ID-based recommendation models: For ID-based models, the classic BPR model, the widely used GNN model SimGCL, and the collaborative filtering method SimRec are selected. These models cover a variety of user-item interaction modeling techniques from matrix factorization to graph-based learning and collaborative filtering.
[0224] LLM-based recommendation models: For LLM-based models, UniTRec and EmbSum are selected, which can utilize text data to enhance recommendation performance, and the P5 model, which is very effective in processing large-scale datasets and integrating various feature representations.
[0225] KG-based recommendation models: For KG-based models, KGAT, KGCL, and CIKGRec are selected. These models are specifically designed to construct and utilize knowledge graphs for recommendations.
[0226] Spatio-temporal-based recommendation models: FIN, SPCS, and CoMAN are selected. These models directly model spatio-temporal data and are specifically designed to utilize it to enhance performance. However, they simplify the representation of spatio-temporal information by using predefined tags or entities.
[0227] 4) Hyperparameter settings
[0228] In the solution provided by this disclosure, the Adam optimizer with a learning rate of 0.001 is used. The pre-trained Qwen2-7B model is utilized to implement the LLM-driven spatio-temporal insight module. Although larger LLMs may produce better results, the Qwen2-7B model already provides satisfactory performance. Therefore, no further experiments have been conducted on larger LLMs. The GNN network adopts a linear graph neural network structure similar to those used in LightGCN and LR-GCCF. To ensure a fair comparison, consistent hyperparameters are maintained across all models, including an embedding dimension of 64.
[0229] 5) Overall performance analysis
[0230] Table 2 Results of SIRec combined with ID-based recommendation models on the dataset of Platform A
[0231]
[0232] Table 3 Results of SIRec combined with ID-based recommendation models on the dataset of Platform B
[0233]
[0234] Table 4 Overall performance of recommendations on the dataset of Platform A
[0235]
[0236] Table 5 Overall performance of recommendations on the dataset of Platform B
[0237]
[0238]
[0239] 6) Performance comparison
[0240] As shown in Tables 2 and 3, the SIRec model proposed in the present invention can always improve the recommendation performance. When combined with ID-based recommendation models, the average improvement is 5%. SIRec integrates spatio-temporal information by simply concatenating spatio-temporal insights with the initial embeddings of these models. This straightforward approach effectively demonstrates the accuracy and robustness of SIRec in spatio-temporal modeling.
[0241] In addition, the consistent performance improvement among different models indicates that SIRec can seamlessly adapt to multiple recommendation frameworks. It is worth noting that when SIRec is combined with models such as SimGCL and SimRec, the performance improvement is more obvious compared to BPR. This is because advanced models like SimGCL and SimRec have more complex architectures and can better capture the collaborative signals in user interactions. These capabilities make it possible to effectively integrate spatio-temporal and collaborative information, resulting in superior recommendation performance. SIRec demonstrates its adaptability among a wide range of ID-based recommendation models. Even when combined with relatively simple models such as BPR, its effectiveness is obvious, further verifying the strength of its spatio-temporal modeling.
[0242] As shown in Tables 4 and 5, the SIRec model proposed by the present invention outperforms other models in terms of both HR and NDCG metrics. Specifically, since SIRec can be seamlessly integrated with any ID-based recommendation model, the best-performing combination, SIRec+SimRec, is selected for comparison with other baselines. It is worth noting that SIRec has achieved performance improvement when combined with other advanced ID-based models. For the sake of simplicity in illustration, SIRec+SimRec is used as an example to demonstrate the effectiveness of SIRec.
[0243] Compared with all baseline models, including existing LLM-based, KG-based, and spatio-temporal-based recommendation models capable of modeling spatio-temporal data, SIRec+SimRec has an average improvement of more than 6%. On the datasets of Platform A and Platform B, SIRec+SimRec has an average performance improvement of 5% and 7% respectively compared to the best baseline model. This difference can be attributed to the nature of the datasets. The dataset of Platform B is sourced from a platform focusing on local life services, reflecting the strong, goal-oriented intention of users to find services under specific spatio-temporal conditions. Therefore, accurately modeling spatio-temporal data significantly enhances the recommendation performance. In contrast, the dataset of Platform A is collected from a short-video platform, where local life services are indirectly promoted through videos, and the characteristic user interaction data is sparser. Users on Platform A interact with local life services based on their interest in short videos, making the benefits of spatio-temporal modeling slightly less obvious. Overall, the excellent performance of SIRec on both datasets demonstrates its effectiveness in modeling spatio-temporal data.
[0244] In addition, SIRec+SimRec maintains a consistent performance improvement at different top-K values (e.g., HR@10, HR@20, NDCG@10, NDCG@20). This consistency indicates that SIRec is highly adaptable to different user needs, whether users need a smaller, highly relevant recommendation list (Top-10) or a more extensive candidate recommendation list (Top-50). The stability of SIRec+SimRec at different K values also shows that SIRec does not overly rely on a specific data distribution or user group, which enables it to generalize better on diverse datasets. This robustness can be attributed to SIRec's ability to independently and effectively integrate spatio-temporal insights into recommendations. By not relying on a specific recommendation structure, SIRec ensures that performance is always enhanced through valuable spatio-temporal insights while reducing significant performance fluctuations. This adaptability and stability further emphasize the reliability and effectiveness of SIRec in capturing spatio-temporal semantics to enhance recommendation performance.
[0245] Figure 6It is a block diagram of a training device for a recommendation model shown according to an exemplary embodiment. As Figure 6 shown, the device includes: a first determination unit 601, a second determination unit 602, and a training unit 603.
[0246] The first determination unit 601 is configured to determine first spatio-temporal constraint information based on sample interaction information, where the sample interaction information includes interaction records between a sample user and a sample item, and the first spatio-temporal constraint information is used to indicate the spatio-temporal connection when the sample user and the sample item have a direct interaction;
[0247] The second determination unit 602 is configured to determine second spatio-temporal constraint information based on the sample interaction information and the first spatio-temporal constraint information, and the second spatio-temporal constraint information is used to indicate the spatio-temporal connection when the sample user and the sample item have an indirect interaction;
[0248] The training unit 603 is configured to train the recommendation model based on the first spatio-temporal constraint information and the second spatio-temporal constraint information.
[0249] In some embodiments, the first determination unit 601 is configured to determine user direct spatio-temporal information and item direct spatio-temporal information based on the sample interaction information; extract the spatio-temporal constraints of the user from the user direct spatio-temporal information based on a first large language model to obtain first direct spatio-temporal constraint information; extract the spatio-temporal constraints of the item from the item direct spatio-temporal information based on the first large language model to obtain second direct spatio-temporal constraint information; and determine the first spatio-temporal constraint information based on the first direct spatio-temporal constraint information and the second direct spatio-temporal constraint information.
[0250] In some embodiments, the first determination unit 601 is configured to encode the first direct spatio-temporal constraint information through an alignment network to obtain a first user spatio-temporal embedding; encode the second direct spatio-temporal constraint information through the alignment network to obtain a first item spatio-temporal embedding; and splice the first user spatio-temporal embedding and the first item spatio-temporal embedding to obtain the first spatio-temporal constraint information.
[0251] In some embodiments, the training unit 603 is further configured to, for any sample user, determine a plurality of positive sample users and a plurality of negative sample users corresponding to the sample user based on the sample interaction information, where the similarity between the positive sample user and the sample user is greater than a similarity threshold; for any sample item, determine a plurality of positive sample items and a plurality of negative sample items corresponding to the sample item based on the sample interaction information, where the similarity between the positive sample item and the sample item is greater than the similarity threshold; determine a contrastive loss based on the plurality of positive sample users, the plurality of negative sample users, the plurality of positive sample items, and the plurality of negative sample items; and train the alignment network based on the contrastive loss.
[0252] In some embodiments, the training unit 603 is further configured to decode the first user spatio-temporal embedding through an alignment network to obtain a second user spatio-temporal embedding; decode the first item spatio-temporal embedding through the alignment network to obtain a second item spatio-temporal embedding; determine a reconstruction loss based on the user direct spatio-temporal information, the item direct spatio-temporal information, the second user spatio-temporal embedding, and the second item spatio-temporal embedding; and train the alignment network based on the reconstruction loss.
[0253] In some embodiments, the first determination unit 601 is further configured to structurally process the user direct spatio-temporal information based on a first prompt template; and structurally process the item direct spatio-temporal information based on a second prompt template.
[0254] In some embodiments, the second determination unit 602 is configured to determine user indirect spatio-temporal information and item indirect spatio-temporal information based on sample interaction information; extract the spatio-temporal constraints of the user from the user indirect spatio-temporal information through a second large language model to obtain first indirect spatio-temporal constraint information; extract the spatio-temporal constraints of the item from the item indirect spatio-temporal information through the second large language model to obtain second indirect spatio-temporal constraint information; and determine second spatio-temporal constraint information based on the first indirect spatio-temporal constraint information, the second indirect spatio-temporal constraint information, and the first spatio-temporal constraint information.
[0255] In some embodiments, the second determination unit 602 is further configured to structurally process the user indirect spatio-temporal information based on a third prompt template; and structurally process the item indirect spatio-temporal information based on a fourth prompt template.
[0256] In some embodiments, the training unit 603 is configured to, based on the sample interaction information, the user embedding, and the item embedding; splice the user embedding, the item embedding, the first spatio-temporal constraint information, and the second spatio-temporal constraint information to obtain an input embedding; input the input embedding into a recommendation model to obtain recommendation result information; and train the recommendation model based on the recommendation result information and the sample interaction information.
[0257] The embodiments of the present disclosure provide a training device for a recommendation model. By obtaining the interaction records between sample users and sample items and mining the spatio-temporal constraint information of direct and indirect interactions to train the recommendation model, the spatio-temporal constraint-based recommendation model can fuse the spatio-temporal information of direct interactions and the spatio-temporal information of indirect interactions, significantly improving the accuracy, timeliness, and scenario adaptability of recommendations and enhancing the recommendation efficiency.
[0258] It should be noted that, for the training device of the recommendation model provided in the above embodiments, only the division of the above functional units is used as an example for illustration. In actual applications, the above functions can be allocated to different functional units according to needs, that is, the internal structure of the electronic device is divided into different functional units to complete all or part of the functions described above. In addition, the training device of the recommendation model provided in the above embodiments and the embodiments of the recommendation model training method belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.
[0259] Regarding the training device of the recommendation model in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0260] Figure 7 is a block diagram of an item recommendation device shown according to an exemplary embodiment. As Figure 7 shown, the device includes: a first determination unit 701, a second determination unit 702, and a recommendation unit 703.
[0261] The first determination unit 701 is configured to determine first spatio-temporal constraint information based on the historical behavior information of the target user, and the first spatio-temporal constraint information is used to indicate the spatio-temporal connection when the target user directly interacts with each item;
[0262] The second determination unit 702 is configured to determine second spatio-temporal constraint information based on the historical behavior information and the first spatio-temporal constraint information, and the second spatio-temporal constraint information is used to indicate the spatio-temporal connection when the target user indirectly interacts with each item;
[0263] The recommendation unit 703 is configured to process the first spatio-temporal constraint information and the second spatio-temporal constraint information based on the recommendation model to obtain at least one target item. The recommendation model is trained based on the above-mentioned recommendation model training method, and the target item is an item to be recommended to the target user.
[0264] The embodiments of the present disclosure provide an item recommendation device. By determining the first spatio-temporal constraint information based on the historical behavior information of the target user to clarify the spatio-temporal connection of direct interaction, and then combining the historical behavior information and the first spatio-temporal constraint information to determine the second spatio-temporal constraint information to indicate the spatio-temporal connection of indirect interaction, and finally using the trained recommendation model to process the two types of spatio-temporal constraint information to obtain the target items to be recommended. This way can more comprehensively and deeply explore the relationship between the target user and the items in the spatio-temporal dimension, so as to accurately screen out suitable recommended items for the target user, improving the accuracy and pertinence of the recommendation and enhancing the recommendation efficiency.
[0265] In the embodiments of the present disclosure, the electronic device may be a terminal or a server. When the electronic device is a terminal, the terminal serves as the execution entity to implement the technical solutions provided in the embodiments of the present disclosure; when the electronic device is a server, the server serves as the execution entity to implement the technical solutions provided in the embodiments of the present disclosure; or, the technical solutions provided in the present disclosure are implemented through the interaction between the terminal and the server. The embodiments of the present disclosure do not limit this.
[0266] Figure 8 It is a block diagram of an electronic device shown according to an exemplary embodiment. Generally, the electronic device 800 includes: a processor 801 and a memory 802.
[0267] The processor 801 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 801 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 801 may also include a main processor and a coprocessor. The main processor is a processor used to process data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, the processor 801 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 801 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computing operations related to machine learning.
[0268] The memory 802 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 802 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 802 is used to store at least one program code, and the at least one program code is used to be executed by the processor 801 to implement the training method of the recommendation model provided in the method embodiments of the present disclosure, or to implement the item recommendation method provided in the method embodiments of the present disclosure.
[0269] In some embodiments, the electronic device 800 may further optionally include: a peripheral device interface 803 and at least one peripheral device. The processor 801, the memory 802, and the peripheral device interface 803 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 803 through a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807, and a power supply 808.
[0270] The peripheral device interface 803 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 801 and the memory 802. In some embodiments, the processor 801, the memory 802, and the peripheral device interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 801, the memory 802, and the peripheral device interface 803 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0271] The radio frequency circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 804 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 804 converts an electrical signal into an electromagnetic signal for transmission, or converts a received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 804 can communicate with other electronic devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: a metropolitan area network, generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 804 may further include a circuit related to NFC (Near Field Communication), and this disclosure does not limit this.
[0272] The display screen 805 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 805 is a touch display screen, the display screen 805 also has the ability to collect touch signals on or above the surface of the display screen 805. The touch signals can be input as control signals to the processor 801 for processing. At this time, the display screen 805 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 805, which is provided on the front panel of the electronic device 800; in other embodiments, there may be at least two display screens 805, which are respectively provided on different surfaces of the electronic device 800 or are in a foldable design; in still other embodiments, the display screen 805 may be a flexible display screen, which is provided on the curved surface or the folding surface of the electronic device 800. Even further, the display screen 805 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 805 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0273] The camera module 806 is used to collect images or videos. Optionally, the camera module 806 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the electronic device, and the rear camera is provided on the back of the electronic device. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera respectively, to achieve functions such as the combination of the main camera and the depth-of-field camera to achieve the background blurring function, the combination of the main camera and the wide-angle camera to achieve panoramic shooting and VR (Virtual Reality) shooting functions or other combined shooting functions. In some embodiments, the camera module 806 may also include a flash. The flash can be a single-color-temperature flash or a dual-color-temperature flash. A dual-color-temperature flash refers to the combination of a warm-light flash and a cold-light flash, which can be used for light compensation under different color temperatures.
[0274] The audio circuit 807 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 801 for processing, or input to the radio frequency circuit 804 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 800. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signals from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 807 may further include a headphone jack.
[0275] The power supply 808 is used to supply power to each component in the electronic device 800. The power supply 808 may be alternating current, direct current, a primary battery or a rechargeable battery. When the power supply 808 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0276] Those skilled in the art can understand that Figure 8 the structure shown in
[0277] does not limit the electronic device 800, and may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements.
[0278] A computer program product includes a computer program, and when the computer program is executed by a processor, it implements the above-mentioned training method of the recommendation model, or implements the above-mentioned item recommendation method.
[0279] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0280] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A training method for a recommendation model, characterized in that, The method includes: Determining first spatio-temporal constraint information based on sample interaction information, where the sample interaction information includes interaction records between sample users and sample items, and the first spatio-temporal constraint information is used to indicate the spatio-temporal connection when the sample user and the sample item have a direct interaction; Determining second spatio-temporal constraint information based on the sample interaction information and the first spatio-temporal constraint information, where the second spatio-temporal constraint information is used to indicate the spatio-temporal connection when the sample user and the sample item have an indirect interaction; Training a recommendation model based on the first spatio-temporal constraint information and the second spatio-temporal constraint information.
2. The training method of the recommendation model according to claim 1, wherein The determining the first spatio-temporal constraint information based on sample interaction information includes: Determining user direct spatio-temporal information and item direct spatio-temporal information based on the sample interaction information; Extracting the spatio-temporal constraints of the user from the user direct spatio-temporal information based on a first large language model to obtain first direct spatio-temporal constraint information; Extracting the spatio-temporal constraints of the item from the item direct spatio-temporal information based on the first large language model to obtain second direct spatio-temporal constraint information; Determining the first spatio-temporal constraint information based on the first direct spatio-temporal constraint information and the second direct spatio-temporal constraint information.
3. The training method of the recommendation model according to claim 2, wherein The determining the first spatio-temporal constraint information based on the first direct spatio-temporal constraint information and the second direct spatio-temporal constraint information includes: Encoding the first direct spatio-temporal constraint information through an alignment network to obtain a first user spatio-temporal embedding; Encoding the second direct spatio-temporal constraint information through the alignment network to obtain a first item spatio-temporal embedding; Concatenating the first user spatio-temporal embedding and the first item spatio-temporal embedding to obtain the first spatio-temporal constraint information.
4. The training method of the recommendation model according to claim 3, wherein The training steps of the alignment network include: For any sample user, determining multiple positive sample users and multiple negative sample users corresponding to the sample user based on the sample interaction information, where the similarity between the positive sample user and the sample user is greater than a similarity threshold; For any sample item, determining multiple positive sample items and multiple negative sample items corresponding to the sample item based on the sample interaction information, where the similarity between the positive sample item and the sample item is greater than a similarity threshold; Determining a contrastive loss based on the multiple positive sample users, the multiple negative sample users, the multiple positive sample items, and the multiple negative sample items; Training the alignment network based on the contrastive loss.
5. The training method of the recommendation model according to claim 4, wherein, The method further includes: Decoding the first user spatio-temporal embedding through the alignment network to obtain a second user spatio-temporal embedding; Decoding the first item spatio-temporal embedding through the alignment network to obtain a second item spatio-temporal embedding; Determining a reconstruction loss based on the user direct spatio-temporal information, the item direct spatio-temporal information, the second user spatio-temporal embedding, and the second item spatio-temporal embedding; Training the alignment network based on the reconstruction loss.
6. The training method of the recommendation model according to claim 2, wherein The method further includes: Structuring the user direct spatio-temporal information based on a first prompt template; Structuring the item direct spatio-temporal information based on a second prompt template.
7. The training method of the recommendation model according to claim 1, characterized in that, Determining the second spatio-temporal constraint information based on the sample interaction information and the first spatio-temporal constraint information includes: Determining user indirect spatio-temporal information and item indirect spatio-temporal information based on the sample interaction information; Extracting the spatio-temporal constraints of the user from the user indirect spatio-temporal information based on the second large language model to obtain the first indirect spatio-temporal constraint information; Extracting the spatio-temporal constraints of the item from the item indirect spatio-temporal information based on the second large language model to obtain the second indirect spatio-temporal constraint information; Determining the second spatio-temporal constraint information based on the first indirect spatio-temporal constraint information, the second indirect spatio-temporal constraint information, and the first spatio-temporal constraint information.
8. The training method of the recommendation model according to claim 7, wherein The method further includes: Structuring the user indirect spatio-temporal information based on a third prompt template; Structuring the item indirect spatio-temporal information based on a fourth prompt template.
9. The training method of the recommendation model according to claim 1, characterized in that Training the recommendation model based on the first spatio-temporal constraint information and the second spatio-temporal constraint information includes: Based on the sample interaction information, user embedding and item embedding; Concatenating the user embedding, the item embedding, the first spatio-temporal constraint information, and the second spatio-temporal constraint information to obtain an input embedding; Inputting the input embedding into the recommendation model to obtain recommendation result information; Training the recommendation model based on the recommendation result information and the sample interaction information.
10. An article recommendation method, characterized in that, The method includes: Determining first spatio-temporal constraint information based on the historical behavior information of the target user, where the first spatio-temporal constraint information is used to indicate the spatio-temporal connection when the target user directly interacts with each item; Determining second spatio-temporal constraint information based on the historical behavior information and the first spatio-temporal constraint information, where the second spatio-temporal constraint information is used to indicate the spatio-temporal connection when the target user indirectly interacts with each item; Processing the first spatio-temporal constraint information and the second spatio-temporal constraint information based on a recommendation model to obtain at least one target item, where the recommendation model is trained based on any one of claims 1-9, and the target item is an item to be recommended to the target user.
11. A training device for a recommendation model, characterized in that, The device includes: A first determination unit configured to determine first spatio-temporal constraint information based on sample interaction information, where the sample interaction information includes interaction records between a sample user and a sample item, and the first spatio-temporal constraint information is used to indicate the spatio-temporal connection when the sample user directly interacts with the sample item; A second determination unit configured to determine second spatio-temporal constraint information based on the sample interaction information and the first spatio-temporal constraint information, where the second spatio-temporal constraint information is used to indicate the spatio-temporal connection when the sample user indirectly interacts with the sample item; A training unit configured to train a recommendation model based on the first spatio-temporal constraint information and the second spatio-temporal constraint information.
12. An item recommendation device, characterized in that, The device includes: A first determination unit configured to determine first spatio-temporal constraint information based on the historical behavior information of the target user, where the first spatio-temporal constraint information is used to indicate the spatio-temporal connection when the target user directly interacts with each item; A second determination unit, configured to determine second spatio-temporal constraint information based on the historical behavior information and the first spatio-temporal constraint information, where the second spatio-temporal constraint information is used to indicate the spatio-temporal connection when the target user indirectly interacts with each item; A recommendation unit, configured to process the first spatio-temporal constraint information and the second spatio-temporal constraint information based on a recommendation model to obtain at least one target item, where the recommendation model is trained based on any one of claims 1-9, and the target item is an item to be recommended to the target user.
13. An electronic device, characterized in that, The electronic device includes: One or more processors; A memory for storing program code executable by the processor; Wherein, the processor is configured to execute the program code to implement the training method of the recommendation model as described in any one of claims 1 to 9, or to implement the item recommendation method as described in claim 10.
14. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the training method of the recommendation model as described in any one of claims 1 to 9, or to execute the item recommendation method as described in claim 10.
15. A computer program product, characterized in that, The computer program product includes a computer program, which when executed by a processor implements the training method of the recommendation model as described in any one of claims 1 to 9, or executes the item recommendation method as described in claim 10.
Citation Information
Cited By
Model training method and device for user recommendation, equipment and medium
CN121030335A