Article recommendation method and device, equipment, storage medium and computer program product
By integrating user behavior maps and item semantic maps, random walks and large language model training, the recommendation system can effectively handle fuzzy queries, improve user experience and recommendation diversity, and solve the problem of user data sparseness.
Patent Information
- Application Number
- CN202510169916.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-05-27
AI Technical Summary
The existing item recommendation system is difficult to effectively handle fuzzy queries, which makes it impossible to meet the user's fuzzy item recommendation needs and has a poor user experience.
By fusing the target object's behavior directed graph and the semantic directed graph of the item set, a fused directed graph is obtained, and randomly walks are performed based on this, and a large language model is trained to determine the semantic embedding representation, and then the items are recommended.
It improves the user's item recommendation experience, enhances the diversity and efficiency of recommendations, and solves the problem of insufficient exposure caused by the sparseness of user browsing data.
Smart Images

Figure CN120047218A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to an item recommendation method, apparatus, computer device, storage medium, and computer program product. Background Art
[0002] With the development of computer technology and Internet technology, the emergence of sequence models has made remarkable progress in the field of item recommendation. That is, in a recommendation system, the modeling of a user's behavior sequence is mainly used to understand and predict the user's behavior pattern. A sequence model is an effective method for modeling a user's behavior sequence. The sequence model can help the background of the recommendation system better understand the user's interests, habits, and behavior trends, so as to provide more personalized services or products.
[0003] However, in the current item recommendation methods, due to the existence of a large number of items in the recommendation system, the item information is usually queried and displayed in the way of precise search, that is, the user needs to give precise information, so that the recommendation platform can screen and recommend for display from the item data of different categories provided by third-party institutions according to the information range given by the user for the user to select from. However, in actual applications, using the above-mentioned item information recommendation method often fails to meet the needs of some users with fuzzy item recommendation requirements, that is, the recommendation platform cannot make relevant guidance or recommendations for the user's fuzzy item recommendation demands, that is, it cannot provide relevant item recommendations for this part of users, resulting in a poor experience for this part of users. Therefore, how to improve the user's item recommendation experience has become an urgent problem to be solved. Summary of the Invention
[0004] Based on this, in view of the above technical problems, it is necessary to provide an item recommendation method, apparatus, computer device, computer-readable storage medium, and computer program product, which can effectively improve the user's item recommendation experience and bring convenience to the user.
[0005] In a first aspect, this application provides an item recommendation method. The method includes: fusing the behavioral directed graph of a target object and the semantic directed graph of an item set to obtain a fused directed graph; the semantic directed graph is constructed based on the semantic correlation weights of the nodes of each item in the item set; taking each node of the fused directed graph as a starting point, performing a preset number of random walks according to the weights of the edges in the fused directed graph to obtain random walk sequences of different starting points; training an initial large language model based on each random walk sequence to obtain a large language model; determining semantic embedding representations matching each item in the item set through the large language model; and recommending a target item in the item set to the target object according to the semantic embedding representations.
[0006] In a second aspect, the present application also provides an item recommendation device. The device includes: a fusion module configured to fuse the behavioral directed graph of a target object and the semantic directed graph of an item set to obtain a fused directed graph; the semantic directed graph is constructed based on the semantic correlation weights of the nodes of each item in the item set; a random walk module configured to start from each node of the fused directed graph and perform a preset number of random walks according to the weights of the edges in the fused directed graph to obtain random walk sequences starting from different said starting points; a training module configured to train an initial large language model based on each of the random walk sequences to obtain a large language model; a determination module configured to determine semantic embedding representations matching each item in the item set through the large language model; and a recommendation module configured to recommend target items in the item set to the target object according to the semantic embedding representations.
[0007] In a third aspect, the present application also provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the following steps are implemented: fusing the behavioral directed graph of a target object and the semantic directed graph of an item set to obtain a fused directed graph; the semantic directed graph is constructed based on the semantic correlation weights of the nodes of each item in the item set; starting from each node of the fused directed graph and performing a preset number of random walks according to the weights of the edges in the fused directed graph to obtain random walk sequences starting from different said starting points; training an initial large language model based on each of the random walk sequences to obtain a large language model; determining semantic embedding representations matching each item in the item set through the large language model; and recommending target items in the item set to the target object according to the semantic embedding representations.
[0008] In a fourth aspect, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, the following steps are implemented: fusing the behavioral directed graph of a target object and the semantic directed graph of an item set to obtain a fused directed graph; the semantic directed graph is constructed based on the semantic correlation weights of the nodes of each item in the item set; starting from each node of the fused directed graph and performing a preset number of random walks according to the weights of the edges in the fused directed graph to obtain random walk sequences starting from different said starting points; training an initial large language model based on each of the random walk sequences to obtain a large language model; determining semantic embedding representations matching each item in the item set through the large language model; and recommending target items in the item set to the target object according to the semantic embedding representations.
[0009] In a fifth aspect, the present application also provides a computer program product. The computer program product includes a computer program which, when executed by a processor, implements the following steps: fusing the behavioral directed graph of a target object and the semantic directed graph of an item set to obtain a fused directed graph; the semantic directed graph is constructed based on the semantic correlation weights of the nodes of each item in the item set; taking each node of the fused directed graph as a starting point, performing a preset number of random walks according to the weights of the edges in the fused directed graph to obtain random walk sequences starting from different said starting points; training an initial large language model based on each of the random walk sequences to obtain a large language model; determining semantic embedding representations matching each item in the item set through the large language model; and recommending a target item in the item set to the target object according to the semantic embedding representations.
[0010] For the above item recommendation method, device, computer device, storage medium and computer program product, by fusing the behavioral directed graph of a target object and the semantic directed graph of an item set, a fused directed graph is obtained; the semantic directed graph is constructed based on the semantic correlation weights of the nodes of each item in the item set; taking each node of the fused directed graph as a starting point, performing a preset number of random walks according to the weights of the edges in the fused directed graph to obtain random walk sequences starting from different said starting points; training an initial large language model based on each of the random walk sequences to obtain a large language model; determining semantic embedding representations matching each item in the item set through the large language model; and recommending a target item in the item set to the target object according to the semantic embedding representations. Since the semantic directed graph in the present application is pre-constructed based on the semantic correlation weights between the nodes of each item in the item set, and the behavioral directed graph is pre-constructed based on the user's behavior sequence, first fusing the behavioral directed graph of the target object and the semantic directed graph of the item set is to integrate the user's behavior sequence features and the semantic correlation between the node contents into the same graph structure. Furthermore, the random walk sequences obtained based on the fused directed graph can not only reflect the items frequently browsed by the user, but also appropriately display those items less visited by the user through semantic correlation. Furthermore, while maintaining good item recommendation efficiency, it can effectively solve the problem of underfitting of the embedding features of the exposure-insufficient nodes caused by the sparsity of user browsing data, thereby effectively improving the diversity of item recommendation, bringing a better item recommendation experience to the user, and bringing convenience to the user. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 It is an application environment diagram of the item recommendation method in an embodiment;
[0012] Figure 2 It is a flowchart of the item recommendation method in an embodiment;
[0013] Figure 3 Schematic diagram of the display interface on the product side for the item recommendation method provided in an embodiment;
[0014] Figure 4 Schematic diagram of the basic structure during model training in an embodiment;
[0015] Figure 5 Schematic flowchart of the step of constructing a semantic directed graph based on the semantic relevance weights determined by the first semantic embedding representations of each item in an embodiment;
[0016] Figure 6 Schematic flowchart of the step of recommending a target item in the item set to a target object according to the semantic embedding representation in an embodiment;
[0017] Figure 7 Schematic diagram of the basic structure of a recommendation model in an embodiment;
[0018] Figure 8 Schematic diagram of the interface display on the product side in an embodiment;
[0019] Figure 9 Block diagram of the structure of an item recommendation device in an embodiment;
[0020] Figure 10 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0021] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0022] It should be noted that in the following description, the terms "first", "second" and "third" only distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first", "second" and "third" can be interchanged in a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0023] The item recommendation method provided by the embodiments of the present application can be applied to, for example Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be a separate device, integrated on the server 104, or placed on the cloud or other network servers. That is, the terminal 102 can interact with the recommendation platform, that is, the server 104. That is, the server 104 can fuse the behavior directed graph of the target object and the semantic directed graph of the item set to obtain a fused directed graph; the semantic directed graph is constructed based on the semantic correlation weights of the nodes of each item in the item set; the server 104 takes each node of the fused directed graph as the starting point, performs a preset number of random walks according to the weights of the edges in the fused directed graph to obtain random walk sequences with different starting points, and trains the initial large language model based on each random walk sequence to obtain a large language model; further, the server 104 can determine the semantic embedding representations matching each item in the item set through the trained large language model, and recommend the target items in the item set to the target object according to the semantic embedding representations. That is, the server 104 can return the target items in the item set recommended to the target object to the terminal 102 so that the terminal 102 recommends the target items to the target object.
[0024] Among them, the terminal 102 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart TV, a smart watch, an Internet of Things device, and a portable wearable device. The Internet of Things device can be a smart vehicle-mounted device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc.
[0025] The server 104 can be an independent physical server or a service node in a blockchain system. A peer-to-peer (Peer To Peer) network is formed among the service nodes in the blockchain system. The Peer To Peer protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP).
[0026] In addition, the server 104 can also be a server cluster composed of multiple physical servers, and can be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0027] The terminal 102 and the server 104 can be connected through communication connection methods such as Bluetooth, USB (Universal Serial Bus), or the network. This application does not make any restrictions here.
[0028] In one embodiment, as Figure 2 shown, an item recommendation method is provided. This method can be executed independently by a server or a terminal, or jointly by a server and a terminal. Taking the terminal in Figure 1 as an example, the method includes the following steps:
[0029] Step 202: Fuse the behavior directed graph of the target object and the semantic directed graph of the item set to obtain a fused directed graph; the semantic directed graph is constructed based on the semantic correlation weights of the nodes of each item in the item set.
[0030] The target object refers to a specific object. For example, the target object in this application can be different registered users in the recommendation application. It can be understood that the target object in this application can refer to different objects used during the model training phase, that is, there can be multiple target objects.
[0031] The behavior directed graph refers to a graph used to reflect the behavior pattern (or behavior trend) of the target object. That is, the behavior directed graph in this application is a graph composed of nodes and edges. In some cases, the behavior directed graph in this application can also be called a behavior network. For example, the behavior directed graph in this application can be graph structure data containing topological structure information, that is, the graph structure in the behavior directed graph is used to reflect the items frequently browsed by the target object.
[0032] The item set refers to a set of items (item identifiers) used during the model training phase. That is, the item set in this application can be a set containing all the items (item identifiers) used in the training phase.
[0033] The semantic directed graph refers to a graph used to reflect the semantic correlation between each item (item node) in the item set. That is, the semantic directed graph in this application is also a graph composed of nodes and edges. In some cases, the semantic directed graph in this application can also be called a semantic network. For example, the semantic directed graph in this application can be graph structure data containing topological structure information, that is, the graph structure in the semantic directed graph is used to reflect the semantic correlation of the nodes of each item in the item set. It can be understood that the semantic correlation between two nodes in the semantic directed graph in this application can be represented by a semantic correlation weight.
[0034] The semantic relevance weight refers to the weight used to reflect the semantic relevance between the embedded representations of two nodes in a semantic directed graph. For example, the semantic relevance weight in the present application can be represented by the semantic relevance weight of the directed edge connecting the two nodes. The semantic relevance weight of the directed edge between two nodes can be determined based on the cosine distance between the semantic embedded representations of the two nodes, or can also be determined based on the product of the semantic embedded representations of the two nodes. No specific limitation is made here.
[0035] The fused directed graph refers to the graph obtained by fusing the behavioral directed graph of the target object and the semantic directed graph of the item set. That is, the fused directed graph in the present application also refers to a graph composed of nodes and edges. In some cases, the fused directed graph in the present application can also be called a fusion network. For example, the fused directed graph in the present application can be graph structure data containing topological structure information, that is, the graph structure in the fused directed graph is used to reflect the preferences in the browsing behavior of the target object. It can be understood that the fused directed graph in the present application can reflect the preferences in the object browsing behavior through the weights of the edges, and by integrating the co-occurrence weights of the behaviors of the target object into the directed edges in the graph, those high-frequency item nodes with higher display probabilities will be more likely to appear in the random walk path. At the same time, through semantic relevance, the semantic nearest neighbor nodes of each item node will fuse the low-weight semantic edges into this directed graph, so that the (item) nodes that are relatively sparse in the behavioral data can also be included in the walk path, ensuring that all nodes can be encoded into the same embedding space.
[0036] The nodes of each item in the item set refer to using the item identifiers of each item in the item set as the nodes in the directed graph. For example, if the three items included in the item set are item A, item B, and item C respectively, then the item identifiers of item A, item B, and item C are used as the nodes in the directed graph to construct different nodes in the directed graph. That is, the nodes of each item in the item set include: item A node, item B node, and item C node.
[0037] Step 204: Starting from each node of the fused directed graph, perform a preset number of random walks according to the weights of the edges in the fused directed graph to obtain random walk sequences with different starting points.
[0038] Among them, the preset number refers to the number of random walks set in advance. For example, the preset number of random walks in the present application can be set to n times. It can be understood that when the terminal in the present application performs a preset number of random walks, it can start from each node of the fused directed graph and perform n random walks of length L, where both the length L and the number n are preset hyperparameters, and the probability of each walk is proportional to the weights of the adjacent edges (each edge) of each node.
[0039] The random walk sequences with different starting points refer to: the sequences obtained by performing random walks with the nodes of the fused directed graph as the starting points. It can be understood that there can be multiple random walk sequences for each node, that is, the random walk sequences with different starting points can be a collection of sequences obtained by performing random walks with the nodes of the fused directed graph as the starting points.
[0040] Step 206: Train the initial large language model based on each random walk sequence to obtain a large language model.
[0041] Among them, the initial large language model refers to an untrained model, and the large language model refers to a trained model. In this application, the trained large language model can be used for item recommendation.
[0042] Specifically, the item recommendation method provided in this application can be widely used in various personalized item recommendation scenarios such as social networking, games, and ticketing. For example, given an item sequence x of n items browsed by a user, 1 …x n The item recommendation method provided in this application can be based on the item sequence x 1 …x n Quickly and accurately recommending one or more target items to the user has wide applications in the fields of social entertainment, item promotion, games, etc. That is, the devices used by different users (operating objects) can interact with the multimedia information platform (or application). When the user (operating object) wants to browse some personalized recommended items, the user can open the multimedia application (Application, APP) on the terminal through a trigger operation, and enter the main page of the multimedia application through a selection operation, that is, the user can log in to the multimedia application (such as a video application) through a trigger operation. Further, the user can initiate an item recommendation request through a trigger operation in the main page displayed by the multimedia application. For example, in the main page of the video application displayed by the terminal, each object (such as a developer) using the video application can view the specific content and related functional information in the video main page, and each object using the video application can trigger the generation of a video recommendation request by clicking the "Recommendation" control. Then, the terminal responds to the video recommendation request triggered by the operating object in the main page of the video application, calls a pre-trained large language model for item recommendation, and determines the semantic embedding representation matching each item in the item set through the large language model, and then recommends the target item in the item set to the user based on the semantic embedding representation.
[0043] Among them, when training a large language model for item recommendation, the terminal can obtain the behavioral directed graph of the target object (i.e., the sample object) and the semantic directed graph of the item set, and fuse the behavioral directed graph of the target object and the semantic directed graph of the item set to obtain a fused directed graph; further, the terminal takes each node in the fused directed graph as a starting point and performs a preset number of random walks according to the weights of the edges in the fused directed graph, and different starting random walk sequences can be obtained; further, the terminal can train the initial large language model based on different starting random walk sequences (i.e., the set of random walk sequences) for a preset number of times to obtain the trained large language model.
[0044] It can be understood that the method provided in this application can be implemented through the interaction between the terminal and the background server of the video application, or through the interaction between the front end and the back end of the terminal. That is, the front end of the terminal is used to display the target items recommended to the target object, and the back end of the terminal is equivalent to the background server, which is used to perform logical processing such as fusing and random walking the behavioral directed graph of the target object and the semantic directed graph of the item set.
[0045] For example, taking the scenario of personalized video recommendation as an example for illustration. Assume that the multimedia application is a video application. As Figure 3 shown, it is a schematic diagram of the display interface on the product side of the item recommendation method provided in this application. That is, when user A (the operating object) wants to browse some personalized recommended videos, user A can open video application A on the terminal through a trigger operation and enter the main page of video application A as shown in Figure 3 . That is, user A can log in to video application A through a trigger operation. Further, user A can trigger the display of the page for starting model training through a trigger operation on the main page shown in Figure 3 . For example, on the main page shown in Figure 3 displayed on the terminal, user A can click the "Train" control shown in the left figure in Figure 3 so that the terminal responds to the above click operation of user A and displays the page for starting model training shown on the right in Figure 3 . After user A completes the training input operation (such as determining the sample sequence for training input), user A can further click the "Start" control on the page for starting model training shown on the right in Figure 3 to initiate a model training request. Then the terminal responds to the click "Start" operation (i.e., the model training request) triggered by user A on this page, obtains the behavioral directed graph of the target object (i.e., the sample object) and the semantic directed graph of the item set , and the behavioral directed graph of the target object and the semantic directed graph of the item set Fuse to obtain a fused directed graph ; Further, the terminal uses each node in the fused directed graph as the starting point, that is, for as the starting point, perform n random walks of length L to obtain a set of sequences of random walks for each node, and use the initial large language model to fit the sequences of random walks. That is, the terminal can perform m training sessions on the initial large language model based on the set of sequences of random walks for each node to obtain the trained large language model. Among them, both the length L and the number n are preset hyperparameters, and the probability of each walk is proportional to the weight of the adjacent edges of each node.
[0046] Step 208: Determine the semantic embedding representations that match each item in the item set through the large language model.
[0047] Among them, the semantic embedding representation refers to a fused semantic embedding representation that combines image features and text features. For example, the semantic embedding representations that match each item in the item set in this application can be: after initially extracting the semantic embedding representations from the image information and text information of each item through a multimodal large language model, then input the initial semantic embedding representations into the embedding layer of the initial large language model as initialization features, and obtain richer semantic embedding representations after model training, and store the trained semantic embedding representations in the database or the embedding layer of the model, so that when applying, the richer semantic embedding representations that match can be directly retrieved from the database or the embedding layer of the model through the item identifiers of each item in the item set.
[0048] Step 210: Recommend the target items in the item set according to the semantic embedding representations.
[0049] Among them, the target item refers to the relevant item in the item set determined based on the semantic embedding representations that match each item in the item set. It can be understood that the target item in this application can be one or multiple, and no specific limitation is made here.
[0050] Specifically, after completing model training, the terminal does not need to obtain the behavior sequence of the target object. The terminal can directly use the trained large language model, that is, determine the semantic embedding representations matching each item in the item set through the trained large language model, and recommend the target item in the item set to the target object according to the semantic embedding representations matching each item in the item set. For example, the terminal can separately find the semantic embedding representations corresponding to the item identifiers of each item in the item set from the embedding layer of the large language model, and further find a preset number of target embedding features whose distances from the semantic embedding representations meet the distance threshold from the database or the embedding layer of the large language model, and use the items corresponding to the item identifiers of the target embedding features as the target items recommended to the target object.
[0051] For example, as Figure 4 shown, it is a schematic diagram of the basic structure during model training. When the Figure 4 recommendation model (large language model) shown in Figure 4 is trained, the terminal can directly utilize the embedding representations of each item ID in the embedding layer of the model structure shown, and through the approximate nearest neighbor algorithm (ANN), find the item corresponding to the item identifier with the closest distance to it for item-to-item (I2I) recommendation.
[0052] In addition, the terminal can also obtain the behavior sequence of other objects. The behavior sequence is composed of the item identifiers of each item interacted with by other objects, and input the behavior sequence into the trained large language model, so that the embedding layer of the large language model sequentially finds the semantic embedding representations corresponding to the item identifiers in the behavior sequence, and determines the sequence embedding feature (i.e., the predicted embedding feature) of the behavior sequence based on the behavior sequence and the found semantic embedding representations; further, the terminal can determine the target item to be recommended based on the predicted sequence embedding feature of the behavior sequence and recommend the target item to other objects.
[0053] In this embodiment, by fusing the behavior directed graph of the target object and the semantic directed graph of the item set, a fused directed graph is obtained; the semantic directed graph is constructed based on the semantic correlation weights of the nodes of each item in the item set; starting from each node of the fused directed graph, a random walk is performed a preset number of times according to the weights of the edges in the fused directed graph to obtain random walk sequences starting from different said starting points; training an initial large language model based on each said random walk sequence to obtain a large language model; determining semantic embedding representations matching each item in the item set through the large language model; and recommending a target item in the item set to the target object according to the semantic embedding representations. Since the semantic directed graph in this application is pre-constructed based on the semantic correlation weights of the content between each item in the item set, and the behavior directed graph is pre-constructed based on the behavior sequences of different users, first fusing the behavior directed graph of the target object and the semantic directed graph of the item set is to integrate the behavior sequence features of the user and the semantic correlation of the content between items into the same graph structure, so that the random walk sequences obtained based on the fused directed graph can not only reflect the items frequently browsed by the user, but also appropriately display those items less visited by the user through semantic correlation, thereby effectively solving the problem of underfitting of the embedding features of the items with insufficient exposure caused by the sparsity of user browsing data while maintaining good item recommendation efficiency, effectively improving the diversity of item recommendations, bringing a better item recommendation experience to the user, and bringing convenience to the user.
[0054] In one embodiment, before fusing the behavior directed graph of the target object and the semantic directed graph of the item set to obtain a fused directed graph, the method further includes:
[0055] Obtain the behavior sequence of the target object; the behavior sequence is composed of the item identifiers of each item interacted with by the target object; each item identifier in the behavior sequence belongs to the identifiers of each item in the item set;
[0056] Construct a behavior directed graph based on the behavior sequence.
[0057] Among them, the behavior sequence refers to a sequence used to reflect the behavior pattern (or behavior trend) of the target object. For example, the behavior sequence in this application can be an item sequence, and the item sequence contains the item identifiers of each item interacted with by the target object. For example, if user A browses item A, item B, and item C in the recommendation application, the behavior sequence obtained based on the item records browsed (interacted with) by user A is S1(SASBSC), user A is the target object, the item sequence S1(SASBSC) is the behavior sequence of user A, and this item sequence S1 is composed of the item identifiers of each item (item A, item B, and item C) browsed (interacted with) by this user A.
[0058] An item identifier is used to identify a unique item. For example, the item identifier in this application can be a serialized code or a string.
[0059] Each item identifier in the behavior sequence belongs to the identifier of each item in the item set, which means that the item set in this application contains the identifiers of all items, and each item identifier included in the behavior sequence can be a subset of the identifiers of each item in the item set.
[0060] Specifically, taking the scenario of personalized video recommendation as an example for illustration. The terminal can, at a preset time interval, obtain the video identifiers of each video interacted with by different users, perform serialization processing on each video identifier to obtain the behavior sequences of different users, and store the behavior sequences of different users in the database, so that the behavior sequences of different users can be directly obtained from the database subsequently. For example, the terminal can obtain the behavior sequence of a target object (different users) from the database and construct a behavior directed graph based on the obtained behavior sequence. Among them, the behavior sequence in this application is composed of the item identifiers of each item interacted with by the target object, and each item identifier in the behavior sequence belongs to the identifier of each item in the item set.
[0061] It can be understood that in some cases, when user A wants to browse some personalized recommended videos, or when the recommendation system wants to recommend some personalized videos to user A, the terminal (or the background of the recommendation system) can also obtain in real time the video identifiers of each video interacted with by the target object, that is, different users, on the same day (or within a preset time period), and perform serialization processing on each video identifier to obtain a set of behavior sequences of the target object, that is, different users, and construct a behavior directed graph based on the set of behavior sequences of different users. Thus, by pre-obtaining the video identifiers interacted with by different users and performing serialization processing on each video identifier, the behavior sequences of different users can be obtained, and a behavior directed graph can be constructed based on the behavior sequences of different users, so that subsequently, the behavior sequence (behavior directed graph) of the user can be directly integrated with the semantic relevance (semantic directed graph) between nodes into the same graph structure (fused directed graph), ensuring that the random walk sequence generated based on the fused directed graph can not only reflect the items frequently browsed by the user, but also appropriately display those items less visited by the user through semantic relevance. Thus, while maintaining a high recommendation efficiency, the problem of underfitting of the embedding features of the exposure-insufficient (item) nodes caused by the sparsity of user browsing data is effectively solved, and then a better item (video) recommendation experience is brought to the user, bringing convenience to the user.
[0062] In one embodiment, the step of constructing a behavior directed graph based on the behavior sequence includes:
[0063] Construct nodes and directed edges corresponding to the behavior sequence based on the preset window value of the sliding window and each item identifier in the behavior sequence;
[0064] Determine the weight of the directed edge based on the construction times of the directed edge;
[0065] Construct a behavior directed graph based on the nodes, directed edges, and weights;
[0066] Among them, the preset window value of the sliding window refers to: the window size value of the preset sliding window, which is used to represent the window size of the sliding window. For example, in this application, the preset window value can be set to W, and W can be an empirical value.
[0067] The weight of the directed edge refers to the number of times the two nodes connected by the directed edge co-occur. The greater the weight on the edge, the stronger the connection between the two nodes. It can be understood that the behavior directed graph in this application can also be called a directed weighted co-occurrence graph of the user's behavior or a directed weighted co-occurrence network of the user's behavior.
[0068] Co-occurrence network: Also known as a co-occurrence network, it is a method of constructing a network. If two nodes appear simultaneously in the same scenario, an edge is connected between the two nodes. The weight on the edge represents the number of times the two nodes co-occur. The greater the weight on the edge, the stronger the connection between the two nodes.
[0069] Specifically, taking the scenario of personalized video recommendation as an example for illustration. After the terminal obtains the behavior sequence of the target object (different users), the terminal can construct nodes and directed edges corresponding to the behavior sequence based on the preset window value of the sliding window and each item identifier in the behavior sequence, and determine the weight of the directed edge based on the construction times of the directed edge. The terminal constructs a behavior directed graph based on the nodes, directed edges, and weights . Among them, is the node set, is the directed edge set, is the weight set. Specifically, , represents the co-occurrence times of the directed edge .
[0070] For example, assume that the window size of the preset co-occurrence window is an empirical value of W, and use this to determine which nodes in the sequence need to establish directed edges. Specifically, the edge construction process of the terminal in this embodiment can be formally expressed as follows: For a subsequence of , where , connects to the subsequent nodes, obtaining directed edges, that is In all the user behavior sequences obtained at the terminal, directed edges within the co-occurrence window size W are established according to the above method, and the number of these co-occurring directed edges is recorded as the weight of the edges in the graph, finally obtaining a directed weighted network based on user behavior. 。
[0071] In this embodiment, since it is proposed in the model training stage to integrate the user behavior sequence and the item nodes with semantic relevance into the same graph structure, it is ensured that the generated random walk sequence can not only reflect the items frequently browsed by the user, but also appropriately display those items less visited by the user through semantic relevance. This method enables the method provided in the embodiment of the present application to maintain better item recommendation efficiency while effectively solving the problem of underfitting of the embedding features of the exposure-deficient nodes caused by the sparsity of user browsing data, effectively improving the efficiency of model training, and at the same time effectively avoiding problems such as poor controllability, large resource consumption, and non-standard generated IDs introduced in the traditional method when using a pre-trained large model, thereby improving the processing efficiency of item recommendation.
[0072] In one of the embodiments, before fusing the directed behavior graph of the target object and the semantic directed graph of the item set to obtain a fused directed graph, the method further includes:
[0073] Extracting features of each item in the item set through a multimodal model to obtain the initial semantic embedding representation of each item;
[0074] Constructing a semantic directed graph based on the semantic relevance weights determined by the initial semantic embedding representations of the items.
[0075] Among them, the initial semantic embedding representation refers to the fused semantic embedding representation of each item extracted in advance through other models. For example, the initial semantic embedding representation in the present application can be the fused semantic embedding representation corresponding to each item finally obtained by pre-using a multimodal large language model to extract and fuse the picture information and text information of each item. That is, the multimodal model in the present application can be a multimodal large language model.
[0076] Specifically, in the model training stage, such as Figure 4As shown in , the terminal fuses the behavior directed graph of the target object and the semantic directed graph of the item set. Before obtaining the fused directed graph, the terminal can obtain the item ID of each item in the item set, and extract features of the items corresponding to the item IDs through the multimodal model to obtain the initial semantic embedding representation of each item, and construct a semantic directed graph based on the semantic relevance weights determined by the initial semantic embedding representations of each item. For example, after the terminal obtains the initial semantic embedding representation of each item node (item ID) through processing by other models (such as a multimodal large language model), the terminal can use the nearest neighbor algorithm to obtain the first k nearest neighboring points of each item node with the closest semantic distance, thereby obtaining a directed edge composed of the k most semantically relevant neighboring points of each item node. Finally, the terminal can obtain a directed weighted graph based on semantic relevance. ,in, is a node set, is a set of directed edges, is a set of semantic relevance weights. , indicating a directed edge . It can be understood that the semantic relevance weight in the implementation of the present application can be the cosine distance between the embedded representation vectors of two nodes. As a result, for newly added item nodes, the method provided in the embodiment of the present application does not need to rely on user interaction information, and can be integrated into the semantic network (i.e., semantic directed graph) through the semantic features of the nodes. This not only can achieve efficient fitting of the wandering sequences of all item nodes in the fused network, but also can obtain the embedded representations of all nodes, so as to make item-to-item recommendations, that is, the efficient model structure of the pre-trained large language model can be used to more effectively capture the user's sequential behavior characteristics and the semantic relevance between items, which can effectively alleviate the cold start problem caused by the sparsity of user browsing data, and can provide users with richer item recommendations, that is, the diversity of item recommendations made through the model is improved.
[0077] In one embodiment, Figure 5 As shown, the step of constructing a semantic directed graph based on the semantic relevance weight determined by the first semantic embedding representation of each item includes:
[0078] Step 502, determining the neighboring points of each node based on the initial semantic embedding representation of each item;
[0079] Step 504, constructing directed edges corresponding to each node based on each node and its adjacent nodes;
[0080] Step 506, determining the semantic relevance weight of the directed edge;
[0081] Step 508: construct a semantic directed graph based on each node, directed edge and the semantic relevance weight of the directed edge.
[0082] Among them, the neighboring points of each node in the present application can be the first k neighboring points with the closest semantic distance to each node calculated according to the nearest neighbor algorithm.
[0083] Specifically, after the terminal obtains the initial semantic embedding representation of each item node (item identifier) through processing by other models (such as a multimodal large language model), the terminal determines the neighboring points of each node based on the initial semantic embedding representation of each item, and constructs the directed edges corresponding to each node based on the neighboring points of each node. For example, the terminal can obtain the first k neighboring points with the closest semantic distance to each item node through the nearest neighbor algorithm, and then construct the directed edges corresponding to each node based on the first k neighboring points of each node and each node, thereby obtaining the directed edges composed of the most semantically relevant neighboring points of each item node.
[0084] Furthermore, the terminal can determine the semantic relevance weight of each directed edge in turn, and construct a semantic directed graph based on the semantic relevance weights of each node, directed edge, and directed edge, that is, a directed weighted graph based on semantic relevance can be finally obtained. ,in, is a node set, is a set of directed edges, is a set of semantic relevance weights. , indicating a directed edge . It can be understood that the semantic relevance weight in the implementation of the present application can be the cosine distance between the embedded representation vectors of two nodes. As a result, for newly added item nodes, the method provided in the embodiment of the present application does not need to rely on user interaction information, but can be integrated into the semantic network (i.e., semantic directed graph) through the semantic features of the nodes. This not only achieves efficient fitting of the wandering sequences of all item nodes in the fused network, but also obtains the embedded representations of all nodes, thereby providing users with richer item recommendations, thereby improving the diversity of item recommendations through the model.
[0085] In one embodiment, the step of determining the semantic relevance weight of the directed edge includes:
[0086] Get the first node and the second node connected by the directed edge;
[0087] Determine a cosine distance between the semantic embedding representation of the first node and the semantic embedding representation of the second node, and use the cosine distance as a semantic relevance weight of the directed edge; or,
[0088] The product of the semantic embedding representation of the first node and the semantic embedding representation of the second node is used as the semantic relevance weight of the directed edge.
[0089] Among them, the first node and the second node are only used to distinguish the two nodes corresponding to connecting a directed edge. For example, the first node in the present application may be the starting node constituting the directed edge, and the second node may be the ending node constituting the directed edge.
[0090] Specifically, after the terminal obtains the initial semantic embedding representation of each item node (item identifier) through the processing of other models (such as multimodal large language models), the terminal can calculate the top k adjacent nodes with the closest semantic distance to each item node through the nearest neighbor algorithm, that is, obtain the top k adjacent nodes closest to the (initial) semantic embedding representation of each item node. Furthermore, the terminal can construct the directed edges corresponding to each node based on each node and the top k adjacent nodes of each node, so as to obtain the directed edges composed of the top k adjacent nodes that are semantically most relevant to each item node. Further, the terminal can sequentially determine the semantic relevance weights of each directed edge. For example, the terminal can obtain the first node A and the second node B connected by the directed edge e, and determine the semantic embedding representation h of the first node A A and the semantic embedding representation h of the second node B B The cosine distance S between them is used as the semantic relevance weight W of the directed edge e AB =S.
[0091] Alternatively, the terminal can obtain the first node A and the second node B connected by the directed edge e, and use the product of the semantic embedding representation h of the first node A A and the semantic embedding representation h of the second node B B as the semantic relevance weight W of the directed edge e AB =( ). Thus, by constructing a semantic network based on the semantic relevance weights between the nodes of each item in the item set, the correlation between the contents of the nodes can be effectively reflected. Furthermore, by fusing the weighted co-occurrence behavior network with the semantic network between items later, the fused network can reflect the preferences in the user's browsing behavior through the weights of the edges, enabling the method provided by the embodiments of the present application to maintain a higher item recommendation efficiency while effectively solving the problem of underfitting of the embedding features of the exposure-insufficient nodes caused by the sparsity of user browsing data.
[0092] In one embodiment, the step of fusing the behavioral directed graph of the target object and the semantic directed graph of the item set to obtain a fused directed graph includes:
[0093] The third node and the first directed edge in the behavioral directed graph of the target object are respectively spliced with the fourth node and the second directed edge in the semantic directed graph of the item set to obtain a spliced directed graph;
[0094] Weight the weight of the first directed edge in the behavior directed graph with the semantic correlation weight of the second directed edge in the semantic directed graph to obtain a splicing weight;
[0095] Based on the spliced directed graph and the splicing weight, obtain a fused directed graph.
[0096] Among them, the third node and the fourth node are only used to distinguish the nodes in different directed graphs. For example, the third node in this application refers to each node in the behavior directed graph, and the fourth node in this application refers to each node in the semantic directed graph.
[0097] The first directed edge and the second directed edge are only used to distinguish the edges in different directed graphs. For example, the first directed edge in this application refers to each edge in the behavior directed graph, and the second directed edge in this application refers to each edge in the semantic directed graph.
[0098] Specifically, in the model training stage, as Figure 4 shown in, after the terminal obtains the behavior directed graph of the target object and the semantic directed graph of the item set, the terminal can fuse the behavior directed graph of the target object and the semantic directed graph of the item set, that is, the terminal can splice the third node and the first directed edge in the behavior directed graph of the target object with the fourth node and the second directed edge in the semantic directed graph of the item set respectively to obtain a spliced directed graph, and weight the weight of the first directed edge in the behavior directed graph with the semantic correlation weight of the second directed edge in the semantic directed graph to obtain a splicing weight , further, the terminal can obtain a fused directed graph based on the spliced directed graph and the splicing weight . Thus, by constructing a semantic network based on the semantic correlation weights between the nodes of each item in the item set, the correlation of the content between the nodes can be effectively reflected. Furthermore, by fusing the weighted co-occurrence behavior network and the semantic network between items, the fused network can reflect the preferences in the user browsing behavior through the weights of the edges, enabling the method provided in the embodiments of this application to maintain a higher item recommendation efficiency while effectively solving the problem of underfitting of the embedding features of the item nodes with insufficient exposure caused by the sparsity of user browsing data.
[0099] In one embodiment, the step of splicing the third node and the first directed edge in the behavior directed graph of the target object with the fourth node and the second directed edge in the semantic directed graph of the item set respectively to obtain a spliced directed graph includes:
[0100] Take the union of the third node in the behavior directed graph and the fourth node in the semantic directed graph, and use the union as the spliced node set;
[0101] Take the union of the first directed edge in the behavior directed graph and the second directed edge in the semantic directed graph, and use the union as the spliced directed edge set;
[0102] Based on the spliced node set and the spliced directed edge set, construct a spliced directed graph.
[0103] Specifically, in the model training stage, as Figure 4 shown, after the terminal obtains the behavior directed graph of the target object and the semantic directed graph of the item set, the terminal can fuse the behavior directed graph of the target object and the semantic directed graph of the item set. That is, the terminal can splice the third node and the first directed edge in the behavior directed graph of the target object with the fourth node and the second directed edge in the semantic directed graph of the item set respectively to obtain a spliced directed graph. That is, the terminal can take the third node in the behavior directed graph and the fourth node in the semantic directed graph of the union , and use the union as the spliced node set ; at the same time, the terminal takes the first directed edge in the behavior directed graph and the second directed edge in the semantic directed graph of the union , and use the union as the spliced directed edge set ; the terminal constructs a spliced directed graph based on the spliced node set and the spliced directed edge set . Thus, by constructing a semantic network based on the semantic correlation weights between the nodes of each item in the item set, the correlation of the content between the nodes can be effectively reflected. Furthermore, when the weighted co-occurrence behavior network and the semantic network between items are fused later, the fused network can reflect the preferences in the user browsing behavior through the weights of the edges. This enables the method provided in the embodiments of the present application to maintain a higher item recommendation efficiency while effectively solving the problem of underfitting of the embedding features of the exposure-insufficient nodes caused by the sparsity of user browsing data.
[0104] In one embodiment, the step of performing a random walk for a preset number of times according to the weights of the edges in the fused directed graph with each node in the fused directed graph as the starting point includes:
[0105] Obtain the preset length of the random walk;
[0106] With each node in the fused directed graph as the starting point, perform a random walk of the preset length according to the weights of the edges in the fused directed graph until the number of times of the random walk of the preset length reaches the preset number, and stop the random walk to obtain random walk sequences with different starting points.
[0107] Among them, the preset length can be preset as L.
[0108] Specifically, after the terminal fuses the behavioral directed graph of the target object and the semantic directed graph of the item set to obtain the fused directed graph, the terminal can obtain the preset length of the random walk, and starting from each node of the fused directed graph, perform a random walk of the preset length according to the weights of the edges in the fused directed graph until the number of times of the random walk of the preset length reaches the preset number of times, and stop the random walk to obtain random walk sequences with different starting points. For example, the terminal starts from each node in the fused directed graph That is, for as the starting point, perform n random walks of length L to obtain a set of sequences of random walks for each node, and use the initial large language model to fit the sequences of random walks. That is, the terminal can train the initial large language model m times based on the set of sequences of random walks for each node to obtain the trained large language model. Among them, both the length L and the number of times n are preset hyperparameters, and the probability of each walk is proportional to the weight of the adjacent edges of each node. Thus, since the random walk sequences obtained by the method provided in this application not only consider the user's behavior sequence, but also incorporate the semantic correlation between nodes, it is very friendly to the user data sparsity problem and cold start problem that are prone to occur in the traditional method. That is, the method provided in the embodiments of this application can effectively solve the problem of underfitting of the embedding features of the exposure-insufficient nodes caused by the sparsity of user browsing data while maintaining a higher item recommendation efficiency, and further bring a better item (video) recommendation experience to the user, bringing convenience to the user.
[0109] In one embodiment, the step of training the initial large language model based on each random walk sequence to obtain the large language model includes:
[0110] Through the embedding layer in the initial large language model, based on the initial semantic embedding representation and the random walk sequence, determine the semantic embedding representation of each item identifier in the random walk sequence;
[0111] Through the decoding layer of the initial large language model, determine the sequence embedding representation of the random walk sequence based on the semantic embedding representation;
[0112] Through the output layer of the initial large language model, determine the output probability value based on the sequence embedding representation until the number of training times reaches the preset number of training times, and stop training to obtain the large language model.
[0113] Specifically, in the model training stage, such as Figure 4As shown in [reference], the terminal can determine the semantic embedding representations of each item identifier in the random walk sequence through the initialized embedding layer in the initial large language model (i.e., the recommendation model), that is, based on the initial semantic embedding representation and the random walk sequence, and determine the sequence embedding representation of the random walk sequence through the decoding layer of the initial large language model based on the semantic embedding representation. For example, the terminal uses the decoding layer of the initial large language model to perform residual connection on the semantic embedding representations of each item identifier in the random walk sequence and the random walk sequence to obtain the residual features after the residual connection, and determines the sequence embedding representation (i.e., the predicted embedding representation) of the random walk sequence based on the residual features; further, the terminal passes through the output layer of the initial large language model, that is Figure 4 the "global item ID probability output layer" in [reference], determines the output probability value P based on the predicted sequence embedding representation, and stops training until the number of training times reaches the preset number of training times n to obtain the trained large language model, that is, the recommendation model.
[0114] In this embodiment, by integrating the user's behavior sequence and the semantic relevance of the content between nodes into the same graph structure, it is ensured that the generated random walk sequence can not only reflect the items frequently browsed by the user, but also appropriately display those items less visited by the user through semantic relevance. Furthermore, training the model using the above random walk sequence can improve the diversity of item recommendations while enhancing the convergence speed of model training.
[0115] In one embodiment, the step of determining the output probability value based on the sequence embedding representation includes:
[0116] Determine the correlation scores between the sequence embedding representation and the semantic embedding representations of the item identifiers of each item in the item set respectively;
[0117] Determine the sum value of each of the correlation scores;
[0118] Determine the ratio between each of the correlation scores and the sum value respectively, and determine the target ratio from the ratios as the output probability value.
[0119] Among them, the correlation score in this application refers to the correlation score between the sequence embedding representation predicted by the model and the semantic embedding representations of each item in the item set.
[0120] The target ratio refers to the maximum value selected from each ratio, or the top k ratios selected after sorting according to the size of the ratios. That is, the target ratio in this application can be one or multiple.
[0121] Specifically, as Figure 4 shown in [reference], when in the model training stage, the terminal passes through the output layer of the initial large language model, that is Figure 4The "global item ID probability output layer" in it determines the output probability value P based on the predicted sequence embedding representation. Training stops until the number of training times reaches the preset number n, and the trained large language model, i.e., the recommendation model, is obtained. Among them, the calculation method for determining the output probability value P based on the predicted sequence embedding representation can be using the following formula (1):
[0122] } (1)
[0123] Among them, represents the probability value corresponding to the t-th item identifier in the predicted input sequence, represents the predicted embedding feature corresponding to the t-th item identifier in the predicted input sequence, represents the feature of each item (0 - X) in the item set in the embedding layer, … respectively represent the predicted embedding feature corresponding to the t-th item identifier in the predicted input sequence and the feature of each item (0 - X) in the item set in the embedding layer the correlation score between them, represents the (global) item set, represents the predicted embedding feature corresponding to the t-th item identifier in the predicted input sequence and the feature of each item (0 - X) in the item set in the embedding layer the sum value of the correlation scores between them. It can be understood that the basic structure during the training of the large language model in this application can also be other structures. For example, as Figure 6 shown, it is a schematic diagram of the basic structure during model training. When the recommendation model (large language model) shown in Figure 6 is trained, the terminal can directly use the embedding representation of each item ID in the embedding layer of the model structure shown in Figure 6 and find the item corresponding to the item identifier with the closest distance to it through the approximate nearest neighbor algorithm (ANN) for item - item (I2I) recommendation.
[0124] In one embodiment, as Figure 7 shown, the steps of recommending a target item in the item set to a target object according to the semantic embedding representation include:[[]]
[0125] Step 702, searching for a preset number of target embedding features whose distance from the semantic embedding representation satisfies the distance threshold;
[0126] Step 704, taking the item corresponding to the item identifier of the target embedding feature as the target item recommended to the target object.
[0127] Specifically, after the training of the model is completed, the terminal does not need to obtain the behavior sequence of the target object. The terminal can directly determine the semantic embedding representations matching each item in the item set through the trained large language model, that is, respectively search for the semantic embedding representations matching the item identifiers of each item in the item set through the embedding layer of the trained large language model, and further search for a preset number of target embedding features in the database or the embedding layer of the trained recommendation model whose distances between the semantic embedding representations matching the item identifiers of each item meet the distance threshold, and use the items corresponding to the item identifiers of the target embedding features as the target items to be recommended. For example, when the recommendation model shown in Figure 4 is trained, the terminal can directly utilize the embedding representations of each item ID in the embedding layer of the model structure shown in Figure 4 and find the items corresponding to the item identifiers with the closest distance through the approximate nearest neighbor algorithm (ANN) for item-item (I2I) recommendation. Thus, by providing different item recommendation methods, relevant items can be recommended to users from different dimensions, improving the diversity of item recommendations, and further enhancing the user's item recommendation experience and bringing convenience to the user.
[0128] In one embodiment, after training the initial large language model based on each random walk sequence to obtain the large language model, the method further includes:
[0129] Obtaining the behavior sequence of other objects; the behavior sequence is composed of the item identifiers of each item interacted with by other objects;
[0130] Inputting the behavior sequence into the large language model so that the embedding layer of the large language model sequentially searches for the semantic embedding representations corresponding to the item identifiers in the behavior sequence;
[0131] Based on the behavior sequence and the obtained semantic embedding representations, determining the sequence embedding features of the behavior sequence;
[0132] Determining the target items to be recommended based on the sequence embedding features of the behavior sequence and recommending the target items to other objects.
[0133] Specifically, as shown in Figure 4 is a schematic diagram of the basic structure of the recommendation model. After the terminal obtains the behavior sequence of other objects, the terminal can input the behavior sequence into the one shown in Figure 4In the large language model shown, the embedding layer of the large language model can sequentially search for semantic embedding representations corresponding to each item identifier in the behavior sequence, so as to obtain the semantic embedding representations corresponding to each item identifier in the behavior sequence of other objects. Further, the terminal can perform decoding processing on the semantic embedding representations corresponding to each item identifier in the behavior sequence and the behavior sequence through a pre-trained recommendation model, that is, the large language model, to obtain the predicted sequence embedding features of the behavior sequence, and determine the item identifier to be recommended based on the predicted sequence embedding features of the behavior sequence through the pre-trained recommendation model, and recommend the item corresponding to the item identifier to other objects as the target item.
[0134] In this embodiment, since the semantic embedding representations corresponding to each item identifier are pre-extracted and input into the model for training and iteration, and the behavior sequence of other objects in this application is composed of the item identifiers of each item interacted by other objects, the semantic embedding representations corresponding to each item identifier can be quickly and accurately determined directly based on the item identifiers in the behavior sequence, so that the sequence embedding features of the behavior sequence can be predicted directly based on the behavior sequence and the determined semantic embedding representations, and the target item to be recommended can be determined based on the predicted sequence embedding features of the behavior sequence, effectively improving the speed of processing the behavior sequence, and the resource consumption is relatively small. Furthermore, the processing efficiency of item recommendation is improved, which can bring a better item recommendation experience to users and bring convenience to users.
[0135] This application also provides an application scenario that applies the above item recommendation method. Specifically, the application of the item recommendation method in this application scenario is as follows:
[0136] In the process of a user interacting with a multimedia information platform, the above item recommendation method can be adopted. For example, Figure 8 As shown, it is a schematic diagram of the interface display on the product side. That is, when the user (the operating object) has browsed some products to be purchased in a shopping application, the shopping application can recommend some relevant personalized products for the user. That is, the background server of the shopping application can obtain, for example, Figure 8The identifiers of the items that the user recently viewed (19 minutes ago) as shown, namely Item 1, Item 2, and Item 3, are serialized to obtain the user's behavior sequence. Further, the background server can construct a directed graph of the user's behavior based on the user's behavior sequence and the behavior sequences of other users. Meanwhile, the terminal can obtain the initial semantic embedding representations of each item in the item set and construct a semantic directed graph based on the semantic correlation weights determined by the initial semantic embedding representations of each item. Furthermore, the terminal can fuse the directed graph of the target object's behavior and the semantic directed graph of the item set to obtain a fused directed graph, and perform a preset number of random walks starting from each node of the fused directed graph according to the weights of each edge in the fused directed graph to obtain random walk sequences starting from different nodes. Further, the terminal can train the initial large language model based on each random walk sequence to obtain a large language model, and determine the semantic embedding representations matching each item in the item set through the trained large language model, and recommend the target item, i.e., the target commodity, in the item set to the target object according to the semantic embedding representations. Thereby, the shopping recommendation experience of the user can be improved, bringing convenience to the user.
[0137] The method provided in the embodiments of the present application can be applied to various scenarios of commodity recommendation, item recommendation, and video recommendation. The following takes the scenario of a user interacting with a multimedia information platform as an example to illustrate the item recommendation method provided in the embodiments of the present application.
[0138] Among them, embedding representation: Generate a distributed representation of the input "object" based on a neural network model. The main role of the embedding representation is to transform the high-dimensional sparse vector of the original object into a low-dimensional and dense vector, so that these low-dimensional and dense vectors can express certain features of the corresponding object, and at the same time the distance between the vectors can reflect the similarity between the objects, thereby facilitating the processing of downstream models, especially deep learning models. These input "objects" are usually words, entities, semantic labels, nodes in a graph, etc. in natural language processing.
[0139] Graph embedding representation: This type of method is also known as vertex representation learning. Generally speaking, graph embedding is a method of mapping nodes in a graph into low-dimensional and dense embedding representations. Graph embedding needs to capture the topological structure features of the graph, such as the relationship between vertices and vertices, and the relationship between subgraphs and edges. Currently, in the vast majority of research, the process of graph embedding representation is equivalent to the learning process of dimensionality reduction representation of nodes in the graph. Note that when the present invention refers to a graph composed of nodes and edges, it is often also called a network, and the two are interchangeable.
[0140] The Transformer is a well-known deep learning model that has been widely used in various fields such as natural language processing (NLP), computer vision (CV), and speech processing. When the Transformer was first proposed, it was a sequence-to-sequence model for machine translation. The Transformer has become the preferred architecture in natural language processing, especially when used as a pre-trained model. The classic Transformer model is a sequence-to-sequence model composed of an encoder and a decoder. Both the encoder and the decoder are made up of a stack of identical Transformer blocks. Each Transformer block mainly consists of a multi-head self-attention module and a feed-forward neural network.
[0141] Batch Normalization: Also known as batch normalization. During the training process of neural networks, to solve the problem of gradient explosion, Batch Normalization (BN) was introduced. The essence of Batch Normalization is to adjust the data distribution to an approximate normal distribution.
[0142] Language Model: A language model is a model used to model natural language, and its purpose is to predict the next word or character of a given text sequence. Language models can be used in various natural language processing tasks, such as semantic extraction of text, text generation, machine translation, speech recognition, etc. Existing research shows that pre-trained language models (PTMs) based on the Transformer can achieve good results in various tasks of natural language processing. Currently, relatively common pre-trained language models include large pre-trained models such as Bert and GPT series models.
[0143] Large Language Model refers to a natural language processing model with a large number of parameters and trained using a vast amount of data. The training process of large language models usually adopts the unsupervised learning method, that is, training the model through a large-scale text corpus, so that the model can learn the probability distribution and language rules of the language. During the training process, large language models usually optimize the model parameters by maximizing the prediction probability of the next word as the objective function. Currently, the most representative large language models are the GPT series models of OpenAI. They adopt the Transformer model structure and are trained on a large-scale corpus, and can generate high-quality natural language texts such as articles and conversations.
[0144] Token: In large language models, a token is the smallest unit in text, usually a word, punctuation mark, or other symbol. In natural language processing, tokenization is one of the basic steps of breaking text into discrete units. In large language models, each token is assigned a unique integer ID, which is usually called the token ID. These token IDs can be used to represent each token in the text and are used for training and evaluating large language models.
[0145] In large language models, the meaning of tokens is very important. Since the training and prediction of the model are both based on tokens, for different tasks and application scenarios, different tokenization methods and token sets need to be selected. For example, in text classification tasks, words are usually used as tokens. In machine translation tasks, subwords are usually used as tokens. In large language models, tokens can also be used to represent the context information in the text. For example, in the BERT model, each token is assigned a position encoding to represent its position in the text. In this way, the model can utilize the context information to better understand the text and generate more accurate predictions. In short, tokens are the basic units in large language models, and their selection and representation methods have an important impact on the performance and effectiveness of the model.
[0146] Scaling Laws refer to the phenomenon in the fields of machine learning and deep learning that the performance of the model (such as accuracy, loss, etc.) improves significantly as the model scale (such as the number of parameters, the amount of training data, computing resources, etc.) increases.
[0147] I2I and U2I: In recommendation systems, the recall method of finding recommended items for users is called U2I recall (i.e., user-to-item, abbreviated as U2I); while the recall method of using an item to recommend similar items is called I2I recall (i.e., item-to-item, abbreviated as I2I).
[0148] Approximate Nearest Neighbor: Approximate Nearest Neighbor (ANN) is a fast nearest neighbor search algorithm for high-dimensional data. It can perform efficient nearest neighbor searches on large-scale datasets and is commonly used in fields such as computer vision and natural language processing.
[0149] In the traditional approach, in the field of recommendation systems, graph neural networks (GNNs) are a widely used algorithmic technique. GNN models mainly obtain embedding representations at the node, edge, or graph level through recursive message passing and aggregation mechanisms between nodes to meet the requirements of different downstream tasks. That is, in traditional recommendation systems, the GNN algorithm mainly explores how to perform message passing on graph-structured data, that is, iteratively update the node's embedding representation by aggregating the embedding representations of neighboring nodes. However, the model structure of this message propagation method is relatively weak in encoding ability compared to large language models.
[0150] Many studies have explored applying large language models (LLMs) to recommendation systems. These methods usually transform the recommendation task into a natural language processing task and utilize the powerful text understanding ability of LLMs for recommendation. That is, combine the text understanding ability of LLMs with the topological structure information of the graph to improve the effect of the model in dealing with graph-related tasks. Among them, the method of using a large language model as a predictor is relatively similar to the method provided in this application. This method usually involves flattening the graph data into a text description so that the LLM can directly process the graph data through the text sequence.
[0151] Currently, there are relatively few methods that directly use the large model structure to perform sequence modeling on item IDs. Although there are also related solutions that propose an autoregressive model using a similar large model architecture and train it by predicting the next item in the browsing item sequence. The autoregressive architecture of this generative model has greatly improved the performance of the recommendation system and shows the scale law effect, that is, an increase in the number of model parameters will significantly improve the performance. For example, after replacing the binary cross-entropy loss function in the SASRec model with the cross-entropy loss function (or its approximate negative sampling cross-entropy loss function), the model effect will be significantly improved. In addition, once the loss function of SASRec is replaced with the cross-entropy loss function based on the output probabilities of all items, the model structure of SASRec will be very similar to that of GPT2 and can be regarded as directly applying the structure of GPT2 to sequence recommendation.
[0152] The problems existing in the traditional approach include:
[0153] Most graph neural network (GNN) models tend to construct a bipartite graph or heterogeneous network of users and items and obtain the embedding representations of relevant nodes through message passing and aggregation mechanisms on these networks. However, in an actual recommendation system, hundreds of millions or even billions of user behavior data may be generated every day. It is unrealistic to store user nodes of such a large scale in the graph model. Therefore, for large-scale recommendation systems, a more practical method is to only model the item nodes in the item network constructed by user behavior.
[0154] When attempting to combine graph data with large language models (LLMs), there are some drawbacks to using the flattening prediction method: By converting graph data into text sequences, this method may lose the original structural information of the graph, resulting in the model output being overly dependent on the pre-trained large language model, emphasizing semantic information rather than truly reflecting the user's behavior patterns in the graph structure. Additionally, when dealing with large-scale graph data, the flattening method may require a large amount of computing resources to generate and process text sequences, thus reducing the prediction efficiency of the model.
[0155] The technical solution provided in this application can address these issues simultaneously through the design of an innovative complete process:
[0156] First, in this application, the new structure of the large language model can fit the paths obtained by probabilistic walks on the directed weighted graph, and this model structure is decoupled from the input topological graph. Therefore, with the progress of large language model technology, the method provided in this application can also continuously improve the ability to fit the random walk paths in the graph.
[0157] Second, in the recommendation scenario, the method provided in this application combines the weighted co-occurrence behavior with the semantic relevance between items. This fused network can reflect the preferences in the user's browsing behavior through the weights of the edges. By incorporating the co-occurrence weights of behaviors into the directed edges in the graph, the high-frequency item nodes with higher display probabilities will be more likely to appear in the random walk paths. At the same time, through semantic relevance, the semantic nearest neighbor nodes of each node will fuse the low-weight semantic edges into this directed graph, so that the nodes that are relatively sparse in the behavioral data can also be included in the walk paths, ensuring that all nodes can be encoded into the same embedding space.
[0158] That is, in this application, in order to effectively solve the problem of underfitting of the embedding features of under-exposed nodes caused by the sparsity of user browsing data, this application proposes a method for I2I recommendation by combining random walks in a directed graph with a large language model structure. It adopts a large language model architecture with stronger encoding ability. First, it uses an autoregressive method to fit the random walk sequence of the graph through the large language model, so as to learn a better graph node embedding representation than traditional GNNs. Secondly, the method provided in this application integrates the user's behavior sequence and semantic relevance into the same graph structure, ensuring that the generated random walk sequence can not only reflect the items frequently browsed by users, but also appropriately display those items less visited by users through semantic relevance. This method enables the method provided in this application to effectively solve the problem of underfitting of the embedding features of under-exposed nodes caused by the sparsity of user browsing data while maintaining the recommendation efficiency. In addition, through the training method proposed in this application, the training speed of the recommendation model based on the large language model structure can be greatly accelerated, so that the model can be more widely applied to various sequence recommendation tasks under limited computing resources.
[0159] That is, this application proposes a new method to use the autoregressive structure of the large language model to fit the random walk sequence in the weighted item network. Through this method, the recommendation system can more effectively capture the sequential behavior characteristics of users and the semantic relevance between items, and can effectively alleviate the cold start problem caused by the sparsity of user browsing data. For newly added item nodes, without relying on user interaction information, they can be integrated into the existing network through the semantic features of the nodes. This can not only achieve efficient fitting of the random walk sequences of all item nodes in the network, but also obtain the embedding representations of all nodes, so as to perform item-to-item recommendation (I2I recommendation), effectively improving the user's item recommendation experience and bringing convenience to users.
[0160] On the product side, the technical solution provided in this application is applicable to various personalized item recommendation scenarios such as social, gaming, ticketing, and shopping.
[0161] On the technical side, the schematic diagram of the principle during the training of the recommendation model proposed in this application can be the process framework as shown in Figure 4 the following. That is, the main training steps of the large language model proposed in this application are as follows:
[0162] 1. First, fuse the co-occurrence network of user behaviors and the semantic nearest neighbor network of item nodes into a graph.
[0163] 2. Then, starting from each node in this fused network, perform multiple probability-based random walks according to the weights on the edges.
[0164] 3. Finally, fit the random walk sequences obtained above through a large language architecture model, so as to obtain the embedding representations of each node in these sequences.
[0165] After completing the above training steps, item-to-item (I2I) recommendations can be made using the nearest neighbor method based on the embedding representations of items.
[0166] The specific methods of the above training steps will be described in detail below:
[0167] The first stage: network fusion
[0168] (a) Directed weighted co-occurrence graph (directed graph) of user behavior
[0169] Suppose there is a sequence of n items browsed by a certain user , where represents the i-th item in the sequence. In this method, the method provided in this application will preset the window size of the co-occurrence window to an empirical value of W to determine which nodes in the sequence need to establish directed edges. Specifically, the edge-building process of the terminal in this embodiment can be formally expressed as follows: for a subsequence of , where , will be connected to the subsequent nodes to obtain directed edges, that is . Among all the user behavior sequences obtained by the terminal, directed edges within the co-occurrence window size W are established according to the above method, and the number of times these co-occurring directed edges appear is recorded as the weight of the edges in the graph. Finally, a directed weighted network based on user behavior is obtained . Among them, is the node set, is the directed edge set, is the weight set. Specifically , represents the co-occurrence times of the directed edge .
[0170] (b) Semantic network of items (directed graph)
[0171] Currently, there are many open-source multi-modal large language models that can fuse image features and text features to generate text based on image information. Briefly, the terminal in this application will input the image information and text information of each item into such a multi-modal large language, and obtain the fused semantic embedding representations of the image and text from the last Transformer decoder of this model.
[0172] After obtaining the semantic embedding representations of each item node, the top k nearest neighbor nodes with the closest semantic distance to each node can be obtained through the nearest neighbor algorithm, so as to obtain the directed edges composed of the neighbor nodes that are semantically most relevant to each node, and thus obtain a directed weighted graph based on semantics , where is the set of nodes, is the set of directed edges, is the set of semantic relevance weights. , indicating the directed edge . It can be understood that the semantic relevance weight in the implementation of this application can be the cosine distance between the embedding representation vectors of two nodes.
[0173] (c) Fusion of item weighted co-occurrence graph and semantic network
[0174] The terminal needs to fuse the two networks obtained in the above two steps into one network, which can be formally expressed as: . Among them, the weight is the linear weighting of the above two weights, that is is a set empirical value used to prevent the behavior weight from being too large.
[0175] The second stage: Random walk in network fusion
[0176] Next, the terminal performs n random walks of length l starting from v∈V in G=(V,E,W), where both the length l and the number of times n are preset hyperparameters, and the probability of each walk is proportional to the weight of the adjacent edges of each node.
[0177] Finally, the terminal can obtain a set of sequences of random walks for each node and fit the sequences of random walks through the next large language model.
[0178] The third stage: Autoregressive training based on the large language model architecture
[0179] In this stage, this application will use the structure of the multi-objective large language model shown in Figure 4 to perform autoregressive training, so that the model learns the characteristics of the sequences of random walks in the graph. That is, the working principle of the large language model proposed in this application is as follows: The large language model will input the representation of the item sequence x 1 x n of n items browsed by the user into the model, where x i represents the i-th item in the sequence. During the training process, the training label is the corresponding x 2 x (n+1) , that is, x iThe next word x corresponding to each position (i+1) . Briefly, the training process is as follows: According to the sequence x 1 x k , based on the embedding representations output by the Transformer layer corresponding to the first k items in x, predict the probability distribution of the next item x (k+1) .
[0180] The main training steps of the model provided in this application are as follows:
[0181] First, obtain the multi-modal embedding features of each item in the item sequence, and use the item multi-modal embedding features for the initialization of the model embedding layer. Compared with the random initialization of the embedding layer, initializing with multi-modal semantic features is more conducive to the convergence of the model. Then, through the "global item ID probability output layer" as shown in Figure 4 , use formula (1) to perform probability prediction.
[0182] The optimization objective of the model during training is to maximize the probability of the next item given the item sequence x<t. Suppose there is an item sequence of length n , where x t represents the item ID at the t-th position. The objective function in this application can be expressed as shown in the following formula (2):
[0183] (2)
[0184] where x<t represents the sequence of item IDs at the first to t-1 positions.
[0185] Brief description of the model structure of the large language model provided in this application:
[0186] The main architecture of the recommendation model based on the large language model used in this application is as shown in Figure 4 , and the main functions of each layer in this model structure include:
[0187] 1. Item sequence input layer: As shown in Figure 4 , it belongs to the ID sequence of n items. If the number of a sequence is less than n, the subsequent unfilled positions need to be filled with masks.
[0188] 2. Embedding Layer: The output of this layer is the item ID embedding representation, so as to obtain the embedding representation vector of each item ID and input it into the subsequent transformer layer.
[0189] Item ID Embedding Layer: This layer is used to query the embedding representation vectors of each item ID that have been pre-extracted by a multimodal large language model. These embeddings are obtained by other large language models through feature extraction of the text and images of the items, that is, the trained multimodal embeddings of each item are extracted.
[0190] Currently, there are many open-source multimodal large language models that can fuse image features and text features to perform text generation based on image information. In short, the technical solution in this application can pre-input the image information and text information of each item into such a multimodal large language model, and obtain the fused embedding representation of the image and text from the last Transformer decoder of this model, and use this fused embedding representation as the multimodal embedding feature of each item to initialize the embedding layer of the model as shown in Figure 4 In order to accelerate the convergence of the model and make the model training process more stable, it is necessary to regularize these fused embedding features of the items through L2 norm (L2 normalization). At the same time, the weights of these multimodal embedding features will be adjusted and changed during the training process of the model. Here, it is mainly to accelerate the convergence of the model.
[0191] 3. Transformer Layer: This layer is mainly a deeper decoder constructed by stacking multiple Transformer decoding blocks with a masking mechanism to improve the performance and generalization ability of the model. Each decoder block mainly includes the following parts:
[0192] a) Masked Multi-Head Self-Attention: This part calculates the multi-head self-attention of the input sequence to capture the dependencies between different positions in the input sequence.
[0193] b) Residual Connection and Layer Normalization: This part accelerates the training of the model and improves the performance of the model by performing a residual connection between the output of the previous part and the input sequence and performing layer normalization on the result of the residual connection.
[0194] c) Feedforward Neural Network: This part usually consists of two connected fully-connected layers. By calculating the fully-connected layers for the representations at each position, the feature representation ability at each position is enhanced. Currently, in the Transformer decoder block, the sparse mixture-of-experts network is often used to replace the traditional Feedforward Neural Network in large language models.
[0195] That is, the terminal inputs the ID sequence of each item browsed by each user into the large language model as shown Figure 4 . After that, the embedding representation vector output by the last Transformer decoder block (i.e., the Nth Transformer decoder block) is extracted from the corresponding position of the last token in the input, and this embedding representation vector is used as the embedding representation (embedding hidden state) predicted by the large language model for the behavior sequence of this user.
[0196] 4. Model Output Layer: This layer also calculates the embedding representation h output by the Transformer layer through the method of embedding weight binding k and the correlation score s i with the feature e of each item i in the embedding layer. Specifically, this process can be formally described as follows:
[0197] The sequence x 1 x n of n items browsed by the user is input into the model, where x i represents the i-th item in the sequence. Correspondingly, the training label is x 2 x (n+1) . With the input being the item sequence x 1 x k , the model will input the embedding representations h k corresponding to the first k items of the input sequence, and calculate the probability distribution of generating the next item I (k+1) by this fully-connected layer. This layer also outputs a dimension of |X| through the method of embedding weight binding, that is, a score r(h k ,e i ) will be output for each item i ∈ X in |X|, where the embedding representations corresponding to the first k items in the Transformer layer output are h k , the feature of item i in the embedding layer is e i , and the correlation score s(h k ,e i ) = hk ×e i 。
[0198] represents the probability of predicting \(x\) after a given sequence of item IDs where \(x < t\), t and is the item ID with the maximum probability calculated by the softmax method, and the specific calculation method is shown in the aforementioned formula (1).
[0199] That is, the goal in the training stage is to maximize the product of these conditional probabilities, i.e., to minimize the negative log-likelihood loss \(L\). During the training process, the model calculates the difference between the predicted item ID and the actual next item ID, and this difference is calculated by the cross-entropy method.
[0200] Applications in recommendation include U2I recommendation and I2I recommendation:
[0201] U2I recommendation: After the model is trained, for any user input sequence \(x\) 1 \(x\) k , the predicted embedding representation \(h\) can be obtained at the output of the Transformer layer. k Only need to use the approximate nearest neighbor algorithm to find the top \(k\) items with the closest distance to \(h\) in the embedding layer of the model for recommendation. k
[0202] I2I recommendation (can be recommended after the model is trained): After the large language model is trained, the initialized embedding representation of each item ID in the embedding layer shown in Figure 4 can be directly used. Through the approximate nearest neighbor algorithm (ANN), find the item with the closest distance to it for item-item (I2I) recommendation.
[0203] It can be understood that the "global ID probability output layer" in this application adopts the method of binding embedding weights in the ID probability output layer. The main purpose of this is to reduce the model parameters, thereby reducing the fitting difficulty of the model. A separate dense layer can also be used to calculate the output probability of each ID. In addition, the two-stage training method of this application can not only be applied to the recommendation model based on the large language model, but also be directly applied to the pre-training fitting of the large language model, thereby accelerating the fitting speed of the large language model. That is Figure 4 the output layer in
[0204] When constructing a semantic correlation network, a semantic network can be constructed only using text information, or a network can be constructed only using the embedding features of images, as long as the correlation of the content between nodes can be reflected.
[0205] The beneficial effects produced by the technical solution of this application include:
[0206] 1. This application proposes a method of using the structure of a large language model to fit the random walk sequence of item IDs in a graph in an autoregressive manner.
[0207] 2. In the fitting process, this application not only considers fitting the user's behavior sequence, but also incorporates the semantic correlation of nodes. Therefore, it is very friendly to the common user data sparsity problem and cold start problem in the traditional method.
[0208] It should be understood that although the steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0209] Based on the same inventive concept, the embodiments of this application also provide an item recommendation device for implementing the above-mentioned item recommendation method. The implementation solutions provided by this device to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the item recommendation device provided below can refer to the limitations on the item recommendation method in the above text, and will not be repeated here.
[0210] In one embodiment, as Figure 9 shown, an item recommendation device is provided, including: a fusion module 902, a random walk module 904, a training module 906, a determination module 908, and a recommendation module 910, where:
[0211] The fusion module 902 is used to fuse the behavior directed graph of the target object and the semantic directed graph of the item set to obtain a fused directed graph; the semantic directed graph is constructed based on the semantic correlation weights of the nodes of each item in the item set.
[0212] A random walk module 904, configured to start from each node of the fused directed graph, perform a preset number of random walks according to the weights of the edges in the fused directed graph, and obtain random walk sequences starting from different said starting points.
[0213] A training module 906, configured to train an initial large language model based on each of the random walk sequences to obtain a large language model.
[0214] A determination module 908, configured to determine semantic embedding representations matching each item in the item set through the large language model.
[0215] A recommendation module 910, configured to recommend a target item in the item set to the target object according to the semantic embedding representation.
[0216] In one embodiment, the device further includes: an acquisition module, configured to acquire a behavior sequence of the target object; the behavior sequence is composed of item identifiers of each item interacted with by the target object; each of the item identifiers in the behavior sequence belongs to the identifiers of each item in the item set; a construction module, configured to perform serialization processing on each of the item identifiers to obtain the behavior sequence of the target object; a storage module, configured to construct the behavior directed graph based on the behavior sequence.
[0217] In one embodiment, the construction module is further configured to construct nodes and directed edges corresponding to the behavior sequence based on a preset window value of a sliding window and each of the item identifiers in the behavior sequence; the determination module is further configured to determine the weight of the directed edge based on the construction times of the directed edge; the construction module is further configured to construct the behavior directed graph based on the nodes, the directed edges, and the weights.
[0218] In one embodiment, the device further includes: a feature extraction module, configured to perform feature extraction on each item in the item set through a multimodal model to obtain initial semantic embedding representations of each of the items; a construction module, configured to construct the semantic directed graph based on semantic correlation weights determined by the initial semantic embedding representations of each of the items.
[0219] In one embodiment, the determination module is further configured to determine adjacent nodes of each of the nodes based on the initial semantic embedding representations of each of the items; the construction module is further configured to construct directed edges corresponding to each of the nodes based on each of the nodes and the adjacent nodes of each of the nodes; the determination module is further configured to determine the semantic correlation weight of the directed edge; the construction module is further configured to construct the semantic directed graph based on each of the nodes, the directed edges, and the semantic correlation weights of the directed edges.
[0220] In one embodiment, the device further includes: an acquisition module, configured to acquire a first node and a second node connected by the directed edge; the determination module is further configured to determine a cosine distance between the semantic embedding representation of the first node and the semantic embedding representation of the second node, and use the cosine distance as the semantic correlation weight of the directed edge; use the product of the semantic embedding representation of the first node and the semantic embedding representation of the second node as the semantic correlation weight of the directed edge.
[0221] In one embodiment, the device further includes: a splicing module, configured to splice a third node and a first directed edge in the behavioral directed graph of the target object with a fourth node and a second directed edge in the semantic directed graph of the item set respectively to obtain a spliced directed graph; a processing module, configured to perform weighted processing on the weight of the first directed edge in the behavioral directed graph and the semantic correlation weight of the second directed edge in the semantic directed graph to obtain a splicing weight; and obtain a fused directed graph based on the spliced directed graph and the splicing weight.
[0222] In one embodiment, the processing module is further configured to take the union of the third node in the behavioral directed graph and the fourth node in the semantic directed graph, and use the union as the spliced node set; take the union of the first directed edge in the behavioral directed graph and the second directed edge in the semantic directed graph, and use the union as the spliced directed edge set; the device further includes: a construction module, configured to construct the spliced directed graph based on the spliced node set and the spliced directed edge set.
[0223] In one embodiment, the device further includes: an acquisition module, configured to acquire a preset length of random walk; the random walk module is further configured to start from each node of the fused directed graph, perform a random walk of the preset length according to the weights of the edges in the fused directed graph, and stop the random walk until the number of times of the random walk of the preset length reaches a preset number of times to obtain random walk sequences starting from different starting points.
[0224] In one embodiment, the determination module is further configured to, through the embedding layer in the initial large language model, determine the semantic embedding representation of each item identifier in the random walk sequence based on the initial semantic embedding representation and the random walk sequence; determine the sequence embedding representation of the random walk sequence based on the semantic embedding representation through the decoding layer of the initial large language model; the device further includes: a training module, configured to determine an output probability value based on the sequence embedding representation through the output layer of the initial large language model, and stop training until the number of training times reaches a preset number of training times to obtain the large language model.
[0225] In one embodiment, the apparatus further includes: a searching module, configured to search for a preset number of target embedding features whose distances from the semantic embedding representation meet a distance threshold; and use the items corresponding to the item identifiers of the target embedding features as target items recommended to the target object.
[0226] In one embodiment, the apparatus further includes: an obtaining module, configured to obtain a behavior sequence of other objects; the behavior sequence is composed of item identifiers of each item interacted with by the other objects; an input module, configured to input the behavior sequence into the large language model, so that the embedding layer of the large language model sequentially searches for semantic embedding representations corresponding to the item identifiers in the behavior sequence; the determining module is further configured to determine a sequence embedding feature of the behavior sequence based on the behavior sequence and the searched semantic embedding representations; determine target items to be recommended based on the sequence embedding feature of the behavior sequence; and the recommending module is further configured to recommend the target items to the other objects.
[0227] Each module in the above item recommendation apparatus can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above respective modules.
[0228] In one embodiment, a computer device is provided. The computer device can be a terminal or a server. In this embodiment, taking the computer device as a terminal as an example for illustration, its internal structure diagram can be as Figure 10As shown in the figure. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements an item recommendation method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0229] Those skilled in the art can understand that Figure 10 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0230] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0231] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0232] In one embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by the processor, the steps in the above method embodiments are implemented.
[0233] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0234] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.
[0235] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0236] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. An item recommendation method, characterized in that: The method comprises: The behavior directed graph of the target object and the semantic directed graph of the item set are merged to obtain a fused directed graph; the semantic directed graph is constructed based on the semantic relevance weights of the nodes of each item in the item set; Taking each node of the fused directed graph as a starting point, performing a preset number of random walks according to the weight of each edge in the fused directed graph, and obtaining a random walk sequence with different starting points; Training the initial large language model based on each of the random walk sequences to obtain a large language model; Determining, by means of the large language model, a semantic embedding representation matching each item in the item set; According to the semantic embedding representation, a target item in the item set is recommended to the target object.
2. The method according to claim 1, characterized in that Before fusing the behavior directed graph of the target object and the semantic directed graph of the item set to obtain the fused directed graph, the method further includes: Acquire a behavior sequence of the target object; the behavior sequence is composed of item identifiers of items that the target object has interacted with; and each item identifier in the behavior sequence belongs to an identifier of each item in the item set; The behavior directed graph is constructed based on the behavior sequence.
3. The method according to claim 2, characterized in that The constructing the behavior directed graph based on the behavior sequence includes: Based on the preset window value of the sliding window and the identifiers of the items in the behavior sequence, constructing nodes and directed edges corresponding to the behavior sequence; Determining a weight of the directed edge based on the number of times the directed edge is constructed; The behavioral directed graph is constructed based on the nodes, the directed edges and the weights.
4. The method according to claim 1, characterized in that: Before fusing the behavior directed graph of the target object and the semantic directed graph of the item set to obtain the fused directed graph, the method further includes: Extracting features of each item in the item set through a multimodal model to obtain an initial semantic embedding representation of each item; The semantic directed graph is constructed based on the semantic relevance weights determined by the initial semantic embedding representations of the items.
5. The method according to claim 4, characterized in that The step of constructing the semantic directed graph based on the semantic relevance weights determined based on the initial semantic embedding representations of the items comprises: Determining adjacent points of each of the nodes based on the initial semantic embedding representation of each of the items; Based on each of the nodes and the adjacent points of each of the nodes, constructing directed edges corresponding to each of the nodes; Determining a semantic relevance weight of the directed edge; The semantic directed graph is constructed based on the nodes, the directed edges and the semantic relevance weights of the directed edges.
6. The method according to claim 5, characterized in that The determining of the semantic relevance weight of the directed edge comprises: Obtain a first node and a second node connected by the directed edge; Determine a cosine distance between the semantic embedding representation of the first node and the semantic embedding representation of the second node, and use the cosine distance as a semantic relevance weight of the directed edge; or, The product of the semantic embedding representation of the first node and the semantic embedding representation of the second node is used as the semantic relevance weight of the directed edge.
7. The method according to claim 1, characterized in that The step of fusing the behavior directed graph of the target object and the semantic directed graph of the item set to obtain a fused directed graph includes: The third node and the first directed edge in the behavior directed graph of the target object are respectively concatenated with the fourth node and the second directed edge in the semantic directed graph of the item set to obtain a concatenated directed graph; Performing weighted processing on the weight of the first directed edge in the behavioral directed graph and the semantic relevance weight of the second directed edge in the semantic directed graph to obtain a splicing weight; A fused directed graph is obtained based on the spliced directed graph and the splicing weight.
8. The method according to claim 7, characterized in that The step of respectively splicing the third node and the first directed edge in the behavior directed graph of the target object with the fourth node and the second directed edge in the semantic directed graph of the item set to obtain a spliced directed graph includes: Taking the union of the third node in the behavior directed graph and the fourth node in the semantic directed graph, and using the union as a concatenated node set; Taking the union of the first directed edge in the behavioral directed graph and the second directed edge in the semantic directed graph, and using the union as a concatenated directed edge set; The spliced directed graph is constructed based on the spliced node set and the spliced directed edge set.
9. The method according to claim 1, characterized in that: The method of taking each node of the fused directed graph as a starting point and performing a preset number of random walks according to the weight of each edge in the fused directed graph to obtain a random walk sequence with different starting points includes: Get the preset length of the random walk; Taking each node of the fused directed graph as a starting point, a random walk of the preset length is performed according to the weight of each edge in the fused directed graph until the number of random walks of the preset length reaches a preset number, and then the random walk is stopped to obtain a random walk sequence with different starting points.
10. The method according to claim 1, characterized in that The initial large language model is trained based on each of the random walk sequences to obtain a large language model, including: Determining, by means of an embedding layer in the initial large language model, the semantic embedding representation of each item identifier in the random walk sequence based on the initial semantic embedding representation and the random walk sequence; Determining, through a decoding layer of the initial large language model, a sequence embedding representation of the random walk sequence based on the semantic embedding representation; The output probability value is determined based on the sequence embedding representation through the output layer of the initial large language model until the number of training times reaches a preset number of training times, and then the training is stopped to obtain the large language model.
11. The method according to claim 1, characterized in that: The recommending the target item in the item set to the target object according to the semantic embedding representation includes: Find a preset number of target embedding features whose distances from the semantic embedding representation meet a distance threshold; The item corresponding to the item identification of the target embedded feature is used as the target item recommended to the target object.
12. The method according to claim 1, characterized in that After the initial large language model is trained based on each of the random walk sequences to obtain the large language model, the method further includes: Obtaining a behavior sequence of other objects; the behavior sequence is composed of item identifiers of each item that the other objects have interacted with; Inputting the behavior sequence into the large language model so that the embedding layer of the large language model sequentially searches for the semantic embedding representation corresponding to each of the item identifiers in the behavior sequence; Determine a sequence embedding feature of the behavior sequence based on the behavior sequence and the found semantic embedding representation; A target item to be recommended is determined based on the sequence embedding feature of the behavior sequence, and the target item is recommended to the other objects.
13. An item recommendation device, characterized in that: The device comprises: A fusion module, used for fusing the behavior directed graph of the target object and the semantic directed graph of the item set to obtain a fused directed graph; the semantic directed graph is constructed based on the semantic relevance weights of the nodes of each item in the item set; A random walk module, used to take each node of the fused directed graph as a starting point, perform a preset number of random walks according to the weight of each edge in the fused directed graph, and obtain a random walk sequence with different starting points; A training module, used for training the initial large language model based on each of the random walk sequences to obtain a large language model; A determination module, configured to determine, through the large language model, a semantic embedding representation matching each item in the item set; A recommendation module is used to recommend a target item in the item set to the target object according to the semantic embedding representation.
14. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
16. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.
Citation Information
Cited By
Intelligent product recommendation method and device based on graph thought and electronic equipment
CN120780920A