A multimedia recommendation method and device, an electronic device and a storage medium
By expanding historical operational resource information with interest recognition models and generating object interest representation information, the problem of low user interest diversity is solved, and the effectiveness of multimedia resource recommendation is improved.
Patent Information
- Application Number
- CN202111370491.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-18
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2041-11-18
AI Technical Summary
Existing technologies suffer from low diversity of user interests and low effectiveness in recommending multimedia resources.
By using an object interest recognition model to extend the historical operation resource information with interest, object interest representation information is generated, and based on this, target interest multimedia resources are identified from the multimedia resources to be recommended and then recommended.
It improves the diversity and generalization of user interests, and enhances the effectiveness of multimedia resource recommendations.
Smart Images

Figure CN114201625B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of information recommendation, and particularly relates to a multimedia recommendation method and device, an electronic device, and a storage medium. BACKGROUND
[0002] According to the preferences of a user when browsing multimedia resources, a personalized recommendation scheme corresponding to different users is determined, and corresponding multimedia resources can be recommended to different users. However, in related technologies, when a user is recommended personally, it is easy to converge to several interest points of the user, which leads to low diversity of the user's interests, and the multimedia resources recommended to the user cannot arouse the user's interest, thereby reducing the effectiveness of multimedia resource recommendation. SUMMARY
[0003] The present disclosure provides a multimedia recommendation method, device, electronic device, and storage medium to at least solve the problem of low diversity of user interests and low effectiveness of multimedia resource recommendation in related technologies. The technical solutions of the present disclosure are as follows.
[0004] According to a first aspect of an embodiment of the present disclosure, a multimedia recommendation method is provided, which includes:
[0005] inputting historical operation resource information corresponding to a to-be-processed object and a multimedia knowledge structure into an object interest recognition model, performing at least one interest expansion on the historical operation resource information based on the multimedia knowledge structure in the object interest recognition model, and obtaining at least one object interest representation information corresponding to the to-be-processed object, wherein the historical operation resource information represents multimedia resources on which the to-be-processed object has performed a preset operation in a preset historical time period, and the multimedia knowledge structure is a graph formed by taking multimedia resource information of a preset multimedia resource and content tag information corresponding to the preset multimedia resource as nodes and taking an association relationship between the multimedia resource information and the content tag information as an edge;
[0006] determining target interest multimedia resources corresponding to the to-be-processed object from to-be-recommended multimedia resources based on the at least one object interest representation information;
[0007] recommending the target interest multimedia resources to the to-be-processed object.
[0008] As an optional embodiment, the object interest recognition model includes a feature extraction layer, a feature expansion layer, and a feature fusion layer, and the inputting of the historical operation resource information corresponding to the to-be-processed object and the multimedia knowledge structure into the object interest recognition model, the performing of at least one interest expansion on the historical operation resource information based on the multimedia knowledge structure in the object interest recognition model, and the obtaining of the at least one object interest representation information corresponding to the to-be-processed object include:
[0009] The historical operation resource information and the multimedia knowledge structure are input into the feature extraction layer for feature extraction to obtain the historical resource feature information of the historical operation resource information and the structural feature information corresponding to the multimedia knowledge structure.
[0010] The historical resource feature information and the structural feature information are input into the feature extension layer. Based on the structural feature information, the historical resource feature information is extended at least once to obtain the associated tag feature information and the associated resource feature information corresponding to the historical resource feature information under the at least one feature extension.
[0011] The associated label feature information and the associated resource feature information are input into the feature fusion layer for feature fusion to obtain the at least one object interest representation information.
[0012] As an optional embodiment, the step of inputting the historical resource feature information and the structural feature information into the feature extension layer, and performing at least one feature extension on the historical resource feature information based on the structural feature information to obtain the associated tag feature information and the associated resource feature information corresponding to the historical resource feature information under the at least one feature extension includes:
[0013] The historical resource feature information and the structural feature information are input into the feature extension layer. In the structural feature information, at least one feature extension is performed with the historical resource feature information as the central node to obtain the associated node associated with the central node at any feature extension. The starting node of any feature extension is the resource feature information in the associated node obtained in the previous feature extension, and the starting node of the first feature extension in any feature extension is the historical resource feature information.
[0014] The tag feature information in the associated node corresponding to the at least one feature expansion is used as the associated tag feature information, and the resource feature information in the associated node corresponding to the at least one feature expansion is used as the associated resource feature information.
[0015] As an optional embodiment, the step of inputting the associated tag feature information and the associated resource feature information into the feature fusion layer for feature fusion to obtain the at least one object interest representation information includes:
[0016] The associated label feature information and associated resource feature information corresponding to each feature expansion are input into the feature fusion layer for feature fusion to obtain the at least one object interest representation information.
[0017] As an optional embodiment, determining the target interest multimedia resource corresponding to the object to be processed from the multimedia resources to be recommended based on the at least one object interest representation information includes:
[0018] Obtain the resource feature information corresponding to the multimedia resource to be recommended;
[0019] Based on the object interest representation information and the resource feature information corresponding to the multimedia resource to be recommended, the resource interest index corresponding to the multimedia resource to be recommended is determined.
[0020] Based on the resource interest index, the target interest multimedia resources are determined from the multimedia resources to be recommended.
[0021] As an optional embodiment, the preset multimedia resources include multiple multimedia resources, and the method further includes:
[0022] Obtain the profile information corresponding to each multimedia resource;
[0023] Based on the portrait information, multimedia resource information for each multimedia resource and at least one content tag information corresponding to each multimedia resource are obtained;
[0024] Using the multimedia resource information of multiple multimedia resources and the content tag information corresponding to multiple multimedia resources as nodes, and constructing edges between the nodes corresponding to the multimedia resource information of each multimedia resource and the nodes corresponding to the content tag information of each multimedia resource, the multimedia knowledge structure is obtained.
[0025] According to a second aspect of the present disclosure, a method for training an object interest recognition model is provided, the method comprising:
[0026] The system acquires a multimedia knowledge structure, positive sample operation resource information, and negative sample operation resource information corresponding to the sample object. The positive sample operation resource information represents multimedia resources for which the sample object has performed a preset operation within a preset sample time period. The negative sample operation resource information represents at least one of the following: multimedia resources for which the sample object has not performed a preset operation within the preset sample time period, sample multimedia resources similar to the positive sample operation resource information, multimedia resources for which the resource interest index corresponding to the sample object meets preset conditions, and multimedia resources for which a negative feedback operation has been performed. The multimedia knowledge structure is a graph constructed with multimedia resource information of the multimedia resource to be recommended and content tag information corresponding to the multimedia resource to be recommended as nodes, and the relationship between the multimedia resource information and the content tag information as edges.
[0027] The multimedia knowledge structure, the positive sample operation resource information, and the negative sample operation resource information are input into the model to be trained. In the model to be trained, based on the multimedia knowledge structure, the positive sample operation resource information and the negative sample operation resource information are extended at least once to obtain at least one object interest representation information corresponding to the sample object and the structural feature information of the multimedia knowledge structure.
[0028] Obtain the first sample resource feature information corresponding to the positive sample operation resource information and the second sample resource feature information corresponding to the negative sample operation resource information;
[0029] Based on the structural feature information, the at least one object interest representation information, the first sample resource feature information, and the second sample resource feature information, the target loss information is determined;
[0030] Based on the target loss information, the model to be trained is trained to obtain the object interest recognition model.
[0031] As an optional embodiment, the method further includes:
[0032] The multimedia resources of the sample object that have not performed the preset operation within the preset historical time period are sampled to obtain the first negative sample operation resource information;
[0033] Multimedia resources similar to the positive sample operation resource information are used as the second negative sample operation resource information.
[0034] The multimedia resources for which negative feedback operations were performed within the preset historical time period are used as the third negative sample operation resource information.
[0035] One or more of the first negative sample operation resource information, the second negative sample operation resource information, and the third negative sample operation resource information are used as the negative sample operation resource information.
[0036] As an optional embodiment, the method further includes:
[0037] If the current training round is not the first training round, obtain the object interest representation information corresponding to the previous training round of the current training round and the resource feature information corresponding to the multimedia resource to be recommended.
[0038] Based on the object interest representation information and the resource feature information corresponding to the multimedia resource to be recommended, the resource interest index corresponding to the multimedia resource to be recommended is determined.
[0039] Based on the resource interest index, the fourth negative sample operation resource information is determined from the multimedia resources to be recommended.
[0040] The step of using one or more of the first negative sample operation resource information, the second negative sample operation resource information, and the third negative sample operation resource information as the negative sample operation resource information includes:
[0041] One or more of the first negative sample operation resource information, the second negative sample operation resource information, the third negative sample operation resource information, and the fourth negative sample operation resource information are used as the negative sample operation resource information.
[0042] As an optional embodiment, determining the target loss information based on the structural feature information, the at least one object interest representation information, the first sample resource feature information, and the second sample resource feature information includes:
[0043] Interest loss information is obtained based on the at least one object interest representation information, the first sample resource feature information, and the second sample resource feature information;
[0044] Based on the structural feature information, node relationship loss information is obtained;
[0045] Based on the at least one object interest representation information, representation loss information is obtained;
[0046] Based on the structural feature information, the first sample resource feature information, and the second sample resource feature information, regularization loss information is obtained;
[0047] The target loss information is determined based on the interest loss information, the node relationship loss information, the representation loss information, and the regularization loss information.
[0048] According to a third aspect of the present disclosure, a multimedia recommendation device is provided, the device comprising:
[0049] The feature extension module is configured to input the historical operation resource information and multimedia knowledge structure corresponding to the object to be processed into the object interest recognition model. In the object interest recognition model, based on the multimedia knowledge structure, the historical operation resource information is extended at least once to obtain at least one object interest representation information corresponding to the object to be processed. The historical operation resource information represents the multimedia resources for which the object to be processed has performed a preset operation within a preset historical time period. The multimedia knowledge structure is a graph with the multimedia resource information of the preset multimedia resources and the content tag information corresponding to the preset multimedia resources as nodes, and the association relationship between the multimedia resource information and the content tag information as edges.
[0050] The target interest resource determination module is configured to determine the target interest multimedia resource corresponding to the object to be processed from the multimedia resources to be recommended based on the at least one object interest representation information.
[0051] The target interest resource recommendation module is configured to recommend multimedia resources of the target interest to the object to be processed.
[0052] As an optional embodiment, the object interest recognition model includes a feature extraction layer, a feature expansion layer, and a feature fusion layer, wherein the feature expansion module includes:
[0053] The feature extraction unit is configured to input the historical operation resource information and the multimedia knowledge structure into the feature extraction layer to perform feature extraction, thereby obtaining historical resource feature information of the historical operation resource information and structural feature information corresponding to the multimedia knowledge structure.
[0054] The feature expansion unit is configured to input the historical resource feature information and the structural feature information into the feature expansion layer, and perform at least one feature expansion on the historical resource feature information based on the structural feature information to obtain the associated tag feature information and the associated resource feature information corresponding to the historical resource feature information under the at least one feature expansion.
[0055] The feature fusion unit is configured to perform feature fusion by inputting the associated label feature information and the associated resource feature information into the feature fusion layer to obtain the at least one object interest representation information.
[0056] As an optional embodiment, the feature extension unit includes:
[0057] The associated node determination unit is configured to input the historical resource feature information and the structural feature information into the feature extension layer, and perform at least one feature extension on the structural feature information with the historical resource feature information as the central node to obtain associated nodes associated with the central node at any feature extension; the starting node of any feature extension is the resource feature information in the associated nodes obtained in the previous feature extension, and the starting node of the first feature extension in any feature extension is the historical resource feature information.
[0058] The association information acquisition unit is configured to use the tag feature information in the association node corresponding to the at least one feature expansion as the association tag feature information, and use the resource feature information in the association node corresponding to the at least one feature expansion as the association resource feature information.
[0059] As an optional embodiment, the feature fusion unit includes:
[0060] The multi-layer feature fusion unit is configured to input the associated label feature information and the associated resource feature information corresponding to each feature expansion into the feature fusion layer for feature fusion to obtain the at least one object interest representation information.
[0061] As an optional embodiment, the target interest resource determination module includes:
[0062] The resource feature acquisition unit is configured to acquire the resource feature information corresponding to the multimedia resource to be recommended.
[0063] The resource interest index determination unit is configured to determine the resource interest index corresponding to the multimedia resource to be recommended based on the object interest representation information and the resource feature information corresponding to the multimedia resource to be recommended.
[0064] The target multimedia resource determination unit is configured to determine the target interest multimedia resource from the multimedia resources to be recommended based on the resource interest index.
[0065] As an optional embodiment, the preset multimedia resources include multiple multimedia resources, and the device further includes:
[0066] The profile information acquisition module is configured to acquire profile information corresponding to each multimedia resource;
[0067] The graph node acquisition module is configured to perform operations based on the profile information to obtain multimedia resource information for each multimedia resource and at least one content tag information corresponding to each multimedia resource.
[0068] The graph construction module is configured to use the multimedia resource information of multiple multimedia resources and the content tag information corresponding to the multiple multimedia resources as nodes, and construct edges between the nodes corresponding to the multimedia resource information of each multimedia resource and the nodes corresponding to the content tag information of each multimedia resource to obtain the multimedia knowledge structure.
[0069] According to a fourth aspect of the present disclosure, an object interest recognition model training apparatus is provided, the apparatus comprising:
[0070] The information acquisition module is configured to acquire multimedia knowledge structure, positive sample operation resource information, and negative sample operation resource information corresponding to the sample object. The positive sample operation resource information represents multimedia resources for which the sample object has performed a preset operation within a preset sample time period. The negative sample operation resource information represents at least one of the following: multimedia resources for which the sample object has not performed a preset operation within the preset sample time period, sample multimedia resources similar to the positive sample operation resource information, multimedia resources for which the resource interest index corresponding to the sample object meets a preset condition, and multimedia resources for which a negative feedback operation has been performed. The multimedia knowledge structure is a graph constructed with multimedia resource information of the multimedia resource to be recommended and content tag information corresponding to the multimedia resource to be recommended as nodes, and the association relationship between the multimedia resource information and the content tag information as edges.
[0071] The training interest extension module is configured to input the multimedia knowledge structure, the positive sample operation resource information, and the negative sample operation resource information into the model to be trained, and perform at least one interest extension on the positive sample operation resource information and the negative sample operation resource information based on the multimedia knowledge structure in the model to be trained, so as to obtain at least one object interest representation information corresponding to the sample object and the structural feature information of the multimedia knowledge structure.
[0072] The sample resource feature acquisition module is configured to acquire the first sample resource feature information corresponding to the positive sample operation resource information and the second sample resource feature information corresponding to the negative sample operation resource information.
[0073] The target loss information determination module is configured to determine target loss information based on the structural feature information, the at least one object interest representation information, the first sample resource feature information, and the second sample resource feature information;
[0074] The model training module is configured to train the model to be trained based on the target loss information to obtain an object interest recognition model.
[0075] As an optional embodiment, the apparatus further includes:
[0076] The first negative sample resource module is configured to sample multimedia resources of the sample object that have not performed a preset operation within a preset historical time period to obtain the first negative sample operation resource information.
[0077] The second negative sample resource module is configured to use multimedia resources similar to the positive sample operation resource information as the second negative sample operation resource information.
[0078] The third negative sample resource module is configured to use multimedia resources that have performed negative feedback operations on the sample object within the preset historical time period as third negative sample operation resource information.
[0079] The negative sample operation resource acquisition module is configured to use one or more of the first negative sample operation resource information, the second negative sample operation resource information, and the third negative sample operation resource information as the negative sample operation resource information.
[0080] As an optional embodiment, the apparatus further includes:
[0081] The previous training information acquisition module is configured to acquire the object interest representation information corresponding to the previous training round and the resource feature information corresponding to the multimedia resource to be recommended when the current training round is not the first training round.
[0082] The training interest index acquisition module is configured to determine the resource interest index corresponding to the multimedia resource to be recommended based on the object interest representation information and the resource feature information corresponding to the multimedia resource to be recommended.
[0083] The fourth negative sample resource acquisition module is configured to determine the fourth negative sample operation resource information from the multimedia resources to be recommended based on the resource interest index.
[0084] The negative sample operation resource acquisition module includes:
[0085] The negative sample operation resource acquisition unit is configured to use one or more of the first negative sample operation resource information, the second negative sample operation resource information, the third negative sample operation resource information, and the fourth negative sample operation resource information as the negative sample operation resource information.
[0086] As an optional embodiment, the target loss information determination module includes:
[0087] The interest loss information determination unit is configured to perform an operation to obtain interest loss information based on the at least one object interest representation information, the first sample resource feature information, and the second sample resource feature information;
[0088] The node relationship loss information determination unit is configured to perform operations based on the structural feature information to obtain node relationship loss information;
[0089] The characterization loss information determination unit is configured to perform characterization loss information based on the at least one object interest characterization information;
[0090] The regularization loss information determination unit is configured to perform a process based on the structural feature information, the first sample resource feature information, and the second sample resource feature information to obtain regularization loss information.
[0091] The target loss information determination unit is configured to determine the target loss information based on the interest loss information, the node relationship loss information, the representation loss information, and the regularization loss information.
[0092] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising:
[0093] processor;
[0094] Memory used to store the processor's executable instructions;
[0095] The processor is configured to execute the instructions to implement the multimedia recommendation method or the object interest recognition model training method described above.
[0096] According to a sixth aspect of the present disclosure, a computer-readable storage medium is provided, wherein when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the multimedia recommendation method or the object interest recognition model training method described above.
[0097] According to a seventh aspect of the present disclosure, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the multimedia recommender or the object interest recognition model training method described above.
[0098] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:
[0099] This method involves inputting historical operation resource information corresponding to the object to be processed into an object interest recognition model. Based on the multimedia knowledge structure within the model, the historical operation resource information is expanded at least once to obtain at least one object interest representation for the object to be processed. Based on this representation, target interest multimedia resources corresponding to the object to be processed are determined from the recommended multimedia resources, and these resources are then recommended to the object. This method, which expands interest resources based on historical operation resource information, can alleviate information cocoons, acquire multimedia resources corresponding to potential user interests, thereby improving the diversity and generalization of user interests and enhancing the effectiveness of multimedia resource recommendations.
[0100] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0101] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0102] Figure 1 This is a schematic diagram illustrating an application scenario of a multimedia recommendation method according to an exemplary embodiment.
[0103] Figure 2 This is a flowchart illustrating a multimedia recommendation method according to an exemplary embodiment.
[0104] Figure 3 This is a flowchart illustrating the construction of a multimedia knowledge structure in a multimedia recommendation method according to an exemplary embodiment.
[0105] Figure 4 This is a schematic diagram of a multimedia knowledge structure in a multimedia recommendation method according to an exemplary embodiment.
[0106] Figure 5 This is a flowchart illustrating interest expansion in a multimedia recommendation method according to an exemplary embodiment.
[0107] Figure 6 This is a flowchart illustrating feature expansion in a feature expansion layer in a multimedia recommendation method according to an exemplary embodiment.
[0108] Figure 7 This is a schematic diagram of a series of consecutive triples in a multimedia recommendation method according to an exemplary embodiment.
[0109] Figure 8This is a schematic diagram illustrating the acquisition of multiple representations of object interests in a multimedia recommendation method according to an exemplary embodiment.
[0110] Figure 9 This is a flowchart illustrating a multimedia recommendation method for obtaining target multimedia resources according to an exemplary embodiment.
[0111] Figure 10 This is a schematic diagram illustrating online recall in a multimedia recommendation method according to an exemplary embodiment.
[0112] Figure 11 This is a flowchart illustrating an object interest recognition model training method according to an exemplary embodiment.
[0113] Figure 12 This is a flowchart illustrating the construction of negative sample operation resource information in an object interest recognition model training method according to an exemplary embodiment.
[0114] Figure 13 This is a flowchart illustrating the determination of the fourth negative sample operation resource information in an object interest recognition model training method according to an exemplary embodiment.
[0115] Figure 14 This is a flowchart illustrating the determination of target loss information in an object interest recognition model training method according to an exemplary embodiment.
[0116] Figure 15 This is a schematic diagram illustrating the calculation of expected click results in an object interest recognition model training method according to an exemplary embodiment.
[0117] Figure 16 This is a block diagram illustrating a multimedia recommendation device according to an exemplary embodiment.
[0118] Figure 17 This is a block diagram illustrating an object interest recognition model training device according to an exemplary embodiment.
[0119] Figure 18 This is a block diagram illustrating a server-side electronic device according to an exemplary embodiment. Detailed Implementation
[0120] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0121] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0122] Figure 1 This is a schematic diagram illustrating an application scenario of a multimedia recommendation method according to an exemplary embodiment, such as... Figure 1 As shown, this application scenario includes a client 110 and a server 120. The server 120 pre-stores at least one object interest representation information corresponding to each user. This object interest representation information can be obtained by expanding the historical operation resource information corresponding to each user based on the object interest recognition model and multimedia knowledge structure. In response to the multimedia recommendation request sent by the client 110, the server 120 recalls the corresponding target multimedia resource based on the interest recommendation index between the resource feature information of the multimedia resource to be recommended and the object interest representation information, and sends the target multimedia resource to the client 110.
[0123] In this embodiment, client 110 includes physical devices such as smartphones, desktop computers, tablets, laptops, digital assistants, and smart wearable devices, and may also include software running on the physical device, such as applications. The operating system running on the physical device in this embodiment may include, but is not limited to, Android, iOS, Linux, Unix, and Windows. Client 110 includes a UI (User Interface) layer, through which client 110 provides the display of target multimedia resources. Additionally, it sends multimedia recommendation requests to server 120 based on API (Application Programming Interface).
[0124] In this embodiment, server 120 may include a standalone server, a distributed server, or a server cluster consisting of multiple servers. Server 120 may include a network communication unit, a processor, and a memory, etc. Specifically, server 120 may expand the historical operation resource information corresponding to each user with interest resources and recall target multimedia resources based on a multimedia knowledge structure.
[0125] Figure 2 This is a flowchart illustrating a multimedia recommendation method according to an exemplary embodiment, such as... Figure 2 As shown, this method, using a server, includes the following steps.
[0126] S210. Input the historical operation resource information and multimedia knowledge structure corresponding to the object to be processed into the object interest recognition model. In the object interest recognition model, based on the multimedia knowledge structure, perform at least one interest expansion on the historical operation resource information to obtain at least one object interest representation information corresponding to the object to be processed. The historical operation resource information represents the multimedia resources that the object to be processed has performed preset operations within a preset historical time period. The multimedia knowledge structure is a graph with the multimedia resource information of the preset multimedia resources and the content tag information corresponding to the preset multimedia resources as nodes and the association relationship between the multimedia resource information and the content tag information as edges.
[0127] As an optional embodiment, the preset operations corresponding to the historical operation resource information include click operations and positive feedback operations. The positive feedback operations are used to represent user preferences and may include like operations, favorite operations, reward operations, etc.
[0128] The object interest recognition model is based on the water wave recommendation model. When performing at least one feature expansion on historical operational resource information, after the first feature expansion, the associated label and associated resource features obtained from the previous feature expansion can be expanded again. When the number of feature expansions reaches a preset number, the associated label and associated resource features obtained from each feature expansion are used as the object interest representation information.
[0129] As an optional embodiment, please refer to Figure 3 ,like Figure 3 As shown, the preset multimedia resources include multiple multimedia resources, and the methods for constructing the multimedia knowledge structure include:
[0130] S310. Obtain the image information corresponding to each multimedia resource;
[0131] S320. Based on the portrait information, obtain the multimedia resource information of each multimedia resource and at least one content tag information corresponding to each multimedia resource;
[0132] S330. Using the multimedia resource information of multiple multimedia resources and the content tag information corresponding to multiple multimedia resources as nodes, and constructing the edges between the nodes corresponding to the multimedia resource information of each multimedia resource and the nodes corresponding to the content tag information of each multimedia resource, a multimedia knowledge structure is obtained.
[0133] As an optional embodiment, when constructing a multimedia knowledge structure, the profile information corresponding to each multimedia resource in the preset multimedia resources can be obtained, and the profile information can be used to extract features to obtain the multimedia resource information of each multimedia resource and at least one content tag information corresponding to each multimedia resource.
[0134] By associating pairs of multimedia resources with the same content tag information, and continuing this process until all multimedia resources and their corresponding content tag information have been processed within the preset multimedia resources, a multimedia knowledge structure can be obtained. In this knowledge structure, each multimedia resource is associated with at least one other multimedia resource based on at least one content tag; that is, any two associated multimedia resources share one or more identical content tags. Furthermore, the same content tag can be associated with two or more multimedia resources within the knowledge structure.
[0135] As an optional embodiment, taking video resources as an example, feature extraction can be performed on the profile information of video resources to obtain various content tag information such as video music, characters in the video, video style, video category, and video text features, as well as video resources. Video resources with the same music can be associated based on the music tag, video resources with the same characters can be associated based on the character tag, video resources with the same video style can be associated based on the video style tag, video resources with the same video category can be associated based on the category tag, or video resources with the same video text can be associated based on the text tag, thereby obtaining the multimedia knowledge structure corresponding to the video resources.
[0136] As an optional embodiment, please refer to Figure 4 The multimedia knowledge structure is a network structure formed by connecting multiple nodes. It includes two types of nodes: the content tag information corresponding to each multimedia resource is designated as the first type of node, and the multimedia resource information of each multimedia resource is designated as the second type of node. Edges are constructed between the first and second types of nodes to obtain the multimedia knowledge structure.
[0137] like Figure 4As shown, multimedia resources X, Y, Z, and W in the preset multimedia resources are respectively second-type nodes X, Y, Z, and W. The three tag information a, b, and c corresponding to multimedia resource X are respectively first-type nodes a, b, and c. The three feature information b, d, and e corresponding to multimedia resource Y are respectively first-type nodes b, d, and e. The three tag information a, c, and f corresponding to multimedia resource Z are respectively first-type nodes a, c, and f. The three feature information b, g, and h corresponding to multimedia resource W are respectively first-type nodes b, g, and h. Therefore, in the multimedia knowledge structure, second-type node X connects to first-type nodes c and a, first-type nodes c and a are both connected to second-type node Z, and second-type node Z is also connected to first-type node f. Second type node X connects to first type node b, first type node b connects to second type node Y and second type node W, second type node Y also connects to first type node d and first type node e, and second type node Z also connects to first type node g and first type node h.
[0138] By associating multimedia resources with the same content tags, a multimedia knowledge structure can be constructed, thereby revealing the potential correlations between different multimedia resources. This improves the visibility of the relationships between multimedia resources and facilitates subsequent expansion based on users' potential interests.
[0139] As an optional embodiment, the object interest recognition model includes a feature extraction layer, a feature expansion layer, and a feature fusion layer. See [link to relevant documentation]. Figure 5 The historical operation resource information and multimedia knowledge structure corresponding to the object to be processed are input into the object interest recognition model. In the object interest recognition model, based on the multimedia knowledge structure, the historical operation resource information is extended at least once to obtain at least one object interest representation information corresponding to the object to be processed, including:
[0140] S510. Input the historical operation resource information and multimedia knowledge structure into the feature extraction layer to extract features, and obtain the historical resource feature information of the historical operation resource information and the structural feature information corresponding to the multimedia knowledge structure.
[0141] S520. Input historical resource feature information and structural feature information into the feature extension layer, perform at least one feature extension on the historical resource feature information based on the structural feature information, and obtain the associated label feature information and the associated resource feature information corresponding to the historical resource feature information under at least one feature extension.
[0142] S530. Input the associated label feature information and associated resource feature information into the feature fusion layer for feature fusion to obtain at least one object interest representation information.
[0143] As an optional embodiment, the historical operation resource information includes multimedia resources clicked by the user within a preset historical time period and multimedia resources for which the user provided positive feedback within the preset historical time period. Inputting the historical operation resource information and the multimedia knowledge structure into the feature extraction layer for feature extraction yields historical resource feature information and structural feature information corresponding to the multimedia knowledge structure. For example, if the historical operation resource information corresponds to the title of a certain film or television work, then the historical resource feature information can be the text feature information corresponding to that title.
[0144] Feature expansion is performed primarily based on historical resource feature information. Based on structural feature information, at least one feature expansion can be performed on the historical resource feature information, yielding associated tag feature information and associated resource feature information corresponding to the historical resource feature information after each expansion. After the first feature expansion, a second feature expansion can be performed based on the associated tag feature information and associated resource feature information obtained in the first expansion, and so on. In each subsequent feature expansion, the associated tag feature information and associated resource feature information obtained in the previous feature expansion can be used for feature expansion.
[0145] In terms of multimedia knowledge structure, the same content tag information between any two multimedia resources can express the potential similarity of multimedia resources. Therefore, by extending the historical operation resource information to interest resources through feature extension, we can explore users' potential interests and improve the effectiveness and generalization of interest resource extension.
[0146] As an optional embodiment, please refer to Figure 6 Historical resource feature information and structural feature information are input into the feature extension layer. Based on the structural feature information, the historical resource feature information is extended at least once to obtain the associated label feature information and the associated resource feature information corresponding to the historical resource feature information under at least one feature extension.
[0147] S610. Input historical resource feature information and structural feature information into the feature extension layer. In the structural feature information, perform at least one feature extension with historical resource feature information as the central node to obtain the associated nodes associated with the central node at any feature extension. The starting node at any feature extension is the resource feature information in the associated nodes obtained in the previous feature extension. The starting node at the first feature extension in any feature extension is the historical resource feature information.
[0148] S620. Use the label feature information in the associated node corresponding to at least one feature expansion as the associated label feature information, and use the resource feature information in the associated node corresponding to at least one feature expansion as the associated resource feature information.
[0149] As an optional embodiment, in the structural feature information, the node corresponding to the historical resource feature information is determined, and this node is used as the central node for at least one feature expansion. During each feature expansion, the starting node of the feature expansion can be updated, but the central node will not be updated. All associated nodes obtained after feature expansion are nodes that are related to the central node.
[0150] In the first feature expansion, the feature expansion is performed starting from the central node, which yields the associated nodes corresponding to the first feature expansion. These associated nodes include tag feature information and resource feature information. Since the central node contains historical resource feature information, the first feature expansion will first obtain the tag feature information corresponding to the historical resource feature information. Then, based on the tag feature information corresponding to the historical resource feature information, the resource feature information corresponding to that tag feature information will be obtained. In other words, the tag feature information corresponding to a circle of historical resource feature information is obtained first, and then the resource feature information corresponding to a circle of tag feature information is obtained.
[0151] During the second feature expansion, the outermost associated nodes obtained from the first feature expansion, i.e. the associated nodes of resource feature information, are used as the starting nodes for feature expansion to obtain the associated nodes corresponding to the second feature expansion. These associated nodes also include label feature information and resource feature information. Similarly, the label feature information corresponding to the resource feature information obtained from the first feature expansion is first expanded, and then the resource feature information corresponding to the label feature information is expanded based on these corresponding label feature information.
[0152] In the third feature expansion, the outermost associated nodes obtained from the second feature expansion (i.e., the associated nodes corresponding to the resource feature information) can be used as the starting node for further feature expansion to obtain the associated nodes corresponding to the third feature expansion. This process continues until the preset number of feature expansions is met.
[0153] During feature expansion, at least one feature expansion results in a node that is directly or indirectly associated with the central node. The label feature information in the associated nodes corresponding to at least one feature expansion is used as the associated label feature information, and the resource feature information in the associated nodes corresponding to at least one feature expansion is used as the associated resource feature information.
[0154] As an optional embodiment, each multimedia resource can correspond to content tag information and the extended multimedia resource corresponding to that content tag information, such as... Figure 7 As shown, the resource feature information of multimedia resources, the tag feature information of content tag information, and the resource feature information of extended multimedia resources can be formed into a triple. Starting from the historical resource feature information at the time of the first feature expansion, the historical resource feature information, the tag feature information corresponding to the historical resource feature information, and the resource feature information of the first feature expansion corresponding to the tag feature information are taken as a triple. The resource feature information of the first feature expansion, the tag feature information corresponding to the resource feature information, and the resource feature information of the second feature expansion corresponding to the tag feature information are taken as a triple. And so on, so that the triple corresponding to the previous feature expansion and the triple corresponding to the next feature expansion are connected to generate a series of connected triples.
[0155] During feature expansion, both multimedia resources and content tag information are expanded simultaneously. Based on the expanded multimedia resources, more content tag information is discovered, and based on the more content tag information, even more multimedia resources are discovered, thereby improving the effectiveness and efficiency of feature expansion.
[0156] As an optional embodiment, the associated label feature information and associated resource feature information are input into the feature fusion layer for feature fusion to obtain at least one object interest representation information, including:
[0157] The associated label feature information and associated resource feature information corresponding to each feature expansion are input into the feature fusion layer for feature fusion to obtain at least one object interest representation information.
[0158] As an optional embodiment, the associated label feature information and associated resource feature information corresponding to each feature expansion can be input into the feature fusion layer for feature fusion to obtain the object interest representation information corresponding to each feature expansion. The number of object interest representation information is the same as the number of feature expansions.
[0159] In the feature fusion layer, the associated label feature information and associated resource feature information corresponding to each feature expansion are fused through the se module (Squeeze-and-Excitationblock, se-block) and sum pooling to obtain object interest representation information.
[0160] As an optional embodiment, please refer to Figure 8 The object identifier to be processed can be input into the object interest recognition model, and feature processing can be performed on the historical operation resource information corresponding to the click operation to obtain click history feature information. Click history feature information can describe user click preferences in the short term. Feature processing on the historical operation resource information corresponding to the click operation can be performed through the se module (Squeeze-and-Excitationblock, se-block) and sum pooling.
[0161] In the feature fusion layer, object interest representation information can be fused with the object identifier and click history feature information to obtain multi-representation information of object interest corresponding to the object to be processed. The object identifier, multi-representation information of object interest, and click history feature information can be combined based on the concat operation, and the combined feature information is passed through a fully connected layer (Dense) to obtain multi-representation information of object interest. Then, based on the multi-representation information of object interest, the target interest multimedia resources corresponding to the object to be processed can be determined from the multimedia resources to be recommended.
[0162] By fusing the associated tag feature information and associated resource feature information obtained from each feature expansion, multi-level object interest representation information can be obtained, which includes both content tag information and multimedia resource information, thereby improving the richness of object interest representation information.
[0163] S220. Based on at least one object interest representation information, determine the target interest multimedia resources corresponding to the object to be processed from the multimedia resources to be recommended;
[0164] As an optional embodiment, in the online sorting and recall step, the object interest representation information is sorted to obtain a multimedia sequence to be recommended, and then multimedia resources are recalled based on the multimedia sequence to be recommended, so as to obtain the target multimedia resources corresponding to the object to be processed.
[0165] As an optional embodiment, please refer to Figure 9 Multimedia filtering is performed on the object's interest representation information to obtain the target multimedia resources corresponding to the object to be processed, including:
[0166] S910. Obtain the resource feature information corresponding to the multimedia resources to be recommended;
[0167] S920. Based on the object interest representation information and the resource feature information corresponding to the multimedia resources to be recommended, determine the resource interest index corresponding to the multimedia resources to be recommended;
[0168] S930. Based on resource interest indicators, determine the target interest multimedia resources from the multimedia resources to be recommended.
[0169] As an optional embodiment, when determining the resource interest index between object interest representation information and resource feature information corresponding to the multimedia resource to be recommended, the dot product between the object interest representation information and the resource feature information corresponding to the multimedia resource to be recommended can be calculated to obtain the resource interest index. If multiple representation information of object interest is obtained based on the combination of object interest representation information, object identifier, and click history feature information, the inner product between the multiple representation information of object interest and the resource feature information corresponding to the multimedia resource to be recommended can be calculated to obtain the resource interest index.
[0170] Based on the magnitude of the resource interest index, the multimedia resources to be recommended are sorted from largest to smallest to obtain a sequence of multimedia resources to be recommended. From the sequence of multimedia resources to be recommended, the top preset number of multimedia resources are selected as target multimedia resources, or the object interest representation information with a resource interest index greater than or equal to a preset threshold is selected as target multimedia resources.
[0171] As an optional embodiment, please refer to Figure 10 ,like Figure 10 As shown, in the online recall part, the resource feature information of the multimedia resources to be recommended obtained offline is input into the feature storage module. In response to the multimedia recommendation request of the object to be processed, the inner product of the object interest representation information and the resource feature information corresponding to the multimedia resources to be recommended is calculated to obtain the resource interest index. Based on the size of the resource interest index, the multimedia resources to be recommended are sorted from smallest to largest to obtain the multimedia sequence to be recommended. Then, based on the multimedia sequence to be recommended, the target multimedia resources are determined from the set of multimedia resources to be recommended. The object interest representation information and the resource feature information corresponding to the multimedia resources to be recommended can be input into the online Artificial Neural Network (ANN) service to determine the target multimedia resources from the multimedia resources to be recommended. The target multimedia resources are then fed back to the client corresponding to the multimedia recommendation request.
[0172] The multimedia recommendation request can be a client-sent request based on input from the user, meaning the server recommends multimedia resources based on the user's active request, or a client-sent request when the user starts an application or a function within the application. In other words, the server proactively recommends multimedia resources to the user.
[0173] Based on the resource interest index calculated from the object interest representation information and resource feature information, the target multimedia resources are recalled from the multimedia resources to be recommended. Information that users are more interested in can be displayed at the top of the recommended information, thereby improving the accuracy of target multimedia resource recommendations and enhancing the user experience.
[0174] S230. Recommend multimedia resources of target interest to the object to be processed.
[0175] As an optional embodiment, please refer to Figure 11 The model training method includes:
[0176] S1110. Obtain the multimedia knowledge structure, positive sample operation resource information, and negative sample operation resource information corresponding to the sample object. The positive sample operation resource information represents the multimedia resources for which the sample object has performed a preset operation within a preset sample time period. The negative sample operation resource information represents at least one of the following: multimedia resources for which the sample object has not performed a preset operation within the preset sample time period, sample multimedia resources similar to the positive sample operation resource information, multimedia resources for which the resource interest index corresponding to the sample object meets preset conditions, and multimedia resources for which a negative feedback operation has been performed. The multimedia knowledge structure is a graph with the multimedia resource information of the multimedia resource to be recommended and the content tag information corresponding to the multimedia resource to be recommended as nodes, and the relationship between the multimedia resource information and the content tag information as edges.
[0177] S1120. Input the multimedia knowledge structure, positive sample operation resource information and negative sample operation resource information into the model to be trained. In the model to be trained, based on the multimedia knowledge structure, perform at least one interest expansion on the positive sample operation resource information and negative sample operation resource information to obtain at least one object interest representation information and the structural feature information of the multimedia knowledge structure corresponding to the sample object.
[0178] S1130. Obtain the first sample resource feature information corresponding to the positive sample operation resource information and the second sample resource feature information corresponding to the negative sample operation resource information;
[0179] S1140. Determine the target loss information based on structural feature information, at least one object interest representation information, first sample resource feature information, and second sample resource feature information;
[0180] S1150. Based on the target loss information, train the model to be trained to obtain the object interest recognition model.
[0181] As an optional embodiment, positive sample operation resource information and negative sample operation resource information corresponding to the sample object are obtained, wherein the negative sample operation resource information can be constructed from one or more negative sample operation resource information. Positive sample operation resource information represents multimedia resources for which the sample object has performed a preset operation within a preset historical time period. The preset operation can be a click operation, meaning the positive sample operation resource information represents multimedia resources that the user has clicked. Negative sample operation resource information represents at least one of the following: multimedia resources for which the sample object has not performed a preset operation within the preset sample time period, sample multimedia resources similar to the positive sample operation resource information, multimedia resources whose resource interest index for the sample object meets preset conditions, and multimedia resources for which a negative feedback operation has been performed. In other words, negative sample operation resource information represents multimedia resources that show the user has received negative preferences to varying degrees.
[0182] Multimedia knowledge structure, positive sample operation resource information, and negative sample operation resource information are input into the model to be trained. In the model, based on the multimedia knowledge structure, at least one feature expansion is performed on the positive and negative sample operation resource information to obtain at least one object interest representation information corresponding to the sample object. The structural feature information of the multimedia knowledge structure extracted in the feature extraction layer of the model to be trained is output. Feature extraction is performed on the positive and negative sample operation resource information to obtain the first sample resource feature information corresponding to the positive sample operation resource information and the second sample resource feature information corresponding to the negative sample operation resource information. Based on the structural feature information, at least one object interest representation information, the first sample resource feature information, and the second sample resource feature information, target loss information is determined. Based on the target loss information, the model to be trained is trained to obtain the object interest recognition model.
[0183] By using positive and negative sample operation resources and target loss information to train the model to be trained, an object interest recognition model can be obtained, which can improve the accuracy of model training.
[0184] As an optional embodiment, please refer to Figure 12 Obtaining the negative sample operation resource information corresponding to the sample object includes:
[0185] S1210. Sample multimedia resources that have not performed preset operations within a preset historical time period to obtain the first negative sample operation resource information;
[0186] S1220. Multimedia resources similar to the positive sample operation resource information are used as the second negative sample operation resource information;
[0187] S1230. Multimedia resources of the sample object that have performed negative feedback operations within a preset historical time period are used as third negative sample operation resource information.
[0188] S1240. Use one or more of the first negative sample operation resource information, the second negative sample operation resource information, and the third negative sample operation resource information as negative sample operation resource information.
[0189] As an optional embodiment, negative sample operation resource information can be constructed in various ways. Specifically, the first negative sample operation resource information is obtained by sampling multimedia resources from which the target operation has not been performed within a preset historical time period. This first negative sample operation resource information enables the model to be trained to perform a coarse screening of object interest representation information.
[0190] The second negative sample operation resource information refers to multimedia resources that are similar to the positive sample operation resource information among multimedia resources that have not performed the target operation. For example, if a user has clicked on movie 1 starring actor A, but has not clicked on movie 2 starring actor A, then movie 1 starring actor A can be used as the positive sample operation resource information, and movie 2 starring actor A can be used as the second negative sample operation resource information. The second negative sample operation resource information enables the model to be trained to identify object interest representation information from similar sample multimedia resources.
[0191] The third negative sample operation resource information refers to multimedia resources for which the sample object has performed negative feedback operations within a preset historical time period. This third negative sample operation resource information contains negative user feedback. For example, if a user reports that video 3 is not interesting, then video 3 can be considered the third negative sample operation resource information. Similarly, if a user dislikes video 4, then video 4 can also be considered the fourth negative sample operation resource information. This third negative sample operation resource information enables the model to filter out non-object interest representation information.
[0192] Based on different negative sample operation resource information, the ability of the model to be trained to identify object interest representation information can be trained with different focuses, which improves the ability of the model to be trained to distinguish between object interest representation information and non-object interest representation information, thereby improving the comprehensiveness and effectiveness of model training.
[0193] As an optional embodiment, please refer to Figure 13 Other methods for obtaining negative sample resources include:
[0194] S1310. If the current training round is not the first training round, obtain the object interest representation information corresponding to the previous training round and the resource feature information corresponding to the multimedia resource to be recommended.
[0195] S1320. Based on the object interest representation information and the resource feature information corresponding to the multimedia resources to be recommended, determine the resource interest index corresponding to the multimedia resources to be recommended;
[0196] S1330. Based on resource interest indicators, determine the fourth negative sample operation resource information from the multimedia resources to be recommended;
[0197] The negative sample operation resource information includes one or more of the first negative sample operation resource information, the second negative sample operation resource information, and the third negative sample operation resource information.
[0198] S1340. Take one or more of the first negative sample operation resource information, the second negative sample operation resource information, the third negative sample operation resource information, and the fourth negative sample operation resource information as negative sample operation resource information.
[0199] As an optional embodiment, the fourth negative sample operation resource information is the negative sample operation resource information updated during each training iteration. This information can only be obtained at the start of the second round of training. The fourth negative sample operation resource information can be multimedia resources in the training multimedia sequence corresponding to the previous training iteration whose resource interest index is less than a preset threshold, or the last preset number of multimedia resources in the sequence. In other words, the fourth negative sample operation resource information consists of multimedia resources obtained in the previous training iteration that have low correlation with the object interest representation information. For example, in the first training iteration, the multimedia resources to be recommended are sorted from largest to smallest according to their resource interest index, resulting in a sequence of multimedia resources to be recommended. The last 300 multimedia resources to be recommended can be used as the fourth negative sample operation resource information for the second training iteration. This fourth negative sample operation resource information enables the trained model to have the ability to rank the multimedia resources to be recommended.
[0200] As an optional embodiment, a first negative sample operation resource information is set to correspond to a first recognition difficulty, a second negative sample operation resource information to correspond to a second recognition difficulty, and a fourth negative sample operation resource information to correspond to a third recognition difficulty. The third negative sample operation resource information is the real negative sample operation resource information. Recognition difficulty refers to the difficulty of identifying the object's interest representation information. Since the first negative sample operation resource information is a multimedia resource that the user has not clicked, the second negative sample operation resource information is a multimedia resource that the user has not clicked and is similar to the positive sample operation resource information, and the fourth negative sample operation resource information is a multimedia resource with low relevance to the sample object among the training object's interest resources, the distance to the positive sample operation resource information gradually decreases from the first negative sample operation resource information to the second negative sample operation resource information, and then to the fourth negative sample operation resource information. Therefore, the third recognition difficulty is greater than the second recognition difficulty, and the second recognition difficulty is greater than the first recognition difficulty. By combining negative sample operation resource information with different recognition difficulties and real negative sample operation resource information according to preset weight information, negative sample operation resource information applied in model training can be obtained. The preset weight information among the first negative sample operation resource information, the second negative sample operation resource information, the fourth negative sample operation resource information, and the third negative sample operation resource information can be 5:2:2:1.
[0201] Based on negative sample operation resource information with different recognition difficulties and real negative sample operation resource information, a multi-level negative sample operation resource information is constructed for training the model to be trained. This can improve the object interest recognition model's ability to recognize object interest representation information, thereby improving the accuracy and transparency of target multimedia resources.
[0202] As an optional embodiment, please refer to Figure 14 Based on structural feature information, at least one object interest representation information, first sample resource feature information, and second sample resource feature information, the target loss information is determined as follows:
[0203] S1410. Based on at least one object interest representation information, first sample resource feature information, and second sample resource feature information, interest loss information is obtained;
[0204] S1420. Based on structural feature information, obtain node relationship loss information;
[0205] S1430. Based on at least one object interest representation information, obtain representation loss information;
[0206] S1440. Based on structural feature information, first sample resource feature information, and second sample resource feature information, the regularization loss information is obtained;
[0207] S1450. Determine the target loss information based on the interest loss information, node relationship loss information, representation loss information, and regularization loss information.
[0208] As an optional embodiment, please refer to Figure 15 ,like Figure 15 As shown, after obtaining at least one object interest representation, this representation is fused with the object identifier and click history features to obtain multi-representation information of object interest corresponding to the sample object. After attention calculation and weighted summation of the multi-representation information, the target interest representation is obtained. Based on the target interest representation and the resource feature information corresponding to the multimedia resource to be recommended, the click probability of the multimedia resource to be recommended is determined, yielding the expected click result for the multimedia resource to be recommended.
[0209] When calculating the loss data, node relationship loss information can be determined based on the node feature information and the connection relationship feature information between nodes in the structural feature information corresponding to the multimedia knowledge structure. Regularization loss information, which can be L2 regularization loss information, can be determined based on the node and connection relationship feature information in the structural feature information, as well as the first and second sample resource feature information. Based on the first and second sample resource feature information, it can be determined whether the sample object clicks on the sample multimedia resource, i.e., the click probability of the sample object on the sample multimedia resource, obtaining the expected click result, which can be the click-through rate (CTR). When the sample resource feature information is the first sample resource feature information, corresponding to the positive sample operation resource information, the expected click result can be determined as clicking. When the sample resource feature information is the second sample resource feature information, corresponding to the negative sample operation resource information, the expected click result can be determined as not clicking. Based on the expected click result, target interest representation information, and the sample resource feature information corresponding to the expected click result, the recommendation click cross-entropy, i.e., interest loss information, can be calculated. Based on at least one object interest representation, the KL divergence loss between each object interest representation and its neighboring object interest representations is calculated to obtain representation loss information, which is used to measure the distance between two adjacent object interest representations.
[0210] As an optional embodiment, the formulas for calculating node relationship loss information, representation loss information, regularization loss information, and interest loss information are as follows:
[0211]
[0212] Where kge_loss represents the node relationship loss information, Ir Let E represent the identity matrix, E represent the eigenvectors corresponding to nodes in the multimedia knowledge structure, and R represent the eigenvectors corresponding to the connections between nodes in the multimedia knowledge structure. T This represents the transpose matrix corresponding to the node. `kl_loss` represents the loss information, where `ue`... num The number of object interest representations is represented by ue, where ue represents at least one object interest representation. The calculation is performed between two adjacent multimedia object interest representations ue. a and ue b The sum of the second norms of the distances between samples. `l2_loss` represents the regularization loss information, where `v` represents the resource features of the first and second samples, and the second norms of `v`, `E`, and `R` are calculated respectively. `base_loss` represents the interest loss data, where `y`... uv This indicates whether the sample object clicked on the recommended multimedia resource; 1 indicates a click, and 0 indicates no click. u represents the target interest representation information. T Train the transpose matrix corresponding to the interest information of the target, σ(u T v) represents the probability distribution calculated between the target interest representation information and the first sample resource feature information, or between the target interest representation information and the second sample resource feature information.
[0213] By maintaining the diversity of training object interest resources based on representation loss information, constraining the multimedia knowledge structure based on regularization loss information and node relationship loss information, and determining the loss information corresponding to the expected click result based on interest loss information, the ability of the model to be trained to identify training object interest resources is improved, thereby improving the accuracy and effectiveness of the object interest recognition model.
[0214] This disclosure proposes a multimedia recommendation method, which includes: inputting historical operation resource information corresponding to an object to be processed into an object interest recognition model; expanding the historical operation resource information into interest resources at least once based on the multimedia knowledge structure in the object interest recognition model to obtain at least one object interest representation information corresponding to the object to be processed; determining the target interest multimedia resource corresponding to the object to be processed from the multimedia resources to be recommended based on the at least one object interest representation information; and recommending the target multimedia resource to the object to be processed. This method expands interest resources based on historical operation resource information, which can alleviate information cocoons, obtain multimedia resources corresponding to potential user interests, thereby improving the diversity and generalization of user interests and enhancing the effectiveness of multimedia resource recommendation.
[0215] Figure 16 This is a block diagram illustrating a multimedia recommendation device according to an exemplary embodiment. (Refer to...) Figure 16 The device includes:
[0216] The feature extension module 1610 is configured to input the historical operation resource information and multimedia knowledge structure corresponding to the object to be processed into the object interest recognition model. In the object interest recognition model, based on the multimedia knowledge structure, the historical operation resource information is extended at least once to obtain at least one object interest representation information corresponding to the object to be processed. The historical operation resource information represents the multimedia resources for which the object to be processed has performed preset operations within a preset historical time period. The multimedia knowledge structure is a graph with the multimedia resource information of the preset multimedia resources and the content tag information corresponding to the preset multimedia resources as nodes and the association relationship between the multimedia resource information and the content tag information as edges.
[0217] The target interest resource determination module 1620 is configured to perform the determination of the target interest multimedia resources corresponding to the object to be processed from the multimedia resources to be recommended based on at least one object interest representation information.
[0218] The target interest resource recommendation module 1630 is configured to recommend target interest multimedia resources to the objects to be processed.
[0219] As an optional embodiment, the object interest recognition model includes a feature extraction layer, a feature expansion layer, and a feature fusion layer. The feature expansion module includes:
[0220] The feature extraction unit is configured to perform feature extraction by inputting historical operation resource information and multimedia knowledge structure into the feature extraction layer, thereby obtaining historical resource feature information of historical operation resource information and structural feature information corresponding to multimedia knowledge structure.
[0221] The feature extension unit is configured to input historical resource feature information and structural feature information into the feature extension layer, perform at least one feature extension on the historical resource feature information based on the structural feature information, and obtain the associated tag feature information and the associated resource feature information corresponding to the historical resource feature information under at least one feature extension.
[0222] The feature fusion unit is configured to perform feature fusion by inputting associated label feature information and associated resource feature information into the feature fusion layer to obtain at least one object interest representation information.
[0223] As an optional embodiment, the feature extension unit includes:
[0224] The associated node determination unit is configured to input historical resource feature information and structural feature information into the feature extension layer. In the structural feature information, at least one feature extension is performed with the historical resource feature information as the central node to obtain the associated nodes associated with the central node at any feature extension. The starting node at any feature extension is the resource feature information in the associated nodes obtained in the previous feature extension, and the starting node at the first feature extension in any feature extension is the historical resource feature information.
[0225] The associated information acquisition unit is configured to use the tag feature information in the associated node corresponding to at least one feature extension as the associated tag feature information, and use the resource feature information in the associated node corresponding to at least one feature extension as the associated resource feature information.
[0226] As an optional embodiment, the feature fusion unit includes:
[0227] The multi-layer feature fusion unit is configured to input the associated label feature information and the associated resource feature information corresponding to each feature expansion into the feature fusion layer for feature fusion to obtain at least one object interest representation information.
[0228] As an optional embodiment, the target interest resource determination module includes:
[0229] The resource feature acquisition unit is configured to acquire resource feature information corresponding to the multimedia resources to be recommended.
[0230] The resource interest index determination unit is configured to determine the resource interest index corresponding to the multimedia resource to be recommended based on the object interest representation information and the resource feature information corresponding to the multimedia resource to be recommended.
[0231] The target multimedia resource determination unit is configured to perform a task based on resource interest indicators to determine target interest multimedia resources from the multimedia resources to be recommended.
[0232] As an optional embodiment, the preset multimedia resources include multiple multimedia resources, and the device further includes:
[0233] The profile information acquisition module is configured to acquire profile information corresponding to each multimedia resource;
[0234] The graph node acquisition module is configured to perform operations based on the profile information to obtain the multimedia resource information of each multimedia resource and at least one content tag information corresponding to each multimedia resource.
[0235] The graph construction module is configured to use the multimedia resource information and content tag information corresponding to multiple multimedia resources as nodes, and construct edges between the nodes corresponding to the multimedia resource information of each multimedia resource and the nodes corresponding to the content tag information of each multimedia resource to obtain the multimedia knowledge structure.
[0236] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0237] Figure 17 This is a block diagram illustrating an object interest recognition model training apparatus according to an exemplary embodiment. (Refer to...) Figure 17 The device includes:
[0238] The information acquisition module 1710 is configured to acquire multimedia knowledge structure, positive sample operation resource information and negative sample operation resource information corresponding to the sample object. Positive sample operation resource information represents multimedia resources for which the sample object has performed a preset operation within a preset sample time period. Negative sample operation resource information represents at least one of the following: multimedia resources for which the sample object has not performed a preset operation within the preset sample time period, sample multimedia resources similar to positive sample operation resource information, multimedia resources for which the resource interest index corresponding to the sample object meets preset conditions, and multimedia resources for which negative feedback operations have been performed. The multimedia knowledge structure is a graph with multimedia resource information of the multimedia resource to be recommended and content tag information corresponding to the multimedia resource to be recommended as nodes, and the relationship between multimedia resource information and content tag information as edges.
[0239] The training interest extension module 1720 is configured to input multimedia knowledge structure, positive sample operation resource information and negative sample operation resource information into the model to be trained, and perform at least one interest extension on the positive sample operation resource information and negative sample operation resource information based on the multimedia knowledge structure in the model to be trained, so as to obtain at least one object interest representation information and structural feature information of multimedia knowledge structure corresponding to the sample object.
[0240] The sample resource feature acquisition module 1730 is configured to acquire the first sample resource feature information corresponding to the positive sample operation resource information and the second sample resource feature information corresponding to the negative sample operation resource information.
[0241] The target loss information determination module 1740 is configured to determine target loss information based on structural feature information, at least one object interest representation information, first sample resource feature information, and second sample resource feature information.
[0242] The model training module is configured to train the model to be trained based on the target loss information to obtain the object interest recognition model.
[0243] As an optional embodiment, the apparatus further includes:
[0244] The first negative sample resource module is configured to sample multimedia resources of the sample object that have not performed a preset operation within a preset historical time period to obtain the first negative sample operation resource information.
[0245] The second negative sample resource module is configured to use multimedia resources similar to the positive sample operation resource information as the second negative sample operation resource information.
[0246] The third negative sample resource module is configured to use multimedia resources that have performed negative feedback operations on the sample object within a preset historical time period as the third negative sample operation resource information.
[0247] The negative sample operation resource acquisition module is configured to use one or more of the first negative sample operation resource information, the second negative sample operation resource information, and the third negative sample operation resource information as negative sample operation resource information.
[0248] As an optional embodiment, the apparatus further includes:
[0249] The previous training information acquisition module is configured to acquire the object interest representation information and the resource feature information corresponding to the multimedia resources to be recommended in the previous training round when the current training round is not the first training round.
[0250] The training interest index acquisition module is configured to perform the determination of resource interest indexes corresponding to the multimedia resources to be recommended based on object interest representation information and resource feature information corresponding to the multimedia resources to be recommended.
[0251] The fourth negative sample resource acquisition module is configured to determine the fourth negative sample operation resource information from the multimedia resources to be recommended based on resource interest indicators.
[0252] The negative sample operation resource acquisition module includes:
[0253] The negative sample operation resource acquisition unit is configured to use one or more of the first negative sample operation resource information, the second negative sample operation resource information, the third negative sample operation resource information, and the fourth negative sample operation resource information as negative sample operation resource information.
[0254] As an optional embodiment, the target loss information determination module includes:
[0255] The interest loss information determination unit is configured to perform an operation based on at least one object interest representation information, first sample resource feature information, and second sample resource feature information to obtain interest loss information.
[0256] The node relationship loss information determination unit is configured to perform operations based on structural feature information to obtain node relationship loss information;
[0257] The characterization loss information determination unit is configured to perform characterization loss information based on at least one object interest characterization information;
[0258] The regularization loss information determination unit is configured to perform regularization loss information based on structural feature information, first sample resource feature information, and second sample resource feature information.
[0259] The target loss information determination unit is configured to determine the target loss information based on interest loss information, node relationship loss information, representation loss information, and regularization loss information.
[0260] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0261] Figure 18 This is a block diagram illustrating an electronic device for training a multimedia recommendation or object interest recognition model according to an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as follows: Figure 18 As shown, the electronic device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a multimedia resource recommendation method or an object interest recognition model training method.
[0262] Those skilled in the art will understand that Figure 18 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the electronic device to which the present disclosure is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0263] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 1804 including instructions, which can be executed by a processor 1820 of an electronic device 1800 to perform the above-described method. Optionally, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0264] In an exemplary embodiment, a computer program product is also provided, including a computer program / instructions, which, when executed by a processor, implement the multimedia recommendation method or the object interest recognition model training method described above.
[0265] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0266] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A multimedia recommendation method, characterized in that, The method includes: The historical operation resource information and multimedia knowledge structure corresponding to the object to be processed are input into the feature extraction layer of the object interest recognition model for feature extraction, so as to obtain the historical resource feature information of the historical operation resource information and the structural feature information corresponding to the multimedia knowledge structure. The historical operation resource information represents the multimedia resources of the object to be processed that have performed preset operations within a preset historical time period. The multimedia knowledge structure is a graph constructed with the multimedia resource information of the preset multimedia resources and the content tag information corresponding to the preset multimedia resources as nodes, and the association relationship between the multimedia resource information and the content tag information as edges. The historical resource feature information and the structural feature information are input into the feature extension layer of the object interest recognition model. In the structural feature information, at least one feature extension is performed with the historical resource feature information as the central node to obtain the tag feature information in the associated nodes associated with the central node at any feature extension. Based on the tag feature information, the resource feature information in the associated nodes is extended. The starting node for any feature extension is the resource feature information in the associated nodes obtained in the previous feature extension. The tag feature information is used as associated tag feature information, and the resource feature information is used as associated resource feature information; The associated label feature information and the associated resource feature information are input into the feature fusion layer of the object interest recognition model for feature fusion to obtain at least one object interest representation information. Based on the at least one object interest representation information, determine the target interest multimedia resources corresponding to the object to be processed from the multimedia resources to be recommended; Recommend the target interest multimedia resources to the object to be processed.
2. The multimedia recommendation method according to claim 1, characterized in that, The starting node for the first feature expansion in any given feature expansion is the historical resource feature information.
3. The multimedia recommendation method according to claim 1, characterized in that, The step of inputting the associated label feature information and the associated resource feature information into the feature fusion layer of the object interest recognition model for feature fusion to obtain at least one object interest representation information includes: The associated label feature information and associated resource feature information corresponding to each feature expansion are input into the feature fusion layer for feature fusion to obtain the at least one object interest representation information.
4. The multimedia recommendation method according to claim 1, characterized in that, The step of determining the target interest multimedia resource corresponding to the object to be processed from the multimedia resources to be recommended based on the at least one object interest representation information includes: Obtain the resource feature information corresponding to the multimedia resource to be recommended; Based on the object interest representation information and the resource feature information corresponding to the multimedia resource to be recommended, the resource interest index corresponding to the multimedia resource to be recommended is determined. Based on the resource interest index, the target interest multimedia resources are determined from the multimedia resources to be recommended.
5. The multimedia recommendation method according to claim 1, characterized in that, The preset multimedia resources include multiple multimedia resources, and the method further includes: Obtain the profile information corresponding to each multimedia resource; Based on the portrait information, multimedia resource information for each multimedia resource and at least one content tag information corresponding to each multimedia resource are obtained; Using the multimedia resource information of multiple multimedia resources and the content tag information corresponding to multiple multimedia resources as nodes, and constructing edges between the nodes corresponding to the multimedia resource information of each multimedia resource and the nodes corresponding to the content tag information of each multimedia resource, the multimedia knowledge structure is obtained.
6. A method for training an object interest recognition model, characterized in that, The method includes: The system acquires a multimedia knowledge structure, positive sample operation resource information, and negative sample operation resource information corresponding to the sample object. The positive sample operation resource information represents multimedia resources for which the sample object has performed a preset operation within a preset sample time period. The negative sample operation resource information represents at least one of the following: multimedia resources for which the sample object has not performed a preset operation within the preset sample time period, sample multimedia resources similar to the positive sample operation resource information, multimedia resources for which the resource interest index corresponding to the sample object meets preset conditions, and multimedia resources for which a negative feedback operation has been performed. The multimedia knowledge structure is a graph constructed with multimedia resource information of the multimedia resource to be recommended and content tag information corresponding to the multimedia resource to be recommended as nodes, and the relationship between the multimedia resource information and the content tag information as edges. The multimedia knowledge structure, the positive sample operation resource information, and the negative sample operation resource information are input into the feature extraction layer of the model to be trained for feature extraction, so as to obtain the resource feature information corresponding to the positive sample operation resource information, the resource feature information corresponding to the negative sample operation resource information, and the current structure feature information corresponding to the multimedia knowledge structure. The resource feature information corresponding to the positive sample operation resource information, the resource feature information corresponding to the negative sample operation resource information, and the current structural feature information are input into the feature expansion layer of the model to be trained. In the current structural feature information, at least one feature expansion is performed using the resource feature information corresponding to the positive sample operation resource information and the resource feature information corresponding to the negative sample operation resource information as sample center nodes. Sample label feature information in the sample association nodes associated with the sample center node is obtained at any feature expansion. Based on the sample label feature information, the sample resource feature information in the sample association nodes is expanded. The starting node for any feature expansion is the sample resource feature information in the sample association nodes obtained in the previous feature expansion. The sample label feature information is used as the sample associated label feature information, and the sample resource feature information is used as the sample associated resource feature information; The sample-associated label feature information and the sample-associated resource feature information are input into the feature fusion layer of the model to be trained for feature fusion to obtain at least one object interest representation information corresponding to the sample object; Obtain the first sample resource feature information corresponding to the positive sample operation resource information and the second sample resource feature information corresponding to the negative sample operation resource information; Based on the structural feature information, the at least one object interest representation information, the first sample resource feature information, and the second sample resource feature information, the target loss information is determined; Based on the target loss information, the model to be trained is trained to obtain the object interest recognition model.
7. The object interest recognition model training method according to claim 6, characterized in that, The method further includes: The multimedia resources of the sample object that have not performed the preset operation within the preset historical time period are sampled to obtain the first negative sample operation resource information; Multimedia resources similar to the positive sample operation resource information are used as the second negative sample operation resource information. The multimedia resources for which negative feedback operations were performed within the preset historical time period are used as the third negative sample operation resource information. One or more of the first negative sample operation resource information, the second negative sample operation resource information, and the third negative sample operation resource information are used as the negative sample operation resource information.
8. The object interest recognition model training method according to claim 7, characterized in that, The method further includes: If the current training round is not the first training round, obtain the object interest representation information corresponding to the previous training round of the current training round and the resource feature information corresponding to the multimedia resource to be recommended. Based on the object interest representation information and the resource feature information corresponding to the multimedia resource to be recommended, the resource interest index corresponding to the multimedia resource to be recommended is determined. Based on the resource interest index, the fourth negative sample operation resource information is determined from the multimedia resources to be recommended. The step of using one or more of the first negative sample operation resource information, the second negative sample operation resource information, and the third negative sample operation resource information as the negative sample operation resource information includes: One or more of the first negative sample operation resource information, the second negative sample operation resource information, the third negative sample operation resource information, and the fourth negative sample operation resource information are used as the negative sample operation resource information.
9. The object interest recognition model training method according to claim 6, characterized in that, The determination of target loss information based on the structural feature information, the at least one object interest representation information, the first sample resource feature information, and the second sample resource feature information includes: Interest loss information is obtained based on the at least one object interest representation information, the first sample resource feature information, and the second sample resource feature information; Based on the structural feature information, node relationship loss information is obtained; Based on the at least one object interest representation information, representation loss information is obtained; Based on the structural feature information, the first sample resource feature information, and the second sample resource feature information, regularization loss information is obtained; The target loss information is determined based on the interest loss information, the node relationship loss information, the representation loss information, and the regularization loss information.
10. A multimedia recommendation device, characterized in that, The device includes: The feature expansion module includes a feature extraction unit, a feature expansion unit, and a feature fusion unit. The feature expansion unit includes an associated node determination unit and an associated information acquisition unit. The feature extraction unit is configured to perform feature extraction by inputting the historical operation resource information and multimedia knowledge structure corresponding to the object to be processed into the feature extraction layer of the object interest recognition model, thereby obtaining the historical resource feature information of the historical operation resource information and the structural feature information corresponding to the multimedia knowledge structure; the historical operation resource information represents the multimedia resources for which the object to be processed has performed preset operations within a preset historical time period, and the multimedia knowledge structure is a graph constructed with the multimedia resource information of the preset multimedia resources and the content tag information corresponding to the preset multimedia resources as nodes, and the association relationship between the multimedia resource information and the content tag information as edges; The associated node determination unit is configured to input the historical resource feature information and the structural feature information into the feature extension layer of the object interest recognition model. In the structural feature information, it performs at least one feature extension with the historical resource feature information as the central node to obtain tag feature information in the associated nodes associated with the central node at any given feature extension. Based on the tag feature information, it extends the resource feature information in the associated nodes. The starting node for each feature extension is the resource feature information in the associated nodes obtained from the previous feature extension. The association information acquisition unit is configured to use the tag feature information as association tag feature information and the resource feature information as association resource feature information; The feature fusion unit is configured to perform feature fusion by inputting the associated label feature information and the associated resource feature information into the feature fusion layer of the object interest recognition model to obtain at least one object interest representation information; The target interest resource determination module is configured to determine the target interest multimedia resource corresponding to the object to be processed from the multimedia resources to be recommended based on the at least one object interest representation information. The target interest resource recommendation module is configured to recommend multimedia resources of the target interest to the object to be processed.
11. The multimedia recommendation device according to claim 10, characterized in that, The starting node for the first feature expansion in any given feature expansion is the historical resource feature information.
12. The multimedia recommendation device according to claim 10, characterized in that, The feature fusion unit includes: The multi-layer feature fusion unit is configured to input the associated label feature information and the associated resource feature information corresponding to each feature expansion into the feature fusion layer for feature fusion to obtain the at least one object interest representation information.
13. The multimedia recommendation device according to claim 10, characterized in that, The target interest resource determination module includes: The resource feature acquisition unit is configured to acquire the resource feature information corresponding to the multimedia resource to be recommended. The resource interest index determination unit is configured to determine the resource interest index corresponding to the multimedia resource to be recommended based on the object interest representation information and the resource feature information corresponding to the multimedia resource to be recommended. The target multimedia resource determination unit is configured to determine the target interest multimedia resource from the multimedia resources to be recommended based on the resource interest index.
14. The multimedia recommendation device according to claim 10, characterized in that, The preset multimedia resources include multiple multimedia resources, and the device further includes: The profile information acquisition module is configured to acquire profile information corresponding to each multimedia resource; The graph node acquisition module is configured to perform operations based on the profile information to obtain multimedia resource information for each multimedia resource and at least one content tag information corresponding to each multimedia resource. The graph construction module is configured to use the multimedia resource information of multiple multimedia resources and the content tag information corresponding to the multiple multimedia resources as nodes, and construct edges between the nodes corresponding to the multimedia resource information of each multimedia resource and the nodes corresponding to the content tag information of each multimedia resource to obtain the multimedia knowledge structure.
15. A training device for an object interest recognition model, characterized in that, The device includes: The information acquisition module is configured to acquire multimedia knowledge structures, positive sample operation resource information, and negative sample operation resource information corresponding to sample objects. The positive sample operation resource information represents multimedia resources for which the sample object has performed a preset operation within a preset sample time period. The negative sample operation resource information represents multimedia resources for which the sample object has not performed a preset operation within the preset sample time period, sample multimedia resources similar to the positive sample operation resource information, multimedia resources for which the resource interest index corresponding to the sample object meets preset conditions, and multimedia resources for which negative feedback operations have been performed. At least one of the multimedia resources; the multimedia knowledge structure is a multimedia resource of the multimedia resources to be recommended. The information and the content tag information corresponding to the multimedia resources to be recommended are used as nodes, and the multimedia resource information is combined with the... The relationships between content tag information form a graph composed of edges; The training interest extension module is configured to perform feature extraction by inputting the multimedia knowledge structure, the positive sample operation resource information, and the negative sample operation resource information into the feature extraction layer of the model to be trained, obtaining resource feature information corresponding to the positive sample operation resource information, resource feature information corresponding to the negative sample operation resource information, and current structure feature information corresponding to the multimedia knowledge structure; and to input the resource feature information corresponding to the positive sample operation resource information, the resource feature information corresponding to the negative sample operation resource information, and the current structure feature information into the feature extension layer of the model to be trained, wherein the resource feature information corresponding to the positive sample operation resource information and the resource feature information corresponding to the negative sample operation resource information are used as the basis for the current structure feature information. The sample center node undergoes at least one feature expansion to obtain sample label feature information in the sample association nodes associated with the sample center node at any given time. Based on the sample label feature information, sample resource feature information in the sample association nodes is expanded. The starting node for each feature expansion is the sample resource feature information in the sample association nodes obtained from the previous feature expansion. The sample label feature information is used as sample association label feature information, and the sample resource feature information is used as sample association resource feature information. The sample association label feature information and the sample association resource feature information are input into the feature fusion layer of the model to be trained for feature fusion to obtain at least one object interest representation information corresponding to the sample object. The sample resource feature acquisition module is configured to acquire the first sample resource feature information corresponding to the positive sample operation resource information and the second sample resource feature information corresponding to the negative sample operation resource information. The target loss information determination module is configured to determine target loss information based on the structural feature information, the at least one object interest representation information, the first sample resource feature information, and the second sample resource feature information; The model training module is configured to train the model to be trained based on the target loss information to obtain an object interest recognition model.
16. The object interest recognition model training device according to claim 15, characterized in that, The device further includes: The first negative sample resource module is configured to sample multimedia resources of the sample object that have not performed a preset operation within a preset historical time period to obtain the first negative sample operation resource information. The second negative sample resource module is configured to use multimedia resources similar to the positive sample operation resource information as the second negative sample operation resource information. The third negative sample resource module is configured to use multimedia resources that have performed negative feedback operations on the sample object within the preset historical time period as third negative sample operation resource information. The negative sample operation resource acquisition module is configured to use one or more of the first negative sample operation resource information, the second negative sample operation resource information, and the third negative sample operation resource information as the negative sample operation resource information.
17. The object interest recognition model training device according to claim 16, characterized in that, The device further includes: The previous training information acquisition module is configured to acquire the object interest representation information corresponding to the previous training round and the resource feature information corresponding to the multimedia resource to be recommended when the current training round is not the first training round. The training interest index acquisition module is configured to determine the resource interest index corresponding to the multimedia resource to be recommended based on the object interest representation information and the resource feature information corresponding to the multimedia resource to be recommended. The fourth negative sample resource acquisition module is configured to determine the fourth negative sample operation resource information from the multimedia resources to be recommended based on the resource interest index. The negative sample operation resource acquisition module includes: The negative sample operation resource acquisition unit is configured to use one or more of the first negative sample operation resource information, the second negative sample operation resource information, the third negative sample operation resource information, and the fourth negative sample operation resource information as the negative sample operation resource information.
18. The object interest recognition model training device according to claim 15, characterized in that, The target loss information determination module includes: The interest loss information determination unit is configured to perform an operation to obtain interest loss information based on the at least one object interest representation information, the first sample resource feature information, and the second sample resource feature information; The node relationship loss information determination unit is configured to perform operations based on the structural feature information to obtain node relationship loss information; The characterization loss information determination unit is configured to perform characterization loss information based on the at least one object interest characterization information; The regularization loss information determination unit is configured to perform a process based on the structural feature information, the first sample resource feature information, and the second sample resource feature information to obtain regularization loss information. The target loss information determination unit is configured to determine the target loss information based on the interest loss information, the node relationship loss information, the representation loss information, and the regularization loss information.
19. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the multimedia recommendation method as described in any one of claims 1 to 5 or the object interest recognition model training method as described in any one of claims 6 to 9.
20. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the multimedia recommendation method as described in any one of claims 1 to 5 or the object interest recognition model training method as described in any one of claims 6 to 9.
21. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the multimedia recommendation method according to any one of claims 1 to 5 or the object interest recognition model training method according to any one of claims 6 to 9.
Citation Information
Patent Citations
Content recommendation method, device and equipment and readable storage medium
CN111680219A