Resource recommendation method, method and device for training deep learning model, and agent
By fusing resource features through a window attention mechanism and recommending target resources based on semantic differences, the problem of fixed user interest boundaries and diminishing resource diversity is solved, thereby improving the diversity and accuracy of resource recommendations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2026-03-27
AI Technical Summary
In existing technologies, the information recommended by users suffers from fixed interest boundaries and diminished resource diversity, making it difficult to meet users' potential needs for expanding their interests, thus reducing the diversity and accuracy of recommended resources.
By fusing resource content feature sequences through a window attention mechanism, multiple resource fusion features are obtained, and target resources are determined from candidate resources based on semantic difference conditions, and target resources that meet the user's interest expansion are recommended.
It improves the diversity and accuracy of resource recommendations, meets users' needs for expanding their interests, and enhances the quality of resource recommendations and user satisfaction.
Smart Images

Figure CN120653846B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, and in particular to the technical field of resource recommendation, intelligent search, big data, etc. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, users can conveniently browse resource information such as news and videos through terminal devices such as smart phones. Related Internet platforms can also recommend resource information to users based on the needs or preferences of users. SUMMARY
[0003] The present disclosure provides a resource recommendation method, a method and apparatus for training a deep learning model, an agent, an electronic device, a storage medium, and a computer program product.
[0004] According to an aspect of the present disclosure, a resource recommendation method is provided, including: obtaining a resource content feature sequence for a target object; fusing a plurality of resource content features in the resource content feature sequence based on a window attention mechanism to obtain a plurality of resource fusion features; and determining a target resource from candidate resources based on the plurality of resource fusion features and recommending the target resource to the target object, wherein a candidate topic of the candidate resource and an initial topic for the resource content feature sequence satisfy a semantic difference condition.
[0005] According to another aspect of the present disclosure, a method for training a deep learning model is provided, including: obtaining a sample resource content feature sequence for a sample object and a label recommendation weight of a sample candidate resource, wherein a sample candidate topic of the sample candidate resource and a sample initial topic for the sample resource content feature sequence satisfy a semantic difference condition, and the label recommendation weight represents a satisfaction degree of the sample object to the sample candidate resource; fusing a plurality of sample resource content features in the sample resource content feature sequence based on a window attention mechanism using a first feature fusion layer of the deep learning model to obtain a plurality of sample resource fusion features; and processing the plurality of sample resource fusion features using a recommendation weight detection network of the deep learning model to obtain a sample recommendation weight for the sample candidate resource; training the deep learning model based on the label recommendation weight and the sample recommendation weight to obtain a trained deep learning model.
[0006] According to another aspect of the present disclosure, a resource recommendation apparatus is provided, including: a first obtaining module configured to obtain a resource content feature sequence for a target object; a first fusing module configured to fuse a plurality of resource content features in the resource content feature sequence based on a window attention mechanism to obtain a plurality of resource fusion features; and a first determining module configured to determine a target resource from candidate resources based on the plurality of resource fusion features and recommend the target resource to the target object, wherein a candidate topic of the candidate resource and an initial topic for the resource content feature sequence satisfy a semantic difference condition.
[0007] According to another aspect of the present disclosure, a method for training a deep learning model is provided, comprising: a second acquisition module configured to acquire a sample resource content feature sequence of a sample object and a label recommendation weight of a sample candidate resource, wherein a semantic difference condition is satisfied between a sample candidate topic of the sample candidate resource and a sample initial topic used for the sample resource content feature sequence, and the label recommendation weight represents a satisfaction degree of the sample object to the sample candidate resource; a second fusion module configured to fuse a plurality of sample resource content features in the sample resource content feature sequence based on a window attention mechanism to obtain a plurality of sample resource fusion features by using a first feature fusion layer of the deep learning model; and a sample recommendation weight obtaining module configured to process the plurality of sample resource fusion features by using a recommendation weight detection network of the deep learning model to obtain a sample recommendation weight of the sample candidate resource; and a training module configured to train the deep learning model based on the label recommendation weight and the sample recommendation weight to obtain a trained deep learning model.
[0008] According to another aspect of the present disclosure, an artificial intelligence agent is provided, comprising: an input module configured to receive input information; a processing module configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, execute the resource recommendation method described above by calling the large model, or execute the method for training a deep learning model described above by calling the large model, and obtain output information; and an output module configured to output the output information obtained by the processing module.
[0009] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.
[0010] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described above.
[0011] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method described above.
[0012] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS
[0013] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0014] Figure 1 This illustration schematically shows an exemplary system architecture to which resource recommendation methods and apparatus can be applied according to embodiments of the present disclosure;
[0015] Figure 2 A flowchart illustrating a resource recommendation method according to an embodiment of the present disclosure is shown schematically;
[0016] Figure 3 The illustration shows a schematic diagram of the principle of fusing multiple resource content features based on a window attention mechanism according to an embodiment of the present disclosure;
[0017] Figure 4 The illustration shows a schematic diagram of the principle of a deep learning model according to an embodiment of the present disclosure;
[0018] Figure 5 A flowchart illustrating a method for training a deep learning model according to an embodiment of the present disclosure is shown schematically.
[0019] Figure 6 The illustration shows a schematic diagram illustrating the principle of training a deep learning model according to an embodiment of the present disclosure;
[0020] Figure 7 A block diagram of a resource recommendation apparatus according to an embodiment of the present disclosure is shown schematically;
[0021] Figure 8 A block diagram of an apparatus for training a deep learning model according to an embodiment of the present disclosure is shown schematically;
[0022] Figure 9 A schematic diagram illustrating the structure of an intelligent agent of artificial intelligence according to embodiments of the present disclosure; and
[0023] Figure 10 A schematic block diagram of an example electronic device is shown that can be used to implement the resource recommendation method and the method for training a deep learning model in the embodiments of this disclosure. Detailed Implementation
[0024] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0025] In the technical solutions of the present disclosure, the acquisition, storage and application of user personal information comply with relevant laws and regulations, necessary security measures are taken, and the public order and good customs are not violated.
[0026] The inventors found that the information recommended to the user has the trend of interest boundary solidification and resource diversity attenuation, which is difficult to meet the potential interest expansion needs of the user, limits the diversity and richness of recommended resources, and reduces the quality of recommended resources.
[0027] Embodiments of the present disclosure provide a resource recommendation method, a method and apparatus for training a deep learning model, an agent, an electronic device, a storage medium, and a computer program product. The resource recommendation method comprises: obtaining a resource content feature sequence for a target object; fusing multiple resource content features in the resource content feature sequence based on a window attention mechanism to obtain multiple resource fusion features; and determining a target resource from candidate resources based on the multiple resource fusion features, and recommending the target resource to the target object, wherein a candidate theme of the candidate resource and an initial theme used for the resource content feature sequence satisfy a semantic difference condition.
[0028] According to embodiments of the present disclosure, by fusing multiple resource content features in the resource content feature sequence based on the window attention mechanism, the resource fusion features can pay more attention to the attention fusion of the multiple resource content features that are locally adjacent based on the window in the resource content feature sequence, so that the resource fusion features can pay more attention to the local interaction interest of the target object in each shorter period. Then, based on the local interaction interest represented by each of the multiple resource fusion features, the target resource that can meet the interest of the target object can be determined from the candidate resources that satisfy the semantic difference with the initial theme, to improve the matching degree between the target resource and the interaction interest of the target object. At the same time, by recommending the target resource that has a theme semantic difference with the initial theme that the target object is currently interested in, the homogenization of resource recommendation is avoided, so that the target resource can be matched with the interest needs of the target object, and the diversity of resources obtained by the target object is expanded, to improve the diversity and accuracy of resource recommendation.
[0029] Figure 1 An exemplary system architecture to which the resource recommendation method and apparatus according to embodiments of the present disclosure can be applied is schematically shown.
[0030] It should be noted that, Figure 1The examples shown are merely illustrative of system architectures applicable to embodiments of this disclosure, intended to help those skilled in the art understand the technical content of this disclosure. They do not imply that embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios. For instance, in another embodiment, an exemplary system architecture to which the resource recommendation method and apparatus can be applied may include a terminal device. However, the terminal device can implement the resource recommendation method and apparatus provided by embodiments of this disclosure without interacting with a server.
[0031] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0032] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (for example only).
[0033] Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0034] Server 105 can be a server that provides various services, such as a backend management server that supports the content browsed by users using terminal devices 101, 102, and 103 (for example only). The backend management server can analyze and process data such as received user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0035] Server 105 can be a cloud server, also known as a cloud computing server or cloud host. It is a host product in the cloud computing service system, which solves the shortcomings of traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"), such as high management difficulty and weak business scalability. Server 105 can also be a server for a distributed system or a server combined with blockchain.
[0036] It should be noted that the resource recommendation method provided in the embodiments of the present disclosure can generally be executed by the terminal device 101, 102, or 103. Accordingly, the resource recommendation apparatus provided in the embodiments of the present disclosure can also be arranged in the terminal device 101, 102, or 103.
[0037] Alternatively, the resource recommendation method provided in the embodiments of the present disclosure can also be generally executed by the server 105. Accordingly, the resource recommendation apparatus provided in the embodiments of the present disclosure can be generally arranged in the server 105. The resource recommendation method provided in the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal device 101, 102, 103 and / or the server 105. Accordingly, the resource recommendation apparatus provided in the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the terminal device 101, 102, 103 and / or the server 105.
[0038] It should be understood that Figure 1 The number of terminal devices, networks, and servers in the above-mentioned system is only illustrative. According to the implementation needs, there can be any number of terminal devices, networks, and servers.
[0039] Figure 2 A flowchart of a resource recommendation method according to an embodiment of the present disclosure is schematically shown.
[0040] As shown in Figure 2 The resource recommendation method includes operations S210-S230.
[0041] In operation S210, a resource content feature sequence for a target object is obtained.
[0042] In operation S220, multiple resource fusion features are obtained by fusing multiple resource content features in the resource content feature sequence based on a window attention mechanism.
[0043] In operation S230, a target resource is determined from candidate resources based on the multiple resource fusion features, and the target resource is recommended to the target object.
[0044] According to an embodiment of the present disclosure, the resource content feature sequence can include multiple resource content features, and the resource content features can represent title content, text content, image content, and the like of the interactive resource. The interactive resource can be a resource on which the target object has performed an interactive operation in a specified period of time, such as a video resource, a news resource, and the like that the target object has browsed in a specified period of time.
[0045] In an embodiment, the resource content feature characterizes a satisfactory consumption resource for the target object. The satisfactory consumption resource indicates an interaction resource with a higher satisfaction degree in the interaction resource in which the target object performs the interaction operation. The satisfaction degree can be determined based on an interaction behavior indicator of the target object for the interaction resource, for example, the satisfaction degree can be determined based on a browsing time length, a like behavior, a comment text preference, and the like of the target object for the interaction resource. Embodiments of the present disclosure do not limit the specific manner of determining the satisfactory consumption resource. By determining the resource content feature based on the satisfactory consumption resource, the resource content feature sequence can more accurately characterize the interaction behavior preference and the resource interest intention of the target object. Thus, based on the resource fusion feature accurately characterizing the resource interest intention of the target object, the target resource that matches the interest intention or the interaction behavior preference of the target object can be determined from the candidate resource with the candidate theme that has a semantic difference from the initial theme, so that the target resource meets the interest expansion demand of the target object.
[0046] It should be noted that the candidate theme in the embodiments of the present disclosure can be represented as a resource theme of the candidate resource, and the initial theme can be represented as a resource theme of the interaction resource in which the target object performs the interaction operation. The interaction operation can be understood as an operation performed by the target object through a smart terminal such as a smart phone, and the interaction behavior can be understood as a type, a frequency, and the like of the interaction attribute of the target object by performing the interaction operation.
[0047] According to embodiments of the present disclosure, based on the window attention mechanism fusing the plurality of resource content features, the local plurality of resource content features in the resource content feature sequence that match the window size can be attention-fused through the size range of the window, so that the resource fusion feature can characterize the local resource content in the resource content feature sequence corresponding to the window size, and the global resource content of the resource content feature sequence can be represented through the resource fusion features corresponding to the plurality of windows. Thus, while reducing the computational overhead of attention calculation, the local resource content range learned by the resource fusion feature can be controlled through the window size, so that the resource fusion feature can more accurately characterize the interaction behavior preference and the interest demand of the target object.
[0048] According to an embodiment of the present disclosure, feature matching can be performed between the plurality of resource fusion features and the resource features of the candidate resources, and a target resource matching the plurality of resource fusion features can be determined from the candidate resources according to the feature matching result. However, this is not a limitation. The target resource can also be determined from the candidate resources based on the plurality of resource fusion features in other manners, for example, a neural network algorithm such as a multi-layer perception can be used to process the plurality of resource fusion features to obtain a recommendation weight of the candidate resources, and the target resource meeting the recommendation weight condition can be determined from one or more candidate resources based on the recommendation weight. The embodiments of the present disclosure do not limit the specific manner of determining the target resource, as long as the plurality of resource fusion features can be used.
[0049] According to an embodiment of the present disclosure, the candidate topic of the candidate resource meets the semantic difference condition with the initial topic used for the sequence of resource content features. In this way, the initial topic represented by the sequence of resource content features can be more accurately detected by the plurality of resource fusion features representing the local interaction interest and the global interaction demand of the target object, so that the target resource with the topic semantic difference is obtained, thereby expanding the resource diversification acquisition demand of the target object. In addition, the plurality of resource fusion features representing the interaction interest of the target object can be used to match the target resource with the interaction behavior preference of the target object, thereby improving the recommendation quality of the target resource.
[0050] In one embodiment, the resource content features are determined based on the interaction resources related to the target object, and the arrangement order of the plurality of resource content features in the sequence of resource content features is determined according to the interaction time attribute of each of the plurality of interaction resources by the target object. For example, the arrangement order of the plurality of interaction resources can be determined based on the time when the target object performs the interaction behavior on the plurality of interaction resources, and the arrangement order of the plurality of resource content features in the sequence of resource content features can be further determined.
[0051] By arranging the plurality of resource content features according to the interaction time attribute of the target object on the interaction resources, the plurality of resource content features in the window range can be fused based on the window attention mechanism, and the resource fusion features can learn the time sequence preference order of the target object performing the interaction behavior on the plurality of interaction resources, so that the resource fusion features can further represent the change of the interaction behavior interest of the target object in the local time period in the specified time period. In this way, the target resource matching the change of the interaction behavior interest of the target object can be determined based on the plurality of resource fusion features, thereby expanding the range of resource recommendation topics while meeting the interest change demand of the target object for resource acquisition.
[0052] According to an embodiment of the present disclosure, the resource content features can include resource title text features, resource body text features, and resource speech text features.
[0053] In an example, text features can be extracted from the resource title text of the interaction resource to obtain resource title text features. The resource title text features can accurately represent the theme and type of the interaction resource content based on fewer vector dimensions, and the resource content features can be determined based on the resource title text features to represent the resource content semantics based on fewer feature vector dimensions.
[0054] In an example, the resource content features can be determined by text feature extraction on the body text of the resource, for example, the resource body text can be processed based on a text encoder to obtain resource body text features. The body text of the resource can comprehensively and in detail represent the resource content of the interaction resource, and thus the resource body text features can comprehensively represent the semantic content of the resource to reduce the loss of resource content information.
[0055] In an example, the resource speech text features can be obtained by feature extraction on resources with speech information such as video resources and speech resources, or can also be determined by text feature extraction on text content representing speech information such as video subtitles and lyric subtitles.
[0056] According to an embodiment of the present disclosure, the resource content features can be obtained by fusing any two or more of the resource title text features, the resource body text features, and the resource speech text features, so as to improve the accuracy and richness of the resource content features in representing the resource content information.
[0057] In an example, the resource content features are semantic vectors obtained by processing the title text of the satisfactory consumption resource for the target object based on a CLIP (Contrastive Language-Image Pretraining) model.
[0058] In an example, the resource content features are determined based on feature extraction on the title text, resource image, and body text of the satisfactory consumption resource for the target object by a CLIP model. The CLIP model maps the resource image and text into a shared semantic space by jointly training an image encoder and a text encoder, and can effectively capture deep semantic information of the text, thereby improving the capture accuracy of the resource fusion features on the interaction behavior interest of the target object.
[0059] According to an embodiment of the present disclosure, the fusing of the plurality of resource content features in the resource content feature sequence based on the window attention mechanism comprises: determining, for a target resource content feature in the resource content feature sequence, an associated resource content feature from the resource-related feature sequence according to a weight window for the target resource content feature; and fusing the target resource content feature and the associated resource content feature based on the attention mechanism to obtain a resource fusion feature related to the target resource content feature.
[0060] According to an embodiment of the present disclosure, the target resource content feature represents the target interactive resource, and the associated resource content feature can represent an associated resource adjacent to the target resource in the interactive resource.
[0061] According to an embodiment of the present disclosure, the window size of the weight window is determined based on the resource satisfaction of the target object to the target interactive resource. The window size can represent the attention range of the weight window. Thus, the window size that matches the satisfaction of the target object to the target resource content feature can be dynamically determined based on the satisfaction of the target object to the target resource content feature. Then, the number of associated resource content features that matches the attention range represented by the window size can be determined from the sequence of resource content features according to the dynamic window size. In this way, the window attention range in the window attention calculation process can be dynamically adjusted based on the satisfaction of the target object to the target interactive resource. The local feature fusion range of the resource fusion feature can be more accurately dynamically adjusted for the interactive resource with higher satisfaction of the target object. The representation accuracy of the resource fusion feature to the interactive interest preference of the target object can be improved. Thus, the matching degree of the interactive interest of the target resource and the target object can be improved.
[0062] In one example, the window size of the weight window is in a positive proportional relationship with the satisfaction of the target resource content feature. Thus, the window attention range can be expanded for the target resource content feature with higher satisfaction of the target object. Thus, a larger number of adjacent associated resource content features in the sequence of resource content features can be queried for the target resource content feature with higher satisfaction. The resource fusion feature obtained based on the window attention mechanism can learn more adjacent resource content features. Thus, the target resource determined based on multiple resource fusion features can expand the theme range of the recommended resource while improving the satisfaction of the target object to the target resource.
[0063] According to an embodiment of the present disclosure, the arrangement order of the plurality of resource content features in the resource content feature sequence can correspond to the interaction time sequence represented by the respective interaction time attributes of the plurality of interaction resources of the target object. By dynamically adjusting the window size of the weight window according to the satisfaction degree of the target resource content feature, the target resource content feature with a higher satisfaction degree can be given a higher attention focus range through the weight window with a dynamically changing attention range, and the determined resource fusion feature can learn the interaction behavior change attribute of the target object in a longer time period based on the position embedding mechanism calculated by attention, so that the resource fusion feature can learn the interaction behavior interest of the resource content feature with a higher satisfaction degree in a longer time period. Therefore, the target resource determined based on the plurality of resource fusion features can further more accurately represent the interaction behavior interest change trend of the target object, so as to improve the matching degree between the target resource and the interest intention of the target object, and improve the resource recommendation quality.
[0064] It should be noted that the target interaction resource involved in the embodiments of the present disclosure can represent a resource that the target object has interacted with in a specified time period, and the target resource can represent a resource currently or in the future recommended to the target object.
[0065] According to an embodiment of the present disclosure, the resource satisfaction degree is determined based on at least one of the following interaction behavior indicators of the target object for the target interaction resource: a browsing time length indicator, a like behavior indicator, a comment behavior indicator, and a collection behavior indicator.
[0066] According to an embodiment of the present disclosure, the browsing time length indicator can be determined based on the browsing time length data of the target object for the target interaction resource. For example, the browsing time length indicator can be determined based on the ratio between the browsing time length and a preset browsing time length, or the browsing time length indicator can also be determined after the browsing time length data is normalized. The specific determination manner of the browsing time length indicator is not limited in the embodiments of the present disclosure, as long as the browsing time length of the target object for the target interaction resource can be represented.
[0067] According to an embodiment of the present disclosure, the like behavior indicator can represent the like interaction behavior data of the target object for the target interaction resource, such as the like frequency and the like duration. The like behavior indicator can be determined by quantifying the like interaction behavior data according to a preset rule, and then the resource satisfaction degree can be determined based on the like behavior indicator.
[0068] According to an embodiment of the present disclosure, the comment behavior indicator can be determined based on the comment word number, the comment sentiment attribute, and other attribute data related to the comment behavior. The collection behavior indicator can be determined based on the collection frequency, the collection duration, and other collection behaviors.
[0069] In an embodiment, the comment behavior attribute data, the collection behavior data, the browsing time length data, and the like, related to the target interactive resource are quantified based on preset quantification rules to obtain quantified browsing time length indicators, like indicators, comment behavior indicators, and collection behavior indicators. At least two of the browsing time length indicators, the like indicators, the comment behavior indicators, and the collection behavior indicators are weighted to obtain a quantified resource satisfaction degree of the target interactive resource. Thus, the weight window size corresponding to the target resource content feature can be obtained based on the quantified resource satisfaction degree.
[0070] According to an embodiment of the present disclosure, the fusion of the target resource content feature and the associated resource content feature based on the attention mechanism can include: determining a query feature based on the target resource content feature; determining a value feature and a key feature based on the weight number of associated resource content features; and fusing the query feature, the value feature, and the key feature based on an attention algorithm.
[0071] According to an embodiment of the present disclosure, determining the query feature based on the target resource content feature can include fusing and calculating the trained query weight and the target resource content feature to obtain the query feature.
[0072] According to an embodiment of the present disclosure, the weight number matches the number of features represented by the window size, which can be represented by the number of associated resource content features, or the number of features can also be understood as the number of associated interactive resources corresponding to the window size. Determining the value feature and the key feature based on the weight number of associated resource content features can include fusing and calculating the value weight and the weight number of associated resource content features to obtain the weight number of value feature vectors; and fusing and calculating the key weight and the weight number of associated resource content features to obtain the weight number of key feature vectors. The attention calculation operation of the attention algorithm is performed on the weight number of key feature vectors and value feature vectors, and the query feature to obtain the resource fusion feature of the target resource content feature of the user. It should be understood that the fusion calculation can include dot product calculation operation.
[0073] In an embodiment, the plurality of resource fusion features can also be arranged according to the interactive time attributes of the plurality of interactive resources, so that the deep learning model can further fully learn the interactive behavior interest trend of the target object by learning the order of the plurality of resource fusion features, so that the target resource can expand the interest theme range while more accurately representing the interest intention of the target object.
[0074] Figure 3 The principle diagram of fusing a plurality of resource content features based on a window attention mechanism according to an embodiment of the present disclosure is schematically shown.
[0075] As Figure 3As shown, the resource content feature sequence 3100 includes a plurality of resource content features A1 to A10 arranged according to the interactive event attributes. The window size of the first weight window C301 corresponding to the first target resource content feature A5 can be determined based on the resource satisfaction degree of the target interactive resource corresponding to the first target resource content feature A5. The plurality of associated resource content features associated with the first target resource content feature A5 in the window range based on the window attention mechanism can be determined as the plurality of first associated resource content features A2, A3 and A4 based on the window size of the first weight window C301. Based on the attention algorithm, A2, A3, A4 and A5 can be fused to determine a query feature (Query) based on A5, a plurality of value features (Value) and a plurality of key features (Key) based on A2, A3 and A4, and a first resource fusion feature B5 corresponding to the first target resource feature A5 based on the attention algorithm to fuse the query feature (Query), the plurality of value features (Value) and the plurality of key features (Key).
[0076] As shown, Figure 3 Based on the resource satisfaction degree of the target interactive resource corresponding to the second target resource content feature A9, the window size of the second weight window C302 corresponding to the second target resource content feature A9 can be determined. The plurality of associated resource content features associated with the second target resource content feature A5 in the window range based on the window attention mechanism can be determined as the plurality of second associated resource content features A7 and A8 based on the window size of the second weight window C302. Based on the attention algorithm, A7, A8 and A9 can be fused to determine a query feature (Query) based on A9, two value features (Value) and two key features (Key) based on A7 and A8, and a second resource fusion feature B9 corresponding to the second target resource feature A9 based on the attention algorithm to fuse the query feature (Query), the two value features (Value) and the two key features (Key).
[0077] It should be understood that the plurality of resource content features A1 to A10 can each correspond to a resource fusion feature, and embodiments of the present disclosure will not be repeated here.
[0078] It should be noted that, Figure 3The window size or the weight window position of the first weight window or the second weight window shown is only illustrative and is not intended to limit the window size or the window position of the weight window. The window size of the weight window can be dynamically adjusted according to the resource satisfaction of the user target resource interaction feature based on actual needs, and the weight window can be dynamically changed in the window size at any window position based on the window attention mechanism, so as to perform attention calculation on the local multiple resource content features in the weight window range to obtain the resource fusion features corresponding to the resource content features in the window attention range corresponding to the resource satisfaction.
[0079] According to an embodiment of the present disclosure, determining the target resource from the candidate resources based on the multiple resource fusion features includes: fusing the initial topic feature representing the initial topic and the multiple resource fusion features to obtain a target fusion feature; determining a recommendation weight for the candidate resource based on the target fusion feature and a candidate topic feature representing a candidate topic; and determining the target resource based on the recommendation weight.
[0080] In one embodiment, the initial topic feature and the multiple resource interaction fusion features are fused based on a cross-attention algorithm to obtain the target fusion feature.
[0081] In one embodiment, the candidate topic feature is obtained by performing semantic encoding on the candidate topic based on an encoder. The candidate topic feature represents the topic semantic attribute of the candidate topic.
[0082] In one embodiment, the initial topic features representing the initial topics of the multiple interaction resources and the multiple resource interaction fusion features are fused based on a cross-attention algorithm to achieve adaptive learning of the resource fusion features and the initial topic features representing the resource content, so that the resource content semantics represented by the resource fusion features can be queried based on the initial topic features, so that the target fusion feature can learn the relationship between the initial topic and the resource content semantics and the interaction behavior interest intention. Therefore, the interest intention of the target object for the candidate resource can be dynamically represented based on the target fusion feature, and the recommendation weight for the candidate resource can be detected based on the candidate topic feature representing the candidate resource and the target fusion feature, so that the recommendation weight can more accurately represent the interaction interest intention of the target object.
[0083] It should be noted that the recommendation weight can be any type of data such as a weight value, a weight vector, etc., and the embodiments of the present disclosure do not limit the specific data type of the recommendation weight. The candidate topic and the candidate topic involved in the embodiments of the present disclosure can represent the topic semantics of the candidate resource, which will not be described here again.
[0084] In one example, the candidate resources with higher recommendation weights can be recalled by determining the recommendation weights of the plurality of candidate resources, and recalling the plurality of candidate resources based on the recommendation weights in a resource recall stage, for example, recalling candidate resources with top pre-set ranking positions of the recommendation weights as the recalled resources. The recalled resources are scored comprehensively by combining the click rate index, the conversion rate index and other recommendation indexes, and the target resource with higher comprehensive score is determined from the recalled resources based on the comprehensive score, and is recommended to the target object.
[0085] In one embodiment, determining the recommendation weight for the candidate resource based on the target fusion feature and the candidate topic feature representing the candidate topic comprises: determining the recommendation weight based on the target fusion feature, the candidate topic feature and the interaction scenario feature for the target object.
[0086] According to embodiments of the present disclosure, the interaction scenario feature is determined based on interaction scenario information for the target object in a specified period. The interaction scenario information can include information related to the interaction scenario of the target object, for example, the time period of the interaction behavior, the interaction behavior frequency, the specified interaction behavior frequency, etc. Embodiments of the present disclosure do not limit the specific type of interaction scenario information. The interaction scenario feature can be determined by feature extraction on the interaction behavior information.
[0087] According to embodiments of the present disclosure, the target fusion feature, the candidate topic feature and the interaction scenario feature for the target object can be processed based on a neural network algorithm to obtain the recommendation weight.
[0088] In one example, the target fusion feature, the candidate topic feature and the interaction scenario feature for the target object can be processed based on a fully connected layer constructed based on a multi-layer perceptron algorithm to obtain the recommendation weight. In this way, the complex semantics in the target fusion feature fused based on the window attention mechanism and the cross attention mechanism can be captured based on the fully connected layer, the semantic understanding ability for longer resource content feature sequences can be improved, and the calculation efficiency can be improved to efficiently and accurately determine the recommendation weight of the candidate resource.
[0089] In one embodiment, based on the target fusion feature and the candidate topic feature representing the candidate topic, the recommendation weight for the candidate resource can also include processing the target fusion feature, the candidate topic feature and the interaction scenario feature based on a neural network algorithm, and other feature information related to resource recommendation such as the object attribute feature representing the object attribute of the target object and the candidate resource attribute feature. In this way, the recommendation weight for the candidate resource can be accurately detected by multi-dimensional feature information to improve the matching degree between the target resource and the target object.
[0090] Figure 4 The principle schematic diagram of the deep learning model according to embodiments of the present disclosure is schematically shown.
[0091] As shown in Figure 4 The deep learning model includes a first feature fusion layer, a second feature fusion layer, and a recommendation weight detection layer. The first feature fusion layer is constructed based on a window attention algorithm, the second feature fusion layer is constructed based on a cross-attention algorithm, and the recommendation weight detection layer is constructed based on a multi-layer perceptron algorithm.
[0092] In this embodiment, the interaction resource for the target object is a plurality of satisfactory consumption resource contents in the satisfactory consumption resource sequence 410 of the target object in a specified period, and the plurality of satisfactory consumption resource contents in the satisfactory consumption resource sequence 410 are arranged based on the interaction behavior time attribute of the target object. The plurality of satisfactory consumption resource contents in the satisfactory consumption resource sequence 410 are processed by the multi-modal encoder constructed based on the CLIP model to obtain the resource content feature sequence 430. The plurality of resource content features of the resource content feature sequence 430 correspond to the arrangement order of the plurality of satisfactory consumption resource contents. The initial theme sequence 420 is obtained by processing the plurality of initial themes in the initial theme sequence 420 by using the feature embedding layer constructed based on the encoder.
[0093] The resource content feature sequence 430 is input into the first feature fusion layer, and the first feature fusion layer fuses the plurality of resource content features in the resource content feature sequence 430 based on the window attention mechanism to obtain the resource fusion feature sequence 450. The resource fusion feature sequence 450 includes a plurality of resource fusion features. The second feature fusion layer is used to perform cross-attention fusion on the resource fusion feature sequence 450 and the initial theme feature sequence 440 to obtain a target fusion feature. The obtained feature set 460 related to the candidate resource includes candidate theme features representing candidate themes of the candidate resource, interaction scene features and object attribute features for the target object, and other types of feature information such as item features representing the candidate resource.
[0094] The recommendation weight for the candidate resource is obtained by inputting the spliced feature set 460 and the target fusion feature into the recommendation weight detection layer. The window attention mechanism is introduced to fuse multiple resource content features so that the resource fusion feature can sufficiently learn the interactive behavior interest intention and the resource content interest intention of the target object. The initial topic feature is cross-attention fused with the multiple resource fusion features to further accurately represent the interest intention of the target object. Therefore, the candidate topic feature, the object attribute feature, and other feature information related to the candidate resource and the target object in the feature set for the candidate resource can be used to more accurately mine the interactive behavior preference and the interest expansion intention of the target object, so that the recommendation weight can represent the potential resource interest demand of the target object. Therefore, the target resource that can meet the interest expansion demand of the target object can be determined from the candidate resource according to the recommendation weight, the “information cocoon” effect can be broken through by recommending the target resource to the target object, the satisfaction of the target object with the resource recommendation is improved, and the activity and long-term interaction stickiness of the target object are improved.
[0095] Based on the resource recommendation method provided in the above embodiments, an embodiment of the present disclosure further provides a method for training a deep learning model.
[0096] Figure 5 A flowchart of the method for training a deep learning model according to an embodiment of the present disclosure is schematically shown.
[0097] As shown in Figure 5 The method for training a deep learning model includes operations S510-S540.
[0098] In operation S510, a sample resource content feature sequence for a sample object and a label recommendation weight of a sample candidate resource are obtained.
[0099] In operation S520, based on a window attention mechanism, a first feature fusion layer of a deep learning model is used to fuse multiple sample resource content features in the sample resource content feature sequence to obtain multiple sample resource fusion features.
[0100] In operation S530, a recommendation weight detection network of the deep learning model is used to process the multiple sample resource fusion features to obtain a sample recommendation weight for the sample candidate resource.
[0101] In operation S540, the deep learning model is trained based on the label recommendation weight and the sample recommendation weight to obtain a trained deep learning model.
[0102] According to an embodiment of the present disclosure, the sample candidate topic of the sample candidate resource and the sample initial topic for the sample resource content feature sequence satisfy a semantic difference condition, so that the sample candidate resource to be recommended can meet the topic interest expansion demand of the target object.
[0103] According to an embodiment of the present disclosure, the sample resource content feature is determined based on resource content of a sample interaction resource, for example, the sample resource content feature can be obtained based on semantic encoding of the resource content of the sample interaction resource by an encoder.
[0104] According to an embodiment of the present disclosure, the label recommendation weight characterizes the satisfaction of the sample object to the sample candidate resource, thereby the deep learning model can be trained based on the label recommendation weight and the sample recommendation weight, so that the deep learning model can detect the matching degree of the candidate resource with semantic difference from the initial theme of the interaction resource to the interest intention of the target object, so that the trained deep learning model can meet the interest intention of the target object and expand the interest theme range of the target object, improve the diversity and accuracy of resource recommendation, and improve the quality of resource recommendation.
[0105] It should be noted that the trained deep learning model determined by the method for training the deep learning model provided by the embodiments of the present disclosure can be applied to the resource recommendation method provided in the above embodiments. For example, the trained deep learning model can be used to process the resource content feature sequence for the target object to obtain a plurality of resource fusion features, determine the recommendation weight for the candidate resource based on the plurality of resource fusion features, and determine the target resource from the candidate resource by using the recommendation weight.
[0106] According to an embodiment of the present disclosure, based on the window attention mechanism, the first feature fusion layer of the deep learning model is used to fuse a plurality of sample resource content features in the sample resource content feature sequence, including: for a sample target resource content feature in the sample resource content feature sequence, determining a sample associated resource content feature from the sample resource related feature sequence according to a sample weight window for the sample target resource content feature; and fusing the sample target resource content feature and the sample associated resource content feature based on the attention mechanism to obtain a sample resource fusion feature related to the sample target resource content feature.
[0107] According to an embodiment of the present disclosure, the window size of the sample weight window is determined based on the resource satisfaction of the sample object to the sample target interaction resource, and the sample target resource content feature characterizes the sample target interaction resource.
[0108] In one example, the resource satisfaction of the sample target interaction resource can be determined based on a sample interaction behavior indicator of the sample object to the sample interaction resource.
[0109] It should be noted that the technical terms involved in the embodiments of the present disclosure, including but not limited to sample target interaction resource, sample resource content feature, etc., have the same or corresponding properties as the technical terms provided by the embodiments of the present disclosure, including but not limited to target interaction resource, resource content feature, etc., and the embodiments of the present disclosure will not be repeated here.
[0110] According to an embodiment of the present disclosure, the recommendation weight detection network based on the recommendation weight detection network includes a second feature fusion layer and a recommendation weight detection layer. The recommendation weight detection network using the deep learning model processes the plurality of sample resource fusion features to obtain the sample recommendation weight for the sample candidate resource, including: based on the attention mechanism, using the second feature fusion layer to fuse the sample initial topic features representing the sample initial topics and the plurality of sample resource fusion features to obtain the sample target fusion features; and using the recommendation weight detection layer to determine the sample recommendation weight based on the sample target fusion features and the sample candidate topic features representing the sample candidate topics.
[0111] Figure 6 The principle schematic diagram of training the deep learning model according to an embodiment of the present disclosure is schematically shown.
[0112] As shown in Figure 6 , the deep learning model includes a first feature fusion layer, a second feature fusion layer and a recommendation weight detection layer. The first feature fusion layer is constructed based on a window attention algorithm, the second feature fusion layer is constructed based on a cross-attention algorithm, and the recommendation weight detection layer is constructed based on a multi-layer perceptron algorithm.
[0113] In this embodiment, the plurality of sample resource content features of the sample resource content feature sequence 610. The initial topics of the plurality of satisfactory consumption resource contents in the satisfactory consumption resource sequence 410 are taken as the initial topic sequence 420. The plurality of initial topics in the initial topic sequence 420 are processed using the feature embedding layer constructed based on the encoder to obtain the initial topic feature sequence 440.
[0114] The plurality of sample resource content features of the sample resource content feature sequence 610 are input into the first feature fusion layer, and the first feature fusion layer fuses the plurality of sample resource content features in the sample resource content feature sequence 610 based on the window attention mechanism to obtain the plurality of sample resource fusion features in the sample resource fusion feature sequence 630. The second feature fusion layer is used to cross-attention fuse the sample resource fusion feature sequence 630 and the sample initial topic feature sequence 620 to obtain the target fusion feature. The sample feature set 640 related to the sample candidate resource obtained includes the sample candidate topic feature representing the sample candidate topic of the sample candidate resource, the sample interaction scene feature and the sample object attribute feature for the sample object, etc. The sample feature set 640 can also include other types of feature information such as the item feature represented by the sample candidate resource.
[0115] The sample recommendation weight of the sample candidate resource is obtained by inputting the spliced sample feature set 640 and target fusion feature into the recommendation weight detection layer. The sample recommendation weight of the sample candidate resource is processed by using a loss function to obtain a weight loss value. The deep learning model is trained by using the weight loss value until the weight loss value meets a convergence condition to obtain a trained deep learning model.
[0116] According to an embodiment of the present disclosure, the label recommendation weight is determined based on the evaluation of the sample object on the sample candidate resource. However, the label recommendation weight can also be determined based on the interaction behavior index of the sample object on the sample candidate resource. The specific setting method of the label recommendation weight is not limited in the embodiments of the present disclosure.
[0117] In one embodiment, the label recommendation weight is determined based on the following operations: obtaining sample interaction behavior change data of the sample object on the sample candidate theme; and determining the label recommendation weight of the sample candidate resource based on the sample interaction behavior change data of the sample candidate theme.
[0118] According to an embodiment of the present disclosure, the sample interaction behavior change data of the sample candidate theme can represent the change of the interaction behavior of the sample object on the resource with the sample candidate theme in a specified period. For example, the sample interaction behavior change data can be represented as a change value of the browsing time of the sample object on one or more resources with the sample candidate theme in a plurality of sub-periods of the specified period, a change value of the number of likes, and other data representing the change of the interaction behavior of the sample object.
[0119] In one embodiment, the sample interaction behavior change data represents the change of the interaction behavior of the sample object on the sample candidate theme after the sample object performs an interaction operation on the sample interaction resource represented by the sample resource content feature. The interaction behavior of the sample object on the sample candidate theme can be understood as the subsequent interaction behavior data of the sample object on the sample candidate resource with the sample candidate theme. The change of the interaction behavior can be understood as the difference information between the previous interaction behavior data generated by the sample object performing an interaction operation on the sample interaction resource with the initial theme and the subsequent interaction behavior data generated by the sample object performing an interaction operation on the sample candidate resource with the sample candidate theme.
[0120] For example, the sample interaction behavior change data can be represented as a ratio between the previous interaction behavior data and the subsequent interaction behavior data, a difference between the previous interaction behavior data and the subsequent interaction behavior data, and the like.
[0121] By determining the label recommendation weights of candidate resources based on sample interaction behavior change data, the label recommendation weights can quantify the changes in the interaction behavior of sample objects towards different sample resources with semantic differences. This allows the label recommendation weights to represent the consumption interest gain generated by recommending the candidate resource to the sample object, thus reflecting the sample object's interaction interest intent towards the candidate resource and the changing trend of the sample object's interaction interest towards sample resources with different themes. Therefore, the label recommendation weights can be used as labels to train a deep learning model, improving the accuracy of the deep learning model in detecting the changing trends of the target object's interaction behavior and the precision in detecting the target object's interest intent towards resources expanding their interest topics. The trained deep learning model can then process resource content feature sequences to identify target resources from the candidate resources that meet the target object's interaction behavior interests and theme expansion intent, improving resource recommendation quality, enhancing the target object's trust in the deployed deep learning model, increasing interaction stickiness, and achieving a positive and efficient interaction between the recommendation system and the target object.
[0122] It should be noted that the specific representation of the sample interaction behavior change data in this embodiment is not limited, as long as it can characterize the change of the interaction behavior of the sample object with the sample candidate topic after performing an interaction operation on the sample interaction resource.
[0123] According to embodiments of this disclosure, the sample interaction behavior change data includes at least one of the following: browsing duration change data, like behavior change data, comment behavior change data, and favorite behavior change data.
[0124] Browsing duration variation data can represent the changes in browsing time for sample objects. Determining tag recommendation weights based on browsing duration variation data can include quantifying the browsing duration variation data to obtain tag recommendation weights.
[0125] In one embodiment, the tag recommendation weight can be determined based on browsing time change data, which can be performed based on the following formula (1).
[0126] (1);
[0127] Wherein, gain represents the tag recommendation weight, topic_avg represents the average browsing time of sample objects on sample interactive resources in the first period, topic_time represents the browsing time of sample objects on sample candidate resources with sample candidate topics in the second period, and λ is a preset smoothing parameter used to reduce the interaction behavior data of active users.
[0128] The label recommendation weight can be quantified by a ratio between the post-browsing duration of the sample candidate resource and the average browsing duration of the previous, which can more accurately represent the change of the sample object's post-browsing duration to represent the sample object's consumption willingness for the sample resource with the sample candidate theme, so that the deep learning model can be trained by the label recommendation weight to capture the sample object's potentially interactive behavior interest and resource theme intention, avoiding the formation of an information cocoon.
[0129] According to an embodiment of the present disclosure, the like behavior change data, the comment behavior change data, and the collection behavior change data can respectively represent the change of the like behavior data, the comment behavior data, and the collection behavior data of the sample object for the sample candidate resource theme after the sample object performs an interactive operation on the sample interactive resource represented by the content feature of the sample resource in the previous period. The label recommendation weight can be determined by quantifying the like behavior change data, the comment behavior change data, or the collection behavior change data of the sample object. For example, the like behavior change data can be processed based on a normalization algorithm to determine the label weight.
[0130] It should be noted that the label recommendation weight can be obtained by quantifying any one or more of the browsing duration change data, the like behavior change data, the comment behavior change data, and the collection behavior change data. For example, one or more initial label recommendation weights can be obtained by quantifying any one or more of the browsing duration change data, the like behavior change data, the comment behavior change data, and the collection behavior change data, and the label recommendation weight can be obtained by weighting calculation of the multiple initial label recommendation weights.
[0131] According to an embodiment of the present disclosure, the label recommendation weight can be determined by diversifying the sample interactive behavior change data, which can make the deep learning model more accurately capture the interactive behavior change trend of the sample object and the preference degree of the sample candidate resource for the expanded sample candidate theme in the training process, thereby improving the resource recommendation quality and diversity of the deep learning model applied to the resource recommendation process.
[0132] Figure 7 A block diagram of a resource recommendation apparatus according to an embodiment of the present disclosure is schematically shown.
[0133] As shown in Figure 7 The resource recommendation apparatus 700 includes a first acquisition module 710, a first fusion module 720, and a first determination module 730.
[0134] The first acquisition module 710 is configured to acquire a resource content feature sequence for a target object.
[0135] The first fusion module 720 is configured to fuse a plurality of resource content features in the resource content feature sequence based on a window attention mechanism to obtain a plurality of resource fusion features.
[0136] The first determination module 730 is configured to determine a target resource from the candidate resources based on the plurality of resource fusion features, and recommend the target resource to the target object, wherein a semantic difference condition is met between a candidate topic of the candidate resource and an initial topic used for the resource content feature sequence.
[0137] According to an embodiment of the present disclosure, the first fusion module 720 includes a first determination unit and a first obtaining unit.
[0138] The first determination unit is configured to determine, for a target resource content feature in the resource content feature sequence, an associated resource content feature from the resource-related feature sequence according to a weight window used for the target resource content feature, wherein a window size of the weight window is determined based on a resource satisfaction of the target object to a target interactive resource, and the target resource content feature characterizes the target interactive resource.
[0139] The first obtaining unit is configured to fuse the target resource content feature and the associated resource content feature based on an attention mechanism to obtain a resource fusion feature related to the target resource content feature.
[0140] According to an embodiment of the present disclosure, the first obtaining unit includes a first determination sub-unit, a second determination sub-unit, and a fusion sub-unit.
[0141] The first determination sub-unit is configured to determine a query feature based on the target resource content feature.
[0142] The second determination sub-unit is configured to determine a value feature and a key feature based on the weight number of associated resource content features, wherein the weight number matches a feature number represented by the window size.
[0143] The fusion sub-unit is configured to fuse the query feature, the value feature, and the key feature based on an attention algorithm.
[0144] According to an embodiment of the present disclosure, the resource satisfaction is determined based on at least one of the following interaction behavior indicators of the target interactive resource: a browsing time length indicator, a like behavior indicator, a comment behavior indicator, and a collection behavior indicator.
[0145] According to an embodiment of the present disclosure, the first determination module 730 includes a target fusion feature obtaining unit, a recommendation weight determination unit, and a target resource determination unit.
[0146] The target fusion feature obtaining unit is configured to fuse an initial topic feature representing an initial topic and the plurality of resource fusion features to obtain a target fusion feature.
[0147] The recommendation weight determination unit is configured to determine a recommendation weight for the candidate resource based on the target fusion feature and the candidate topic feature representing the candidate topic.
[0148] The target resource determination unit is configured to determine the target resource based on the recommendation weight.
[0149] According to an embodiment of the present disclosure, the recommendation weight determination unit comprises a recommendation weight determination subunit.
[0150] The recommendation weight determination subunit is configured to determine the recommendation weight based on the target fusion feature, the candidate topic feature, and an interaction scene feature for the target object, wherein the interaction scene feature is determined based on interaction scene information for the target object in a specified time period.
[0151] According to an embodiment of the present disclosure, the resource content feature comprises at least one of a resource title text feature, a resource body text feature, and a resource voice text feature.
[0152] According to an embodiment of the present disclosure, the resource content feature is determined based on an interaction resource related to the target object, and the arrangement order of the plurality of resource content features in the resource content feature sequence is determined according to respective interaction time attributes of the plurality of interaction resources with respect to the target object.
[0153] Figure 8 A block diagram of an apparatus for training a deep learning model according to an embodiment of the present disclosure is schematically shown.
[0154] As shown in Figure 8 The apparatus 800 for training a deep learning model comprises a second acquisition module 810, a second fusion module 820, a sample recommendation weight obtaining module 830, and a training module 840.
[0155] The second acquisition module 810 is configured to acquire a sample resource content feature sequence for a sample object and a label recommendation weight of a sample candidate resource, wherein a semantic difference condition is met between a sample candidate topic of the sample candidate resource and a sample initial topic for the sample resource content feature sequence, and the label recommendation weight represents a satisfaction degree of the sample object with respect to the sample candidate resource.
[0156] The second fusion module 820 is configured to fuse a plurality of sample resource content features in the sample resource content feature sequence by using a first feature fusion layer of the deep learning model based on a window attention mechanism to obtain a plurality of sample resource fusion features.
[0157] The sample recommendation weight obtaining module 830 is configured to process the plurality of sample resource fusion features by using a recommendation weight detection network of the deep learning model to obtain a sample recommendation weight for the sample candidate resource.
[0158] The training module 840 is configured to train the deep learning model based on the label recommendation weight and the sample recommendation weight, to obtain a trained deep learning model.
[0159] According to an embodiment of the present disclosure, the label recommendation weight is determined based on the following operation: obtaining sample interaction behavior change data of a sample object on a sample candidate topic, wherein the sample interaction behavior change data represents a change in interaction behavior of the sample object on the sample candidate topic after performing an interaction operation on a sample interaction resource represented by the sample resource content feature; and determining the label recommendation weight of the sample candidate resource based on the sample interaction behavior change data of the sample candidate topic.
[0160] According to an embodiment of the present disclosure, the sample interaction behavior change data includes at least one of the following: browsing time length change data, like behavior change data, comment behavior change data, and collection behavior change data.
[0161] According to an embodiment of the present disclosure, the second fusion module 820 includes a sample associated resource content feature determination unit and a sample resource fusion feature determination unit.
[0162] The sample associated resource content feature determination unit is configured to, for a sample target resource content feature in the sample resource content feature sequence, determine a sample associated resource content feature from the sample resource related feature sequence according to a sample weight window for the sample target resource content feature, wherein a window size of the sample weight window is determined based on a resource satisfaction of the sample object on a sample target interaction resource, and the sample target resource content feature represents the sample target interaction resource.
[0163] The sample resource fusion feature determination unit is configured to fuse the sample target resource content feature and the sample associated resource content feature based on an attention mechanism, to obtain a sample resource fusion feature related to the sample target resource content feature.
[0164] According to an embodiment of the present disclosure, the recommendation weight detection network includes a second feature fusion layer and a recommendation weight detection layer. The sample recommendation weight obtaining module 830 includes a sample target fusion feature obtaining unit and a sample recommendation weight determination unit.
[0165] The sample target fusion feature obtaining unit is configured to fuse, based on an attention mechanism, a sample initial topic feature representing a sample initial topic and a plurality of sample resource fusion features by using the second feature fusion layer, to obtain a sample target fusion feature.
[0166] The sample recommendation weight determination unit is configured to determine the sample recommendation weight based on the sample target fusion feature and a sample candidate topic feature representing a sample candidate topic by using the recommendation weight detection layer.
[0167] Figure 9A structural block diagram of an agent of artificial intelligence is shown schematically.
[0168] In embodiments of the present disclosure, as shown in Figure 9 The AI agent 900 can include an input module 910, a processing module 920, and an output module 930.
[0169] The input module 910 is configured to receive input information.
[0170] The processing module 920 is configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, execute the resource recommendation method provided by embodiments of the present disclosure by calling the large language model, or obtain output information by executing the method for training a deep learning model provided by embodiments of the present disclosure by calling the large model.
[0171] The output module 930 is configured to output the output information obtained by the processing module.
[0172] According to embodiments of the present disclosure, the input module 910 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the outside world (e.g., users or external environment), and converting them into formats that the AI agent 900 can understand and process. The input module 910 is the first link for the AI agent 900 to interact with the outside world, which enables the AI agent 900 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to these information.
[0173] In examples, the input module 910 can input the resource content feature sequence or the sample resource content feature sequence, the label recommendation weight, etc. described in the foregoing.
[0174] In examples, the processing module 920 is the core support for the AI agent 900 to process complex tasks. The processing module 920 can execute the resource recommendation method and the method for training a deep learning model described in the foregoing.
[0175] In examples, the performance of the processing module 920 can be closely related to the large model on which the AI agent 900 is based. In order to fully exert the capabilities of the large model, the internal structure of the processing module 920 can be designed to be highly configurable and extensible in order to cope with various different types of tasks and demands in real scenarios.
[0176] In examples, after obtaining the resource content feature sequence, the processing module 920 can process the resource content feature sequence using the large model to obtain a plurality of resource fusion features, process the plurality of resource fusion features using the large model to obtain a target resource, and pass the target resource to the output module 930.
[0177] It can be understood that the large model can be a large language model. Although the large language model has excellent language understanding and generation capabilities, it is limited in the tasks that can be solved without the aid of any tools, like a human. When the AI agent 900 is endowed with the ability to call tools, it can implement tasks such as completing mathematical operations with a calculator, completing data analysis with Python, and completing weather forecasts with a search engine.
[0178] In an example, the output module 930 can output the target resource or the trained deep learning model described above.
[0179] The AI agent 900 according to the embodiments of the present disclosure can simply and effectively improve the intelligent degree and improve the flexibility and versatility.
[0180] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0181] According to the embodiments of the present disclosure, an electronic device includes at least one processor, and a memory connected with the at least one processor in communication. The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method provided by the embodiments of the present disclosure.
[0182] According to the embodiments of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to perform the method provided by the embodiments of the present disclosure.
[0183] According to the embodiments of the present disclosure, a computer program product includes a computer program, and the computer program, when executed by a processor, implements the method provided by the embodiments of the present disclosure.
[0184] Figure 10 A schematic block diagram of an example electronic device that can be used to implement the resource recommendation method and the method of training a deep learning model according to the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.
[0185] As Figure 10As shown, the device 1000 includes a computing unit 1001 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0186] A plurality of components in the device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, and the like; an output unit 1007, such as various types of displays, speakers, and the like; a storage unit 1008, such as a magnetic disk, an optical disk, and the like; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, and the like. The communication unit 1009 allows the device 1000 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0187] The computing unit 1001 can be various general-purpose and / or special-purpose processing components having processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, and the like. The computing unit 1001 performs various methods and processes described above, such as the resource recommendation method or the method of training a deep learning model. For example, in some embodiments, the resource recommendation method or the method of training a deep learning model can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the resource recommendation method or the method of training a deep learning model described above can be performed. Alternatively, in other embodiments, the computing unit 1001 can be configured to perform the resource recommendation method or the method of training a deep learning model by any other appropriate means, such as by means of firmware.
[0188] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a load programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0189] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, or entirely on a remote machine or server.
[0190] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical conductors, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0191] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0192] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0193] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server is generally established by computer programs running on the respective computers and having a client-server relationship to each other. The servers can be cloud servers, servers of a distributed system, or servers combined with a blockchain.
[0194] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, without departing from the desired results of the technology disclosed herein, which are not limited herein.
[0195] The specific embodiments described above are not intended to be limiting, and persons skilled in the art will appreciate that various modifications, combinations, sub-combinations and alternatives can be made to the specific embodiments without departing from the spirit and principles of the disclosure. Accordingly, the disclosure is not limited to the specific embodiments described above.
Claims
1. A resource recommendation method, comprising: obtaining a sequence of resource content features for a target object; fusing a plurality of resource content features in the sequence of resource content features based on a window attention mechanism to obtain a plurality of resource fusion features; and fusing initial topic features representing an initial topic with the plurality of resource fusion features to obtain target fusion features; determining a recommendation weight for a candidate resource based on the target fusion features, candidate topic features representing a candidate topic, and an interaction scenario feature for the target object, the interaction scenario feature being determined based on interaction scenario information for the target object in a specified time period, the candidate topic of the candidate resource and the initial topic for the sequence of resource content features satisfying a semantic difference condition; determining a target resource from the candidate resource based on the recommendation weight and recommending the target resource to the target object. The fusing of the plurality of resource content features in the sequence of resource content features based on the window attention mechanism comprises:
2. The method of claim 1, wherein, for a target resource content feature in the sequence of resource content features, determining associated resource content features from the sequence of resource content features according to a weight window for the target resource content feature, wherein a window size of the weight window is determined based on a resource satisfaction degree of the target object to a target interaction resource, the target resource content feature representing the target interaction resource; and fusing the target resource content feature and the associated resource content features based on an attention mechanism to obtain a resource fusion feature related to the target resource content feature. The fusing of the target resource content feature and the associated resource content features based on the attention mechanism comprises:
3. The method of claim 2, wherein, determining a query feature based on the target resource content feature; determining a value feature and a key feature based on a weight number of the associated resource content features, wherein the weight number matches a feature number represented by the window size; and fusing the query feature, the value feature, and the key feature based on an attention algorithm. The resource satisfaction degree is determined based on at least one of the following interaction behavior indicators for the target interaction resource:
4. The method of claim 2, wherein, a browsing time length indicator, a like behavior indicator, a comment behavior indicator, and a collection behavior indicator. The resource content features include at least one of the following:
5. The method of claim 1, wherein, a resource title text feature, a resource body text feature, and a resource speech text feature. The resource content features are determined based on interaction resources related to the target object, and an arrangement order of a plurality of the resource content features in the sequence of resource content features is determined according to respective interaction time attributes of a plurality of the interaction resources for the target object.
6. The method of claim 1 or 2, wherein, 7. A method for training a deep learning model, comprising: obtaining a sequence of sample resource content features for a sample object and a label recommendation weight of a sample candidate resource, wherein a sample candidate topic of the sample candidate resource and a sample initial topic for the sequence of sample resource content features satisfy a semantic difference condition, the label recommendation weight representing a satisfaction degree of the sample object to the sample candidate resource; The deep learning model is used to fuse multiple sample resource content features in the sample resource content feature sequence based on a window attention mechanism to obtain multiple sample resource fusion features; The deep learning model is used to process the multiple sample resource fusion features by using a recommendation weight detection network to obtain a sample recommendation weight for the sample candidate resource, wherein the sample recommendation weight is determined by using the recommendation weight detection network to perform the following operations: The sample initial topic feature representing the sample initial topic is fused with the multiple sample resource fusion features to obtain a sample target fusion feature; The sample recommendation weight is determined based on the sample target fusion feature, a sample candidate topic feature representing the sample candidate topic, and a sample interaction scene feature for the sample object, wherein the sample interaction scene feature is determined based on interaction scene information for the sample object in a specified time period; The deep learning model is trained based on the label recommendation weight and the sample recommendation weight to obtain a trained deep learning model.
8. The method of claim 7, wherein, The label recommendation weight is determined based on the following operations: Sample interaction behavior change data of the sample object for the sample candidate topic is obtained, wherein the sample interaction behavior change data represents the change in interaction behavior of the sample object for the sample candidate topic after performing an interaction operation on a sample interaction resource represented by the sample resource content feature; Based on the sample interaction behavior change data for the sample candidate topic, a label recommendation weight for the sample candidate resource is determined.
9. The method of claim 8, wherein, The sample interaction behavior change data includes at least one of the following: Browsing time change data, like behavior change data, comment behavior change data, and collection behavior change data.
10. The method of claim 7, wherein, Based on a window attention mechanism, the deep learning model is used to fuse multiple sample resource content features in the sample resource content feature sequence, including: For a sample target resource content feature in the sample resource content feature sequence, a sample associated resource content feature is determined from the sample resource related feature sequence according to a sample weight window for the sample target resource content feature, wherein the window size of the sample weight window is determined based on the resource satisfaction of the sample object for a sample target interaction resource, and the sample target resource content feature represents the sample target interaction resource; and Based on an attention mechanism, the sample target resource content feature and the sample associated resource content feature are fused to obtain a sample resource fusion feature related to the sample target resource content feature.
11. A resource recommendation device, comprising: a first acquisition module configured to obtain a resource content feature sequence for a target object; a first fusion module configured to fuse multiple resource content features in the resource content feature sequence based on a window attention mechanism to obtain multiple resource fusion features; and a recommendation weight detection network configured to process the multiple resource fusion features to obtain a sample recommendation weight for the sample candidate resource, wherein the sample recommendation weight is determined by using the recommendation weight detection network to perform the following operations: a sample initial topic feature representing the sample initial topic is fused with the multiple sample resource fusion features to obtain a sample target fusion feature; the sample recommendation weight is determined based on the sample target fusion feature, a sample candidate topic feature representing the sample candidate topic, and a sample interaction scene feature for the sample object, wherein the sample interaction scene feature is determined based on interaction scene information for the sample object in a specified time period; the deep learning model is trained based on the label recommendation weight and the sample recommendation weight to obtain a trained deep learning model. the label recommendation weight is determined based on the following operations: sample interaction behavior change data of the sample object for the sample candidate topic is obtained, wherein the sample interaction behavior change data represents the change in interaction behavior of the sample object for the sample candidate topic after performing an interaction operation on a sample interaction resource represented by the sample resource content feature; based on the sample interaction behavior change data for the sample candidate topic, a label recommendation weight for the sample candidate resource is determined. the sample interaction behavior change data includes at least one of the following: browsing time change data, like behavior change data, comment behavior change data, and collection behavior change data. based on a window attention mechanism, the deep learning model is used to fuse multiple sample resource content features in the sample resource content feature sequence, including: for a sample target resource content feature in the sample resource content feature sequence, a sample associated resource content feature is determined from the sample resource related feature sequence according to a sample weight window for the sample target resource content feature, wherein the window size of the sample weight window is determined based on the resource satisfaction of the sample object for a sample target interaction resource, and the sample target resource content feature represents the sample target interaction resource; and based on an attention mechanism, the sample target resource content feature and the sample associated resource content feature are fused to obtain a sample resource fusion feature related to the sample target resource content feature. The first determining module is configured to determine a target resource from candidate resources based on the resource fusion features, and recommend the target resource to a target object, wherein a candidate theme of the candidate resource meets a semantic difference condition with an initial theme used for the resource content feature sequence; The first determining module includes: A target fusion feature obtaining unit is configured to fuse an initial theme feature representing the initial theme and the resource fusion features to obtain a target fusion feature; A recommendation weight determining unit is configured to determine a recommendation weight for the candidate resource based on the target fusion feature and a candidate theme feature representing the candidate theme; A target resource determining unit is configured to determine the target resource based on the recommendation weight; The recommendation weight determining unit includes: A recommendation weight determining sub-unit is configured to determine the recommendation weight based on the target fusion feature, the candidate theme feature, and an interaction scene feature for the target object, wherein the interaction scene feature is determined based on interaction scene information for the target object in a specified time period.
12. The apparatus of claim 11, wherein, The first fusion module includes: A first determining unit is configured to determine, for a target resource content feature in the resource content feature sequence, an associated resource content feature from the resource-related feature sequence according to a weight window for the target resource content feature, wherein a window size of the weight window is determined based on a resource satisfaction degree of the target object to a target interaction resource, and the target resource content feature represents the target interaction resource; and A first obtaining unit is configured to fuse the target resource content feature and the associated resource content feature based on an attention mechanism to obtain a resource fusion feature related to the target resource content feature.
13. An apparatus for training a deep learning model, comprising: A second obtaining module is configured to obtain a sample resource content feature sequence for a sample object and a label recommendation weight of a sample candidate resource, wherein a sample candidate theme of the sample candidate resource meets a semantic difference condition with a sample initial theme used for the sample resource content feature sequence, and the label recommendation weight represents a satisfaction degree of the sample object to the sample candidate resource; A second fusion module is configured to fuse, based on a window attention mechanism, a plurality of sample resource content features in the sample resource content feature sequence by using a first feature fusion layer of the deep learning model to obtain a plurality of sample resource fusion features; A sample recommendation weight obtaining module is configured to process the plurality of sample resource fusion features by using a recommendation weight detection network of the deep learning model to obtain a sample recommendation weight of the sample candidate resource, wherein the sample recommendation weight is determined by performing the following operations by using the recommendation weight detection network: fuse a sample initial theme feature representing the sample initial theme and the plurality of sample resource fusion features to obtain a sample target fusion feature; determine the sample recommendation weight based on the sample target fusion feature, a sample candidate topic feature representing the sample candidate topic, and a sample interaction scenario feature of the sample object, the sample interaction scenario feature being determined based on interaction scenario information of the sample object in a specified time period; a training module configured to train the deep learning model based on the label recommendation weight and the sample recommendation weight to obtain a trained deep learning model.
14. An artificial intelligence agent product, comprising: an input module configured to receive input information; a processing module configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, execute the method of any one of claims 1 to 6 by invoking the large model, or execute the method of any one of claims 7 to 10 by invoking the large model, and obtain output information; an output module configured to output the output information obtained by the processing module.
15. An electronic device, comprising: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of any one of claims 1 to 10.
16. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to execute the method of any one of claims 1 to 10.
17. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1 to 10.
17. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1 to 10.
Citation Information
Patent Citations
Multimedia resource recommendation method and device, model training method and device and storage medium
CN116956183A