Resource recommendation method, method and device for training deep learning model, and intelligent agent

By integrating resource content features through the window attention mechanism, the problems of interest boundary solidification and diversity attenuation in user recommendations are solved, and the diversity and accuracy of resource recommendations are improved.

CN120653846AActive Publication Date: 2025-09-16BEIJING BAIDU NETCOM SCI & TECH CO LTD

Patent Information

Application Number
CN202510865661.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-09-16
Estimated Expiration
2045-06-25

AI Technical Summary

Technical Problem

In the existing technology, the resources recommended by users have the problems of solidified interest boundaries and reduced resource diversity, which makes it difficult to meet the user's potential interest expansion needs, resulting in reduced diversity and accuracy of recommended resources.

Method used

The window attention mechanism is used to fuse resource content feature sequences. By focusing on multiple locally adjacent resource content features, the target resource is determined to meet the semantic difference condition and recommend target resources that match the user's interests.

Benefits of technology

It improves the diversity and accuracy of resource recommendations, meets users' interest expansion needs, and improves the quality of resource recommendations and user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653846A_ABST
    Figure CN120653846A_ABST
Patent Text Reader

Abstract

The invention provides a resource recommendation method, a method and device for training a deep learning model and an intelligent agent, and relates to the technical field of artificial intelligence, in particular to the technical fields of resource recommendation, intelligent search, big data and the like. The resource recommendation method comprises the steps of obtaining a resource content feature sequence for a target object; fusing a plurality of resource content features in the resource content feature sequence based on a window attention mechanism to obtain a plurality of resource fusion features; and determining a target resource from the candidate resources based on the plurality of resource fusion features, and recommending the target resource to the target object, a candidate topic of the candidate resource and an initial topic for the resource content feature sequence satisfying a semantic difference condition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, particularly to resource recommendation, intelligent search, big data and other technical fields. Background Art

[0002] With the rapid development of artificial intelligence technology, users can easily browse news, videos and other resource information through terminal devices such as smartphones. Relevant Internet platforms can also recommend resource information to users based on their needs or preferences. Summary of the Invention

[0003] The present disclosure provides a resource recommendation method, a method for training a deep learning model, an apparatus, an intelligent agent, an electronic device, a storage medium, and a computer program product.

[0004] According to one aspect of the present disclosure, a resource recommendation method is provided, comprising: obtaining a resource content feature sequence for a target object; fusing a plurality of resource content features in the resource content feature sequence based on a window attention mechanism to obtain a plurality of resource fusion features; and determining a target resource from candidate resources based on the plurality of resource fusion features, and recommending the target resource to the target object, wherein a semantic difference condition is satisfied between a candidate topic of the candidate resource and an initial topic for the resource content feature sequence.

[0005] According to another aspect of the present disclosure, a method for training a deep learning model is provided, comprising: obtaining a sample resource content feature sequence for a sample object and a label recommendation weight for a sample candidate resource, wherein a semantic difference condition is satisfied between a sample candidate topic of the sample candidate resource and a sample initial topic for the sample resource content feature sequence, and the label recommendation weight represents the satisfaction of the sample object with the sample candidate resource; based on a window attention mechanism, utilizing a first feature fusion layer of a deep learning model to fuse a plurality of sample resource content features in the sample resource content feature sequence to obtain a plurality of sample resource fusion features; and utilizing a recommendation weight detection network of the deep learning model to process the plurality of sample resource fusion features to obtain a sample recommendation weight for the sample candidate resource; and training the deep learning model based on the label recommendation weight and the sample recommendation weight to obtain a trained deep learning model.

[0006] According to another aspect of the present disclosure, a resource recommendation device is provided, comprising: a first acquisition module for acquiring a resource content feature sequence for a target object; a first fusion module for fusing multiple resource content features in the resource content feature sequence based on a window attention mechanism to obtain multiple resource fusion features; and a first determination module for determining a target resource from candidate resources based on the multiple resource fusion features, and recommending the target resource to the target object, wherein a semantic difference condition is satisfied between a candidate topic of the candidate resource and an initial topic for the resource content feature sequence.

[0007] According to another aspect of the present disclosure, a method for training a deep learning model is provided, including: a second acquisition module, used to obtain a sample resource content feature sequence for a sample object and a label recommendation weight of a sample candidate resource, wherein a semantic difference condition is satisfied between the sample candidate topic of the sample candidate resource and the sample initial topic for the sample resource content feature sequence, and the label recommendation weight represents the satisfaction of the sample object with the sample candidate resource; a second fusion module, used to fuse multiple sample resource content features in the sample resource content feature sequence using a first feature fusion layer of a deep learning model based on a window attention mechanism to obtain multiple sample resource fusion features; and a sample recommendation weight acquisition module, used to process the multiple sample resource fusion features using a recommendation weight detection network of the deep learning model to obtain a sample recommendation weight for the sample candidate resource; a training module, used to train the deep learning model based on the label recommendation weight and the sample recommendation weight to obtain a trained deep learning model.

[0008] According to another aspect of the present disclosure, an artificial intelligence agent is provided, comprising: an input module for receiving input information; a processing module for determining a target task based on the input information received by the input module, determining a large model based on the target task, and obtaining output information by calling the large model to execute the above-mentioned resource recommendation method, or by calling the large model to execute the above-mentioned method for training a deep learning model; and an output module for outputting the output information obtained by the processing module.

[0009] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above-mentioned method.

[0010] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the above method.

[0011] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the above method when executed by a processor.

[0012] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0014] Figure 1 Schematically illustrates an exemplary system architecture to which the resource recommendation method and apparatus according to an embodiment of the present disclosure can be applied;

[0015] Figure 2 The following schematically shows a flow chart of a resource recommendation method according to an embodiment of the present disclosure;

[0016] Figure 3 Schematically illustrates a schematic diagram of the principle of fusing multiple resource content features based on a window attention mechanism according to an embodiment of the present disclosure;

[0017] Figure 4 Schematically illustrates a schematic diagram of the principle of a deep learning model according to an embodiment of the present disclosure;

[0018] Figure 5 Schematically shows a flow chart of a method for training a deep learning model according to an embodiment of the present disclosure;

[0019] Figure 6 Schematically illustrates a schematic diagram of the principle of training a deep learning model according to an embodiment of the present disclosure;

[0020] Figure 7 A block diagram of a resource recommendation device according to an embodiment of the present disclosure is schematically shown;

[0021] Figure 8 A block diagram schematically illustrates an apparatus for training a deep learning model according to an embodiment of the present disclosure;

[0022] Figure 9 A block diagram schematically illustrates a structure of an artificial intelligence agent according to an embodiment of the present disclosure; and

[0023] Figure 10 A schematic block diagram of an example electronic device that can be used to implement the resource recommendation method and the method for training a deep learning model according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0024] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0025] In the technical solution disclosed herein, the acquisition, storage and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.

[0026] The inventors found that the information recommended to users has changing trends such as solidified interest boundaries and diminished resource diversity, which makes it difficult to meet users' potential interest expansion needs, limits the diversity and richness of recommended resources, and reduces the quality of recommended resources.

[0027] Embodiments of the present disclosure provide a resource recommendation method, a method for training a deep learning model, an apparatus, an intelligent agent, an electronic device, a storage medium, and a computer program product. The resource recommendation method includes: obtaining a resource content feature sequence for a target object; fusing multiple resource content features in the resource content feature sequence based on a window attention mechanism to obtain multiple resource fusion features; and determining a target resource from candidate resources based on the multiple resource fusion features, and recommending the target resource to the target object, wherein a candidate topic of the candidate resource satisfies a semantic difference condition with an initial topic used for the resource content feature sequence.

[0028] According to an embodiment of the present disclosure, by fusing multiple resource content features in a resource content feature sequence based on a window attention mechanism, the resource fusion feature can pay more attention to the attention fusion of multiple locally adjacent resource content features determined based on the window in the resource content feature sequence, so that the resource fusion feature can pay more attention to the local interactive interests of the target object in each shorter time period, and then based on the local interactive interests represented by each of the multiple resource fusion features, the target resource that can meet the interest of the target object can be determined from the candidate resources that meet the semantic difference with the initial topic, so as to improve the matching degree between the target resource and the interactive interest of the target object, and at the same time, by recommending the initial topic that the target object is currently interested in with a topic semantic difference to avoid homogenized resource recommendation, the target resource can be matched with the interest needs of the target object, and the diversity of the resources acquired by the target object can be expanded to improve the diversity and accuracy of resource recommendation.

[0029] Figure 1 An exemplary system architecture to which the resource recommendation method and apparatus according to an embodiment of the present disclosure can be applied is schematically shown.

[0030] It should be noted that Figure 1The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure. This does not mean that the embodiments of the present disclosure cannot be applied to other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the resource recommendation method and apparatus may be applied may include a terminal device, but the terminal device may implement the resource recommendation method and apparatus provided by the embodiments of the present disclosure without interacting with a server.

[0031] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0032] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).

[0033] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0034] Server 105 may be a server that provides various services, such as a background management server (for example only) that supports content browsed by users using terminal devices 101, 102, and 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal device.

[0035] Server 105 can be a cloud server, also known as a cloud computing server or cloud host. This is a host product within a cloud computing service system that addresses the management difficulties and limited scalability of traditional physical hosts and VPS (Virtual Private Server) services. Server 105 can also be a server for a distributed system or a server integrated with blockchain.

[0036] It should be noted that the resource recommendation method provided in the embodiment of the present disclosure can generally be executed by the terminal device 101, 102, or 103. Accordingly, the resource recommendation apparatus provided in the embodiment of the present disclosure can also be set in the terminal device 101, 102, or 103.

[0037] Alternatively, the resource recommendation method provided in the embodiment of the present disclosure may also be generally executed by the server 105. Accordingly, the resource recommendation apparatus provided in the embodiment of the present disclosure may generally be provided in the server 105. The resource recommendation method provided in the embodiment of the present disclosure may also be performed by a server or server cluster that is different from the server 105 and that is capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the resource recommendation apparatus provided in the embodiment of the present disclosure may also be provided in a server or server cluster that is different from the server 105 and that is capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.

[0038] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0039] Figure 2 The flowchart of the resource recommendation method according to an embodiment of the present disclosure is schematically shown.

[0040] like Figure 2 As shown, the resource recommendation method includes operations S210 to S230.

[0041] In operation S210 , a resource content feature sequence for a target object is acquired.

[0042] In operation S220 , multiple resource content features in the resource content feature sequence are fused based on a window attention mechanism to obtain multiple resource fusion features.

[0043] In operation S230 , a target resource is determined from the candidate resources based on the plurality of resource fusion features, and the target resource is recommended to the target object.

[0044] According to embodiments of the present disclosure, a resource content feature sequence may include multiple resource content features, each of which may represent the resource content of an interactive resource, such as its title, text, and image content. An interactive resource may be a resource with which a target subject has performed interactive operations within a specified time period, such as a video resource or news resource that the target subject has viewed within a specified time period.

[0045] In one embodiment, resource content feature characterization is used for satisfactory consumption resources of the target object. Satisfied consumption resources represent interactive resources with higher satisfaction among the interactive resources on which the target object has performed interactive operations. The satisfaction can be determined based on the interactive behavior indicators of the target object for the interactive resources. For example, the satisfaction can be determined based on the interactive behaviors of the target object such as browsing time, like behavior, comment text preference, etc. on the interactive resources. The embodiments of the present disclosure do not limit the specific method of determining satisfactory consumption resources. By determining resource content features based on satisfactory consumption resources, the resource content feature sequence can be made to more accurately characterize the interactive behavior preferences and resource interest intentions of the target object. Therefore, based on the resource fusion features that accurately characterize the resource interest intentions of the target object, the target resources that match the interest intentions or interactive behavior preferences of the target object can be determined from the candidate resources of the candidate topics that have semantic differences with the initial topics, so that the target resources meet the interest expansion needs of the target object.

[0046] It should be noted that in the embodiments of the present disclosure, candidate topics can be represented as resource topics of candidate resources, and initial topics can be represented as resource topics of interactive resources on which the target object has performed interactive operations. Interactive operations can be understood as operations performed by the target object via a smart terminal such as a smartphone, and interactive behaviors can be understood as interactive attributes such as the type and frequency of interactive operations performed by the target object.

[0047] According to an embodiment of the present disclosure, by fusing multiple resource content features based on a window attention mechanism, the window size range can be used to fuse the local resource content features in the resource content feature sequence that match the window size, so that the resource fusion feature can represent the local resource content of the target object in the resource content feature sequence corresponding to the window size, and the resource fusion features corresponding to multiple windows can be used to represent the global resource content of the resource content feature sequence. This can reduce the computational overhead of attention calculation while controlling the local resource content range learned by the resource fusion feature through the window size, so that the resource fusion feature can more accurately represent the target object's interactive behavior preferences and interest needs.

[0048] According to the embodiments of the present disclosure, feature matching can be performed based on multiple resource fusion features and resource features for candidate resources, and target resources that match the multiple resource fusion features can be determined from the candidate resources based on the feature matching results. However, this is not limited to this, and target resources can be determined from candidate resources based on multiple resource fusion features in other ways. For example, multiple resource fusion features can be processed based on neural network algorithms such as multi-layer perceptrons to obtain recommendation weights for candidate resources, and target resources that meet the recommendation weight conditions can be determined from one or more candidate resources based on the recommendation weights. The embodiments of the present disclosure do not limit the specific method for determining the target resource, as long as multiple resource fusion features can be utilized.

[0049] According to the embodiments of the present disclosure, the candidate topics of the candidate resources meet the semantic difference condition with the initial topics used for the resource content feature sequence. Therefore, by using multiple resource fusion features that characterize the local interactive interests and global interactive needs of the target object, it is possible to more accurately detect the target resources that can be characterized by the initial topics of the resource content feature sequence and have thematic semantic differences, thereby achieving the expansion of the target object's resource diversification acquisition needs. And by using multiple resource fusion features that can characterize the target object's interactive interests, the target resource can be matched with the target object's interactive behavior preferences, thereby improving the recommendation quality of the target resource.

[0050] In one embodiment, resource content features are determined based on interactive resources associated with a target object, and the order of the resource content features in the resource content feature sequence is determined based on the time attributes of the target object's interactions with each of the multiple interactive resources. For example, the order of the multiple interactive resources, and thus the order of the multiple resource content features in the resource content feature sequence, can be determined based on the time at which the target object interacts with the multiple interactive resources.

[0051] By arranging multiple resource content features according to the target subject's interaction time attributes with interactive resources, the window attention mechanism can be used to fuse multiple temporally adjacent resource content features within the window range. This allows the resource fusion feature to learn the temporal preference order of the target subject's interaction behaviors with multiple interactive resources, further characterizing the target subject's changes in interaction behavior interest within a specific time period. This allows the target resources that match the target subject's changes in interaction behavior interest to be determined based on multiple resource fusion features, expanding the scope of resource recommendation topics while meeting the target subject's changing interest in resource acquisition.

[0052] According to an embodiment of the present disclosure, resource content features may include resource title text features, resource body text features, and resource voice text features.

[0053] In one example, text features can be extracted from the interactive resource title text to obtain resource title text features. Resource title text features can accurately characterize the subject and type of interactive resource content using fewer vector dimensions. Resource title text features can be used as resource content features to represent resource content semantics using fewer feature vector dimensions.

[0054] In one example, resource body text features can be determined by extracting text features from the resource body text. For example, the resource body text can be processed using a text encoder to obtain the resource body text features. The resource body text can more comprehensively and detailedly represent the resource content of the interactive resource, thereby enabling the resource body text features to fully represent the semantic content of the resource, thereby reducing resource content information loss.

[0055] In an example, the resource speech text features may be obtained by extracting features from resources with speech information, such as video resources and speech resources, or may be determined by extracting text features from text content representing speech information, such as video subtitles and lyrics subtitles.

[0056] According to an embodiment of the present disclosure, resource content features can be obtained by integrating any two or more of resource title text features, resource body text features, and resource voice text features, so as to improve the accuracy and richness of resource content features in representing resource content information.

[0057] In one example, the resource content feature is a semantic vector obtained by processing the title text of the satisfactory consumption resource for the target object using a CLIP (Contrastive Language-Image Pretraining) model.

[0058] In one example, resource content features are determined based on feature extraction of the title text, resource image, and body text of the satisfactory consumption resources for the target object using the CLIP model. The CLIP model maps resource images and text into a shared semantic space by jointly training the image encoder and text encoder, enabling the resource content features to effectively capture the deep semantic information of the text, thereby improving the accuracy of the resource fusion features in capturing the interactive behavior interests of the target object.

[0059] According to an embodiment of the present disclosure, fusing multiple resource content features in a resource content feature sequence based on a window attention mechanism includes: for a target resource content feature in the resource content feature sequence, determining an associated resource content feature from a resource-related feature sequence according to a weight window for the target resource content feature; and fusing the target resource content feature and the associated resource content feature based on the attention mechanism to obtain a resource fusion feature related to the target resource content feature.

[0060] According to an embodiment of the present disclosure, the target resource content feature represents the target interactive resource, and the associated resource content feature may represent the associated resources adjacent to the target resource in the interactive resource.

[0061] According to an embodiment of the present disclosure, the window size of the weight window is determined based on the target object's resource satisfaction with the target interactive resource. The window size can represent the attention range of the weight window, thereby dynamically determining the window size that matches the satisfaction based on the target object's satisfaction with the target resource content features, and then determining the number of associated resource content features that match the attention range represented by the window size from the resource content feature sequence for the dynamic window size, so as to dynamically adjust the window attention range in the window attention calculation process based on the target object's satisfaction with the target interactive resource, so that the resource fusion feature can more accurately and dynamically adjust the local feature fusion range for the interactive resource with higher target object satisfaction, so as to improve the representation accuracy of the resource fusion feature for the target object's interactive interest preference, and thereby improve the matching degree between the target resource and the target object's interactive interest.

[0062] In one example, the window size of the weight window is directly proportional to the satisfaction level of the target resource content features, thereby expanding the window attention range for target resource content features with higher satisfaction levels of the target objects, thereby querying a larger number of adjacent related resource content features in the resource content feature sequence for target resource content features with higher satisfaction levels, so that the resource fusion features obtained by fusion based on the window attention mechanism can learn a larger number of adjacent resource content features, thereby enabling the target resources determined based on multiple resource fusion features to expand the subject scope of the recommended resources while improving the target objects' satisfaction with the target resources.

[0063] According to the embodiments of the present disclosure, the arrangement order of multiple resource content features in the resource content feature sequence can correspond to the interaction timing characterized by the target object's respective interaction time attributes for multiple interactive resources. By dynamically adjusting the window size of the weight window based on the satisfaction level of the target resource content feature, the attention range of the target resource content feature with higher satisfaction can be improved through the weight window with dynamically changing attention range, and the position embedding mechanism based on attention calculation can be used to enable the determined resource fusion feature to learn the target object's interactive behavior change attributes over a longer period of time, so that the resource fusion feature can learn the interactive behavior interest of the resource content feature with higher satisfaction over a longer period of time. In this way, the target resource determined based on multiple resource fusion features can further more accurately characterize the target object's interactive behavior interest change trend, so as to improve the matching degree between the target resource and the target object's interest intention, and improve the quality of resource recommendation.

[0064] It should be noted that the target interaction resources involved in the embodiments of the present disclosure may represent resources with which the target object has interacted in a specified time period, and the target resources may represent resources recommended to the target object currently or in the future.

[0065] According to an embodiment of the present disclosure, resource satisfaction is determined based on at least one of the following interaction behavior indicators of the target object for the target interactive resource: browsing time indicator, like behavior indicator, comment behavior indicator, and collection behavior indicator.

[0066] According to embodiments of the present disclosure, the browsing duration indicator can be determined based on the target object's browsing duration data for the target interactive resource. For example, the browsing duration indicator can be determined based on the ratio between the browsing duration and a preset browsing duration. Alternatively, the browsing market indicator can be determined by normalizing the browsing duration data. Embodiments of the present disclosure do not specify a specific method for determining the browsing duration indicator; as long as it can represent the target object's browsing duration for the target interactive resource, it will suffice.

[0067] According to embodiments of the present disclosure, the like behavior indicator can represent the target subject's like interaction behavior data for the target interactive resource, such as the frequency and duration of likes. The like interaction behavior data can be quantified using preset rules to determine the like behavior indicator, and then the resource satisfaction can be determined based on the like behavior indicator.

[0068] According to an embodiment of the present disclosure, the comment behavior index can be determined based on attribute data related to comment behavior, such as the number of words in the comment, the emotional attribute of the comment, etc. The collection behavior index can be determined based on collection behaviors, such as the number of collections and the duration of collections.

[0069] In one embodiment, the comment behavior attribute data, favorite behavior data, browsing time data, and like interaction behavior data related to the target interactive resource are quantified based on preset quantization rules to obtain quantified browsing time indicators, like behavior indicators, comment behavior indicators, and favorite behavior indicators. By performing a weighted calculation on at least two of the browsing time indicators, like behavior indicators, comment behavior indicators, and favorite behavior indicators, a quantized resource satisfaction score for the target interactive resource is obtained. Based on the quantified resource satisfaction score, the window size of the weight window corresponding to the content characteristics of the target resource can be obtained.

[0070] According to an embodiment of the present disclosure, fusing target resource content features and associated resource content features based on an attention mechanism may include: determining query features based on target resource content features; determining value features and key features based on a weighted number of associated resource content features; and fusing query features, value features, and key features based on an attention algorithm.

[0071] According to an embodiment of the present disclosure, determining the query feature based on the target resource content feature may include performing a fusion calculation based on the trained query weight and the target resource content feature to obtain the query feature.

[0072] According to an embodiment of the present disclosure, the number of weights matches the number of features represented by the window size, and the number of features can be expressed as the number of associated resource content features, or the number of features can also be understood as the number of associated interactive resources corresponding to the window size. Determining the value features and key features based on the weighted number of associated resource content features can include performing a fusion calculation based on the value weight and the weighted number of associated resource content features to obtain the weighted number of value feature vectors; and performing a fusion calculation based on the key weight and the weighted number of associated resource content features to obtain the weighted number of key feature vectors. By performing an attention calculation operation on the weighted number of key feature vectors and the value feature vector, as well as the query feature through an attention algorithm, a resource fusion feature of the user's target resource content feature can be obtained. It should be understood that the fusion calculation can include a dot product calculation operation.

[0073] In one embodiment, multiple resource fusion features can also be arranged according to the interaction time attributes of the respective interaction resources, so that the deep learning model can further fully learn the target object's interaction behavior interest change trend by learning the order of multiple resource fusion features, so that the target resource can expand the scope of interest topics while more accurately representing the target object's interest intentions.

[0074] Figure 3 The schematic diagram schematically shows the principle of fusing multiple resource content features based on the window attention mechanism according to an embodiment of the present disclosure.

[0075] like Figure 3As shown, resource content feature sequence 3100 includes multiple resource content features A1 to A10 arranged according to interaction event attributes. The window size of a first weight window C301 for the first target resource content feature A5 can be determined based on the resource satisfaction of the target interaction resource corresponding to the first target resource content feature A5. The window size of the first weight window C301 determines the window range of the window attention mechanism based on which the multiple associated resource content features associated with the first target resource content feature A5 are A2, A3, and A4, respectively. A2, A3, A4, and A5 are fused based on an attention algorithm to determine a query feature (Query) based on A5, multiple value features (Value) and multiple key features (Key) based on A2, A3, and A4, and the query feature (Query), multiple value features (Value), and multiple key features (Key) are fused based on the attention algorithm to obtain a first resource fused feature B5 corresponding to the first target resource feature A5.

[0076] like Figure 3 As shown, based on the resource satisfaction of the target interactive resource corresponding to the second target resource content feature A9, the window size of the second weight window C302 for the second target resource content feature A9 is determined. The window size of the second weight window C302 determines that within the window range based on the window attention mechanism, multiple associated resource content features associated with the second target resource content feature A5 are multiple second associated resource content features A7 and A8. Based on the attention algorithm, A7, A8, and A9 are fused, and the query feature (Query) is determined based on A9, and two value features (Value) and two key features (Key) are determined based on A7 and A8. The query feature (Query), two value features (Value), and two key features (Key) are fused based on the attention algorithm to obtain the second resource fusion feature B9 corresponding to the second target resource feature A9.

[0077] It should be understood that the plurality of resource content features A1 to A10 may each correspond to a resource fusion feature, and the embodiments of the present disclosure will not be described in detail herein.

[0078] It should be noted that Figure 3The window size or weight window position of the first weight window or the second weight window shown is only for illustration and is not used to limit the window size or window position of the weight window. The window size of the weight window can be dynamically adjusted based on actual needs according to the resource satisfaction of the user's target resource interaction characteristics. It can also be based on the window attention mechanism to perform attention calculations on multiple local resource content features in the weight window range at any window position according to the weight window with dynamically changing window size, so as to obtain resource fusion features corresponding to the resource content features in the window attention range corresponding to the resource satisfaction.

[0079] According to an embodiment of the present disclosure, determining a target resource from candidate resources based on multiple resource fusion features includes: fusing an initial topic feature representing an initial topic with multiple resource fusion features to obtain a target fusion feature; determining a recommendation weight for the candidate resource based on the target fusion feature and the candidate topic feature representing the candidate topic; and determining the target resource based on the recommendation weight.

[0080] In one embodiment, the initial topic features and multiple resource interaction fusion features are fused based on the cross-attention algorithm to obtain the target fusion features.

[0081] In one embodiment, the candidate topic features are obtained by semantically encoding the candidate topic through an encoder, and represent the topic semantic attributes of the candidate topic.

[0082] In one embodiment, the initial topic features representing the initial topics of multiple interactive resources and the multiple resource interaction fusion features are fused based on the cross-attention algorithm to achieve adaptive learning of resource fusion features and initial topic features representing resource content, so that the resource content semantics represented by the resource fusion features can be queried based on the initial topic features, so that the target fusion features can learn the relationship between the initial topic and the resource content semantics and the interest intention of the interactive behavior. In this way, the interest intention of the target object for the candidate resource can be dynamically represented based on the target fusion features. By detecting the recommendation weight for the candidate resource based on the candidate topic features representing the candidate resource and the target fusion features, the recommendation weight can be made to more accurately represent the interactive interest intention of the target object.

[0083] It should be noted that the recommendation weight can be any type of data such as a weight value, a weight vector, etc., and the embodiments of this disclosure do not limit the specific data type of the recommendation weight. The candidate topics and candidate topics involved in the embodiments of this disclosure can represent the topic semantics of the candidate resources, and the embodiments of this disclosure will not be repeated here.

[0084] In one example, recommendation weights for multiple candidate resources can be determined. During the resource recall phase, these candidate resources are recalled based on the recommendation weights, with candidate resources with higher recommendation weights, such as those ranked higher in a preset order, being selected as the recalled resources. The recalled resources are then comprehensively scored by combining other recommendation metrics, such as click-through rate and conversion rate. Based on the comprehensive score, target resources with higher comprehensive scores are identified from the recalled resources and recommended to the target audience.

[0085] In one embodiment, determining a recommendation weight for a candidate resource based on the target fusion feature and a candidate topic feature representing the candidate topic includes determining the recommendation weight based on the target fusion feature, the candidate topic feature, and an interaction scenario feature for the target object.

[0086] According to embodiments of the present disclosure, interaction scene features are determined based on interaction scene information for a target object during a specified time period. This interaction scene information may include information related to the target object's interaction scene, such as the time period of the interaction behavior, the frequency of the interaction behavior, or a specified frequency of the interaction behavior. Embodiments of the present disclosure do not limit the specific type of interaction scene information. Interaction scene features may be determined by extracting features from the interaction behavior information.

[0087] According to an embodiment of the present disclosure, target fusion features, candidate topic features, and interaction scene features for a target object may be processed based on a neural network algorithm to obtain a recommendation weight.

[0088] In one example, a fully connected layer built using a multi-layer perceptron algorithm can process target fusion features, candidate topic features, and interaction scene features for the target object to obtain recommendation weights. This allows the fully connected layer to capture the complex semantics of the target fusion features derived from the fusion of window attention and cross attention mechanisms, improving semantic understanding capabilities for longer sequences of resource content features and increasing computational efficiency, enabling efficient and accurate determination of recommendation weights for candidate resources.

[0089] In one embodiment, based on the target fusion features and candidate topic features representing the candidate topics, the recommendation weight for the candidate resource can also be determined by processing the target fusion features, candidate topic features, and interaction scene features based on a neural network algorithm, as well as object attribute features representing the object attributes of the target object, candidate resource attribute features, and other feature information related to resource recommendation. In this way, the recommendation weight for the candidate resource can be accurately detected through multi-dimensional feature information to improve the matching degree between the target resource and the target object.

[0090] Figure 4 The schematic diagram schematically shows the principle of the deep learning model according to an embodiment of the present disclosure.

[0091] like Figure 4 As shown in the figure, the deep learning model includes a first feature fusion layer, a second feature fusion layer, and a recommendation weight detection layer. The first feature fusion layer is built based on the window attention algorithm, the second feature fusion layer is built based on the cross attention algorithm, and the recommendation weight detection layer is built based on the multi-layer perceptron algorithm.

[0092] In this embodiment, the interactive resources for the target object are multiple satisfactory consumption resource contents in the satisfactory consumption resource sequence 410 of the target object in a specified time period. The multiple satisfactory consumption resource contents in the satisfactory consumption resource sequence 410 are arranged based on the time attribute of the target object's interactive behavior. The multiple satisfactory consumption resource contents in the satisfactory consumption resource sequence 410 are processed by a multimodal encoder constructed based on the CLIP model to obtain a resource content feature sequence 430. The multiple resource content features of the resource content feature sequence 430 correspond to the arrangement order of the multiple satisfactory consumption resource contents. The initial topics of the multiple satisfactory consumption resource contents in the satisfactory consumption resource sequence 410 are respectively used as the initial topic sequence 420. The multiple initial topics in the initial topic sequence 420 are processed using a feature embedding layer constructed based on the encoder to obtain an initial topic feature sequence 440.

[0093] The resource content feature sequence 430 is input into the first feature fusion layer. The first feature fusion layer fuses multiple resource content features in the resource content feature sequence 430 based on the window attention mechanism to obtain a resource fusion feature sequence 450. The resource fusion feature sequence 450 includes multiple resource fusion features. The second feature fusion layer performs cross-attention fusion on the resource fusion feature sequence 450 and the initial topic feature sequence 440 to obtain a target fusion feature. The feature set 460 obtained related to the candidate resource includes features such as candidate topic features representing the candidate resource's candidate topic, interactive scene features for the target object, and object attribute features. The feature set 460 may also include other types of feature information such as item features representing the candidate resource.

[0094] By concatenating feature set 460 and the target fusion feature and inputting them into the recommendation weight detection layer, a recommendation weight for the candidate resource is obtained. By introducing a window attention mechanism to fuse multiple resource content features, the resource fusion feature fully learns the target object's interactive behavior interest intentions and resource content interest intentions. By cross-attentionally fusing the initial topic feature with multiple resource fusion features, the target fusion feature further accurately represents the target object's interest intentions. Thus, by using feature information related to the candidate resource and the target object, such as candidate topic features and object attribute features in the feature set for the candidate resource, the target object's interactive behavior preferences and interest expansion intentions can be more accurately mined, allowing the recommendation weight to represent the target object's potential resource interest needs. Therefore, based on the recommendation weight, the target resource that can meet the target object's interest expansion needs can be determined from the candidate resources. By recommending the target resource to the target object, the "information cocoon" effect can be broken, the target object's satisfaction with resource recommendations can be improved, and the target object's activity and long-term interaction stickiness can be enhanced.

[0095] Based on the resource recommendation method provided in the above embodiment, an embodiment of the present disclosure also provides a method for training a deep learning model.

[0096] Figure 5 A flowchart of a method for training a deep learning model according to an embodiment of the present disclosure is schematically shown.

[0097] like Figure 5 As shown, the method for training a deep learning model includes operations S510 to S540.

[0098] In operation S510 , a sample resource content feature sequence for a sample object and a tag recommendation weight of a sample candidate resource are obtained.

[0099] In operation S520 , based on the window attention mechanism, a first feature fusion layer of the deep learning model is used to fuse multiple sample resource content features in the sample resource content feature sequence to obtain multiple sample resource fusion features.

[0100] In operation S530 , a recommendation weight detection network of a deep learning model is used to process the fusion features of multiple sample resources to obtain sample recommendation weights for sample candidate resources.

[0101] In operation S540 , a deep learning model is trained based on the tag recommendation weight and the sample recommendation weight to obtain a trained deep learning model.

[0102] According to an embodiment of the present disclosure, the sample candidate topics of the sample candidate resources and the sample initial topics used for the sample resource content feature sequence meet the semantic difference condition, thereby enabling the sample candidate resources to be recommended to meet the target object's subject interest expansion needs.

[0103] According to an embodiment of the present disclosure, the sample resource content features are determined based on the resource content of the sample interaction resource. For example, the sample resource content features can be obtained by semantically encoding the resource content of the sample interaction resource using an encoder. The sample interaction resource can be resource information about the interaction behavior performed by the sample object.

[0104] According to the embodiments of the present disclosure, the label recommendation weight represents the satisfaction of the sample object with the sample candidate resources. Therefore, a deep learning model can be trained based on the label recommendation weight and the sample recommendation weight, so that the deep learning model can detect the matching degree of the candidate resources with semantic differences from the initial topics of the interactive resources for the interest intention of the target object, with respect to the interactive resources with which the target object has interacted. This can enable the trained deep learning model to meet the interest intention of the target object, expand the scope of the target object's interest topics, improve the diversity and accuracy of resource recommendations, and improve the quality of resource recommendations.

[0105] It should be noted that the trained deep learning model determined by the method for training a deep learning model provided in the embodiments of this disclosure can be applied to the resource recommendation method provided in the above embodiments. For example, the trained deep learning model can be used to process a sequence of resource content features for a target object to obtain multiple resource fusion features. Based on the multiple resource fusion features, recommendation weights for candidate resources can be determined, and the recommendation weights can be used to determine the target resource from the candidate resources.

[0106] According to an embodiment of the present disclosure, based on the window attention mechanism, multiple sample resource content features in the sample resource content feature sequence are fused using the first feature fusion layer of the deep learning model, including: for the sample target resource content feature in the sample resource content feature sequence, determining the sample associated resource content feature from the sample resource related feature sequence according to the sample weight window used for the sample target resource content feature; and fusing the sample target resource content feature and the sample associated resource content feature based on the attention mechanism to obtain a sample resource fusion feature related to the sample target resource content feature.

[0107] According to an embodiment of the present disclosure, the window size of the sample weight window is determined based on the resource satisfaction of the sample object to the sample target interactive resource, and the sample target resource content feature characterizes the sample target interactive resource.

[0108] In one example, the resource satisfaction of the sample target interaction resource may be determined based on a sample interaction behavior indicator of the sample object with respect to the sample interaction resource.

[0109] It should be noted that the technical terms involved in the embodiments of the present disclosure, including but not limited to sample target interaction resources, sample resource content features, etc., have the same or corresponding attributes as the technical terms provided in the embodiments of the present disclosure, including but not limited to target interaction resources, resource content features, etc., and the embodiments of the present disclosure will not be repeated here.

[0110] According to an embodiment of the present disclosure, a recommendation weight detection network includes a second feature fusion layer and a recommendation weight detection layer. The recommendation weight detection network of a deep learning model processes multiple sample resource fusion features to obtain a sample recommendation weight for a sample candidate resource, including: based on an attention mechanism, using the second feature fusion layer to fuse the sample initial topic feature representing the sample initial topic with multiple sample resource fusion features to obtain a sample target fusion feature; and using the recommendation weight detection layer to determine the sample recommendation weight based on the sample target fusion feature and the sample candidate topic feature representing the sample candidate topic.

[0111] Figure 6 The schematic diagram schematically shows the principle of training a deep learning model according to an embodiment of the present disclosure.

[0112] like Figure 6 As shown in the figure, the deep learning model includes a first feature fusion layer, a second feature fusion layer, and a recommendation weight detection layer. The first feature fusion layer is built based on the window attention algorithm, the second feature fusion layer is built based on the cross attention algorithm, and the recommendation weight detection layer is built based on the multi-layer perceptron algorithm.

[0113] In this embodiment, the sample resource content feature sequence 610 includes multiple sample resource content features. The initial topics of the multiple satisfactory consumption resource contents in the satisfactory consumption resource sequence 410 are used as the initial topic sequence 420. The initial topics in the initial topic sequence 420 are processed using a feature embedding layer constructed based on the encoder to obtain an initial topic feature sequence 440.

[0114] The multiple sample resource content features of the sample resource content feature sequence 610 are input into the first feature fusion layer. The first feature fusion layer fuses the multiple sample resource content features in the sample resource content feature sequence 610 based on the window attention mechanism to obtain the multiple sample resource fusion features in the sample resource fusion feature sequence 630. The second feature fusion layer performs cross-attention fusion on the sample resource fusion feature sequence 630 and the sample initial topic feature sequence 620 to obtain the target fusion feature. The obtained sample feature set 640 related to the sample candidate resource includes features such as sample candidate topic features representing the sample candidate topic of the sample candidate resource, sample interaction scene features for the sample object, and sample object attribute features. The sample feature set 640 may also include other types of feature information such as item features representing the sample candidate resource.

[0115] The sample feature set 640 and the target fusion feature are concatenated and input into the recommendation weight detection layer to obtain the sample recommendation weights for the sample candidate resources. The sample recommendation weights for the sample candidate resources are processed using a loss function to obtain a weight loss value. The deep learning model is trained using the weight loss value until the weight loss value meets the convergence condition, resulting in a trained deep learning model.

[0116] According to an embodiment of the present disclosure, the tag recommendation weight is determined based on the sample subject's evaluation of the sample candidate resource. However, this is not limited to this. The tag recommendation weight can also be determined based on the sample subject's interaction behavior indicators with the sample candidate resource. The embodiments of the present disclosure do not limit the specific method for setting the tag recommendation weight.

[0117] In one embodiment, the tag recommendation weight is determined based on the following operations: obtaining sample interaction behavior change data of the sample object to the sample candidate topic; and determining the tag recommendation weight of the sample candidate resource based on the sample interaction behavior change data for the sample candidate topic.

[0118] According to embodiments of the present disclosure, sample interaction behavior change data for a sample candidate topic may represent changes in the sample subject's interaction behavior with resources having the sample candidate topic during a specified time period. For example, the sample interaction behavior change data may be represented by data characterizing changes in the sample subject's interaction behavior, such as changes in browsing time or number of likes, for one or more resources having the sample candidate topic during multiple sub-periods of the specified time period.

[0119] In one embodiment, the sample interaction behavior change data represents changes in the sample object's interaction behavior with respect to a sample candidate topic after the sample object performs an interaction operation on a sample interaction resource characterized by the sample resource content characteristics. The sample object's interaction behavior with respect to the sample candidate topic can be understood as subsequent interaction behavior data of the sample object with respect to a sample candidate resource having the sample candidate topic. Changes in interaction behavior can be understood as the difference between the previous interaction behavior data generated by the sample object performing an interaction operation on the sample interaction resource having the initial topic and the subsequent interaction behavior data generated by the sample object performing an interaction operation on the sample candidate resource having the sample candidate topic.

[0120] For example, the sample interaction behavior change data can be expressed as a ratio between the previous interaction behavior data and the subsequent interaction behavior data, a difference between the previous interaction behavior data and the subsequent interaction behavior data, etc.

[0121] By determining the tag recommendation weights of sample candidate resources based on the sample interaction behavior change data, the tag recommendation weights can be used to quantify the changes in the sample object's interaction behavior for different sample resources with semantic differences. This allows the tag recommendation weights to represent the consumer interest gain generated by recommending the sample candidate resource to the sample object, thereby achieving the use of the tag recommendation weights to represent the sample object's interaction interest intention for the sample candidate resource, as well as the sample object's interaction interest change trend for sample resources with different themes. Therefore, the tag recommendation weights can be used as labels to train a deep learning model to improve the deep learning model's detection accuracy of the target object's interaction behavior change trend, as well as the target object's detection accuracy of the resource interest intention for expanding the topic of interest. The trained deep learning model can then determine the target resource from the candidate resources that satisfies the target object's interaction behavior interest and topic expansion intention by processing the resource content feature sequence, thereby improving the quality of resource recommendations, enhancing the target object's trust and interaction stickiness in deploying the trained deep learning model, and achieving a benign and efficient interaction between the recommendation system and the target object.

[0122] It should be noted that the embodiments of the present disclosure do not limit the specific representation method of the sample interaction behavior change data, as long as it can represent the changes in the interaction behavior of the sample object with the sample candidate resource having the sample candidate topic after performing an interaction operation on the sample interaction resource.

[0123] According to an embodiment of the present disclosure, the sample interaction behavior change data includes at least one of the following: browsing time change data, like behavior change data, comment behavior change data, and collection behavior change data.

[0124] The browsing time change data may represent changes in browsing time of the sample object. Determining the tag recommendation weight based on the browsing time change data may include quantifying the browsing time change data to obtain the tag recommendation weight.

[0125] In one embodiment, the tag recommendation weight can be determined based on the browsing time change data, which can be performed based on the following formula (1).

[0126] (1);

[0127] Where gain represents the tag recommendation weight, topic_avg represents the average browsing time of sample subjects on sample interactive resources in the previous period, topic_time represents the browsing time of sample subjects on sample candidate resources with sample candidate topics in the later period, and λ is a preset smoothing parameter used to reduce the interactive behavior data of active users.

[0128] By quantifying the tag recommendation weight by the ratio of the subsequent browsing time of the sample candidate resources to the average browsing time before, we can more accurately represent the changes in the browsing time of the sample objects in the later period to express the consumption willingness of the sample objects for the sample resources with the sample candidate topics. Therefore, the deep learning model can be trained by the tag recommendation weight to enable the deep learning model to capture the potential interactive behavior interests and resource topic intentions of the sample objects, avoiding the formation of information cocoons.

[0129] According to an embodiment of the present disclosure, the like behavior change data, comment behavior change data, and favorite behavior change data can respectively represent the changes in the like behavior data, comment behavior data, and favorite behavior data of the sample candidate resource theme after the sample object performs an interactive operation on the sample interactive resource represented by the sample resource content characteristics in the previous period. The tag recommendation weight can be determined by quantifying the like behavior change data, comment behavior change data, or favorite behavior change data of the sample object. For example, the like behavior change data can be processed based on a normalization algorithm to determine the tag weight.

[0130] It should be noted that the tag recommendation weight can be obtained by quantifying any one or more of the browsing time change data, the like behavior change data, the comment behavior change data, and the collection behavior change data. For example, any one or more of the browsing time change data, the like behavior change data, the comment behavior change data, and the collection behavior change data can be quantified to obtain one or more initial tag recommendation weights, and the tag recommendation weight can be obtained by weighted calculation of multiple initial tag recommendation weights.

[0131] According to the embodiments of the present disclosure, by determining the label recommendation weights through diversified sample interaction behavior change data, the deep learning model can more accurately capture the interaction behavior change trends of sample objects and the preference degree of sample candidate resources for expanded sample candidate topics during the training process, thereby improving the resource recommendation quality and diversity of the deep learning model applied to the resource recommendation process.

[0132] Figure 7 The block diagram of the resource recommendation device according to an embodiment of the present disclosure is schematically shown.

[0133] like Figure 7 As shown, the resource recommendation device 700 includes: a first acquisition module 710 , a first fusion module 720 and a first determination module 730 .

[0134] The first acquisition module 710 is configured to acquire a resource content feature sequence for a target object.

[0135] The first fusion module 720 is used to fuse multiple resource content features in the resource content feature sequence based on the window attention mechanism to obtain multiple resource fusion features.

[0136] The first determination module 730 is used to determine a target resource from candidate resources based on multiple resource fusion features and recommend the target resource to the target object, where the candidate topic of the candidate resource meets a semantic difference condition with the initial topic used for the resource content feature sequence.

[0137] According to an embodiment of the present disclosure, the first fusion module 720 includes: a first determining unit and a first obtaining unit.

[0138] The first determination unit is used to determine the associated resource content features from the resource-related feature sequence according to the weight window used for the target resource content features in the resource content feature sequence, wherein the window size of the weight window is determined based on the resource satisfaction of the target object with the target interactive resource, and the target resource content features represent the target interactive resource.

[0139] The first obtaining unit is used to fuse the target resource content features and the related resource content features based on the attention mechanism to obtain the resource fusion features related to the target resource content features.

[0140] According to an embodiment of the present disclosure, the first obtaining unit includes: a first determining subunit, a second determining subunit, and a fusing subunit.

[0141] The first determining subunit is configured to determine query features based on target resource content features.

[0142] The second determining subunit is configured to determine the value feature and the key feature based on the weighted number of associated resource content features, wherein the weighted number matches the number of features represented by the window size.

[0143] The fusion subunit is used to fuse query features, value features, and key features based on the attention algorithm.

[0144] According to an embodiment of the present disclosure, resource satisfaction is determined based on at least one of the following interaction behavior indicators for the target interaction resource: a browsing time indicator, a like behavior indicator, a comment behavior indicator, and a collection behavior indicator.

[0145] According to an embodiment of the present disclosure, the first determination module 730 includes: a target fusion feature obtaining unit, a recommendation weight determining unit, and a target resource determining unit.

[0146] The target fusion feature obtaining unit is used to fuse the initial topic feature representing the initial topic with multiple resource fusion features to obtain the target fusion feature.

[0147] The recommendation weight determination unit is used to determine the recommendation weight for the candidate resource based on the target fusion feature and the candidate topic feature representing the candidate topic.

[0148] The target resource determination unit is used to determine the target resource based on the recommendation weight.

[0149] According to an embodiment of the present disclosure, the recommendation weight determination unit includes: a recommendation weight determination subunit.

[0150] The recommendation weight determination subunit is used to determine the recommendation weight based on the target fusion feature, the candidate topic feature and the interaction scene feature for the target object, wherein the interaction scene feature is determined based on the interaction scene information for the target object in a specified time period.

[0151] According to an embodiment of the present disclosure, the resource content feature includes at least one of the following: a resource title text feature, a resource body text feature, and a resource voice text feature.

[0152] According to an embodiment of the present disclosure, resource content features are determined based on interactive resources related to a target object, and the arrangement order of multiple resource content features in a resource content feature sequence is determined based on the target object's interaction time attributes with each of the multiple interactive resources.

[0153] Figure 8 A block diagram of an apparatus for training a deep learning model according to an embodiment of the present disclosure is schematically shown.

[0154] like Figure 8 As shown, the device 800 for training a deep learning model includes: a second acquisition module 810, a second fusion module 820, a sample recommendation weight acquisition module 830 and a training module 840.

[0155] The second acquisition module 810 is used to obtain the sample resource content feature sequence for the sample object and the label recommendation weight of the sample candidate resource, wherein the sample candidate topic of the sample candidate resource and the sample initial topic used for the sample resource content feature sequence meet the semantic difference condition, and the label recommendation weight represents the satisfaction of the sample object with the sample candidate resource.

[0156] The second fusion module 820 is configured to fuse multiple sample resource content features in the sample resource content feature sequence using the first feature fusion layer of the deep learning model based on the window attention mechanism to obtain multiple sample resource fusion features.

[0157] The sample recommendation weight obtaining module 830 is used to use the recommendation weight detection network of the deep learning model to process the fusion features of multiple sample resources to obtain the sample recommendation weights for the sample candidate resources.

[0158] The training module 840 is used to train the deep learning model based on the label recommendation weight and the sample recommendation weight to obtain a trained deep learning model.

[0159] According to an embodiment of the present disclosure, the tag recommendation weight is determined based on the following operations: obtaining sample interaction behavior change data of the sample object with respect to the sample candidate topic, wherein the sample interaction behavior change data represents the change in the interaction behavior of the sample object with respect to the sample candidate topic after the sample object performs an interaction operation on the sample interaction resource represented by the content feature of the sample resource; and determining the tag recommendation weight of the sample candidate resource based on the sample interaction behavior change data for the sample candidate topic.

[0160] According to an embodiment of the present disclosure, the sample interaction behavior change data includes at least one of the following: browsing time change data, like behavior change data, comment behavior change data, and collection behavior change data.

[0161] According to an embodiment of the present disclosure, the second fusion module 820 includes: a sample-associated resource content feature determination unit and a sample resource fusion feature determination unit.

[0162] The sample-associated resource content feature determination unit is used to determine the sample-associated resource content feature from the sample resource-related feature sequence according to a sample weight window used for the sample target resource content feature in the sample resource content feature sequence, wherein the window size of the sample weight window is determined based on the resource satisfaction of the sample object with the sample target interactive resource, and the sample target resource content feature represents the sample target interactive resource.

[0163] The sample resource fusion feature determination unit is used to fuse the sample target resource content features and the sample associated resource content features based on the attention mechanism to obtain the sample resource fusion features related to the sample target resource content features.

[0164] According to an embodiment of the present disclosure, based on the recommendation weight detection network including a second feature fusion layer and a recommendation weight detection layer, the sample recommendation weight acquisition module 830 includes: a sample target fusion feature acquisition unit and a sample recommendation weight determination unit.

[0165] The sample target fusion feature acquisition unit is used to fuse the sample initial topic feature representing the sample initial topic with multiple sample resource fusion features based on the attention mechanism to obtain the sample target fusion feature.

[0166] The sample recommendation weight determination unit is used to determine the sample recommendation weight based on the sample target fusion feature and the sample candidate topic feature representing the sample candidate topic by using the recommendation weight detection layer.

[0167] Figure 9The structural block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is schematically shown.

[0168] In the embodiments of the present disclosure, Figure 9 As shown, the AI ​​agent 900 may include an input module 910 , a processing module 920 and an output module 930 .

[0169] Input module 910, for receiving input information;

[0170] Processing module 920, configured to determine a target task based on input information received by the input module, determine a large model based on the target task, and obtain output information by invoking the large language model to execute the resource recommendation method provided in accordance with an embodiment of the present disclosure, or by invoking the large model to execute the method for training a deep learning model provided in accordance with an embodiment of the present disclosure;

[0171] The output module 930 is used to output the output information obtained by the processing module.

[0172] According to an embodiment of the present disclosure, the input module 910 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the outside world (e.g., a user or the external environment) and converting it into a format that can be understood and processed by the AI ​​agent 900. The input module 910 is the primary link for the AI ​​agent 900 to interact with the outside world. It enables the AI ​​agent 900 to efficiently and accurately obtain the necessary "sensory" information from the outside world and respond to this information.

[0173] In an example, the input module 910 may input the resource content feature sequence or sample resource content feature sequence, tag recommendation weight, etc. described above.

[0174] In this example, the processing module 920 is the core support for the AI ​​agent 900 to handle complex tasks. The processing module 920 can execute the resource recommendation method and the method for training the deep learning model described above.

[0175] In this example, the performance of processing module 920 may be closely related to the large model underlying AI agent 900. To fully leverage the capabilities of the large model, the internal structure of processing module 920 may be designed to be highly configurable and extensible to handle a variety of different tasks and requirements in real-world scenarios.

[0176] In the example, after the AI ​​agent 900 obtains the resource content feature sequence, the processing module 920 can use the large model to process the resource content feature sequence to obtain multiple resource fusion features. The large model processes multiple resource fusion features to obtain the target resource and passes the target resource to the output module 930.

[0177] It's understandable that a large model can be a large language model. Although large language models have excellent language understanding and generation capabilities, like humans, they can only perform limited tasks without tools. Once AI agent 900 is empowered with tool-based capabilities, it can perform tasks such as using a calculator to perform mathematical operations, using Python to perform data analysis, and using search engines to perform weather forecasts.

[0178] In an example, the output module 930 may output the target resource or the trained deep learning model described above.

[0179] The AI ​​agent 900 according to the embodiment of the present disclosure can simply and effectively improve the level of intelligence, and enhance flexibility and versatility.

[0180] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0181] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided by the embodiment of the present disclosure.

[0182] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the method provided by the embodiment of the present disclosure.

[0183] According to an embodiment of the present disclosure, a computer program product includes a computer program. When the computer program is executed by a processor, the method provided by the embodiment of the present disclosure is implemented.

[0184] Figure 10 A schematic block diagram of an example electronic device that can be used to implement the resource recommendation method and the method for training a deep learning model of an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0185] like Figure 10As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. RAM 1003 may also store various programs and data required for the operation of device 1000. Computing unit 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to bus 1004.

[0186] Various components in device 1000 are connected to I / O interface 1005, including an input unit 1006, such as a keyboard, mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, optical disk, etc.; and a communication unit 1009, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0187] The computing unit 1001 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as the resource recommendation method or the method for training a deep learning model. For example, in some embodiments, the resource recommendation method or the method for training a deep learning model can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the resource recommendation method or the method for training a deep learning model described above can be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to execute the resource recommendation method or the method for training a deep learning model in any other appropriate manner (for example, by means of firmware).

[0188] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0189] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0190] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0191] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0192] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0193] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0194] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0195] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A resource recommendation method, comprising: Get the resource content feature sequence for the target object; fusing multiple resource content features in the resource content feature sequence based on a window attention mechanism to obtain multiple resource fusion features; as well as A target resource is determined from candidate resources based on a plurality of resource fusion features, and the target resource is recommended to a target object, wherein a candidate topic of the candidate resource satisfies a semantic difference condition with an initial topic used for the resource content feature sequence.

2. The method according to claim 1, wherein The fusing of multiple resource content features in the resource content feature sequence based on the window attention mechanism includes: For a target resource content feature in the resource content feature sequence, determining an associated resource content feature from the resource-related feature sequence according to a weight window for the target resource content feature, wherein a window size of the weight window is determined based on the target object's resource satisfaction with a target interactive resource, the target resource content feature representing the target interactive resource; and The target resource content feature and the associated resource content feature are fused based on the attention mechanism to obtain a resource fusion feature related to the target resource content feature.

3. The method according to claim 2, wherein: The fusing of the target resource content features and the associated resource content features based on the attention mechanism includes: Determining query characteristics based on the target resource content characteristics; Determining a value feature and a key feature based on a weighted number of the associated resource content features, wherein the weighted number matches the number of features represented by the window size; and The query feature, the value feature, and the key feature are fused based on an attention algorithm.

4. The method according to claim 2, wherein: The resource satisfaction is determined based on at least one of the following interaction behavior indicators for the target interaction resource: Browsing time indicator, like behavior indicator, comment behavior indicator, and collection behavior indicator.

5. The method according to claim 1 or 2, wherein: Determining a target resource from candidate resources based on the plurality of resource fusion features includes: fusing an initial topic feature representing the initial topic with a plurality of resource fusion features to obtain a target fusion feature; determining a recommendation weight for the candidate resource based on the target fusion feature and a candidate topic feature representing the candidate topic; and The target resource is determined based on the recommendation weight.

6. The method according to claim 5, wherein: The determining of the recommendation weight for the candidate resource based on the target fusion feature and the candidate topic feature representing the candidate topic includes: The recommendation weight is determined based on the target fusion feature, the candidate topic feature, and an interaction scene feature for the target object, wherein the interaction scene feature is determined based on interaction scene information for the target object in a specified time period.

7. The method according to claim 1, wherein The resource content characteristics include at least one of the following: Resource title text features, resource body text features, resource voice text features.

8. The method according to claim 1 or 2, wherein: The resource content feature is determined based on the interactive resources related to the target object, and the arrangement order of the multiple resource content features in the resource content feature sequence is determined according to the interaction time attributes of the target object to each of the multiple interactive resources.

9. A method for training a deep learning model, comprising: Obtaining a sample resource content feature sequence and a tag recommendation weight of a sample candidate resource for a sample object, wherein a sample candidate topic of the sample candidate resource and a sample initial topic for the sample resource content feature sequence satisfy a semantic difference condition, and the tag recommendation weight represents the sample object's satisfaction with the sample candidate resource; Based on the window attention mechanism, the first feature fusion layer of the deep learning model is used to fuse multiple sample resource content features in the sample resource content feature sequence to obtain multiple sample resource fusion features; Using the recommendation weight detection network of the deep learning model to process the fusion features of the plurality of sample resources, to obtain sample recommendation weights for the sample candidate resources; The deep learning model is trained based on the tag recommendation weight and the sample recommendation weight to obtain a trained deep learning model.

10. The method according to claim 9, wherein: The tag recommendation weight is determined based on the following operations: Acquiring sample interaction behavior change data of the sample object with respect to the sample candidate topic, wherein the sample interaction behavior change data represents a change in the interaction behavior of the sample object with respect to the sample candidate topic after the sample object performs an interaction operation on the sample interaction resource represented by the sample resource content feature; Based on the sample interaction behavior change data for the sample candidate topic, a tag recommendation weight of the sample candidate resource is determined.

11. The method according to claim 10, wherein: The sample interaction behavior change data includes at least one of the following: Changes in browsing time, likes, comments, and favorites.

12. The method according to claim 9, wherein Based on the window attention mechanism, using the first feature fusion layer of the deep learning model to fuse multiple sample resource content features in the sample resource content feature sequence includes: determining, for a sample target resource content feature in the sample resource content feature sequence, a sample associated resource content feature from the sample resource associated feature sequence according to a sample weight window for the sample target resource content feature, wherein a window size of the sample weight window is determined based on resource satisfaction of the sample subject with a sample target interactive resource, the sample target resource content feature characterizing the sample target interactive resource; and The sample target resource content feature and the sample associated resource content feature are fused based on the attention mechanism to obtain a sample resource fusion feature related to the sample target resource content feature.

13. The method according to claim 9, wherein: Based on the recommendation weight detection network including a second feature fusion layer and a recommendation weight detection layer; the recommendation weight detection network using the deep learning model processes the plurality of sample resource fusion features to obtain the sample recommendation weight for the sample candidate resource, including: Based on the attention mechanism, the second feature fusion layer is used to fuse the sample initial topic feature representing the sample initial topic with the plurality of sample resource fusion features to obtain a sample target fusion feature; and The recommendation weight detection layer is utilized to determine the sample recommendation weight based on the sample target fusion feature and the sample candidate topic feature characterizing the sample candidate topic.

14. A resource recommendation device, comprising: A first acquisition module is used to acquire a resource content feature sequence for a target object; A first fusion module is configured to fuse multiple resource content features in the resource content feature sequence based on a window attention mechanism to obtain multiple resource fusion features; as well as The first determination module is used to determine a target resource from candidate resources based on multiple resource fusion features and recommend the target resource to a target object, wherein a candidate topic of the candidate resource meets a semantic difference condition with an initial topic used for the resource content feature sequence.

15. The device according to claim 14, wherein The first fusion module includes: a first determining unit configured to determine, for a target resource content feature in the resource content feature sequence, an associated resource content feature from the resource-related feature sequence according to a weight window for the target resource content feature, wherein a window size of the weight window is determined based on a resource satisfaction score of the target object with respect to a target interactive resource, the target resource content feature representing the target interactive resource; and The first obtaining unit is configured to fuse the target resource content feature and the associated resource content feature based on an attention mechanism to obtain a resource fusion feature related to the target resource content feature.

16. A device for training a deep learning model, comprising: A second acquisition module is configured to acquire a sample resource content feature sequence and a tag recommendation weight of a sample candidate resource for a sample object, wherein a sample candidate topic of the sample candidate resource satisfies a semantic difference condition with a sample initial topic for the sample resource content feature sequence, and the tag recommendation weight represents the sample object's satisfaction with the sample candidate resource; a second fusion module, configured to fuse multiple sample resource content features in the sample resource content feature sequence using the first feature fusion layer of the deep learning model based on a window attention mechanism to obtain multiple sample resource fusion features; a sample recommendation weight obtaining module, configured to process the fusion features of the plurality of sample resources using the recommendation weight detection network of the deep learning model to obtain the sample recommendation weights for the sample candidate resources; A training module is used to train the deep learning model based on the label recommendation weight and the sample recommendation weight to obtain a trained deep learning model.

17. An artificial intelligence agent, comprising: An input module, used for receiving input information; a processing module, configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and obtain output information by calling the large model to execute the method of any one of claims 1 to 8, or by calling the large model to execute the method of any one of claims 9 to 13; An output module is used to output the output information obtained by the processing module.

18. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 13.

19. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 13.

20. A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Media resource recommendation method, training method of recommendation model and related equipment

    CN116450860A

  • Multimedia resource recommendation method and device, model training method and device and storage medium

    CN116956183A

  • Method for extracting target frame, electronic device and computer program product

    CN118570691A

  • Resource recommendation method and device based on large model, intelligent agent, electronic equipment and storage medium

    CN119089047A

  • Method, apparatus, electronic device, and storage medium for recommending multimedia resource

    US20200288205A1

Cited By

  • Resource recommendation method and device, equipment, medium and product

    CN121980076A