Data processing method and related device
By separating user information, scene information, and media information, generating a fused representation, and conducting comparative learning training, the problem of poor cross-domain transferability of project content representation is solved, achieving a more efficient cross-domain recommendation effect.
Patent Information
- Application Number
- CN202410637403.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-21
- Publication Date
- 2025-11-21
AI Technical Summary
Existing technologies encode user information, scene information, and media information together when extracting project features, resulting in poor cross-domain transferability of project content representation and making it unsuitable for recommendation tasks in other scenarios.
By separating user information and scene information from media information, user perspective information and project media information are generated, and fusion representation learning is performed. Media content representations are extracted using sparse hybrid expert models and fine-grained interactive graphic pre-trained models, and comparative learning training is performed to generate a pre-trained library for recommendation tasks.
It improves the accuracy and efficiency of cross-domain transfer of media content representation, and enhances the model's versatility and training accuracy in multiple scenarios.
Smart Images

Figure CN120994896A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and more particularly to a data processing method and related apparatus. Background Technology
[0002] In recent years, intelligent recommendation technology has been widely used in internet products. For example, news apps recommend news to users, shopping apps recommend products, and music apps recommend songs. Recommendation algorithms can be abstractly modeled as click-through rate (CTR) prediction algorithms, which predict the probability that a user will click on a particular item in a specific environment. The specific process involves extracting features of items clicked by the user over a period of time, and training a CTR prediction model based on these features. Once trained, the model can be used to predict the probability of an item being clicked and recommend items with high CTR rates to the user.
[0003] When recommending content to users, it's common practice to recommend content from the same domain (scenario). For example, recommending short videos to users in the short video domain and news articles in the text and image domain. Therefore, during training, features of clicked items within a single domain are often extracted to train a model for recommendation within that domain. However, current methods encode domain information, user information, and item content information together when extracting item features for each domain. This results in poor cross-domain transferability of item content representations, making them unsuitable for recommendation tasks in other scenarios. Summary of the Invention
[0004] This application provides a data processing method and related apparatus, which separates user information and scene information, allowing the media content representation part to focus on learning content representation, thereby improving the cross-domain transfer accuracy of content representation.
[0005] The first aspect of this application provides a data processing method applicable to a data processing device. The method includes: acquiring multiple first items clicked by a user within a first time period; generating first-view information of the user in a first scene based on the multiple first items, where the first scene is the scene to which the multiple first items belong, and the first-view information includes user information and first scene information; acquiring multiple second items not clicked by the user in the first scene within the first time period; extracting first media information and second media information for each first item; fusing the first-view information and the first media information to obtain a first fusion representation for each first item; fusing the first-view information and the second media information to obtain a second fusion representation for each second item; performing comparative learning training using the multiple first fusion representations as positive samples and the multiple second fusion representations as negative samples to obtain a pre-training library; extracting a media content representation portion from the pre-training library, where the media content representation portion includes multiple third media information, which are obtained after training from the multiple first media information; and performing a recommendation task based on the media content representation portion.
[0006] Acquire user behavior data for a first time period, including multiple first items clicked and multiple second items not clicked during that period. Behavioral data can be obtained from user behavior logs, specifically including clicking items and subsequent actions such as browsing, commenting, downloading, and purchasing. The first scenario can be a short video scenario, an image and text information scenario, or a video scenario. Image and text information scenarios include content with images and / or text; short video scenarios include short videos where content can be switched by swiping up and down or left and right; and video scenarios include video content with titles. Correspondingly, the first and second items in the first scenario can be short videos, images, articles, or products. The first time period can be 3 days, 7 days, or one month, with no specific limitation.
[0007] Project features include user information, domain (i.e., scene) specific information, and media information, separating user and scene information from media information. Specifically, by clicking on multiple first projects, the user's perspective information within the first scene to which the first project belongs is generated. This perspective information includes both user and scene information. Then, media information from the first and second projects is extracted. Media information refers to the project's content information, including video, audio, images, and text.
[0008] The user's first-person perspective information in the first scene is then fused with the project's media information to obtain the project's fused representation. Multiplying the first-person perspective information with the first media information of the first project yields the first fused representation of the first project, and multiplying the first-person perspective information with the second media information of the second project yields the second fused representation of the second project. The first fused representations are used as positive samples, and the second fused representations as negative samples for comparative learning training. This process narrows the distance between first fused representations and widens the distance between second fused representations and first fused representations in the representation space. After training convergence, the multiple first fused representations and multiple second fused representations trained through comparative learning are stored in a pre-training database (pre-training library).
[0009] Then, the media content representation is extracted from the pre-training library. This media content representation includes multiple third media information extracted from multiple first fusion representations after training. In other words, the third media information is the first media information after training; that is, the third media information is the first media information after narrowing the distance. The media content representation may also include multiple fourth media information, which is extracted from the second fusion representation after training. In other words, the fourth media information is the second media information after training; that is, the fourth media information is the second media information after widening the distance.
[0010] The media content representation section is used to perform online recommendation tasks, that is, to filter out projects with similar content representations from the media content representation section based on the user's request access, project characteristics and context information, and display them in the recommendation list.
[0011] In the first aspect of this application, the project features are decoupled by acquiring perspective information and media information separately, and user information and scene-specific information are separated. This allows the media content representation part to focus on learning the content itself when conducting comparative learning training on the fusion representation, thereby improving the accuracy and efficiency of cross-domain transfer of content representation (i.e., media information).
[0012] In one possible implementation of the first aspect, the first scenario includes a second scenario and a third scenario, and the multiple first items include a third item and a fourth item, wherein the third item belongs to the second scenario and the fourth item belongs to the third scenario. The above steps: generating user perspective information in the first scenario based on the multiple first items include: generating user second perspective information in the second scenario based on the third item; and generating user third perspective information in the third scenario based on the fourth item.
[0013] In this possible implementation, the first scenario includes multiple scenarios, and the multiple first items are items clicked by the user in multiple scenarios. This means that cross-scenario co-occurrence information can be used to train the model, thereby increasing the amount of data used for training and improving the model's training accuracy. Furthermore, compared to items in other scenarios, items within each scenario can better improve the model's training performance in that scenario. Therefore, using items from multiple scenarios can improve the model's training accuracy across multiple corresponding scenarios, enhancing the model's generality.
[0014] When a first scenario encompasses multiple scenarios, and multiple first items belong to multiple scenarios, generating the user's perspective information within the first scenario means generating the user's perspective information within that scenario based on the first items in each scenario. For example, if the first scenario includes a second and a third scenario, and multiple first items include a third and a fourth item, where the third item belongs to the second scenario and the fourth item belongs to the third scenario, then generating the user's first-perspective information in the first scenario involves generating the user's second-perspective information in the second scenario based on the third item, and generating the user's third-perspective information in the third scenario based on the fourth item. The first-perspective information includes both the second-perspective and third-perspective information.
[0015] In one possible implementation of the first aspect, the above steps: performing a recommendation task based on the media content representation part, including: obtaining the fifth item clicked by the user at the current moment; extracting the fifth media information of the fifth item; filtering out target media information from multiple third media information, wherein the similarity between the target media information and the fifth media information is less than a preset value; and recommending the item corresponding to the target media information to the user.
[0016] After a user clicks on the fifth item, the system recommends items based on the media information and content representation of that item. Specifically, it extracts the fifth item's media information and compares its similarity with multiple third-party media information items included in the content representation section. Target media information items with a similarity less than a preset value are then selected, and the items corresponding to these target media information items are the items to be recommended.
[0017] The media content representation section may also include fourth media information. The specific screening process may involve selecting media information with a similarity less than a preset value from the third and fourth media information, and then determining the target media information that is close in distance in the representation space from the screening results.
[0018] This possible implementation method limits the process of performing recommendation tasks based on media content assurance, thus improving the feasibility of the solution.
[0019] In one possible implementation of the first aspect, the above steps: fusing first-view information and first media information to obtain a first fused representation of each first item include: converting the first-view information into a first matrix; converting the first media information into a first vector; and multiplying the first matrix and the first vector to obtain a first fused representation of each first item.
[0020] An encoder maps first-view information into a view matrix (i.e., the first matrix), and an encoder converts first-media information into a vector (i.e., the first vector). Then, the first-view information and the first-media information are fused; specifically, the first matrix is multiplied by the first vector. The result of this matrix-vector multiplication is still a vector, which is the first fused representation. This possible implementation limits the fusion method of view information and media information, improving the feasibility of the solution.
[0021] In one possible implementation of the first aspect, the above steps of generating user first-view information in the first scene based on multiple first items include: generating user first-view information in the first scene based on multiple first items using a sparse MOE model.
[0022] In this possible implementation, the user behavior sequence is input into a sparse mixture of experts (Sparse MOE) model. The Sparse MOE model generates the user's perspective information in the first scene, where the user behavior sequence represents the first item the user clicked. Sparse MOEs are machine learning models used to solve multi-objective prediction problems, aiming to improve the performance and generalization ability of neural networks by integrating multiple expert networks. Its main idea is to distribute the input data to multiple expert networks for processing, and then perform a weighted average of the outputs of each expert network to obtain the final prediction result. Sparse MOEs can reduce the number of parameters, thereby improving training speed and generalization ability. Furthermore, they can adapt to different data distributions because they can use different expert networks to process different subsets of data, and can also handle high-dimensional sparse data.
[0023] In one possible implementation of the first aspect, the above steps: extracting first media information for each first item and second media information for each second item include: extracting first media information for each first item and second media information for each second item through the fine-grained interactive graphic pre-trained model FILIP.
[0024] First, multi-modal data for the first and second items, such as text, image, and video data, are initially extracted. Then, this multi-modal data is input into a fine-grained interactive language-image pre-training (FILIP) model. The FILIP model extracts representation vectors of media information, specifically the first and second media information. The FILIP model can solve the fine-grained matching problem in image-text matching, achieving more refined alignment through a cross-modal post-interaction mechanism. By modifying only the contrastive loss, the FILIP model successfully utilizes the fine-grained expressive power between image patches and text words, while simultaneously gaining the ability to pre-compute image and text representations offline during inference, maintaining efficiency for large-scale training and inference.
[0025] In one possible implementation of the first aspect, the first scenario includes short video scenario, graphic and text information scenario, and video scenario.
[0026] This possible implementation method limits the specific form of the first scenario, thus improving the feasibility of the solution.
[0027] A second aspect of this application provides a data processing apparatus, including an acquisition unit, a generation unit, an extraction unit, a fusion unit, a training unit, and an execution unit. The acquisition unit is used to acquire multiple first items clicked by a user in a first time period; the generation unit is used to generate first-view information of the user in a first scene based on the multiple first items, where the first scene is the scene to which the multiple first items belong, and the first-view information includes user information and first scene information; the acquisition unit is also used to acquire multiple second items not clicked by the user in the first scene during the first time period; the extraction unit is used to extract first media information of each first item and second media information of each second item; the fusion unit is used to fuse the first-view information and the first media information to obtain a first fusion representation of each first item; the fusion unit is also used to fuse the first-view information and the second media information to obtain a second fusion representation of each second item; the training unit is used to perform comparative learning training using multiple first fusion representations as positive samples and multiple second fusion representations as negative samples to obtain a pre-training library; the extraction unit is also used to extract a media content representation part from the pre-training library, where the media content representation part includes multiple third media information, which are obtained after training from multiple first media information; and the execution unit is used to perform a recommendation task based on the media content representation part.
[0028] In one possible implementation of the second aspect, the first scenario includes a second scenario and a third scenario, and the multiple first items include a third item and a fourth item, wherein the third item belongs to the second scenario and the fourth item belongs to the third scenario. The generation unit is specifically used to: generate second-view information of the user in the second scenario based on the third item; and generate third-view information of the user in the third scenario based on the fourth item.
[0029] In one possible implementation of the second aspect, the execution unit is specifically used to: obtain the fifth item clicked by the user at the current moment; extract the fifth media information of the fifth item; filter out the target media information from multiple third media information, wherein the similarity between the target media information and the fifth media information is less than a preset value; and recommend the item corresponding to the target media information to the user.
[0030] In one possible implementation of the second aspect, the fusion unit is specifically used to: convert the first perspective information into a first matrix; convert the first media information into a first vector; and multiply the first matrix and the first vector to obtain a first fusion representation of each first item.
[0031] In one possible implementation of the second aspect, the generation unit is specifically used to: generate first-person perspective information of the user in the first scene based on multiple first items using a sparse hybrid expert (SMOE) model.
[0032] In one possible implementation of the second aspect, the extraction unit is specifically used to: extract the first media information of each first item and the second media information of each second item through the fine-grained interactive graphic pre-trained model FILIP.
[0033] In one possible implementation of the second aspect, the first scenario includes short video scenarios, graphic and text information scenarios, and video conferencing scenarios.
[0034] The data processing apparatus provided in the second aspect of this application is used to perform the method described in the first aspect or any possible implementation thereof.
[0035] A third aspect of this application provides a data processing apparatus, including a processor and a memory. The memory is used to store instructions, and the processor is used to retrieve the instructions stored in the memory to execute the method described in the first aspect or any possible implementation thereof.
[0036] A fourth aspect of this application provides a computer-readable storage medium comprising instructions that, when executed on a computer, cause the computer to perform the method described in the first aspect or any possible implementation thereof.
[0037] The fifth aspect of this application provides a computer program product containing instructions that, when the computer program product is run on a computer, cause the computer to perform the method described in the first aspect or any possible implementation thereof.
[0038] The sixth aspect of this application provides a chip system including at least one processor and a communication interface, the communication interface and the at least one processor being interconnected via a line, the at least one processor being configured to run a computer program or instructions to perform the method described in the first aspect or any possible implementation thereof. Attached Figure Description
[0039] Figure 1 This is a schematic diagram illustrating an application scenario of the data processing method provided in the embodiments of this application;
[0040] Figure 2 A schematic diagram of an embodiment of the data processing method provided in this application;
[0041] Figure 3a This is a schematic diagram illustrating the generation of viewpoint information in an embodiment of this application;
[0042] Figure 3b This is another schematic diagram illustrating the generation of viewpoint information in an embodiment of this application;
[0043] Figure 4a This is a schematic diagram illustrating the extraction of media information in an embodiment of this application;
[0044] Figure 4b This is another schematic diagram illustrating the extraction of media information in an embodiment of this application;
[0045] Figure 5 This is a schematic diagram illustrating the acquisition of fusion characterization in an embodiment of this application;
[0046] Figure 6 This is a schematic diagram of comparative learning in an embodiment of this application;
[0047] Figure 7 A schematic diagram of another embodiment of the data processing method provided in this application;
[0048] Figure 8 A schematic diagram illustrating the recall effect of the data processing method provided in the embodiments of this application;
[0049] Figure 9 A schematic diagram of a deletion experiment for the data processing method provided in the embodiments of this application;
[0050] Figure 10 A schematic diagram of the structure of the data processing apparatus provided in the embodiments of this application;
[0051] Figure 11 Another schematic diagram of the data processing apparatus provided in the embodiments of this application. Detailed Implementation
[0052] This application provides a data processing method that can improve the accuracy and efficiency of cross-domain migration of media content representation. This application also provides corresponding apparatus, computer-readable storage media, and computer program products, which will be described below.
[0053] The embodiments of this application are described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. As those skilled in the art will understand, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0054] The terms “domain” and “scene,” “model” and “network,” “media information” and “content representation,” etc., used in the specification, claims, and accompanying drawings of this application are interchangeable. Unless otherwise specified, ordinal numbers such as “first,” “second,” etc., are used to distinguish multiple objects and are not used to limit the order, sequence, priority, or importance of the multiple objects. It should be understood that such terms are interchangeable where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0055] To facilitate understanding, the relevant terms and concepts mainly involved in the embodiments of this application will be introduced below.
[0056] 1. Machine Learning Systems
[0057] Based on input features and labels, the parameters of a machine learning model are trained using optimization methods such as gradient descent, and the trained model is then used to predict the labels of unknown data.
[0058] 2. Recommendation System
[0059] The recommendation system uses machine learning algorithms to analyze and learn from users' historical click behavior, then predicts new user requests and returns a list of recommended items.
[0060] 3. Click-through rate
[0061] This refers to the probability that a user will click on a specific displayed item in a particular environment.
[0062] 4. Content Representation / Multimodal Representation
[0063] A technique for representing multimedia information as vectors, such as input images, and extracting intermediate layer vectors through a neural network.
[0064] 5. Portability
[0065] A model or representation trained in one scenario can still be used in another scenario.
[0066] Please refer to the following: Figure 1 This is a schematic diagram of an application scenario of the data processing method provided in the embodiments of this application.
[0067] like Figure 1 The diagram illustrates a recommendation scenario. A recommendation scenario involves suggesting content that a user might be interested in, including short videos, text and image information, videos, and products. Recommendation technology can be simplified to predicting the click-through rate of items. If an item has a high probability of being clicked, it is recommended to the user; otherwise, it is not recommended. Specifically, user actions on the front-end page generate behavioral data, such as clicking on items and subsequent browsing, commenting, and downloading. This behavioral data is logged. The user behavior data in the logs is then used for offline model training, producing a prediction model after training convergence. This prediction model is then deployed in an online service environment and provides recommendation results based on user requests, item characteristics, and contextual information, recommending items with a high probability of being clicked to increase user engagement and activity. Recommended items are displayed as a list on the front-end page. After recommending items to users, user feedback is collected—whether the user clicked on the recommended item—and the prediction model is adjusted based on this feedback data.
[0068] When training predictive models, user behavior data within a specific scenario is typically used to train the model's recommendation capabilities for that scenario. When recommending items, items relevant to the current scenario are also recommended. For example, data on users browsing short videos can be used to train the model's recommendation capabilities in a short video scenario, and data on users browsing text and image information can be used to train the model's recommendation capabilities in a text and image information scenario. When making recommendations, if a user is browsing text and image news, the recommended items will also be text and image information; if the user is watching short videos, short videos will also be recommended to the user. It is understood that a scenario can also be called a domain; for example, a short video scenario can also be called a short video domain. This application uses scenarios as examples for description, and no specific limitations are made.
[0069] Using user behavior data within a specific scenario involves extracting features of items clicked by the user within that scenario. However, current methods encode scenario information, user information, and item content information together when extracting features for each scenario item. This results in poor cross-domain transferability of item content information, making it unsuitable for recommendation tasks in other scenarios. For example, if a user frequently clicks on sports news in a text and image information scenario, user information and scenario information are encoded together into the sports news item during feature extraction, making sports news unusable for recommendation in a short video scenario.
[0070] In view of this, embodiments of this application provide a data processing method that decouples user information, scene information, and media information in project features, improving the model's learning effect on media information, thereby enhancing the accuracy and effectiveness of cross-scene project transfer. Furthermore, embodiments of this application use cross-scene behavioral data from the same time period to train the model together, thus making fuller use of co-occurrence information, increasing the amount of data used for training, and further improving the model's generalization and accuracy.
[0071] Please refer to the following: Figure 2 This is a schematic diagram of one embodiment of the data processing method provided in this application. Figure 2 As shown, this embodiment includes steps 201 to 209.
[0072] 201. Obtain the first items that the user clicked in the first time period.
[0073] User actions in various scenarios are recorded in the user behavior log. These actions include clicking on items and subsequent actions such as browsing, commenting, downloading, or purchasing. The user behavior log is analyzed to extract the sequence of first items clicked by the user within a specific time period (i.e., multiple first items). The first time period can be 3 days, 7 days, or one month; its length is not limited. It's understood that the module storing user behavior data can be called a user behavior log, user operation log, or other names; the specific name is not limited here.
[0074] The first item can be short videos, images, articles, and products, etc. In one possible scenario, the first item covers multiple scenarios. For example, the information flow service provided by a mobile browser includes text and image information, short videos, and video content. Text and image information includes content containing images and / or text, short videos are short videos that can be switched by swiping up and down or left and right, and video content is video content with titles. Text and image information mainly consists of image and text modalities, while short videos and video content mainly consist of video and audio modalities. Within the first time period, a user can browse multiple service scenarios. After clicking to browse text and image news, they can continue to browse video content and short videos. Therefore, the multiple first items clicked within the first time period cover text and image information scenarios, video scenarios, and short video scenarios.
[0075] The first item clicked within a certain period can be considered an item that the user is most interested in, and multiple first items clicked are called co-occurring items. These multiple first items are grouped into positive samples, which are then used as labels during subsequent comparative learning training.
[0076] 202. Generate first-person perspective information of the user in the first scene based on multiple first projects. The first scene is the scene to which multiple first projects belong. The first-person perspective information includes user information and first scene information.
[0077] The items a user clicks on in each scenario reflect the user's perspective within that scenario. Therefore, by acquiring multiple first items, we can describe the user's first-person perspective information within the first scenario to which the first item belongs. This first-person perspective information includes the user's own information and the information of the first scenario; essentially, it is the user information and scenario information (i.e., domain information) decoupled from the features of the first items.
[0078] In one possible approach, multiple first items encompass multiple use cases; in other words, a first scenario includes multiple scenarios. For example, a first scenario includes a second scenario and a third scenario, and multiple first items include a third item and a fourth item, where the third item belongs to the second scenario and the fourth item belongs to the third scenario. Specifically, generating the user's first-view information in the first scenario based on the multiple first items clicked by the user involves generating the user's second-view information in the second scenario based on the third item, and generating the user's third-view information in the third scenario based on the fourth item. The first-view information includes both second-view and third-view information. In other words, when multiple scenarios are included, the view information of the scenario to which the clicked item belongs is generated.
[0079] The process of generating perspective information based on multiple first-projects can be combined with Figure 3a and Figure 3b To understand. For example Figure 3aAs shown, in the first scenario, the user clicks on multiple first items in the first time period, including item 1, item 2, item 3, item 4, and item 5. These multiple first items are converted into corresponding embeddings using a shared item encoder. An embedding is a low-dimensional vector representing an item, also known as an item vector mapping. The vector mapping for item 1 is emb1, for item 2 it's emb2, for item 3 it's emb3, for item 4 it's emb4, and for item 5 it's emb5. Then, a domain encoder obtains the user's projection matrix in the first scenario. This projection matrix is then used to transform the embeddings, ultimately yielding the user's viewpoint information in the first scenario. Since the viewpoint information is obtained by transforming the item embeddings, it can also be called a viewpoint matrix.
[0080] Figure 3b This describes the specific process of generating first-person perspective information when the first item encompasses multiple scenes. Items 1 and 2 clicked by the user in the first time period belong to the second scene, while items 3, 4, and 5 clicked in the first time period belong to the third scene. Multiple first items obtain corresponding embeddings by sharing an item encoder. Then, the routing layer routes these embeddings to their respective scene encoders: embedding 1 and embedding 2 are routed to the second scene encoder, and embeddings 3, 4, and 5 are routed to the third scene encoder, thus obtaining the prediction matrices for the second and third scenes. After performing matrix transformation on the item embeddings using the prediction matrices for the corresponding scenes, the perspective information for each scene is obtained: the second-person perspective information for the second scene and the third-person perspective information for the third scene.
[0081] In one possible approach, the user's perspective information in the first scene is generated using a sparse mixture of experts (Sparse MOE), as described above. Figure 3a and Figure 3bThe shared item encoder, scene encoder, and routing layer are modules in the sparse hybrid expert model. Specifically, the user behavior sequence is input into the sparse hybrid expert model to obtain the user's viewpoint matrix in the corresponding scene; the user behavior sequence represents the items the user has clicked. The sparse hybrid expert model is a machine learning model used to solve multi-objective prediction problems, aiming to improve the performance and generalization ability of neural networks by integrating multiple expert networks. Its main idea is to distribute the input data to multiple expert networks for processing, and then perform a weighted average of the outputs of each expert network to obtain the final prediction result. The sparse hybrid expert model can reduce the number of parameters, thereby improving training speed and generalization ability. Furthermore, it can adapt to different data distributions because it can use different expert networks to process different subsets of data, and it can also handle high-dimensional sparse data.
[0082] Besides sparse mixture of expert models, perspective information can also be generated through other models, such as mixture of experts (MOE) and soft mixture of experts (soft MOE), etc., without being limited here.
[0083] 203. Obtain multiple second items that the user did not click in the first time period under the first scenario.
[0084] Retrieve multiple second items that the user did not click in the first scenario within the first time period. When the first scenario includes only one scenario, the second item is also only an item within that single scenario. For example, if the first scenario is a text and image information scenario, then the second item is the article within that scenario. When the first scenario includes multiple scenarios, the second item is the item within each of those scenarios. For example, if the first scenario includes a second and a third scenario, where the second scenario is a text and image information scenario and the third scenario is a short video scenario, then the second item includes the article within the text and image information scenario and the short video within the short video scenario.
[0085] Unclicked items indicate that users may be less interested in them. Multiple second items are combined into negative samples for use as labels in subsequent comparative learning training.
[0086] 204. Extract the first media information for each first project and the second media information for each second project.
[0087] Extract the primary media information of clicked items and the secondary media information of unclicked items. Media information constitutes the content information of the items, including text, images, audio, and video. Media information can vividly and intuitively represent the content of the items, and combining media information allows for more accurate item recommendations to users.
[0088] The first scenario can be a single scenario. For example, if the first scenario is a text and image information scenario, then the modalities of the first and second items are images and / or text. The first scenario can also include multiple scenarios. For example, if the first scenario includes a second and a third scenario, where the second scenario is a text and image information scenario and the third scenario is a short video scenario, then the modalities of the media information in the first and second items include images, text, audio, and video.
[0089] For multiple first-item clicks by a user within a certain period, the representations of these first-item clicks can be considered similar. Therefore, after extracting media information, representation alignment is performed, allowing for the recommendation of related content to the user during subsequent recommendations. For example, if a user clicks on basketball news and then football news, these two are considered similar, and their corresponding representations are aligned in space after media information extraction. After training the model using the aligned media information, the model can recommend related content to the user. For instance, when a user browses basketball news, the recommendation list will display both other basketball news and football news. When the first scene includes only one scene, intra-domain alignment of media details is sufficient after media information extraction. When the first scene includes multiple scenes, in addition to intra-domain alignment of the extracted media information, cross-domain alignment is also required.
[0090] The process of extracting media information is as follows Figure 4a and Figure 4b As shown. It is understandable that the process of extracting information from the first media and the second media is similar. Figure 4a and Figure 4b Let's take First Media Information as an example.
[0091] like Figure 4a As shown, the first scene is a text and image information scene. The first clicked item includes item 1, item 2, and item 3. The media information modalities of the three items include images and text. Multimodal data is extracted from item 1 to obtain text 1 and image 1. Multimodal data is extracted from item 2 to obtain text 2 and image 2. Multimodal data is extracted from item 3 to obtain text 3 and image 3. Each data is input into the corresponding modality's representation encoder, i.e., text is input into the text encoder, and images are input into the image encoder. After processing by the modality encoder, the data is input into the domain adapter to obtain a representation vector containing only media information. The domain MOE adapter can be a scene hybrid expert model adapter. After extracting the media information of each item, information alignment within the scene is performed.
[0092] Figure 4bThe first scenario includes both short video and image / text information scenarios, where the user clicks on three items: Item 1, Item 2, and Item 3. Item 1 belongs to the short video scenario, with modalities of audio and video. Items 2 and 3 belong to the image / text information scenario, with modalities of text and image. Multimodal data is extracted from Item 1 to obtain Audio 1 and Video 1. Multimodal data is extracted from Item 2 to obtain Text 1 and Image 1. Multimodal data is extracted from Item 3 to obtain Text 2 and Image 2. Each data point is input into its corresponding modality's representation encoder: audio is input to the voice encoder, video to the video encoder, text to the text encoder, and images to the image encoder. After processing by the modality encoder, the data is then input into the corresponding scene adapters: Audio 1 and Video 1 are input to the short video scene adapter, and Text 1, Image 1, Text 2, and Image 2 are input to the image / text information scene adapter, resulting in a representation vector containing only media information. After extracting the media information for each item, both intra-domain and cross-domain alignment are performed.
[0093] In one possible approach, multimodal data (text data, video data, etc.) for the first and second items are initially extracted. This multimodal data is then input into a fine-grained interactive language-image pre-training (FILIP) model. The FILIP model extracts representation vectors containing only media information; that is, it extracts the first and second media information. FILIP can solve the fine-grained matching problem in image-text matching by achieving finer alignment through a cross-modal post-interaction mechanism. This mechanism uses the maximum similarity at the token level between visual tokens and text tokens to guide the objective function of contrastive learning. By modifying only the contrastive loss, the FILIP model successfully utilizes the fine-grained expressive power between image patches and text words, while simultaneously gaining the ability to pre-compute image and text representations offline during inference, maintaining efficiency for large-scale training and inference.
[0094] 205. By integrating first-person perspective information and first-media information, the first integrated representation of each first project is obtained.
[0095] Clicking on a project allows you to obtain the perspective of users within the project's domain. This user perspective is the perspective information, specifically including user information and scene information of the project's domain. Media information includes the project's content information. Therefore, fusing perspective information and media information yields a complete representation of the project. Fusing the first perspective information and the first media information results in a first fused representation for the first project.
[0096] Optionally, the specific process of fusion is as follows: the first viewpoint information is mapped into a viewpoint matrix (i.e., the first matrix) by an encoder, and the first media information is converted into a vector (i.e., the first vector) by an encoder. Then, the first matrix and the first vector are multiplied together. The result of matrix multiplication by vector is still a vector, which is the first fused representation.
[0097] 206. By integrating first-person perspective information and second-media information, a second fused representation of each second item is obtained.
[0098] Similar to step 205, the first-view information and the second-media information are fused to obtain a second fused representation for representing the second item.
[0099] Optionally, the specific process of fusion involves mapping the first-view information into a first matrix using an encoder, and then converting the second media information into a second vector using the same encoder. The first matrix is then multiplied by the second vector; the result of this matrix-vector multiplication is still a vector, which is the second fused representation.
[0100] The process of obtaining the first fusion representation and the second fusion representation can be combined Figure 5 To understand, Figure 5 Taking the acquisition of the first fusion characterization as an example. Figure 5 The dashed box below illustrates the process of extracting media information described in step 204. After extracting the representation vector of the media information for each project, it is fused with the viewpoint information of the first scene to obtain the fused representation of each project. When the first scene includes multiple scenes, the media information of each project is fused with the viewpoint information of its respective scene. For example, project 1 belongs to the second scene, and the viewpoint information of the second scene is second-viewpoint information. Projects 2 and 3 belong to the third scene, and the viewpoint information of the third scene is third-viewpoint information. The fusion process is as follows: the media information of project 1 is fused with the second-viewpoint information to obtain the fused representation of project 1. The media information of projects 2 and 3 is fused with the third-viewpoint information to obtain the fused representations of projects 2 and 3. After obtaining the fused representation, the representation vectors also need to be aligned in the representation space. When the first scene includes only one scene, only intra-domain alignment is required. When the first scene includes multiple scenes, both intra-domain alignment and cross-domain alignment are required.
[0101] The process of obtaining fusion representations can also be combined with Figure 3a and Figure 3b To understand. For example Figure 3a and Figure 3b As shown, the dashed boxes connecting the viewpoint information represent the project's media information. After performing matrix transformation to obtain the viewpoint matrix, it is fused with the media information to obtain the fused representation of the project.
[0102] The process of acquiring the fused representation can be represented by the formula E = M × P, where E is the fused representation, M is the media information, and P is the viewpoint information. That is, the representation vector of the project's media information is multiplied by the viewpoint matrix of the corresponding scene to obtain the fused representation vector of the project; the multiplication of the vector and the matrix still results in a vector. Traditional multimodal representation learning methods directly learn E, but in this embodiment, E is decomposed into the product of M and P, allowing the media information representation extraction network to focus on extracting and learning content representations. In other words, the learned content representations do not carry user information or domain information, thereby improving the cross-domain transferability of media content and enabling its use in cross-domain recommendation tasks.
[0103] 207. Use the first fusion representation as a positive sample and the second fusion representation as a negative sample for comparative learning training to obtain a pre-training library.
[0104] After obtaining the first fusion representation of the first item and the second fusion representation of the second item, comparative learning training is performed using the previously constructed positive and negative category labels. Specifically, during the comparative learning process, multiple first fusion representations corresponding to multiple first items are used as positive samples, and multiple second fusion representations corresponding to multiple second items are used as negative samples. This reduces the distance between the first fusion representations and increases the distance between the second fusion representations and the first fusion representations. After training convergence, the multiple first fusion representations and multiple second fusion representations trained through comparative learning are stored in a pre-training database (pre-training library). During the training process, the parameters of the decoupling module and the media content representation module are learned simultaneously. The decoupling module is the module that generates viewpoint information, and the media content representation module is the module that acquires media information.
[0105] The process of comparative learning can be combined with Figure 6 If we understand that both fusion representation 1 and fusion representation 2 are the first fusion representation, then we can shorten the distance between fusion representation 1 and fusion representation 2 in space.
[0106] 208. Extract the media content representation part from the pre-training library. The media content representation part includes multiple third media information, which are obtained after training from multiple first media information.
[0107] The media content representation is extracted from the pre-training library. This media content representation includes multiple third media information extracted from multiple first fusion representations after training. In other words, the third media information is the first media information after training; that is, the third media information is the first media information after narrowing the distance. The media content representation may also include multiple fourth media information, which is extracted from the second fusion representation after training. In other words, the fourth media information is the second media information after training; that is, the fourth media information is the second media information after widening the distance.
[0108] 209. Perform recommendation tasks based on the media content representation part.
[0109] The media content representation component is used to perform online recommendation tasks. Based on user requests, item characteristics, and contextual information, it identifies items with similar multimodal representation distances and displays them in the recommendation list. Because the media content representation component focuses on content information, it can perform cross-domain recommendations. For example, if a user is browsing sports news in a text-based context in a browser, in addition to recommending sports news in text-based contexts, it will also recommend text-based news in short video contexts or video contexts.
[0110] In one possible approach, the recommendation task involves obtaining the fifth item clicked by the user at the current moment, and then recommending items based on the media information and media content representation of the fifth item. Specifically, the fifth media information of the fifth item is extracted and compared with the similarity of multiple third media information items included in the media content representation. Target media information with a similarity less than a preset value is selected, and the item corresponding to this target media information is the item to be recommended.
[0111] Optionally, the media content representation section may also include fourth media information. The specific screening process may involve selecting media information with a similarity less than a preset value from the third and fourth media information, and then determining the target media information that is close in distance in the representation space from the screening results.
[0112] Recommendation tasks can be project recall tasks and project ranking tasks. Project recall tasks involve identifying a small subset of projects with similar characteristics from a massive pool of candidate projects, such as filtering from tens of thousands of projects to a few hundred. Project ranking tasks, on the other hand, involve selecting an even smaller number of projects from the small subset selected by the project recall task through refined sorting, such as selecting a few dozen projects from the hundreds selected by the recall task.
[0113] Furthermore, recommendation tasks can be combined with project coarse ranking and project re-ranking tasks. Coarse ranking is a step between the recall and fine ranking tasks; for example, filtering 500 recalled projects down to 100. Fine ranking then processes these 100 projects further. Re-ranking, on the other hand, is a step following fine ranking, continuing to process the projects selected during fine ranking. This includes removing content already read by users, removing duplicate content, repackaging, and inserting advertisements and promotional content.
[0114] In this embodiment, the project features are decoupled by acquiring perspective information and media information separately, and user information and scene-specific information are separated. This allows the media content representation part to focus on the content itself, thereby improving the accuracy and efficiency of cross-domain transfer of content representation (i.e., media information). When transferring to a new scene, there is no need to fine-tune the model, which can save training costs.
[0115] Optionally, this embodiment uses cross-scene co-occurrence information (i.e., items clicked in different scenes within the same time period) for training, increasing the amount of data used for training, alleviating the problem of sparsity in collaborative signals, and improving the model's accuracy and training effect. Furthermore, cross-domain training can avoid introducing excessive domain information, thereby improving the transfer effect of media information. Although this embodiment improves the transfer effect of content representation by decoupling user information, scene information, and media information, items in each scene still perform better for training in that scene. Therefore, using items from multiple scenes for training can better improve the model's accuracy in the corresponding scene and enhance the model's generality. In addition, the improved cross-domain transferability of content representation in this embodiment can further improve the efficiency of using co-occurrence information, thereby improving training efficiency.
[0116] In summary, the following is a combination of... Figure 7 The data processing methods provided in the embodiments of this application will be described in general terms. Figure 7 Taking the use of co-occurring signals across different scenarios for training as an example. Figure 7 As shown, the user clicked on item 1 in the second scene and also clicked on items 2 and 3 in the third scene within the first time period. The three items are input into the user domain decoupling module to separate user information and scene information, thus generating the corresponding scene's perspective information. Specifically, the three items are first input into the all-domain ID embedding layer for vector mapping, and the output data is then processed by the user behavior encoder. After processing, each item goes to its respective scene encoder; that is, item 1 is input into the scene encoder for the second scene, and items 2 and 3 are input into the scene encoder for the third scene, thus obtaining the prediction matrix for the second scene and the prediction matrix for the third scene. Finally, matrix transformation processing is performed to obtain the second-viewpoint information for the second scene and the third-viewpoint information for the third scene.
[0117] Next, media information from the three projects is extracted. Specifically, multimodal data for the three projects is extracted first. The media information modalities for Project 1 are audio and video; Audio 1 and Video 1 are extracted from Project 1. The media information modalities for Projects 2 and 3 are text and image; Text 1 and Image 1 are extracted from Project 2, and Text 2 and Image 2 are extracted from Project 3. Each modal information is processed by its corresponding modality representation encoder: audio is input to the audio encoder, video to the video encoder, text to the text encoder, and image to the image encoder. After processing by each encoder, the media information is then processed by the corresponding scene adapter: the media information for Project 1 is input to the second scene adapter, and the media information for Projects 2 and 3 is input to the third scene adapter, thus obtaining the vector representation of the media information for each project.
[0118] Then, the perspective information and media information are fused to obtain the fused representation of each project. Specifically, the media information of project 1 is multiplied by the second perspective information to obtain the fused representation of project 1. The media information of project 2 is multiplied by the third perspective information to obtain the fused representation of project 2. The media information of project 3 is multiplied by the third perspective information to obtain the fused representation of project 3.
[0119] After obtaining the fused representations of each project, cross-domain comparative learning is performed to narrow the gap between Project 1, Project 2, and Project 3 in the representation space. After training convergence, the media content representation is used for project recommendation tasks, such as recall and ranking tasks.
[0120] The data processing method provided in this application has undergone recall and deletion experiments on a publicly available dataset. The following will first combine... Figure 8 Explain the effectiveness of the recall. Figure 8 UniEmbedding is the data processing method provided in this application embodiment, while UniSRec and Item2Vec are existing methods. Testset is the test dataset, including HR@10, HR@50, MRR@10, MRR@50, NDCG@10, and NDCG@50. Old seq represents the old sequence, and new seq represents the new sequence. A larger number indicates better recall. Figure 8 As can be seen, compared with UniSRec, UniEmbedding has better recall performance on both new and old sequences in various datasets, and the improvement is significant. Compared with Item2Vec, UniEmbedding also achieved better recall performance on all datasets except HR@50.
[0121] Figure 9This is a pruning experiment, testing recall performance after removing some features from UniEmbedding. The "UniEmbedding" column indicates no pruning was performed; "UniEmbedding (item only)" indicates the case where only viewpoint information is used for contrastive learning; and "UniEmbedding (single domain)" indicates the case where only co-occurrence information from a single scene is used for training. Figure 9 It can be seen that, compared with UniEmbedding, when only viewpoint information is used, the recall performance on both new and old sequences in each dataset is significantly reduced. Furthermore, when only single-scene co-occurrence information is used, the recall performance on both new and old sequences in each dataset also shows a significant decline.
[0122] The data processing method provided in this application has also achieved good results after practical application. In browser video scenarios, the average number of visits per user increased by 1.55%, and the average usage time per user increased by 32.53 seconds over 4 days. In browser text and image information scenarios, the average number of visits per user increased by 1.89%, and the average usage time per user increased by 3.57 seconds. In browser short video scenarios, the average number of visits per user increased by 1.1%, and the average usage time per user increased by 13.99 seconds over 7 days.
[0123] The embodiments of this application have been described above from the perspective of methodology. The relevant devices in the embodiments of this application will be introduced below from the perspective of specific device implementation.
[0124] Please see Figure 10 This application provides a schematic diagram of a data processing device 1000. The data processing device 1000 includes an acquisition unit 1001, a generation unit 1002, an extraction unit 1003, a fusion unit 1004, a training unit 1005, and an execution unit 1006.
[0125] The acquisition unit 1001 is used to acquire multiple first items that the user clicked in the first time period.
[0126] The generation unit 1002 is used to generate first-view information of a user in a first scene based on multiple first items. The first scene is the scene to which the multiple first items belong. The first-view information includes user information and information of the first scene.
[0127] The acquisition unit 1001 is also used to acquire multiple second items that the user did not click in the first time period under the first scenario.
[0128] Extraction unit 1003 is used to extract the first media information of each first item and the second media information of each second item.
[0129] The fusion unit 1004 is used to fuse first-view information and first media information to obtain the first fusion representation of each first item.
[0130] The fusion unit 1004 is also used to fuse first-view information and second-media information to obtain a second fusion representation of each second item.
[0131] Training unit 1005 is used to perform comparative learning training by using multiple first fusion representations as positive samples and multiple second fusion representations as negative samples to obtain a pre-training library.
[0132] The extraction unit 1003 is also used to extract the media content representation part in the pre-training library. The media content representation part includes multiple third media information, which are obtained after training from multiple first media information.
[0133] Execution unit 1006 is used to perform recommendation tasks based on the media content representation part.
[0134] Optionally, the first scene includes a second scene and a third scene, and the multiple first items include a third item and a fourth item, wherein the third item belongs to the second scene and the fourth item belongs to the third scene. The generation unit 1002 is specifically used to generate the user's second perspective information in the second scene based on the third item; and to generate the user's third perspective information in the third scene based on the fourth item.
[0135] Optionally, the execution unit 1006 is specifically used to obtain the fifth item clicked by the user at the current moment; extract the fifth media information of the fifth item; filter out the target media information from multiple third media information, wherein the similarity between the target media information and the fifth media information is less than a preset value; and recommend the item corresponding to the target media information to the user.
[0136] Optionally, the fusion unit 1004 is specifically used to convert the first perspective information into a first matrix; convert the first media information into a first vector; and multiply the first matrix and the first vector to obtain the first fusion representation of each first item.
[0137] Optionally, the generation unit 1002 is specifically used to generate first-person perspective information of the user in the first scene based on multiple first items through the SparseMOE model of the Sparse Hybrid Expert.
[0138] Optionally, the extraction unit 1003 is specifically used to extract the first media information of each first item and the second media information of each second item through the fine-grained interactive graphic pre-trained model FILIP.
[0139] Optionally, the first scenario includes short video scenarios, text and image information scenarios, and video conferencing scenarios.
[0140] Each module in the data processing device 1000 performs as described above. Figures 2 to 5 as well as Figure 7 The operation of the data processing device in the illustrated embodiment will not be described in detail here.
[0141] Please refer to the following: Figure 11 This is a possible structural diagram of a data processing device 1100 provided in an embodiment of this application, including a processor 1101, a communication interface 1102, a memory 1103, and a bus 1104. The processor 1101, the communication interface 1102, and the memory 1103 are interconnected via the bus 1104. In the embodiments of this application, the processor 1101 is used to control and manage the operation of the data processing device; for example, the processor 1101 is used to execute... Figure 2 The steps performed by the data processing device in the illustrated method embodiment are as follows: Communication interface 1102 is used to support communication by the data processing device. Memory 1103 is used to store the program code and data of the data processing device.
[0142] The processor 1101 can be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, etc. The bus 1104 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 11 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0143] This application also provides a computer-readable storage medium, which includes instructions that, when executed on a computer, cause the computer to perform the aforementioned actions. Figures 2 to 5 as well as Figure 7 The method in the illustrated embodiment.
[0144] This application also provides a computer program product containing instructions, which, when run on a computer, causes the computer to perform the aforementioned... Figures 2 to 5 as well as Figure 7 The method in the illustrated embodiment.
[0145] This application also provides a chip system, which includes at least one processor and a communication interface. The communication interface and the at least one processor are interconnected via a circuit. The at least one processor is used to run computer programs or instructions to perform the aforementioned... Figures 2 to 5 as well as Figure 7 The method in the illustrated embodiment.
[0146] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0147] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0148] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.
[0149] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0150] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0151] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
Claims
1. A data processing method, characterized in that, include: Get the first items that the user clicked in the first time period; The user's first perspective information in a first scene is generated based on the plurality of first items. The first scene is the scene to which the plurality of first items belong. The first perspective information includes the user's information and the information of the first scene. Obtain multiple second items that the user did not click during the first time period in the first scenario; Extract the first media information for each first project and the second media information for each second project; By fusing the first perspective information and the first media information, a first fused representation of each first item is obtained; By fusing the first perspective information and the second media information, a second fused representation of each second item is obtained; Multiple first fusion representations are used as positive samples and multiple second fusion representations are used as negative samples for comparative learning training to obtain a pre-training library; Extract the media content representation part from the pre-training library. The media content representation part includes multiple third media information, which are obtained after training from multiple first media information. The recommendation task is performed based on the media content representation.
2. The method according to claim 1, characterized in that, The first scenario includes a second scenario and a third scenario, and the plurality of first items includes a third item and a fourth item, wherein the third item belongs to the second scenario, and the fourth item belongs to the third scenario. Generating the user's perspective information in the first scenario based on the plurality of first items includes: Based on the third item, generate the user's second perspective information in the second scenario; The user's third-person perspective information in the third scene is generated based on the fourth item.
3. The method according to claim 1 or 2, characterized in that, The step of performing the recommendation task based on the media content representation includes: Get the fifth item that the user clicked at the current moment; Extract the fifth media information of the fifth item; Target media information is selected from the plurality of third media information, wherein the similarity between the target media information and the fifth media information is less than a preset value; Recommend the items corresponding to the target media information to the user.
4. The method according to any one of claims 1 to 3, characterized in that, The fusion of the first perspective information and the first media information to obtain the first fused representation of each first item includes: Convert the first perspective information into a first matrix; Convert the first media information into a first vector; Multiplying the first matrix by the first vector yields the first fusion representation of each first item.
5. The method according to any one of claims 1 to 4, characterized in that, The generation of the user's first-view information in the first scene based on the plurality of first items includes: Based on the multiple first items, the user's first-view information in the first scene is generated through the Sparse MOE model.
6. The method according to any one of claims 1 to 5, characterized in that, The extraction of the first media information for each first item and the second media information for each second item includes: The first media information of each first item and the second media information of each second item are extracted using the fine-grained interactive graphic pre-trained model FILIP.
7. The method according to any one of claims 1 to 6, characterized in that, The first scenario includes short video scenarios, image and text information scenarios, and video conferencing scenarios.
8. A data processing apparatus, characterized in that, include: The acquisition unit is used to acquire multiple first items that the user clicked in the first time period; A generation unit is configured to generate first-view information of the user in a first scene based on the plurality of first items, wherein the first scene is the scene to which the plurality of first items belong, and the first-view information includes information of the user and information of the first scene; The acquisition unit is further configured to acquire multiple second items that the user did not click during the first time period in the first scenario; An extraction unit is used to extract the first media information of each first item and the second media information of each second item; A fusion unit is used to fuse the first perspective information and the first media information to obtain a first fusion representation of each first item; The fusion unit is further configured to fuse the first perspective information and the second media information to obtain a second fusion representation of each second item; The training unit is used to perform comparative learning training by using multiple first fusion representations as positive samples and multiple second fusion representations as negative samples to obtain a pre-training library. The extraction unit is further configured to extract the media content representation part from the pre-training library, the media content representation part including multiple third media information, the multiple third media information being obtained after training multiple first media information; An execution unit is used to perform recommendation tasks based on the media content representation portion.
9. The apparatus according to claim 8, characterized in that, The first scenario includes a second scenario and a third scenario, and the plurality of first items include a third item and a fourth item, wherein the third item belongs to the second scenario, the fourth item belongs to the third scenario, and the generation unit is specifically used for: Based on the third item, generate the user's second perspective information in the second scenario; The user's third-person perspective information in the third scene is generated based on the fourth item.
10. The apparatus according to claim 8 or 9, characterized in that, The execution unit is specifically used for: Get the fifth item that the user clicked at the current moment; Extract the fifth media information of the fifth item; Target media information is selected from the plurality of third media information, wherein the similarity between the target media information and the fifth media information is less than a preset value; Recommend the items corresponding to the target media information to the user.
11. The apparatus according to any one of claims 8 to 10, characterized in that, The fusion unit is specifically used for: Convert the first perspective information into a first matrix; Convert the first media information into a first vector; Multiplying the first matrix by the first vector yields the first fusion representation of each first item.
12. The apparatus according to any one of claims 8 to 11, characterized in that, The generation unit is specifically used for: Based on the multiple first items, the user's first-view information in the first scene is generated through the Sparse MOE model.
13. The apparatus according to any one of claims 8 to 12, characterized in that, The extraction unit is specifically used for: The first media information of each first item and the second media information of each second item are extracted using the fine-grained interactive graphic pre-trained model FILIP.
14. The apparatus according to any one of claims 8 to 13, characterized in that, The first scenario includes short video scenarios, image and text information scenarios, and video conferencing scenarios.
15. A data processing apparatus, characterized in that, include: Processor and memory; The memory is used to store instructions; The processor is configured to execute instructions stored in the memory to implement the method according to any one of claims 1 to 7.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by one or more processors, it implements the method as described in any one of claims 1 to 7.
17. A computer program product containing instructions, characterized in that, When the computer program product is run on a computer, it causes the computer to perform the method as described in any one of claims 1 to 7.