Multimedia data processing method, device, equipment and storage medium

By obtaining the attribute information of the target object and candidate multimedia data, predicting the operation probability and commercial value, the problem of low accuracy of multimedia data push is solved, and accurate recommendations and resource utilization are improved.

CN114331495BActive Publication Date: 2025-08-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111462968.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-02
Publication Date
2025-08-22
Estimated Expiration
2041-12-02

AI Technical Summary

Technical Problem

In the prior art, multimedia data push methods cannot achieve accurate push, resulting in low push accuracy and waste of resources.

Method used

By obtaining the attribute information of the target object and candidate multimedia data, predict the operation probability and commercial value of the target object to the multimedia data, and select the appropriate multimedia data for pushing.

Benefits of technology

It realizes accurate recommendation of multimedia data, improves push accuracy, and avoids waste of resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114331495B_ABST
    Figure CN114331495B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses a multimedia data processing method, apparatus, device and storage medium, wherein the method relates to the field of artificial intelligence and blockchain technology, and the method comprises: obtaining target object attribute information of a target object, and candidate multimedia data P in a multimedia data set to be pushed i According to the target object attribute information and the candidate media attribute information, the predicted target object for the candidate multimedia data P i Execute the target operation information corresponding to the operation; determine the candidate multimedia data P according to the candidate media attribute information i The target operation information and the media asset factor are used to select candidate multimedia data for pushing to the target object from the multimedia data set as the target multimedia data, and the target multimedia data is pushed to the terminal corresponding to the target object. This application can improve the accuracy of multimedia data push.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of artificial intelligence and blockchain technology, and in particular to a multimedia data processing method, apparatus, device, and storage medium. Background Art

[0002] With the development of Internet technology, online information push technology has been applied in various scenarios. For example, in product promotion scenarios, multimedia data about products is pushed to users via web pages, search engines, browsers, or other multimedia platforms. Currently, random push is the main method used to push multimedia data to users on multimedia platforms. However, in practice, this method cannot achieve precise push, resulting in low accuracy of multimedia data push. Summary of the Invention

[0003] The technical problem to be solved by the embodiments of the present application is to provide a multimedia data processing method, apparatus, device and storage medium, which can improve the push accuracy of multimedia data.

[0004] An embodiment of the present application provides a method for processing multimedia data, including:

[0005] Obtain the target object attribute information of the target object and the candidate multimedia data P in the multimedia data set to be pushed i candidate media attribute information; i is a positive integer less than or equal to M, and M is the number of candidate multimedia data in the multimedia data set;

[0006] According to the target object attribute information and the candidate media attribute information, a prediction is made to reflect the target object's response to the candidate multimedia data P i Target operation information corresponding to the execution operation;

[0007] According to the candidate media attribute information, the candidate multimedia data P is determined i The media asset factor is used to reflect the media asset factor of the candidate multimedia data P i The actual amount of assets that the object to which the candidate multimedia data belongs needs to spend when the conversion operation is executed;

[0008] According to the target operation information and the media asset factor, candidate multimedia data for pushing to the target object is selected from the multimedia data set as the target multimedia data, and the target multimedia data is pushed to the terminal corresponding to the target object. In one aspect, an embodiment of the present application provides a multimedia data processing device, comprising:

[0009] The acquisition module is used to obtain the target object attribute information of the target object and the candidate multimedia data P in the multimedia data set to be pushed. i candidate media attribute information; i is a positive integer less than or equal to M, and M is the number of candidate multimedia data in the multimedia data set;

[0010] The prediction module is configured to predict, based on the target object attribute information and the candidate media attribute information, a value that reflects the target object's response to the candidate multimedia data P i Target operation information corresponding to the execution operation;

[0011] A determination module is used to determine the candidate multimedia data P according to the candidate media attribute information. i The media asset factor is used to reflect the media asset factor of the candidate multimedia data P i The actual amount of assets that the object to which the candidate multimedia data belongs needs to spend when the conversion operation is executed;

[0012] A selection module is used to select candidate multimedia data for pushing to the target object from the multimedia data set according to the target operation information and the media asset factor, as target multimedia data, and push the target multimedia data to the terminal corresponding to the target object.

[0013] In one aspect, the present application provides a computer device, comprising: a processor and a memory;

[0014] The memory is used to store a computer program, and the processor is used to call the computer program to execute the steps in the method.

[0015] On one hand, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the steps in the method are executed.

[0016] On one hand, an embodiment of the present application provides a computer program product, including a computer program / instruction, which implements the steps in the above method when executed by a processor.

[0017] In this application, the target operation information is specifically used to reflect the probability of the target object performing an operation (such as a shallow conversion operation, a deep conversion operation, or a non-conversion operation) on the candidate multimedia data. The target operation information can, to a certain extent, reflect the target object's interest in the candidate multimedia data. For example, the target operation information indicates that the probability of the target object performing a deep conversion operation on the candidate multimedia data is relatively high, indicating that the target object has a high interest in the candidate multimedia data. The media asset factor is used to reflect the probability of the candidate multimedia data P being used to represent the target object's interest in the candidate multimedia data. i When the conversion operation is executed, the actual amount of assets that the object to which the candidate multimedia data belongs needs to spend, that is, the media asset factor can, to a certain extent, reflect the commercial value that the candidate multimedia data brings to the multimedia platform. By selecting the candidate multimedia data to be pushed to the target object from the multimedia data set according to the target operation information and the media asset factor, the target multimedia data is pushed to the terminal corresponding to the target object. In other words, by comprehensively considering the user's interest characteristics in the candidate multimedia data and the commercial value of the candidate multimedia data, multimedia data can be recommended to the user, which can achieve accurate recommendation and improve the accuracy of multimedia data recommendation; it can not only avoid recommending multimedia data that the user is not interested in to the user, resulting in invalid multimedia data recommendation and wasting the resources of the multimedia platform, but also avoid recommending multimedia data with relatively low commercial value to the user, resulting in relatively low resource utilization of the multimedia platform, thereby improving the resource utilization of the multimedia platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1a This is a schematic diagram of a scenario for automatic expansion of intelligent targeted advertising provided by this application;

[0020] Figure 1b This is a schematic diagram of a preferred scenario of a smart targeted advertising system provided by this application;

[0021] Figure 2 This is a schematic diagram of the architecture of a multimedia data processing system provided by this application;

[0022] Figure 3 This is a schematic diagram of a scenario of interaction between various devices in a multimedia data processing system provided by the present application;

[0023] Figure 4 This is a flowchart of the first multimedia data processing method provided by this application;

[0024] Figure 5 This is a flowchart of the second multimedia data processing method provided by this application;

[0025] Figure 6 This is a schematic diagram of a target operation identification model and a target asset identification model provided by this application;

[0026] Figure 7 This is a schematic diagram of a scenario for obtaining target operation information corresponding to an operation performed by a target object on candidate multimedia data, provided by the present application;

[0027] Figure 8 This is a schematic diagram of a scenario for obtaining the media asset factors of candidate multimedia data provided by the present application;

[0028] Figure 9 This is a schematic diagram of the structure of an asset expert network provided by this application;

[0029] Figure 10 This is a schematic diagram of a scenario for constructing training sample data provided by an embodiment of the present application;

[0030] Figure 11 This is a schematic diagram of a scenario for retrieving target multimedia data provided by an embodiment of the present application;

[0031] Figure 12 This is a schematic diagram of a scenario for obtaining the commercial value of candidate multimedia data provided by an embodiment of the present application;

[0032] Figure 13a This is a schematic diagram of a scenario for obtaining target multimedia data provided by an embodiment of the present application;

[0033] Figure 13b This is a schematic diagram of a process for obtaining target multimedia data provided by an embodiment of the present application;

[0034] Figure 14 This is a schematic diagram of a multimedia data processing structure provided by an embodiment of the present application;

[0035] Figure 15 It is a structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0036] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0037] First, the nouns involved in the embodiments of this application are introduced:

[0038] Multimedia data: Data on which users can perform operations, which may include one or more of images, text, videos, and web links. When multimedia data is displayed on a multimedia platform, users can perform operations on the multimedia data. The operations performed here may include shallow conversion operations, deep conversion operations, and non-conversion operations. Deep conversion operations refer to operations that can bring assets to the publisher of the multimedia data this time. Shallow conversion operations may refer to operations with a first probability of bringing assets to the publisher of the multimedia data in the future. Non-conversion operations refer to operations that will not bring assets to the publisher of the multimedia data, or operations with a second probability of bringing assets to the publisher of the multimedia data in the future, where the first probability is greater than the second probability. Shallow conversion operations, deep conversion operations, and non-conversion operations can be specifically determined based on the type of multimedia data. For example, if the multimedia data is a product promotion advertisement, when the product promotion advertisement is displayed on a multimedia platform, if the user clicks on the product promotion advertisement, the terminal jumps to the webpage corresponding to the product promotion advertisement in response to the click operation, and the webpage includes an application for purchasing the product. If the user's terminal already has the app installed, they confirm and launch it. Alternatively, if the user doesn't have the app installed, they download and install it, and then place an order for the promoted product on the app. Therefore, the click action described above can be considered a non-conversion action, the download action (or confirmation and launch action) can be considered a shallow conversion action, and the order action can be considered a conversion action.

[0039] Publisher of multimedia data: The entity to which multimedia data belongs. For example, if the multimedia data is an advertisement, the publisher can be an advertiser, specifically a business selling or promoting its products and services online. Advertisers publish advertising campaigns and pay the multimedia platform based on the total number of marketing effects and the price per effect specified in the advertising campaign.

[0040] oCPA (optimized click per action) advertising: This is an advertising format that uses a new bidding method. Specifically, it refers to the expected cost price (target_cpa) set by the advertiser for a specific conversion action. The platform is responsible for controlling the bid for each impression to ensure that the average cost per conversion is within 1.2 of the advertiser's expected cost price (target_cpa). The expected cost price can also be called the media's expected asset information.

[0041] Actual cost of advertising: For the conversion actions set by the advertiser, the average consumption (i.e. expenditure) of each conversion action of oCPA advertising is the cost.

[0042] eCPM: effective Cost per Mille (total charge per thousand impressions), an indicator of bidding ranking on multimedia platforms. Ads with high eCPM mean that they can bring more revenue to multimedia platforms and get priority exposure.

[0043] GMV: Gross Merchandise Volume (GMV) is an indicator that measures the total transaction volume of advertisers on multimedia platforms. The calculation formula is the product of the advertiser's conversion number (the number of conversion operations performed) and the conversion bid. A higher GMV means the higher the value the multimedia platform brings to the advertiser.

[0044] Recall / retrieval: In multimedia platforms, selecting a set of advertisements that match the current user's interests from a massive advertisement library is called recall, also known as retrieval.

[0045] Smart Targeted Advertising: Unlike traditional ad targeting, which requires advertisers to select target audiences based on prior knowledge and provide corresponding audience packages, smart targeted advertising allows advertisers to simply provide basic audience targeting (such as gender, age, and region). The advertising system then automatically selects user groups that meet these basic targeting requirements using model strategies. Only when the model identifies high-quality user requests will the advertiser's ads be recalled and bid for exposure. Smart Targeted Advertising comes in two forms: automatic expansion and system optimization.

[0046] 1) Automatic expansion is usually used with precise narrow targeting. Here, precise narrow targeting refers to the original narrow targeting selected by the advertiser, such as Figure 1aA and B and C and D (A / B / C / D refers to the traditional tag targeting conditions introduced earlier). When enabling automatic expansion, advertisers can also set an unbreakable portion of the original targeting. For example, if it's A and B, the expansion strategy will replace the target audience with E. The final ad targeting audience is based on the original targeting of A and B and C and D, with A and B and E added to achieve the expansion effect.

[0047] 2) System optimization is usually used with broad targeting. Broad targeting here refers to the original targeting selected by the advertiser, which is slightly broad, such as Figure 1b In A and B, the preferred strategy population is replaced by F, and the final advertising targeted population is Aand B and F, achieving the effect of preferred adjustment.

[0048] Typically, multimedia platforms include a large amount of multimedia data to be pushed. However, the area (i.e., ad space) for displaying multimedia data in a multimedia platform is limited. Therefore, only a portion of the multimedia data can be selected for exposure each time. Only the exposed multimedia data has the opportunity to be converted, and thus, can bring assets (i.e., commercial value) to the object to which the multimedia data belongs and the multimedia data platform. Currently, a random push method is mainly used to push the multimedia data that needs to be pushed in the multimedia platform to the user, which cannot achieve accurate push, resulting in a relatively low accuracy of multimedia data push. In other words, this push method has the same push probability for each multimedia data, so it is easy to push multimedia data that the user is not interested in, or multimedia data with relatively low commercial value to the user, wasting the exposure opportunity of the multimedia data, resulting in a relatively low accuracy of multimedia data push. Based on this, the present application provides a multimedia data processing method, which includes: a computer device can obtain target object attribute information of a target object, and candidate media attribute information of candidate multimedia data in a multimedia data set to be pushed, and predict target operation information corresponding to the target object performing an operation on the candidate multimedia data based on the target object attribute information and the candidate media attribute information, that is, the target operation information is used to reflect the probability of the target object performing an operation (such as a shallow conversion operation, a deep conversion operation, or a non-conversion operation) on the candidate multimedia data. The target operation information can reflect the interest of the target object in the candidate multimedia data to a certain extent. For example, the target operation information indicates that the probability of the target object performing a deep conversion operation on the candidate multimedia data is relatively high, indicating that the target object has a relatively high interest in the candidate multimedia data. Further, based on the candidate multimedia attribute information, the media asset factor of the candidate multimedia data is determined. The media asset factor is used to reflect the probability of the candidate multimedia data P iWhen a conversion operation is executed, the actual amount of assets required by the object to which the candidate multimedia data belongs, that is, the media asset factor, can, to a certain extent, reflect the commercial value that the candidate multimedia data brings to the multimedia platform. Then, based on the target operation information and the media asset factor, the candidate multimedia data to be pushed to the target object is selected from the multimedia data set as the target multimedia data, and the target multimedia data is pushed to the terminal corresponding to the target object. In other words, by comprehensively considering the user's interest characteristics in the candidate multimedia data and the commercial value of the candidate multimedia data, multimedia data can be recommended to the user, achieving precise recommendations and improving the accuracy of multimedia data recommendations.

[0049] In order to facilitate a clearer understanding of the present application, a multimedia data processing system for implementing the multimedia data processing method of the present application is first introduced. Figure 2 As shown, the multimedia data processing system includes Figure 2 As shown, the multimedia data processing system includes a server 10 and a terminal cluster. The terminal cluster may include one or more terminals. The number of terminals is not limited here. Figure 2 As shown, the terminal cluster can specifically include terminal 1, terminal 2, ..., terminal n; it can be understood that terminal 1, terminal 2, terminal 3, ..., terminal n can all be connected to the server 10 through the network, so that each terminal can exchange data with the server 10 through the network connection.

[0050] The server 10 may be a multimedia data management device, for example, a device that provides backend services for a multimedia platform, such as a social application, a video application, a multimedia webpage, etc. Specifically, the server 10 may be configured to recommend multimedia data to a user based on media attribute information of the multimedia data and object attribute information of the user.

[0051] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of at least two physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal can specifically refer to a vehicle-mounted terminal with multimedia data processing capabilities, a smart phone, a smart speaker, a screen speaker, a smart watch, etc., but is not limited to this. Each terminal and server can be directly or indirectly connected through wired or wireless communication. At the same time, the number of terminals and servers can be one or at least two, and this application does not limit this.

[0052] It should be noted that the server in this application may refer to a node device in a blockchain network. When the node device receives new candidate multimedia data, it obtains the candidate media data information of the candidate multimedia data. When the terminal receives a request from the target object to obtain multimedia data, it sends the request to the node device. The node device can obtain the target object attribute information of the target object, and retrieve the candidate media data to be pushed to the target object from the candidate media attribute information based on the target object attribute information, and push the target media data to the target object as the target media data.

[0053] based on Figure 2 The multimedia data processing system shown can be used to implement the multimedia data processing method in this application. Figure 3 Taking multimedia data as advertisement as an example, the multimedia data processing method may include an offline processing process and an online processing process:

[0054] The offline processing process primarily involves obtaining ad feature vectors and training the model. This process is implemented by the offline service module in the server, which includes seven modules: log parsing, feature construction, sample construction and model training, delivery database, ad data stream, and ad feature vectors. Log parsing, feature construction, sample construction, and model training primarily extract the various features (including ad features, user features, and ad competition environment features) required for the candidate asset and operation recognition models from the raw logs or feature database of the multimedia platform (i.e., the advertising system). Based on these features, training samples for the candidate asset and operation recognition models are constructed using appropriate sample construction methods. Furthermore, the candidate asset and operation recognition models are trained using the training samples to obtain the target asset and operation recognition models. The delivery database, ad data stream, and ad feature vector modules primarily capture the latest real-time ad status and use the target asset and operation recognition models to calculate the latest ad feature vectors (i.e., candidate media attribute information).

[0055] The online processing process mainly involves obtaining user feature vectors and multimedia data to be recommended. This process is implemented by the online service module in the terminal, which includes five modules: fine ranking, coarse ranking, recall, model service, and advertising library. When a user requests multimedia data, the target operation recognition model is used to calculate the latest user feature vector. Based on the user feature vector, the corresponding ad feature vector is retrieved from the advertising library. The retrieved ad undergoes subsequent coarse and fine ranking on the multimedia platform until the bid is successful and the ad is exposed on the current user request.

[0056] Further, see Figure 4 , is a flowchart of a multimedia data processing method provided by an embodiment of the present application. Figure 4 As shown, this method can be Figure 2 The terminal can also be executed by Figure 2 It can also be executed by the server in Figure 2 The terminal and server in the embodiment are executed together. In this application, the devices used to execute the method can be collectively referred to as computer devices. The multimedia data processing method may include the following steps S101 to S104:

[0057] S101, obtaining target object attribute information of the target object and candidate multimedia data P in the multimedia data set to be pushed i candidate media attribute information; i is a positive integer less than or equal to M, and M is the number of candidate multimedia data in the multimedia data set.

[0058] In this application, the user who sends the multimedia data acquisition request can be called the target object. The computer device can obtain the log data of the multimedia platform, parse the log data, and obtain the attribute information of the target object and the candidate multimedia data P in the multimedia data set to be pushed. i The candidate media attribute information. The target object attribute information includes one or more of basic portrait features, historical behavior statistical features, behavior sequence features, behavior interest mining features, etc. The basic portrait features include one or more of age, gender, province, occupation, consumption status, marital status, and education level. Historical behavior statistical features include one or more of the number of clicks in the historical time period (such as the past month, the past week, the past three months), the number of clicks on the multimedia data details page, the number of video clicks, the number of times multimedia data is set to a type of multimedia data that is not of interest, etc., and the average number of multimedia data exposures. Behavioral sequence features include multimedia data, apps, and other information on exposure, clicks, and conversions within the historical time period of the target object; behavioral interest mining features include label features such as multimedia data types of long-term and short-term interest, keywords, etc. mined from the original behavior sequence of the target object. The candidate media attribute information includes context features and basic attribute information. The context features include multimedia data bit information (such as multimedia data bit ID, multimedia data bit material specifications, etc.), device information (device operating system, device networking type), multimedia data bit context information, etc.; basic attribute information includes media expected asset information, multimedia data ID, creative ID, product ID, multimedia data master ID, multimedia data type, creative content keywords, multimedia data keywords and other features. Here, the media expected asset information is used to reflect the candidate multimedia data P iWhen the operation is converted and executed, the expected amount of assets that the object to which the candidate multimedia data belongs needs to spend is the amount of assets that the candidate multimedia data P i Before exposure, candidate multimedia data P i The object's quotation (i.e. bid) to the multimedia data platform.

[0059] It should be noted that the target object attribute information can be called the user embedding (i.e., vector), and the candidate object attribute information can be called the multimedia data embedding. The user embedding and multimedia data embedding are separate: the multimedia data embedding can be calculated offline in advance, while the user embedding is calculated online in real time. For example, when a target object requests multimedia data, the user embedding of the target object can be calculated in real time. The user embedding is then used to retrieve the neighboring multimedia data embeddings, and the candidate multimedia data corresponding to these neighboring multimedia data embeddings is pushed to the target object. This improves the efficiency of obtaining the target multimedia data.

[0060] Among them, when the above embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of the target object's object attribute information, the sample attribute information of the sample object, and the log data of the target object and the sample object regarding multimedia data must comply with the relevant laws, regulations and standards of the relevant countries and regions. In other words, the computer device can only obtain the target object's object attribute information, the sample attribute information of the sample object, and the log data of the target object and the sample object regarding multimedia data when the computer device obtains the user's authorization information for the above information; that is, the target object's object attribute information, the sample attribute information of the sample object, and the log data of the target object and the sample object regarding multimedia data are obtained only after the user's authorization.

[0061] For example, a computer device displays a permission prompt interface in the multimedia interface of a multimedia platform. The permission prompt interface is used to prompt the user that object attribute information of the target object, sample attribute information of the sample object, and log data of the target object and the sample object about multimedia data are currently being collected. After the user issues a confirmation operation on the permission prompt interface, the step of obtaining the object attribute information of the target object, sample attribute information of the sample object, and log data of the target object and the sample object about multimedia data is started, otherwise the step ends.

[0062] S102: predicting the target object's attributes for the candidate multimedia data P based on the target object's attributes and the candidate media attributes. iTarget operation information corresponding to the execution operation.

[0063] In this application, the computer device can associate and identify the target object attribute information and the candidate media attribute information to obtain a value reflecting the target object's response to the candidate multimedia data P. i The target operation information corresponding to the execution operation is used to reflect the target object's operation on the candidate multimedia data P i The probability corresponding to the execution operation, that is, the probability can reflect the target object's response to the candidate multimedia data P to a certain extent. i For example, the target object has a certain interest in the candidate multimedia data P i The probability of executing the conversion operation is relatively high, indicating that the target object is interested in the candidate multimedia data P i The interest of the target object in the candidate multimedia data P is relatively high; if the target object i The probability of executing the conversion operation is relatively low, indicating that the target object has a relatively low probability of executing the conversion operation on the candidate multimedia data P. i The interest level is relatively low.

[0064] S103: Determine the candidate multimedia data P according to the candidate media attribute information. i The media asset factor is used to reflect the media asset factor of the candidate multimedia data P i The actual amount of assets that the object to which the candidate multimedia data belongs needs to spend when the conversion operation is executed.

[0065] In this application, the computer device can determine the candidate multimedia data P according to the candidate media attribute information. i The media asset factor is used to reflect the media asset factor of the candidate multimedia data P i When the conversion operation is executed, the actual asset amount that the object to which the candidate multimedia data belongs needs to spend is the actual asset amount when the candidate multimedia data P i When the conversion operation is executed, the object to which the candidate multimedia data belongs needs to pay the asset amount to the multimedia platform. Therefore, the media asset factor can reflect the candidate multimedia data P to a certain extent. i For example, the higher the media asset factor is, the better the candidate multimedia data P is. i The actual amount of assets that the object needs to pay to the multimedia platform is relatively high, that is, the candidate multimedia data P i On the contrary, the lower the media asset factor is, the higher the commercial value of the candidate multimedia data P is. i The actual amount of assets that the object needs to pay to the multimedia platform is relatively low, that is, the candidate multimedia data P i The commercial value is relatively low.

[0066] S104 : Select candidate multimedia data for pushing to the target object from the multimedia data set according to the target operation information and the media asset factor, as target multimedia data, and push the target multimedia data to the terminal corresponding to the target object.

[0067] In the present application, since the multimedia platform includes a large number of candidate multimedia data to be pushed, it is impossible to push all the candidate multimedia data to the target object at one time. Therefore, the computer device can select the candidate multimedia data to be pushed to the target object from the multimedia data set based on the target operation information and the media asset factor, and push the target multimedia data to the terminal corresponding to the target object; that is, by comprehensively considering the target object's interest characteristics in the candidate multimedia data and the commercial value of the candidate multimedia data, pushing multimedia data to the target object can achieve accurate recommendation and improve the accuracy of multimedia data recommendation.

[0068] In the present application, when a computer device receives a request from a target object to obtain multimedia data, it can use the target object attribute information of the target object and the candidate media attribute information of the candidate multimedia data in the multimedia data set to be pushed. Based on the target object attribute information and the candidate media attribute information, it predicts the target operation information corresponding to the operation performed by the target object on the candidate multimedia data. That is, the target operation information is specifically used to reflect the probability of the target object performing an operation (such as a shallow conversion operation, a deep conversion operation, or a non-conversion operation) on the candidate multimedia data. The target operation information can reflect the interest of the target object in the candidate multimedia data to a certain extent. For example, the target operation information indicates that the probability of the target object performing a deep conversion operation on the candidate multimedia data is relatively high, indicating that the target object has a relatively high interest in the candidate multimedia data. Further, based on the candidate multimedia attribute information, the media asset factor of the candidate multimedia data is determined. The media asset factor is used to reflect the probability of the candidate multimedia data P iWhen the conversion operation is executed, the actual amount of assets that the object to which the candidate multimedia data belongs needs to spend, that is, the media asset factor can, to a certain extent, reflect the commercial value that the candidate multimedia data brings to the multimedia platform. Then, based on the target operation information and the media asset factor, the candidate multimedia data for pushing to the target object is selected from the multimedia data set as the target multimedia data, and the target multimedia data is pushed to the terminal corresponding to the target object. In other words, by comprehensively considering the user's interest characteristics in the candidate multimedia data and the commercial value of the candidate multimedia data, multimedia data can be recommended to the user, which can achieve accurate recommendation and improve the accuracy of multimedia data recommendation; it can not only avoid recommending multimedia data that the user is not interested in to the user, resulting in invalid multimedia data recommendation and wasting the resources of the multimedia platform, but also avoid recommending multimedia data with relatively low commercial value to the user, resulting in relatively low resource utilization of the multimedia platform, thereby improving the resource utilization of the multimedia platform.

[0069] Further, see Figure 5 , is a flowchart of a multimedia data processing method provided by an embodiment of the present application. Figure 5 As shown, the method can be executed by the terminal in Figure 1, or by the server in Figure 1, or by both the terminal and the server in Figure 1. In this application, the devices used to execute the method can be collectively referred to as computer devices. The multimedia data processing method can include the following steps S201 to S206:

[0070] S201, obtaining target object attribute information of the target object and candidate multimedia data P in the multimedia data set to be pushed i candidate media attribute information; i is a positive integer less than or equal to M, and M is the number of candidate multimedia data in the multimedia data set.

[0071] S202: Extract the target object attribute information related to the candidate multimedia data P i The key object attribute information associated with the operation attribute.

[0072] In this application, the computer device can extract the target object attribute information based on historical experience or through a model (such as a target operation recognition model and a target asset recognition model) and the candidate multimedia data P i The key object attribute information associated with the operation attributes, such as the key object attribute information may include one or more of consumption status, occupation, historical behavior statistical characteristics, behavior sequence characteristics, behavior interest mining characteristics, etc.

[0073] Optionally, the key object attribute information includes the key object attribute information related to the candidate multimedia data P iThe first key object attribute information associated with the conversion operation attribute of the candidate multimedia data P i The second key object attribute information associated with the non-conversion operation attribute, the first key object attribute information and the second key object attribute information may specifically refer to the second key object attribute information associated with the candidate multimedia data P i For example, the candidate multimedia data P i For product promotion advertisements, the first key object attribute information includes consumption status, historical behavior statistical features, behavior sequence features, and behavior interest mining features, and the second key object attribute information includes one or more of occupation, historical behavior statistical features, and behavior sequence features. Optionally, the first key object attribute information may specifically include information related to the candidate multimedia data P i The first conversion operation attribute associated with the key object attribute information, and the candidate multimedia data P i The second conversion operation attribute is associated with any one or both of the key object attribute information. The first conversion operation attribute, the second conversion operation attribute, and the non-conversion operation attribute can be determined based on the type of multimedia data. For example, if the multimedia data is a product promotion advertisement, the first conversion operation attribute refers to the order operation attribute, i.e., the deep conversion operation attribute; the second conversion operation attribute refers to the download operation attribute, i.e., the shallow conversion operation attribute; and the non-conversion operation attribute can refer to the click operation attribute.

[0074] For example, Figure 6 As shown, the target operation recognition model includes expert network 1, expert network 2 and expert network 3, and gated network 1, gated network 2 and gated network 3 corresponding to expert network 1, expert network 2 and expert network 3 respectively. The expert network 1 and gated network 1 in the target operation recognition model are used to extract the target object attribute information related to the candidate multimedia data P from the target object attribute information. i Specifically, the expert network 1 in the target operation recognition model extracts the second key object attribute information associated with the candidate multimedia data P from the target object attribute information. i The key object attribute information associated with the non-conversion operation attribute of the target operation recognition model is determined by the gating network 1 for the key object attribute information output by the expert network 1. The confidence is used to weight the key object attribute information output by the expert network 1 to obtain the second key object attribute information. Similarly, the expert network 2 and the gating network 2 in the target operation recognition model are used to extract the key object attribute information related to the candidate multimedia data P from the target object attribute information. i The shallow key object attribute information associated with the shallow transformation operation attribute of the target operation recognition model is used to extract the shallow key object attribute information related to the candidate multimedia data P from the target object attribute information. iThe deep key object attribute information associated with the deep conversion operation attribute is determined, and the shallow key object attribute information and the deep key object attribute information are determined as the first key object attribute information.

[0075] S203, extract the candidate multimedia data P from the candidate media attribute information i The key media attribute information associated with the operation attribute.

[0076] In this application, the computer device can extract the candidate multimedia data P from the candidate multimedia attribute information. i The key media attribute information associated with the operation attribute, such as the key media attribute information including one or more of multimedia data bit context information, multimedia data type, creative content keywords, multimedia data keywords, etc.

[0077] Optionally, the key media attribute information includes the key media attribute information related to the candidate multimedia data P i The first key media attribute information associated with the conversion operation attribute of the candidate multimedia data P i The second key media attribute information associated with the non-conversion operation attribute of the candidate multimedia data P i The key media attribute information associated with the first conversion operation attribute of the candidate multimedia data P i The first key media attribute information and the second key media attribute information may be any one or two of the key media attribute information associated with the second conversion operation attribute of the candidate multimedia data P. i The type is determined.

[0078] For example, Figure 6 As shown, the target asset recognition model includes expert network 1, expert network 2, expert network 3, asset expert network, and gated network 1, gated network 2, gated network 3 and gated network 4 corresponding to expert network 1, expert network 2, expert network 3 and asset expert network respectively. The target asset recognition model includes expert network 1 and gated network 1 for extracting the candidate multimedia attribute information related to the candidate multimedia data P i The target asset recognition model includes an expert network 2 and a gated network 2 for extracting the second key media attribute information related to the candidate multimedia data P from the candidate multimedia attribute information. i The shallow key media attribute information associated with the shallow transformation operation attribute, the target asset recognition model includes an expert network 3 and a gated network 3 for extracting the candidate multimedia attribute information related to the candidate multimedia data P iThe deep key media attribute information associated with the deep conversion operation attribute is determined, and the shallow key media attribute information and the deep key media attribute information are determined as the first key media attribute information.

[0079] S204: using the target operation recognition model's operation recognition network, based on the key object attribute information and the key media attribute information, predicting the target object's operation on the candidate multimedia data P. i Target operation information corresponding to the execution operation.

[0080] In this application, the computer device can use the operation recognition network of the target operation failure model to predict the target object for the candidate multimedia data P based on the key object attribute information and key media attribute information. i By extracting the target operation information corresponding to the candidate multimedia data P from the target object attribute information and the candidate media attribute information respectively, i The key feature information associated with the operation attribute is analyzed only for the key feature information, and there is no need to analyze the invalid information in the target object attribute information and the candidate media attribute information, which can save resources.

[0081] Optional, such as Figure 7 As shown, the key object attribute information includes the candidate multimedia data P i The first key object attribute information associated with the conversion operation attribute of the candidate multimedia data P i The second key object attribute information associated with the non-conversion operation attribute of the candidate multimedia data P i The first key media attribute information associated with the conversion operation attribute of the candidate multimedia data P i The above step S204 includes: using the operation recognition network of the target operation recognition model, based on the first key object attribute information and the first key media attribute information, predicting the target object for the candidate multimedia data P i Execute the conversion operation information corresponding to the conversion operation; predict the target object for the candidate multimedia data P according to the second key object attribute information and the second key media attribute information i Execute the non-conversion operation information corresponding to the non-conversion operation; determine the conversion operation information and the non-conversion operation information as a function of reflecting the target object's response to the candidate multimedia data P i Target operation information corresponding to the execution operation.

[0082] The computer device may use the operation recognition network of the target operation recognition model to predict, based on the first key object attribute information and the first key media attribute information, an operation recognition network for reflecting the target object's operation on the candidate multimedia data P. i The conversion probability corresponding to the conversion operation is executed, and the conversion probability is determined as the conversion operation information. Then, the operation recognition network is used to predict the second key object attribute information and the second key media attribute information to reflect the target object for the candidate multimedia data P i The non-conversion operation probability corresponding to the non-conversion operation is executed, and the non-conversion operation probability is determined as non-conversion operation information. Further, the conversion operation information and the non-conversion operation information can be determined as a function of reflecting the target object's response to the candidate multimedia data P. i By respectively obtaining the target object for the candidate multimedia data P i The operation information corresponding to the execution of the conversion operation and the execution of the non-conversion operation can be mined to find the target object and the candidate multimedia data P i More detailed association information between them can further improve the recommendation accuracy of multimedia data.

[0083] For example, Figure 6 As shown, the operation recognition network includes CTR (click-through rate) twin towers, CVR (conversion rate) twin towers, and deep CVR twin towers. The CTR (click-through rate) twin towers are used to predict the target object for the candidate multimedia data P according to the second key object attribute information and the second key media attribute information. i The non-conversion operation probability corresponding to the non-conversion operation is performed, that is, the CTR double tower is used to predict the probability of the target object for the candidate multimedia data P according to the second key object attribute information and the second key media attribute information. i The probability of executing the click operation corresponding to the click operation is predicted by using the CVR dual tower based on the shallow object attribute information and shallow media attribute information to reflect the target object for the candidate multimedia data P i The download operation probability corresponding to the download operation is predicted by using the CTR dual tower based on the deep object attribute information and the deep media attribute information to reflect the target object for the candidate multimedia data P i The probability of executing the order operation corresponding to the order operation.

[0084] S205: Determine the candidate multimedia data P according to the candidate media attribute information. i The media asset factor is used to reflect the media asset factor of the candidate multimedia data P i The actual amount of assets that the object to which the candidate multimedia data belongs needs to spend when the conversion operation is executed.

[0085] Optionally, the above step S205 includes: extracting the candidate multimedia data P from the candidate media attribute information. i The key asset attribute information associated with the asset attribute of the target asset identification model is obtained by performing cross-correlation identification on the key asset attribute information and obtaining cross-correlation information; the target asset identification model and the target operation identification model are independent of each other;

[0086] Perform deep correlation identification on the key asset attribute information to obtain deep correlation information; and determine the media asset factor based on the cross correlation information and the deep correlation information.

[0087] For example, Figure 8 As shown, the asset identification network of the target asset identification model includes a cross sub-network and a deep sub-network. The computer device can obtain the media asset factor through the asset identification network. The asset identification network can be referred to as a DCN (Deep and Cross Network), or of course, it can also refer to other networks. Specifically, the computer device extracts key asset attribute information associated with the asset attributes of the candidate multimedia data Pi from the candidate multimedia attribute information, and uses the cross sub-network of the asset identification network to perform cross-correlation relationship identification on the key asset attribute information to obtain cross-correlation relationship information. The cross-correlation relationship information is used to reflect the correlation relationship between the key asset attribute information of each two dimensions in the key asset attribute information. Furthermore, the deep sub-network of the asset identification network is used to perform deep correlation relationship identification on the key asset attribute information to obtain deep correlation relationship information. The deep correlation relationship information is used to reflect the correlation relationship between the key asset attribute information of two or more dimensions in the key asset attribute information. The media asset factor is determined based on the cross-correlation relationship information and the deep correlation relationship information. By mining the cross-correlation and deep correlation between each key asset attribute information, detailed information about asset attributes can be mined, providing more information for determining media asset factors and improving the accuracy of obtaining media asset factors.

[0088] Optionally, the above extracts the candidate multimedia data P from the candidate media attribute information. i The key asset attribute information associated with the asset attribute of the target asset recognition model includes: using the asset expert network of the target asset recognition model to extract the key asset attribute information associated with the candidate multimedia data P from the candidate media attribute information. i The attribute information associated with the asset attribute of the target asset is used as the candidate asset attribute information; the asset gating network of the target asset recognition model is used to determine the confidence of the candidate asset attribute information; the confidence is used to weight the candidate asset attribute information to obtain the confidence associated with the candidate multimedia data Pi Key asset attribute information associated with the asset attributes.

[0089] The target asset identification model includes one or more asset expert networks, one asset expert network corresponds to one asset gated network, and the computer device can obtain key asset attribute information through the asset expert network and the asset gated network. Specifically, for example, Figure 6 As shown in Figure 2, when the target asset recognition model includes an asset expert network and an asset gating network (i.e. Figure 6 When the gated network 4) in the target asset recognition model is used, the computer device can use the asset expert network of the target asset recognition model to extract the candidate multimedia data P from the candidate media attribute information. i The attribute information associated with the asset attribute of the target asset is used as the candidate asset attribute information. Then, the asset gating network of the target asset recognition model is used to determine the confidence of the candidate asset attribute information. The confidence is used to reflect the accuracy of the candidate asset attribute information. Further, the confidence is used to weight the candidate asset attribute information to obtain the confidence associated with the candidate multimedia data P. i When the target asset recognition model includes at least two asset expert networks and at least two asset gating networks, the computer device can respectively use the asset expert networks of the target asset recognition model to extract the key asset attribute information associated with the candidate multimedia data P from the candidate media attribute information. i The attribute information associated with the asset attribute of the target asset recognition model is used as the candidate asset attribute information. Then, each asset gating network of the target asset recognition model is used to determine the confidence of the corresponding candidate asset attribute information. Further, the corresponding candidate asset attribute information is weighted using the confidence to obtain the weighted candidate asset attribute information. The weighted candidate asset attribute information is fused to obtain the candidate multimedia data P i Key asset attribute information associated with the asset attributes. Key asset attribute information is obtained through an asset expert network and an asset gating network to improve the accuracy of obtaining key asset attribute information. It should be noted that the independence of the target asset identification model and the target operation identification model in this application can mean that the output factors of the model are independently fitted and independently modeled and optimized, which can improve the accuracy of the target asset identification model and the target operation identification model.

[0090] For example, the asset expert network can refer to a PNN structure, and the asset gating network structure can adopt a softmax (logistic regression) structure. Figure 9As shown, the asset gating network may include a mapping layer, a physical layer, a hidden layer 1, and a hidden layer 2. The mapping layer is used to convert the attribute information of each dimension in the candidate media attribute information into a feature vector of the same length, and the physical layer is used to identify the association between each feature vector. The hidden layer 1 and the hidden layer 2 are used to extract the candidate multimedia data P from the candidate media attribute information based on the association. i Key asset attribute information associated with the asset attributes.

[0091] Optionally, obtain sample object attribute information of the sample object, sample media attribute information of the sample multimedia data, and annotated operation information of the sample object regarding the sample multimedia data; use a candidate operation recognition model to predict the sample object attribute information and the sample media attribute information to obtain predicted operation information corresponding to the operation performed by the sample object on the sample multimedia data; adjust the candidate operation recognition model based on the annotated operation information and the predicted operation information, and determine the adjusted candidate operation recognition model as the target operation recognition model.

[0092] The computer device may use sample object attribute information of the sample object, sample media attribute information of the sample multimedia data, and annotated operation information of the sample object with respect to the sample multimedia data. The annotated operation information of the sample object with respect to the sample multimedia data may be determined based on log data of the sample object with respect to the sample multimedia data, thereby avoiding errors caused by manual annotation and automatically generating the annotated operation information, thereby improving the efficiency and accuracy of obtaining the annotated operation information. Alternatively, the computer device may obtain annotated operation information obtained by annotating the sample multimedia data with multiple objects in combination with the object attribute information of the sample object, correct the multiple annotated operation information, and obtain the annotated operation information of the sample multimedia data. A candidate operation recognition model may be used to predict the sample object attribute information and the sample media attribute information to obtain predicted operation information corresponding to the operation performed by the sample object on the sample multimedia data. If the annotated operation information is identical to or relatively close to the predicted operation information, then the candidate operation recognition model has a relatively high operation prediction accuracy. If the annotated operation information differs significantly from the predicted operation information, then the candidate operation recognition model has a relatively low operation prediction accuracy. Therefore, the computer device can adjust the candidate operation recognition model based on the annotated operation information and the predicted operation information, and determine the adjusted candidate operation recognition model as the target operation recognition model. The target operation recognition model is obtained by training the candidate operation recognition model based on the sample object attribute information and sample media attribute information of the sample object, thereby improving the operation prediction accuracy of the target operation recognition model. The target operation recognition model and the target asset recognition model are trained and optimized separately to improve the accuracy of the target operation recognition model.

[0093] Optionally, the marked operation information includes first marked conversion operation information, second marked conversion operation information and marked non-conversion operation information; the predicted operation information includes first predicted conversion operation information, second marked predicted conversion operation information and predicted non-conversion operation information; the first marked conversion operation information and the first pre-conversion operation information are used to reflect the probability that the target object performs the first conversion operation on the sample multimedia data, the second marked conversion operation information and the second predicted conversion operation information are used to reflect the probability that the target object performs the second conversion operation on the sample multimedia data, and the marked non-conversion operation information and the predicted non-conversion operation information are used to reflect the probability that the target object performs the non-conversion operation on the sample multimedia data. The above-mentioned adjustment of the candidate operation recognition model based on the labeled operation information and the predicted operation information includes: determining the first conversion operation prediction error of the candidate operation recognition model based on the first labeled conversion operation information and the first predicted conversion operation information; determining the second conversion operation prediction error of the candidate operation recognition model based on the second labeled conversion operation information and the second predicted conversion operation information; determining the non-conversion operation prediction error of the candidate operation recognition model based on the labeled non-conversion operation information and the predicted non-conversion operation information; adjusting the candidate operation recognition model based on the first conversion operation prediction error, the second conversion operation prediction error and the non-conversion operation prediction error to obtain an adjusted candidate operation recognition model.

[0094] The computer device may determine a first conversion operation prediction error for the candidate operation recognition model based on the first labeled conversion operation information and the first predicted conversion operation information. A lower first conversion operation prediction error indicates a higher recognition accuracy for the first conversion operation of the candidate operation recognition model; a higher first conversion operation prediction error indicates a lower recognition accuracy for the first conversion operation of the candidate operation recognition model. Then, based on the second labeled conversion operation information and the second predicted conversion operation information, a second conversion operation prediction error for the candidate operation recognition model may be determined. A lower second conversion operation prediction error indicates a higher recognition accuracy for the second conversion operation of the candidate operation recognition model; a higher second conversion operation prediction error indicates a lower recognition accuracy for the second conversion operation of the candidate operation recognition model. Furthermore, based on the labeled non-conversion operation information and the predicted non-conversion operation information, a non-conversion operation prediction error for the candidate operation recognition model may be determined. A lower non-conversion operation prediction error indicates a higher recognition accuracy for the non-conversion operation of the candidate operation recognition model; a higher non-conversion operation prediction error indicates a lower recognition accuracy for the non-conversion operation of the candidate operation recognition model. The computer device can adjust the candidate operation recognition model based on the first conversion operation prediction error, the second conversion operation prediction error, and the non-conversion operation prediction error to obtain an adjusted candidate operation recognition model. If the prediction errors of the candidate operation recognition model are all in a convergence state, the training of the candidate operation recognition model can be terminated to obtain an adjusted candidate operation recognition model. Alternatively, the computer device determines the operation prediction error of the candidate operation recognition model based on the first conversion operation prediction error, the second conversion operation prediction error, and the non-conversion operation prediction error, and adjusts the candidate operation recognition model based on the operation prediction error to obtain an adjusted candidate operation recognition model. The candidate operation recognition model information is trained by the operation prediction errors in multiple dimensions to obtain a target operation recognition model, thereby improving the operation prediction accuracy of the target operation recognition model.

[0095] Optionally, the aforementioned adjustment of the candidate operation recognition model based on the first conversion operation prediction error, the second conversion operation prediction error, and the non-conversion operation prediction error to obtain an adjusted candidate operation recognition model includes: the computer device may perform a weighted sum of the first conversion operation prediction error, the second conversion operation prediction error, and the non-conversion operation prediction error to obtain an operation prediction error of the candidate operation recognition model; if the operation prediction error is not in a converged state, the candidate operation recognition model is adjusted based on the operation prediction error to obtain an adjusted candidate operation recognition model. The operation prediction error, i.e., the total operation prediction error, is obtained by performing a weighted sum of the operation prediction errors in multiple dimensions, and the candidate operation recognition model information is trained based on the total operation prediction error to obtain a target operation recognition model, thereby improving the operation prediction accuracy of the target operation recognition model.

[0096] Optional, Figure 10 As shown, when recommending multimedia data to a target object, the prediction space is the entire multimedia data set (e.g., the entire advertising library). If the training samples only select the exposed multimedia data, this will lead to inconsistency between the model training space and the prediction space, i.e., the "sample selection bias" problem, which will result in a relatively low accuracy of the trained model. Therefore, the problem of inconsistency between the model training space and the prediction space can be avoided by adopting a multi-level negative sample construction scheme. Specifically, the computer device can filter out first sample multimedia data from the sample multimedia data set, which is the operation performed by the sample object in a historical time period, i.e., the first sample multimedia data is the multimedia data exposed to the sample object in the historical time period. Furthermore, the annotated operation information corresponding to the first sample multimedia data can be determined based on the log data of the sample object regarding the first sample multimedia data (i.e., the log data in the historical time period); second sample multimedia data on which the sample object did not perform an operation in the historical time period can be filtered out from the sample multimedia data set; the second sample multimedia data includes one or both of multimedia data that was not exposed to the sample object and multimedia data that was exposed to the sample object and on which the sample object did not perform an operation, i.e., the second sample multimedia data is a negative sample. The labeled operation information indicating that the sample object did not perform an operation on the second sample multimedia data is determined as the labeled operation information for the second sample multimedia data, and the first sample multimedia data and the second sample multimedia data are determined as the sample multimedia data corresponding to the sample object. Negative samples are constructed using the second sample multimedia data on which the sample object did not perform an operation within a historical time period, providing rich and comprehensive training data for the model training process. This ensures that the model's prediction space is consistent with the training space, thereby improving the accuracy of model training.

[0097] It should be noted that the above-mentioned sample multimedia data set and the multimedia data set may be the same or different, but the types of sample multimedia data included in the sample multimedia data set may cover the types of candidate multimedia data included in the multimedia data set, and the number of sample multimedia data included in the sample multimedia data set may be greater than the number of candidate multimedia data included in the multimedia data set, which is conducive to improving the accuracy of model training.

[0098] Optionally, the above-mentioned filtering out, from the sample multimedia data set, the second sample multimedia data on which the sample object has not performed an operation within the historical time period, includes: randomly selecting, from the sample multimedia data set, sample multimedia data that has not been recommended to the sample object within the historical time period as first candidate sample multimedia data; selecting, from the sample multimedia data set, sample multimedia data that has been recommended to the sample object within the historical time period and has not been operated upon by the sample object as second candidate sample multimedia data; and determining the first candidate sample multimedia data and the second candidate sample multimedia data as the second sample multimedia data on which the sample object has not performed an operation within the historical time period.

[0099] The computer device may randomly select, from the sample multimedia data set, sample multimedia data that was not recommended (i.e., not exposed) to the sample subject within the historical time period as first candidate sample multimedia data, and select, from the sample multimedia data set, sample multimedia data that was recommended to the sample subject within the historical time period and was not operated on by the sample subject as second candidate sample multimedia data, i.e., the second candidate sample multimedia data is multimedia data that was exposed to the sample subject and was not operated on by the sample subject. That is, both the first candidate sample multimedia data and the second candidate sample multimedia data are negative sample multimedia data, and the first candidate sample multimedia data and the second candidate sample multimedia data are determined to be second sample multimedia data that was not operated on by the sample subject within the historical time period. Multi-level (i.e., multiple types) of negative sample multimedia data are constructed using the first candidate sample multimedia data and the second candidate sample multimedia data, so that the prediction space of the model is consistent with the training space, thereby improving the accuracy of model training.

[0100] Optionally, obtain the media expected asset information of the sample multimedia data; the media expected asset information is used to reflect the expected asset information of the candidate multimedia data P iThe expected amount of assets that the object to which the candidate multimedia data belongs needs to expend when the operation is converted and executed; determining the media annotated asset information of the sample multimedia data based on the predicted operation information and the media expected asset information; using a candidate asset recognition model, predicting the predicted operation information and the sample media attribute information to obtain the media predicted asset information of the sample multimedia data; adjusting the candidate asset recognition model based on the media annotated asset information and the media predicted asset information, and determining the adjusted candidate asset recognition model as the target asset recognition model.

[0101] The computer device can obtain the media expected asset information of the sample multimedia data; the media expected asset information is used to reflect the expected asset information of the candidate multimedia data P i The expected amount of assets that the object to which the candidate multimedia data belongs will need to expend when the operation is converted and executed. Then, based on the predicted operation information and the expected media asset information, the media annotated asset information of the sample multimedia data is determined. A candidate asset recognition model is used to predict the predicted operation information and the sample media attribute information to obtain the predicted media asset information of the sample multimedia data. If the media annotated asset information is identical or relatively close to the predicted media asset information, the candidate asset recognition model has a relatively high asset prediction accuracy. If the media annotated asset information and the predicted media asset information differ significantly, the candidate asset recognition model has a relatively low operation prediction accuracy. Therefore, based on the media annotated asset information and the predicted media asset information, the candidate asset recognition model is adjusted, and the adjusted candidate asset recognition model is determined as the target asset recognition model. By training the candidate asset recognition model based on the expected media asset information and the annotated media asset information, a target asset recognition model is obtained, thereby improving the asset prediction accuracy of the target asset recognition model. The target operation recognition model and the target asset recognition model are trained and optimized separately to improve the accuracy of the target asset recognition model. At the same time, the media annotation asset information of the sample multimedia data is automatically generated according to the predicted operation information and the media expected asset information, without the need for human intervention, thereby improving the generation efficiency of the media annotation asset information.

[0102] For example, the function of the traditional CTCVR model can be expressed by the following formula (1):

[0103] CTCVR=sigmoid(user*ad) (1)

[0104] In formula (1), user represents object attribute information, i.e., user feature vector, and ad refers to media attribute information, i.e., advertisement feature vector. When the commercial value of media data (eCPM) needs to be further considered, the CTCVR model can be transformed into the following formula (2) under the traditional architecture:

[0105] ECPM=bid*sigmoid(user*ad) (2)

[0106] Wherein, in formula (2), bid represents the expected amount of assets that the object to which the sample multimedia data belongs needs to spend when the sample multimedia data is converted. In this application, based on the ANN retrieval architecture, the candidate asset identification model can be expressed as follows:

[0107]

[0108] Wherein, in formula (3), Gann is the media asset factor (i.e., the advertising bid factor), and Gann-weighted-ad can be expressed using the following formula (4):

[0109] Gann-weighted-ad=Gann*ad (4)

[0110] in, Figure 11 As shown, the traditional ANN retrieval architecture refers to the use of target object attribute information

[0111] (ie user embedding) retrieves candidate media attribute information (ie ad embedding). The ANN retrieval architecture in this application uses user embedding to retrieve G_ann weighted ad embedding, thereby fitting the eCPM.

[0112] Optionally, the candidate asset identification model is adjusted based on the media labeled asset information and the media predicted asset information, including: determining the asset prediction error of the candidate asset identification model based on the media labeled asset information and the media predicted asset information; if the asset prediction error is not in a convergence state, adjusting the candidate asset identification model based on the asset prediction error to obtain an adjusted candidate asset identification model.

[0113] The computer device can determine an asset prediction error of the candidate asset identification model based on the media-labeled asset information and the media-predicted asset information. A relatively low asset prediction error indicates that the candidate asset identification model has a relatively high asset prediction accuracy. Conversely, a relatively high asset prediction error indicates that the candidate asset identification model has a relatively low asset prediction accuracy. Therefore, if the asset prediction error is not in a convergent state, it indicates that the asset prediction accuracy of the candidate asset identification model is relatively low. Therefore, the candidate asset identification model is adjusted based on the asset prediction error to obtain an adjusted candidate asset identification model.

[0114] It should be noted that the media prediction asset information here includes deep prediction asset information (i.e., deep commercial value) and shallow media prediction asset information (i.e., shallow commercial value). The media prediction asset information can be expressed using the following formula (5):

[0115]

[0116] In formula (5), shallow-predict refers to shallow commercial value, and deep-predict refers to deep commercial value. The loss function of the candidate asset identification model can adopt Huber Loss in regression loss, as shown in the following formula (6):

[0117] loss aux =HuberLoss(predict,ecpm) (6)

[0118] Wherein, predict in formula (6) represents shallow-predict or deep-predict, and ecpm represents the media-annotated asset information, which may include shallow-annotated asset information and deep-annotated asset information. That is, when predict represents shallow-predict, ecpm represents shallow-annotated asset information; when predict represents deep-predict, ecpm represents deep-annotated asset information, which can be expressed by the following formula (7):

[0119] eCPM=bid×pCTR×pCVR×λ (7)

[0120] In formula (7), pCTR and pCVR represent the estimated click-through rate and conversion rate output by the candidate operation model, respectively. Huber Loss is defined as follows (8):

[0121]

[0122] Among them, L in formula (8) δ (a) represents the asset prediction error of the candidate asset recognition model, and a represents the residual between the media-annotated asset information and the media-predicted asset information. The starting point for using Huber Loss here is to prevent the high eCPM sample loss from causing unstable model convergence.

[0123] It should be noted that the total prediction error Loss of the candidate operation identification model and the candidate asset identification model can be expressed by the following formula (9):

[0124] Loss = αloss ctr +βloss shallow_cvr+γloss deep_cvr +θloss aux (9)

[0125] In formula (9), loss aux Represents the asset prediction error of the candidate asset identification model, loss ctr Represents the non-conversion operation prediction error of the candidate operation recognition model, loss shallow-cvr Represents the deep conversion operation prediction error of the candidate operation recognition model, loss deep-cvr 、loss shallow-cvr 、loss ctr Both can be expressed using the cross entropy loss function in formula (10):

[0126] loss x =-∑[y i *logP i +(1-y i )*log(1-P i )] (10)

[0127] Where Pi in formula (10) is the predicted operation information of the candidate operation recognition model, yi represents the labeled operation information, and x represents the input of the candidate operation recognition model, namely, the sample object attribute information and the sample media attribute information.

[0128] S206 : Select candidate multimedia data for pushing to the target object from the multimedia data set according to the target operation information and the media asset factor, as target multimedia data, and push the target multimedia data to the terminal corresponding to the target object.

[0129] Optionally, the above step S206 includes: determining the candidate multimedia data P according to the target operation information and the media asset factor. i The media estimation asset information is used to reflect the media estimation asset information of the candidate multimedia data P i When the target object performs an operation, the amount of assets that the object to which the candidate multimedia data belongs needs to spend; i The estimated media asset information is used to select candidate multimedia data for pushing to the target object from the multimedia data set as the target multimedia data.

[0130] The computer device can determine the candidate multimedia data P according to the target operation information and the media asset factor. i The media estimation asset information is used to reflect the media estimation asset information of the candidate multimedia data P iWhen the target object performs an operation, the asset amount that the object to which the candidate multimedia data belongs needs to spend, that is, the media estimated asset information is used to reflect the amount of assets that the candidate multimedia data P i When the target object performs an operation, the estimated amount of assets that the object to which the candidate multimedia data belongs needs to pay to the multimedia platform. The higher the estimated amount of assets, the higher the amount of assets that the object to which the candidate multimedia data belongs needs to pay to the multimedia platform, that is, the higher the asset value of the candidate multimedia data P i The higher the commercial value of the candidate multimedia data, the lower the estimated asset amount is, indicating that the object to which the candidate multimedia data belongs needs to pay a lower asset amount to the multimedia platform, that is, the candidate multimedia data P i The lower the commercial value of the multimedia data set, the lower the value. The computer device can select candidate multimedia data from the multimedia data set whose estimated asset amount indicated by the media estimated asset information is greater than the asset amount threshold as the target multimedia data, and push the target multimedia data to the terminal corresponding to the target object. Recommending multimedia data based on the media estimated asset information helps maximize the commercial value of the multimedia platform and improve the utilization value and resource utilization of the multimedia platform.

[0131] Optionally, the conversion operation information includes information for reflecting the target object's response to the candidate multimedia data P i Execute the first conversion operation corresponding to the first conversion operation, and the first conversion operation information for reflecting the target object for the candidate multimedia data P i Execute the second conversion operation information corresponding to the second conversion operation, the first conversion operation is given to the candidate multimedia data P i The asset amount brought by the object to which it belongs is greater than the asset amount brought by the second conversion operation to the candidate multimedia data P i The first conversion operation information is used to reflect the target object's conversion of the candidate multimedia data P i The first conversion probability (ie, deep conversion probability) corresponding to the first conversion operation is executed, and the second conversion operation information is used to reflect the target object's response to the candidate multimedia data P i The second conversion probability (ie, shallow conversion probability) corresponding to the second conversion operation is executed. The candidate multimedia data P is determined based on the target operation information and the media asset factor. i The method comprises: determining first media estimated asset information according to the media asset factor, the first conversion operation information, and the non-conversion operation information; determining second media estimated asset information according to the media asset factor, the second conversion operation information, and the non-conversion operation information; and determining the first media estimated asset information and the second media estimated asset information as the candidate multimedia data P. i Media estimated asset information.

[0132] The computer device may determine first media estimated asset information by multiplying the media asset factor, the first conversion probability corresponding to the first conversion operation information, and the non-conversion probability corresponding to the non-conversion operation information; and determine second media estimated asset information by multiplying the media asset factor, the second conversion probability corresponding to the second conversion operation information, and the non-conversion probability corresponding to the non-conversion operation information; and determine the first media estimated asset information and the second media estimated asset information as the candidate multimedia data P. i By analyzing the target object for the candidate multimedia data P i The shallow conversion probability, media asset factor, deep conversion probability and non-conversion probability corresponding to the execution operation are used to obtain the candidate multimedia data P i The media estimated asset information can be used to analyze the multi-dimensional information of the target object and improve the acquisition of the candidate multimedia data P i The accuracy of media estimation asset information can be improved, and the recommendation accuracy of multimedia data can be further improved.

[0133] For example, Figure 12 As shown, the first conversion operation information includes deep conversion probability, such as order probability, the second conversion operation information includes shallow conversion probability, such as download probability, and non-conversion operation information includes click probability. The computer device can determine the product of deep conversion probability, click probability and media asset factor as deep commercial value (i.e., first media estimated asset information), and determine the product of shallow conversion probability, click probability and media asset factor as shallow commercial value (i.e., second media estimated asset information).

[0134] Optionally, the above-mentioned candidate multimedia data P i The media estimation asset information of the multimedia data set is selected from the multimedia data set for pushing candidate multimedia data to the target object as the target multimedia data, including: obtaining a multimedia data network; the multimedia data network includes a method for reflecting the candidate multimedia data P in the multimedia data set i , and edges formed by connecting nodes corresponding to candidate multimedia data with associated relationships; traversing the nodes in the multimedia data network according to the node path of the multimedia data network and the media estimation asset information to obtain candidate multimedia data with a neighboring relationship with the target object attribute information; determining the candidate multimedia data with a neighboring relationship with the target object attribute information in the multimedia data set as target multimedia data for pushing to the target object.

[0135] The computer may obtain a multimedia data network, which includes a multimedia data network for reflecting the candidate multimedia data P in the multimedia data set. iThe nodes of the media data network and the edges formed by connecting the nodes corresponding to the candidate multimedia data with the associated relationship; if the nodes in the media data network include candidate media attribute information of the candidate multimedia data, the edges of the multimedia data network are formed by connecting the nodes whose matching degree between the candidate multimedia attribute information is greater than the matching degree threshold. The media data network may include one or more sub-networks. When the media data network includes a sub-network, the sub-network may include M nodes, one node corresponding to one candidate multimedia data. The computer device may traverse the nodes in the multimedia data network according to the node path of the multimedia data network and the media estimated asset information to obtain the candidate multimedia data with a neighboring relationship with the target object attribute information. For example, the candidate multimedia data are sorted in descending order of the estimated asset amount indicated by the media estimated asset information, and the top k candidate multimedia data ranked in the sub-network are determined as the candidate multimedia data with a neighboring relationship with the target object attribute information. Alternatively, based on the node path of the multimedia data network, when k candidate multimedia data items are traversed and the estimated asset amount indicated by the media estimated asset information is greater than the asset amount threshold, the traversal ends and the candidate multimedia data items with the estimated asset amount indicated by the k media estimated asset information being greater than the asset amount threshold are determined as candidate multimedia data items having a close neighbor relationship with the target object attribute information. Then, the candidate multimedia data items in the multimedia data set that have a close neighbor relationship with the target object attribute information are determined as target multimedia data to be pushed to the target object. Through the multimedia data network, analysis of the entire set of candidate multimedia data items can be avoided, saving resources and improving the efficiency of acquiring target media data.

[0136] For example, Figure 13a As shown, the multimedia data network includes three layers, namely layer0, layer1, and layer2. Layer0 includes candidate media attribute information of all candidate multimedia data in the multimedia data set, and layer2 only includes candidate media attribute information of one candidate multimedia data in the multimedia data set. The first media estimated asset information (deep commercial value) can be used as a search indicator to retrieve candidate multimedia data corresponding to the candidate media attribute information having a close neighbor relationship with the target object attribute information from the multimedia data network, and such candidate multimedia data is used as a deep optimization target advertising library. The second media estimated asset information (shallow commercial value) can be used as a search indicator to retrieve candidate multimedia data corresponding to the candidate media attribute information having a close neighbor relationship with the target object attribute information from the multimedia data network, and such candidate multimedia data is used as a shallow optimization target advertising library.

[0137] For example, Figure 13bAs shown, after obtaining the deep optimization target database and the shallow optimization target database, the computer device can process the candidate multimedia data in the deep optimization target database and the shallow optimization target database through rough sorting, multi-way merging, and fine sorting to obtain the target multimedia data, and push the target multimedia data to the terminal corresponding to the target object. In the rough sorting process, the click-through rate and conversion rate of each advertisement are first estimated to calculate the eCPM score of each advertisement: eCPM = bid (ad bid) * liteCTR (estimated click-through rate) * liteCVR (estimated conversion rate). Then, the top N advertisements are selected based on the order of the advertisements' eCPM. During the precise ranking process, the multimedia platform first obtains the pCTR (precisely estimated click-through rate) and pCVR (precisely estimated conversion rate) of all ads, and then calculates the precise ranking bidding scores for the top N ads: eCPM2 = bid (ad bid) * pCTR (precisely estimated click-through rate) * pCVR (precisely estimated conversion rate). Based on the high or low eCPM2, the top 1-2 ads are selected and presented to users.

[0138] Optionally, when the media data network includes at least two subnetworks, the number of nodes in each subnetwork may be the same or different, and the candidate multimedia data corresponding to the nodes in the subnetworks may be partially or completely different. In this case, the computer device may randomly select a target subnetwork from the at least two subnetworks. When k candidate multimedia data items with estimated asset amounts indicated by media estimated asset information exceeding a threshold are obtained through traversal in the target subnetwork, the traversal ends and the k candidate multimedia data items with estimated asset amounts indicated by media estimated asset information exceeding the threshold are determined as candidate multimedia data items having a close neighbor relationship with the target object attribute information. Alternatively, the estimated asset amounts indicated by the media estimated asset information in the target subnetwork are ranked among the first k candidate multimedia data items and determined as candidate multimedia data items having a close neighbor relationship with the target object attribute information. This avoids analyzing the entire amount of candidate multimedia data, saves resources, and improves the efficiency of acquiring target media data.

[0139] Optionally, when the media data network may include at least two sub-networks, the multimedia data network includes sub-network N j and subnetwork N j+1 , the subnetwork N j The number of nodes in the subnetwork is less than N j+1 The number of nodes in the subnetwork N j The set of candidate multimedia data corresponding to the nodes in the subnetwork N j+1 The subset of the set consisting of candidate multimedia data corresponding to the nodes in the subnetwork N KThe number of nodes in is the same as the corresponding number of candidate multimedia data in the multimedia data set, j is a positive integer less than K, and K is the number of sub-networks in the multimedia data network.

[0140] The above-mentioned traversing the nodes in the multimedia data network according to the node path of the multimedia data network and the media estimated asset information to obtain the candidate multimedia data having a neighbor relationship with the target object attribute information includes: according to the node path of the multimedia data network, from the subnetwork N j The candidate multimedia data whose estimated asset amount indicated by the media estimated asset information is the maximum estimated asset amount is determined as the designated candidate multimedia data; the sub-network N j+1 The node used to reflect the specified candidate multimedia data is determined as the initial traversal node; if the subnetwork N j+1 The number of nodes in the subnetwork N is less than the number of candidate multimedia data in the multimedia data set, then the initial traversal node is used as the traversal starting point, and the subnetwork N j+1 The candidate multimedia data whose estimated asset amount indicated by the media estimated asset information is determined to be the maximum estimated asset amount will be transmitted from the sub-network N j+1 The determined candidate multimedia data is updated to the designated candidate multimedia data; if the sub-network N j+1 If the number of nodes in is equal to the corresponding number of candidate multimedia data in the multimedia data set, then among the nodes connected to the initial traversal node, the nodes whose asset quantity indicated by the media estimated asset information is greater than the asset quantity threshold are traversed as candidate multimedia data having a neighboring relationship with the object attribute information.

[0141] Computer devices can be connected from subnet N1 to N K In the order of the node path of the multimedia data network, from the sub-network N j The candidate multimedia data whose estimated asset amount indicated by the media estimated asset information is the maximum estimated asset amount is determined as the designated candidate multimedia data; the sub-network N j+1 The node used to reflect the specified candidate multimedia data is determined as the initial traversal node. j+1 The number of nodes in is less than the number of candidate multimedia data in the multimedia data set, indicating that there is a sub-network that has not been traversed. Therefore, the initial traversal node is used as the traversal starting point, and the sub-network N j+1 The candidate multimedia data whose estimated asset amount indicated by the media estimated asset information is determined to be the maximum estimated asset amount will be transmitted from the sub-network N j+1 The determined candidate multimedia data is updated to the designated candidate multimedia data. j+1If the number of nodes in the traversal equals the number of candidate multimedia data in the multimedia data set, indicating that all subnetworks have been traversed, then among the nodes connected to the initial traversal node, nodes whose estimated media asset amount is greater than the asset amount threshold are traversed and selected as candidate multimedia data with a close neighbor relationship with the object attribute information. This avoids analyzing the entire set of candidate multimedia data, saves resources, and improves the efficiency of acquiring target media data.

[0142] In this application, the target operation information is specifically used to reflect the probability of the target object performing an operation (such as a shallow conversion operation, a deep conversion operation, or a non-conversion operation) on the candidate multimedia data. The target operation information can, to a certain extent, reflect the target object's interest in the candidate multimedia data. For example, the target operation information indicates that the probability of the target object performing a deep conversion operation on the candidate multimedia data is relatively high, indicating that the target object has a high interest in the candidate multimedia data. The media asset factor is used to reflect the probability of the candidate multimedia data P being used to represent the target object's interest in the candidate multimedia data. i When the conversion operation is executed, the actual amount of assets that the object to which the candidate multimedia data belongs needs to spend, that is, the media asset factor can, to a certain extent, reflect the commercial value that the candidate multimedia data brings to the multimedia platform. By selecting the candidate multimedia data to be pushed to the target object from the multimedia data set according to the target operation information and the media asset factor, the target multimedia data is pushed to the terminal corresponding to the target object. In other words, by comprehensively considering the user's interest characteristics in the candidate multimedia data and the commercial value of the candidate multimedia data, multimedia data can be recommended to the user, which can achieve accurate recommendation and improve the accuracy of multimedia data recommendation; it can not only avoid recommending multimedia data that the user is not interested in to the user, resulting in invalid multimedia data recommendation and wasting the resources of the multimedia platform, but also avoid recommending multimedia data with relatively low commercial value to the user, resulting in relatively low resource utilization of the multimedia platform, thereby improving the resource utilization of the multimedia platform.

[0143] See Figure 14 , is a structural diagram of a multimedia data processing device provided in an embodiment of the present application. The multimedia data processing device can be a computer program (including program code) running on a computer device, for example, the multimedia data processing device is an application software; the device can be used to execute the corresponding steps of the method provided in an embodiment of the present application. Figure 14 As shown, the multimedia data processing device may include: an acquisition module 141 , a prediction module 142 , a determination module 143 and a selection module 144 .

[0144] The acquisition module is used to obtain the target object attribute information of the target object and the candidate multimedia data P in the multimedia data set to be pushed. i candidate media attribute information; i is a positive integer less than or equal to M, and M is the number of candidate multimedia data in the multimedia data set;

[0145] The prediction module is configured to predict, based on the target object attribute information and the candidate media attribute information, a value that reflects the target object's response to the candidate multimedia data P i Target operation information corresponding to the execution operation;

[0146] A determination module is used to determine the candidate multimedia data P according to the candidate media attribute information. i The media asset factor is used to reflect the media asset factor of the candidate multimedia data P i The actual amount of assets that the object to which the candidate multimedia data belongs needs to spend when the conversion operation is executed;

[0147] A selection module is used to select candidate multimedia data for pushing to the target object from the multimedia data set according to the target operation information and the media asset factor, as target multimedia data, and push the target multimedia data to the terminal corresponding to the target object.

[0148] Obtain the target object attribute information of the target object and the candidate multimedia data P in the multimedia data set to be pushed i candidate media attribute information; i is a positive integer less than or equal to M, and M is the number of candidate multimedia data in the multimedia data set;

[0149] According to the target object attribute information and the candidate media attribute information, a prediction is made to reflect the target object's response to the candidate multimedia data P i Target operation information corresponding to the execution operation;

[0150] According to the candidate media attribute information, the candidate multimedia data P is determined i The media asset factor is used to reflect the media asset factor of the candidate multimedia data P i The actual amount of assets that the object to which the candidate multimedia data belongs needs to spend when the conversion operation is executed;

[0151] According to the target operation information and the media asset factor, candidate multimedia data for being pushed to the target object is selected from the multimedia data set as target multimedia data, and the target multimedia data is pushed to a terminal corresponding to the target object.

[0152] Optionally, the prediction module predicts the target object for the candidate multimedia data P according to the target object attribute information and the candidate media attribute information. i The target operation information corresponding to the execution operation includes:

[0153] Extract the target object attribute information from the candidate multimedia data P i Key object attribute information associated with the operation attribute;

[0154] Extract the candidate multimedia data P from the candidate media attribute information i Key media attribute information associated with the operation attributes;

[0155] The operation recognition network using the target operation recognition model predicts the target object's operation for the candidate multimedia data P based on the key object attribute information and the key media attribute information. i Target operation information corresponding to the execution operation.

[0156] Optionally, the key object attribute information includes the key object attribute information related to the candidate multimedia data P i The first key object attribute information associated with the conversion operation attribute, and the candidate multimedia data P i The second key object attribute information associated with the non-conversion operation attribute; the key media attribute information includes the second key object attribute information associated with the candidate multimedia data P i The first key media attribute information associated with the conversion operation attribute, and the first key media attribute information associated with the candidate multimedia data P i Second key media attribute information associated with the non-conversion operation attribute;

[0157] Optionally, the prediction module uses an operation recognition network of a target operation recognition model to predict, based on the key object attribute information and the key media attribute information, the target object for the candidate multimedia data P i The target operation information corresponding to the execution operation includes:

[0158] The operation recognition network of the target operation recognition model is used to predict the target object's operation for the candidate multimedia data P based on the first key object attribute information and the first key media attribute information. i Conversion operation information corresponding to executing the conversion operation;

[0159] According to the second key object attribute information and the second key media attribute information, a prediction is made for reflecting the target object's response to the candidate multimedia data P i Execute non-conversion operation information corresponding to the non-conversion operation;

[0160] The conversion operation information and the non-conversion operation information are determined as the information used to reflect the target object's response to the candidate multimedia data P i Target operation information corresponding to the execution operation.

[0161] Optionally, the determination module determines the candidate multimedia data P according to the candidate media attribute information. i Media asset factors include:

[0162] Extract the candidate multimedia data P from the candidate media attribute information i Key asset attribute information associated with the asset attributes;

[0163] An asset identification network of a target asset identification model is used to perform cross-correlation identification on the key asset attribute information to obtain cross-correlation information; the target asset identification model and the target operation identification model are independent of each other;

[0164] Performing deep correlation identification on the key asset attribute information to obtain deep correlation information;

[0165] The media asset factor is determined according to the cross-association relationship information and the deep-association relationship information.

[0166] Optionally, the determination module extracts the candidate multimedia data P from the candidate media attribute information. i The key asset attribute information associated with the asset attributes includes:

[0167] The asset expert network of the target asset recognition model is used to extract the candidate multimedia data P from the candidate media attribute information. i The attribute information associated with the asset attribute is used as the candidate asset attribute information;

[0168] Determining the confidence level of the candidate asset attribute information using an asset gating network of the target asset identification model;

[0169] The confidence level is used to weight the candidate asset attribute information to obtain the candidate multimedia data P i Key asset attribute information associated with the asset attributes.

[0170] Optionally, the selection module selects candidate multimedia data for pushing to the target object from the multimedia data set according to the target operation information and the media asset factor, including:

[0171] Determine the candidate multimedia data P according to the target operation information and the media asset factor iThe media estimation asset information is used to reflect the media estimation asset information when the candidate multimedia data P i An estimated amount of assets that the object to which the candidate multimedia data belongs needs to spend when the target object performs an operation;

[0172] According to the candidate multimedia data P i The media estimation asset information is used to select candidate multimedia data for pushing to the target object from the multimedia data set as target multimedia data.

[0173] Optionally, the conversion operation information includes information for reflecting the target object's response to the candidate multimedia data P i Execute the first conversion operation corresponding to the first conversion operation, and the first conversion operation information for reflecting the target object for the candidate multimedia data P i Execute the second conversion operation information corresponding to the second conversion operation, the first conversion operation is performed on the candidate multimedia data P i The asset amount brought by the object to which it belongs is greater than the asset amount brought by the second conversion operation to the candidate multimedia data P i The amount of assets brought by the object;

[0174] Optionally, the selection module determines the candidate multimedia data P according to the target operation information and the media asset factor. i Estimated media asset information, including:

[0175] determining first media estimated asset information according to the media asset factor, the first conversion operation information, and the non-conversion operation information;

[0176] determining second media estimated asset information based on the media asset factor, the second conversion operation information, and the non-conversion operation information;

[0177] The first media estimated asset information and the second media estimated asset information are determined as the candidate multimedia data P i Media estimated asset information.

[0178] Optionally, the selection module selects the candidate multimedia data P i The method further comprises: selecting candidate multimedia data for pushing to the target object from the multimedia data set as target multimedia data, including:

[0179] Acquire a multimedia data network; the multimedia data network includes a method for reflecting the candidate multimedia data P in the multimedia data set i , and the edges formed by connecting the nodes corresponding to the candidate multimedia data with associated relationships;

[0180] Traversing nodes in the multimedia data network according to the node path of the multimedia data network and the media estimation asset information to obtain candidate multimedia data having a neighbor relationship with the target object attribute information;

[0181] The candidate multimedia data in the multimedia data set that has a neighbor relationship with the target object attribute information is determined as the target multimedia data for being pushed to the target object.

[0182] Optionally, the multimedia data network includes a subnetwork N j and subnetwork N j+1 , the sub-network N j The number of nodes in the subnetwork is less than N j+1 The number of nodes in the subnetwork N j The set of candidate multimedia data corresponding to the nodes in the subnetwork N j+1 A subset of the set consisting of candidate multimedia data corresponding to the nodes in , j is a positive integer less than K, and K is the number of subnetworks in the multimedia data network;

[0183] The selection module traverses the nodes in the multimedia data network according to the node path of the multimedia data network and the media estimation asset information to obtain candidate multimedia data having a neighbor relationship with the target object attribute information, including:

[0184] According to the node path of the multimedia data network, from the sub-network N j determining, as designated candidate multimedia data, the candidate multimedia data for which the estimated asset amount indicated by the media estimated asset information is the maximum estimated asset amount;

[0185] The sub-network N j+1 The node reflecting the specified candidate multimedia data is determined as the initial traversal node;

[0186] If the sub-network N j+1 The number of nodes in the subnetwork N is less than the number of candidate multimedia data in the multimedia data set, then the initial traversal node is used as the traversal starting point, and the subnetwork N j+1 The candidate multimedia data whose estimated asset amount indicated by the media estimated asset information is the maximum estimated asset amount is determined to be from the sub-network N j+1 The determined candidate multimedia data is updated as designated candidate multimedia data;

[0187] If the sub-network N j+1If the number of nodes in is equal to the corresponding number of candidate multimedia data in the multimedia data set, then among the nodes connected to the initial traversal node, the nodes whose asset amount indicated by the media estimated asset information is greater than the asset amount threshold are traversed as candidate multimedia data having a neighbor relationship with the object attribute information.

[0188] Optionally, an acquisition module is used to obtain sample object attribute information of a sample object, sample media attribute information of sample multimedia data, and annotated operation information of the sample object regarding the sample multimedia data; a candidate operation recognition model is used to predict the sample object attribute information and the sample media attribute information to obtain predicted operation information corresponding to the operation performed by the sample object on the sample multimedia data; the candidate operation recognition model is adjusted according to the annotated operation information and the predicted operation information, and the adjusted candidate operation recognition model is determined as the target operation recognition model.

[0189] Optionally, the labeled operation information includes first labeled conversion operation information, second labeled conversion operation information, and labeled non-conversion operation information; the predicted operation information includes first predicted conversion operation information, second labeled prediction operation information, and predicted non-conversion operation information; the acquisition module adjusts the candidate operation recognition model according to the labeled operation information and the predicted operation information, including:

[0190] determining a first conversion operation prediction error of the candidate operation recognition model based on the first labeled conversion operation information and the first predicted conversion operation information;

[0191] determining a second conversion operation prediction error of the candidate operation recognition model according to the second labeled conversion operation information and the second predicted conversion operation information;

[0192] determining a non-conversion operation prediction error of the candidate operation recognition model based on the marked non-conversion operation information and the predicted non-conversion operation information;

[0193] The candidate operation recognition model is adjusted according to the first conversion operation prediction error, the second conversion operation prediction error, and the non-conversion operation prediction error to obtain an adjusted candidate operation recognition model.

[0194] Optionally, the acquisition module adjusts the candidate operation recognition model according to the first conversion operation prediction error, the second conversion operation prediction error, and the non-conversion operation prediction error to obtain an adjusted candidate operation recognition model, including:

[0195] performing a weighted summation of the first conversion operation prediction error, the second conversion operation prediction error, and the non-conversion operation prediction error to obtain an operation prediction error of the candidate operation recognition model;

[0196] If the operation prediction error is not in a converged state, the candidate operation recognition model is adjusted according to the operation prediction error to obtain an adjusted candidate operation recognition model.

[0197] Optionally, the acquiring module acquires sample media attribute information of the sample multimedia data and the annotation operation information of the sample object on the sample multimedia data, including:

[0198] Filtering first sample multimedia data of operations performed by the sample object within a historical time period from the sample multimedia data set;

[0199] determining, according to log data of the sample object regarding the first sample multimedia data, annotation operation information corresponding to the first sample multimedia data;

[0200] Filtering out second sample multimedia data from the sample multimedia data set, for which the sample object has not performed any operation within a historical time period;

[0201] determining, as the marking operation information of the second sample multimedia data, the marking operation information indicating that the sample object has not performed an operation on the second sample multimedia data;

[0202] The first sample multimedia data and the second sample multimedia data are determined as sample multimedia data corresponding to the sample object.

[0203] Optionally, the acquiring module filters out, from the sample multimedia data set, second sample multimedia data on which the sample object has not performed any operation within a historical time period, including:

[0204] Randomly selecting, from the sample multimedia data set, sample multimedia data that has not been recommended to the sample object within the historical time period as first candidate sample multimedia data;

[0205] Selecting, from the sample multimedia data set, sample multimedia data that is recommended to the sample object within the historical time period and has not been operated by the sample object as second candidate sample multimedia data;

[0206] The first candidate sample multimedia data and the second candidate sample multimedia data are determined to be second sample multimedia data on which no operation is performed by the sample object within a historical time period.

[0207] Optionally, an acquisition module is used to acquire media expected asset information of the sample multimedia data; the media expected asset information is used to reflect the expected asset information of the candidate multimedia data P i When the operation is converted and executed, the expected amount of assets that the object to which the candidate multimedia data belongs needs to spend; based on the predicted operation information and the media expected asset information, the media annotated asset information of the sample multimedia data is determined; using a candidate asset recognition model, the predicted operation information and the sample media attribute information are predicted to obtain the media predicted asset information of the sample multimedia data; based on the media annotated asset information and the media predicted asset information, the candidate asset recognition model is adjusted, and the adjusted candidate asset recognition model is determined as the target asset recognition model.

[0208] Optionally, the acquisition module adjusts the candidate asset recognition model according to the media annotated asset information and the media predicted asset information, including:

[0209] determining an asset prediction error of the candidate asset identification model based on the media annotated asset information and the media predicted asset information;

[0210] If the asset prediction error is not in a converged state, the candidate asset identification model is adjusted according to the asset prediction error to obtain an adjusted candidate asset identification model.

[0211] In this application, the target operation information is specifically used to reflect the probability of the target object performing an operation (such as a shallow conversion operation, a deep conversion operation, or a non-conversion operation) on the candidate multimedia data. The target operation information can, to a certain extent, reflect the target object's interest in the candidate multimedia data. For example, the target operation information indicates that the probability of the target object performing a deep conversion operation on the candidate multimedia data is relatively high, indicating that the target object has a high interest in the candidate multimedia data. The media asset factor is used to reflect the probability of the candidate multimedia data P being used to represent the target object's interest in the candidate multimedia data. iWhen the conversion operation is executed, the actual amount of assets that the object to which the candidate multimedia data belongs needs to spend, that is, the media asset factor can, to a certain extent, reflect the commercial value that the candidate multimedia data brings to the multimedia platform. By selecting the candidate multimedia data to be pushed to the target object from the multimedia data set according to the target operation information and the media asset factor, the target multimedia data is pushed to the terminal corresponding to the target object. In other words, by comprehensively considering the user's interest characteristics in the candidate multimedia data and the commercial value of the candidate multimedia data, multimedia data can be recommended to the user, which can achieve accurate recommendation and improve the accuracy of multimedia data recommendation; it can not only avoid recommending multimedia data that the user is not interested in to the user, resulting in invalid multimedia data recommendation and wasting the resources of the multimedia platform, but also avoid recommending multimedia data with relatively low commercial value to the user, resulting in relatively low resource utilization of the multimedia platform, thereby improving the resource utilization of the multimedia platform.

[0212] See Figure 15 , is a structural diagram of a computer device provided in an embodiment of the present application. As shown in Figure 15, the above-mentioned computer device 1000 may include: a processor 1001, a network interface 1004 and a memory 1005. In addition, the above-mentioned computer device 1000 may also include: a media content interface 1003, and at least one communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. Among them, the media content interface 1003 may include a display screen (Display), a keyboard (Keyboard), and the optional media content interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface, a wireless interface (such as W I -F I Interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The memory 1005 may optionally be at least one storage device away from the aforementioned processor 1001. Figure 15 As shown, the memory 1005 as a computer-readable storage medium may include an operating system, a network communication module, a media content interface module, and a device control application.

[0213] exist Figure 15 In the computer device 1000 shown, the network interface 1004 can provide network communication functions; the media content interface 1003 is mainly used to provide an input interface for media content; and the processor 1001 can be used to call the device control application stored in the memory 1005 to achieve:

[0214] Obtain the target object attribute information of the target object and the candidate multimedia data P in the multimedia data set to be pushed i candidate media attribute information; i is a positive integer less than or equal to M, and M is the number of candidate multimedia data in the multimedia data set;

[0215] According to the target object attribute information and the candidate media attribute information, a prediction is made to reflect the target object's response to the candidate multimedia data P i Target operation information corresponding to the execution operation;

[0216] According to the candidate media attribute information, the candidate multimedia data P is determined i The media asset factor is used to reflect the media asset factor of the candidate multimedia data P i The actual amount of assets that the object to which the candidate multimedia data belongs needs to spend when the conversion operation is executed;

[0217] According to the target operation information and the media asset factor, candidate multimedia data for being pushed to the target object is selected from the multimedia data set as target multimedia data, and the target multimedia data is pushed to a terminal corresponding to the target object.

[0218] Optionally, the processor 1001 may be configured to call a device control application stored in the memory 1005 to implement:

[0219] Extract the target object attribute information from the candidate multimedia data P i Key object attribute information associated with the operation attribute;

[0220] Extract the candidate multimedia data P from the candidate media attribute information i Key media attribute information associated with the operation attributes;

[0221] The operation recognition network using the target operation recognition model predicts the target object's operation for the candidate multimedia data P based on the key object attribute information and the key media attribute information. i Target operation information corresponding to the execution operation.

[0222] Optionally, the key object attribute information includes the key object attribute information related to the candidate multimedia data P i The first key object attribute information associated with the conversion operation attribute, and the candidate multimedia data P i The second key object attribute information associated with the non-conversion operation attribute; the key media attribute information includes the second key object attribute information associated with the candidate multimedia data P iThe first key media attribute information associated with the conversion operation attribute, and the first key media attribute information associated with the candidate multimedia data P i The processor 1001 can be used to call the device control application stored in the memory 1005 to implement the operation recognition network using the target operation recognition model, based on the key object attribute information and the key media attribute information, to predict the target object for the candidate multimedia data P i The target operation information corresponding to the execution operation includes:

[0223] The operation recognition network of the target operation recognition model is used to predict the target object's operation for the candidate multimedia data P based on the first key object attribute information and the first key media attribute information. i Conversion operation information corresponding to executing the conversion operation;

[0224] According to the second key object attribute information and the second key media attribute information, a prediction is made for reflecting the target object's response to the candidate multimedia data P i Non-conversion operation information corresponding to executing non-conversion operations;

[0225] The conversion operation information and the non-conversion operation information are determined as the information used to reflect the target object's response to the candidate multimedia data P i Target operation information corresponding to the execution operation.

[0226] Optionally, the processor 1001 may be configured to call a device control application stored in the memory 1005 to determine the candidate multimedia data P according to the candidate media attribute information. i Media asset factors include:

[0227] Extract the candidate multimedia data P from the candidate media attribute information i Key asset attribute information associated with the asset attributes;

[0228] An asset identification network of a target asset identification model is used to perform cross-correlation identification on the key asset attribute information to obtain cross-correlation information; the target asset identification model and the target operation identification model are independent of each other;

[0229] Performing deep correlation identification on the key asset attribute information to obtain deep correlation information;

[0230] The media asset factor is determined according to the cross-association relationship information and the deep-association relationship information.

[0231] Optionally, the processor 1001 may be configured to call a device control application stored in the memory 1005 to extract the media attribute information related to the candidate multimedia data P from the candidate media attribute information. i The key asset attribute information associated with the asset attributes includes:

[0232] The asset expert network of the target asset recognition model is used to extract the candidate multimedia data P from the candidate media attribute information. i The attribute information associated with the asset attribute is used as the candidate asset attribute information;

[0233] Determining the confidence level of the candidate asset attribute information using an asset gating network of the target asset identification model;

[0234] The confidence level is used to weight the candidate asset attribute information to obtain the candidate multimedia data P i Key asset attribute information associated with the asset attributes.

[0235] Optionally, the processor 1001 may be configured to call a device control application stored in the memory 1005 to select candidate multimedia data for pushing to the target object from the multimedia data set based on the target operation information and the media asset factor, including:

[0236] Determine the candidate multimedia data P according to the target operation information and the media asset factor i The media estimation asset information is used to reflect the media estimation asset information when the candidate multimedia data P i An estimated amount of assets that the object to which the candidate multimedia data belongs needs to spend when the target object performs an operation;

[0237] According to the candidate multimedia data P i The media estimation asset information is used to select candidate multimedia data for pushing to the target object from the multimedia data set as target multimedia data.

[0238] Optionally, the conversion operation information includes information for reflecting the target object's response to the candidate multimedia data P i Execute the first conversion operation corresponding to the first conversion operation, and the first conversion operation information for reflecting the target object for the candidate multimedia data P i Execute the second conversion operation information corresponding to the second conversion operation, the first conversion operation is performed on the candidate multimedia data P i The asset amount brought by the object to which it belongs is greater than the asset amount brought by the second conversion operation to the candidate multimedia data P i The amount of assets brought by the object;

[0239] Optionally, the processor 1001 may be configured to call a device control application stored in the memory 1005 to determine the candidate multimedia data P according to the target operation information and the media asset factor. i Estimated media asset information, including:

[0240] determining first media estimated asset information according to the media asset factor, the first conversion operation information, and the non-conversion operation information;

[0241] determining second media estimated asset information based on the media asset factor, the second conversion operation information, and the non-conversion operation information;

[0242] The first media estimated asset information and the second media estimated asset information are determined as the candidate multimedia data P i Media estimated asset information.

[0243] Optionally, the processor 1001 may be configured to call a device control application stored in the memory 1005 to implement the control of the candidate multimedia data P i The method further comprises: selecting candidate multimedia data for pushing to the target object from the multimedia data set as target multimedia data, including:

[0244] Acquire a multimedia data network; the multimedia data network includes a method for reflecting the candidate multimedia data P in the multimedia data set i , and the edges formed by connecting the nodes corresponding to the candidate multimedia data with associated relationships;

[0245] Traversing nodes in the multimedia data network according to the node path of the multimedia data network and the media estimation asset information to obtain candidate multimedia data having a neighbor relationship with the target object attribute information;

[0246] The candidate multimedia data in the multimedia data set that has a neighbor relationship with the target object attribute information is determined as the target multimedia data for being pushed to the target object.

[0247] Optionally, the multimedia data network includes a subnetwork N j and subnetwork N j+1 , the sub-network N j The number of nodes in the subnetwork is less than N j+1 The number of nodes in the subnetwork N j The set of candidate multimedia data corresponding to the nodes in the subnetwork N j+1wherein j is a positive integer less than K, and K is the number of subnetworks in the multimedia data network; optionally, the processor 1001 may be configured to call a device control application stored in the memory 1005 to traverse the nodes in the multimedia data network according to the node path of the multimedia data network and the media estimation asset information, and obtain candidate multimedia data having a neighbor relationship with the target object attribute information, including:

[0248] According to the node path of the multimedia data network, from the sub-network N j determining, as designated candidate multimedia data, the candidate multimedia data for which the estimated asset amount indicated by the media estimated asset information is the maximum estimated asset amount;

[0249] The sub-network N j+1 The node reflecting the specified candidate multimedia data is determined as the initial traversal node;

[0250] If the sub-network N j+1 The number of nodes in the subnetwork N is less than the number of candidate multimedia data in the multimedia data set, then the initial traversal node is used as the traversal starting point, and the subnetwork N j+1 The candidate multimedia data whose estimated asset amount indicated by the media estimated asset information is the maximum estimated asset amount is determined to be from the sub-network N j+1 The determined candidate multimedia data is updated as designated candidate multimedia data;

[0251] If the sub-network N j+1 If the number of nodes in is equal to the corresponding number of candidate multimedia data in the multimedia data set, then among the nodes connected to the initial traversal node, the nodes whose asset amount indicated by the media estimated asset information is greater than the asset amount threshold are traversed as candidate multimedia data having a neighbor relationship with the object attribute information.

[0252] Optionally, the processor 1001 may be configured to call a device control application stored in the memory 1005 to obtain sample object attribute information of the sample object, sample media attribute information of the sample multimedia data, and annotation operation information of the sample object with respect to the sample multimedia data;

[0253] Using a candidate operation recognition model, predicting the sample object attribute information and the sample media attribute information to obtain predicted operation information reflecting the operation performed by the sample object on the sample multimedia data;

[0254] The candidate operation recognition model is adjusted according to the marked operation information and the predicted operation information, and the adjusted candidate operation recognition model is determined as the target operation recognition model.

[0255] Optionally, the annotated operation information includes first annotated conversion operation information, second annotated conversion operation information, and annotated non-conversion operation information; the predicted operation information includes first predicted conversion operation information, second predicted conversion operation information, and predicted non-conversion operation information; optionally, the processor 1001 can be used to call a device control application stored in the memory 1005 to adjust the candidate operation recognition model based on the annotated operation information and the predicted operation information, including:

[0256] determining a first conversion operation prediction error of the candidate operation recognition model based on the first labeled conversion operation information and the first predicted conversion operation information;

[0257] determining a second conversion operation prediction error of the candidate operation recognition model according to the second labeled conversion operation information and the second predicted conversion operation information;

[0258] determining a non-conversion operation prediction error of the candidate operation recognition model based on the marked non-conversion operation information and the predicted non-conversion operation information;

[0259] The candidate operation recognition model is adjusted according to the first conversion operation prediction error, the second conversion operation prediction error, and the non-conversion operation prediction error to obtain an adjusted candidate operation recognition model.

[0260] Optionally, the processor 1001 may be configured to call a device control application stored in the memory 1005 to adjust the candidate operation recognition model based on the first conversion operation prediction error, the second conversion operation prediction error, and the non-conversion operation prediction error, to obtain an adjusted candidate operation recognition model, including:

[0261] performing a weighted summation of the first conversion operation prediction error, the second conversion operation prediction error, and the non-conversion operation prediction error to obtain an operation prediction error of the candidate operation recognition model;

[0262] If the operation prediction error is not in a converged state, the candidate operation recognition model is adjusted according to the operation prediction error to obtain an adjusted candidate operation recognition model.

[0263] Optionally, the processor 1001 may be configured to call a device control application stored in the memory 1005 to obtain sample media attribute information of the sample multimedia data and annotation operation information of the sample object on the sample multimedia data, including:

[0264] Filtering first sample multimedia data of operations performed by the sample object within a historical time period from the sample multimedia data set;

[0265] determining, according to log data of the sample object regarding the first sample multimedia data, annotation operation information corresponding to the first sample multimedia data;

[0266] Filtering out second sample multimedia data from the sample multimedia data set, for which the sample object has not performed any operation within a historical time period;

[0267] determining, as the marking operation information of the second sample multimedia data, the marking operation information indicating that the sample object has not performed an operation on the second sample multimedia data;

[0268] The first sample multimedia data and the second sample multimedia data are determined as sample multimedia data corresponding to the sample object.

[0269] Optionally, the processor 1001 may be configured to call a device control application stored in the memory 1005 to filter out, from the sample multimedia data set, second sample multimedia data on which the sample object has not performed an operation within a historical time period, including:

[0270] Randomly selecting, from the sample multimedia data set, sample multimedia data that has not been recommended to the sample object within the historical time period as first candidate sample multimedia data;

[0271] Selecting, from the sample multimedia data set, sample multimedia data that is recommended to the sample object within the historical time period and has not been operated by the sample object as second candidate sample multimedia data;

[0272] The first candidate sample multimedia data and the second candidate sample multimedia data are determined to be second sample multimedia data on which no operation is performed by the sample object within a historical time period.

[0273] Optionally, the processor 1001 may be configured to call a device control application stored in the memory 1005 to obtain the media expected asset information of the sample multimedia data; the media expected asset information is used to reflect the media expected asset information of the candidate multimedia data P i The expected amount of assets that the object to which the candidate multimedia data belongs needs to spend when the operation is converted and executed;

[0274] Determining media annotated asset information of the sample multimedia data based on the predicted operation information and the expected media asset information;

[0275] Using a candidate asset recognition model, predicting the predicted operation information and the sample media attribute information to obtain predicted media asset information of the sample multimedia data;

[0276] The candidate asset recognition model is adjusted according to the media annotated asset information and the media predicted asset information, and the adjusted candidate asset recognition model is determined as the target asset recognition model.

[0277] Optionally, the processor 1001 may be configured to call a device control application stored in the memory 1005 to adjust the candidate asset recognition model based on the media annotated asset information and the media predicted asset information, including:

[0278] determining an asset prediction error of the candidate asset identification model based on the media annotated asset information and the media predicted asset information;

[0279] If the asset prediction error is not in a converged state, the candidate asset identification model is adjusted according to the asset prediction error to obtain an adjusted candidate asset identification model.

[0280] In this application, the target operation information is specifically used to reflect the probability of the target object performing an operation (such as a shallow conversion operation, a deep conversion operation, or a non-conversion operation) on the candidate multimedia data. The target operation information can, to a certain extent, reflect the target object's interest in the candidate multimedia data. For example, the target operation information indicates that the probability of the target object performing a deep conversion operation on the candidate multimedia data is relatively high, indicating that the target object has a high interest in the candidate multimedia data. The media asset factor is used to reflect the probability of the candidate multimedia data P being used to represent the target object's interest in the candidate multimedia data. iWhen the conversion operation is executed, the actual amount of assets that the object to which the candidate multimedia data belongs needs to spend, that is, the media asset factor can, to a certain extent, reflect the commercial value that the candidate multimedia data brings to the multimedia platform. By selecting the candidate multimedia data to be pushed to the target object from the multimedia data set according to the target operation information and the media asset factor, the target multimedia data is pushed to the terminal corresponding to the target object. In other words, by comprehensively considering the user's interest characteristics in the candidate multimedia data and the commercial value of the candidate multimedia data, multimedia data can be recommended to the user, which can achieve accurate recommendation and improve the accuracy of multimedia data recommendation; it can not only avoid recommending multimedia data that the user is not interested in to the user, resulting in invalid multimedia data recommendation and wasting the resources of the multimedia platform, but also avoid recommending multimedia data with relatively low commercial value to the user, resulting in relatively low resource utilization of the multimedia platform, thereby improving the resource utilization of the multimedia platform.

[0281] It should be understood that the computer device 1000 described in the embodiment of the present application can execute the above Figure 4 And the previous article Figure 5 The description of the multimedia data processing method in the corresponding embodiment can also be performed by performing the description of the multimedia data processing device in the embodiment corresponding to FIG10 , which will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated here either.

[0282] In addition, it should be noted that: the embodiment of the present application further provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the multimedia data processing device mentioned above, and the computer program includes program instructions. When the processor executes the program instructions, the computer program can execute the multimedia data processing device mentioned above. Figure 4 and Figure 5 The description of the multimedia data processing method described in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated here. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application.

[0283] As an example, the above program instructions may be deployed on a computer device for execution, or deployed on at least two computer devices at one location for execution, or executed on at least two computer devices distributed at at least two locations and interconnected via a communication network. The at least two computer devices distributed at at least two locations and interconnected via a communication network may constitute a blockchain network.

[0284] The computer-readable storage medium may be the multimedia data processing device provided in any of the aforementioned embodiments or the central storage unit of the computer device, such as a hard disk or a memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device. Furthermore, the computer-readable storage medium may also include both the central storage unit of the computer device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the computer device. The computer-readable storage medium may also be used to temporarily store data that has been output or is to be output.

[0285] The terms "first," "second," and the like in the description, claims, and drawings of the embodiments of the present application are used to distinguish between contents in different media, rather than to describe a specific order. In addition, the terms "including" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not limited to the listed steps or modules, but may optionally include steps or modules that are not listed, or may optionally include other steps and units inherent to these processes, methods, apparatuses, products, or devices.

[0286] The present application also provides a computer program product, including a computer program / instruction, which implements the above-mentioned Figure 4 and Figure 5 The description of the multimedia data processing method described in the corresponding embodiment will not be repeated here. In addition, the description of the beneficial effects of using the same method will not be repeated here. For technical details not disclosed in the embodiments of the computer program product involved in this application, please refer to the description of the method embodiment of this application.

[0287] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0288] The methods and related devices provided in the embodiments of the present application are described with reference to the method flow charts and / or structural diagrams provided in the embodiments of the present application. Specifically, each process and / or box in the method flow charts and / or structural diagrams, as well as the combination of processes and / or boxes in the flow charts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable multimedia data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable multimedia data processing device generate a device for implementing the functions specified in one or more processes in the flow chart and / or one or more boxes in the structural diagram. These computer program instructions can also be stored in a computer-readable memory that can guide the computer or other programmable multimedia data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory generate an article of manufacture including an instruction device, which implements the functions specified in one or more processes in the flow chart and / or one or more boxes in the structural diagram. These computer program instructions can also be loaded onto a computer or other programmable multimedia data processing device so that a series of operating steps are executed on the computer or other programmable device to produce computer-implemented processing, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the structural diagram.

[0289] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A multimedia data processing method, characterized in that: include: Obtain the target object attribute information of the target object and the candidate multimedia data P in the multimedia data set to be pushed i candidate media attribute information; i is a positive integer less than or equal to M, and M is the number of candidate multimedia data in the multimedia data set; According to the target object attribute information and the candidate media attribute information, a prediction is made to reflect the target object's response to the candidate multimedia data P i Target operation information corresponding to the execution operation; According to the candidate media attribute information, the candidate multimedia data P is determined i Media asset factor; The media asset factor is used to reflect the i When the conversion operation is executed, the actual asset amount that the object to which the candidate multimedia data belongs needs to pay to the multimedia platform, wherein the multimedia platform is a platform for pushing multimedia data; Determine the candidate multimedia data P according to the target operation information and the media asset factor i Estimated media asset information; The media estimation asset information is used to reflect the candidate multimedia data P i An estimated amount of assets that the object to which the candidate multimedia data belongs needs to pay to the multimedia platform when the operation is performed by the target object; According to the candidate multimedia data P i The media estimated asset information of the target object is used, and the candidate multimedia data whose estimated asset amount indicated by the media estimated asset information in the multimedia data set is greater than the asset amount threshold is used as the target multimedia data, and the target multimedia data is pushed to the terminal corresponding to the target object.

2. The method according to claim 1, wherein The target object attribute information and the candidate media attribute information are used to predict the target object for the candidate multimedia data P i The target operation information corresponding to the execution operation includes: Extract the target object attribute information from the candidate multimedia data P i Key object attribute information associated with the operation attribute; Extract the candidate multimedia data P from the candidate media attribute information i Key media attribute information associated with the operation attributes; The operation recognition network using the target operation recognition model predicts the target object's operation for the candidate multimedia data P based on the key object attribute information and the key media attribute information. i Target operation information corresponding to the execution operation.

3. The method according to claim 2, wherein The key object attribute information includes the key object attribute information related to the candidate multimedia data P i The first key object attribute information associated with the conversion operation attribute, and the candidate multimedia data P i The second key object attribute information associated with the non-conversion operation attribute; the key media attribute information includes the second key object attribute information associated with the candidate multimedia data P i The first key media attribute information associated with the conversion operation attribute, and the first key media attribute information associated with the candidate multimedia data P i Second key media attribute information associated with the non-conversion operation attribute; The operation recognition network using the target operation recognition model predicts the target object's operation on the candidate multimedia data P based on the key object attribute information and the key media attribute information. i The target operation information corresponding to the execution operation includes: The operation recognition network of the target operation recognition model is used to predict the target object's operation for the candidate multimedia data P based on the first key object attribute information and the first key media attribute information. i Conversion operation information corresponding to executing the conversion operation; According to the second key object attribute information and the second key media attribute information, a prediction is made for reflecting the target object's response to the candidate multimedia data P i Execute non-conversion operation information corresponding to the non-conversion operation; The conversion operation information and the non-conversion operation information are determined as the information used to reflect the target object's response to the candidate multimedia data P i Target operation information corresponding to the execution operation.

4. The method according to claim 3, wherein The candidate multimedia data P is determined based on the candidate media attribute information. i Media asset factors include: Extract the candidate multimedia data P from the candidate media attribute information i Key asset attribute information associated with the asset attributes; An asset identification network of a target asset identification model is used to perform cross-correlation identification on the key asset attribute information to obtain cross-correlation information; the target asset identification model and the target operation identification model are independent of each other; Performing deep correlation identification on the key asset attribute information to obtain deep correlation information; The media asset factor is determined according to the cross-association relationship information and the deep-association relationship information.

5. The method according to claim 4, wherein The candidate multimedia data P is extracted from the candidate media attribute information. i The key asset attribute information associated with the asset attributes includes: The asset expert network of the target asset recognition model is used to extract the candidate multimedia data P from the candidate media attribute information. i The attribute information associated with the asset attribute is used as the candidate asset attribute information; Determining the confidence level of the candidate asset attribute information using an asset gating network of the target asset identification model; The confidence level is used to weight the candidate asset attribute information to obtain the candidate multimedia data P i Key asset attribute information associated with the asset attributes.

6. The method according to claim 3, wherein The conversion operation information includes information for reflecting the target object's response to the candidate multimedia data P i Execute the first conversion operation corresponding to the first conversion operation, and the first conversion operation information for reflecting the target object for the candidate multimedia data P i Execute the second conversion operation information corresponding to the second conversion operation, the first conversion operation is performed on the candidate multimedia data P i The asset amount brought by the object to which it belongs is greater than the asset amount brought by the second conversion operation to the candidate multimedia data P i The amount of assets brought by the object; The candidate multimedia data P is determined based on the target operation information and the media asset factor. i Estimated media asset information, including: determining first media estimated asset information according to the media asset factor, the first conversion operation information, and the non-conversion operation information; determining second media estimated asset information based on the media asset factor, the second conversion operation information, and the non-conversion operation information; The first media estimated asset information and the second media estimated asset information are determined as the candidate multimedia data P i Media estimated asset information.

7. The method according to claim 1, wherein The candidate multimedia data P i The estimated media asset information of the multimedia data set is used as target multimedia data, wherein the candidate multimedia data in the multimedia data set has an estimated asset amount indicated by the estimated media asset information greater than an asset amount threshold. Acquire a multimedia data network; the multimedia data network includes a method for reflecting the candidate multimedia data P in the multimedia data set i , and the edges formed by connecting the nodes corresponding to the candidate multimedia data with associated relationships; traversing nodes in the multimedia data network according to the node path of the multimedia data network and the estimated media asset information to obtain candidate multimedia data having a neighbor relationship with the target object attribute information; the candidate multimedia data having a neighbor relationship with the target object attribute information is the candidate multimedia data for which the estimated asset amount indicated by the estimated media asset information in the multimedia data set is greater than an asset amount threshold; The candidate multimedia data in the multimedia data set that has a neighbor relationship with the target object attribute information is determined as the target multimedia data for being pushed to the target object.

8. The method according to claim 7, wherein The multimedia data network includes a subnetwork N j and subnetwork N j+1 , the sub-network N j The number of nodes in the subnetwork is less than N j+1 The number of nodes in the subnetwork N j The set of candidate multimedia data corresponding to the nodes in the subnetwork N j+1 A subset of the set consisting of candidate multimedia data corresponding to the nodes in , j is a positive integer less than K, and K is the number of subnetworks in the multimedia data network; The traversing nodes in the multimedia data network according to the node path of the multimedia data network and the media estimation asset information to obtain candidate multimedia data having a neighbor relationship with the target object attribute information includes: According to the node path of the multimedia data network, from the sub-network N j determining, as designated candidate multimedia data, the candidate multimedia data for which the estimated asset amount indicated by the media estimated asset information is the maximum estimated asset amount; The sub-network N j+1 The node reflecting the specified candidate multimedia data is determined as the initial traversal node; If the sub-network N j+1 The number of nodes in the subnetwork N is less than the number of candidate multimedia data in the multimedia data set, then the initial traversal node is used as the traversal starting point, and the subnetwork N j+1 The candidate multimedia data whose estimated asset amount indicated by the media estimated asset information is the maximum estimated asset amount is determined to be from the sub-network N j+1 The determined candidate multimedia data is updated as designated candidate multimedia data; If the sub-network N j+1 If the number of nodes in is equal to the corresponding number of candidate multimedia data in the multimedia data set, then among the nodes connected to the initial traversal node, the nodes whose asset amount indicated by the media estimated asset information is greater than the asset amount threshold are traversed as candidate multimedia data having a neighbor relationship with the object attribute information.

9. The method according to claim 2, wherein The method further comprises: Acquire sample object attribute information of the sample object, sample media attribute information of the sample multimedia data, and annotation operation information of the sample object with respect to the sample multimedia data; Using a candidate operation recognition model, predicting the sample object attribute information and the sample media attribute information to obtain predicted operation information reflecting the operation performed by the sample object on the sample multimedia data; The candidate operation recognition model is adjusted according to the marked operation information and the predicted operation information, and the adjusted candidate operation recognition model is determined as the target operation recognition model.

10. The method according to claim 9, wherein The marked operation information includes first marked conversion operation information, second marked conversion operation information and marked non-conversion operation information; the predicted operation information includes first predicted conversion operation information, second predicted conversion operation information and predicted non-conversion operation information; The adjusting the candidate operation recognition model according to the marked operation information and the predicted operation information includes: determining a first conversion operation prediction error of the candidate operation recognition model based on the first labeled conversion operation information and the first predicted conversion operation information; determining a second conversion operation prediction error of the candidate operation recognition model according to the second labeled conversion operation information and the second predicted conversion operation information; determining a non-conversion operation prediction error of the candidate operation recognition model based on the marked non-conversion operation information and the predicted non-conversion operation information; The candidate operation recognition model is adjusted according to the first conversion operation prediction error, the second conversion operation prediction error, and the non-conversion operation prediction error to obtain an adjusted candidate operation recognition model.

11. The method according to claim 10, wherein The adjusting the candidate operation recognition model according to the first conversion operation prediction error, the second conversion operation prediction error, and the non-conversion operation prediction error to obtain an adjusted candidate operation recognition model includes: performing a weighted summation of the first conversion operation prediction error, the second conversion operation prediction error, and the non-conversion operation prediction error to obtain an operation prediction error of the candidate operation recognition model; If the operation prediction error is not in a converged state, the candidate operation recognition model is adjusted according to the operation prediction error to obtain an adjusted candidate operation recognition model.

12. The method according to claim 9, wherein The acquiring of sample media attribute information of the sample multimedia data and the annotation operation information of the sample object on the sample multimedia data includes: Filtering first sample multimedia data of operations performed by the sample object within a historical time period from the sample multimedia data set; determining, according to log data of the sample object regarding the first sample multimedia data, annotation operation information corresponding to the first sample multimedia data; Filtering out second sample multimedia data from the sample multimedia data set, for which the sample object has not performed any operation within a historical time period; determining, as the marking operation information of the second sample multimedia data, the marking operation information indicating that the sample object has not performed an operation on the second sample multimedia data; The first sample multimedia data and the second sample multimedia data are determined as sample multimedia data corresponding to the sample object.

13. The method according to claim 12, wherein: The step of filtering out, from the sample multimedia data set, second sample multimedia data on which the sample object has not performed any operation within a historical time period includes: Randomly selecting, from the sample multimedia data set, sample multimedia data that has not been recommended to the sample object within the historical time period as first candidate sample multimedia data; Selecting, from the sample multimedia data set, sample multimedia data that is recommended to the sample object within the historical time period and has not been operated by the sample object as second candidate sample multimedia data; The first candidate sample multimedia data and the second candidate sample multimedia data are determined to be second sample multimedia data on which no operation is performed by the sample object within a historical time period.

14. The method according to claim 9, wherein The method further comprises: Acquire the media expected asset information of the sample multimedia data; the media expected asset information is used to reflect the expected asset information of the candidate multimedia data P i The expected amount of assets that the object to which the candidate multimedia data belongs needs to spend when the operation is converted and executed; Determining media annotated asset information of the sample multimedia data based on the predicted operation information and the expected media asset information; Using a candidate asset recognition model, predicting the predicted operation information and the sample media attribute information to obtain predicted media asset information of the sample multimedia data; The candidate asset recognition model is adjusted according to the media annotated asset information and the media predicted asset information, and the adjusted candidate asset recognition model is determined as the target asset recognition model.

15. The method according to claim 14, wherein The adjusting the candidate asset identification model according to the media annotated asset information and the media predicted asset information includes: determining an asset prediction error of the candidate asset identification model based on the media annotated asset information and the media predicted asset information; If the asset prediction error is not in a converged state, the candidate asset identification model is adjusted according to the asset prediction error to obtain an adjusted candidate asset identification model.

16. A multimedia data processing device, characterized in that: include: The acquisition module is used to obtain the target object attribute information of the target object and the candidate multimedia data P in the multimedia data set to be pushed. i candidate media attribute information; i is a positive integer less than or equal to M, and M is the number of candidate multimedia data in the multimedia data set; The prediction module is configured to predict, based on the target object attribute information and the candidate media attribute information, a value that reflects the target object's response to the candidate multimedia data P i Target operation information corresponding to the execution operation; A determination module is used to determine the candidate multimedia data P according to the candidate media attribute information. i Media asset factor; The media asset factor is used to reflect the i When the conversion operation is executed, the actual asset amount that the object to which the candidate multimedia data belongs needs to pay to the multimedia platform, wherein the multimedia platform is a platform for pushing multimedia data; A selection module is used to determine the candidate multimedia data P based on the target operation information and the media asset factor. i Estimated media asset information; The media estimation asset information is used to reflect the candidate multimedia data P i When the target object performs an operation, the estimated asset amount that the object to which the candidate multimedia data belongs needs to pay to the multimedia platform; i The media estimated asset information of the target object is used, and the candidate multimedia data whose estimated asset amount indicated by the media estimated asset information in the multimedia data set is greater than the asset amount threshold is used as the target multimedia data, and the target multimedia data is pushed to the terminal corresponding to the target object.

17. A computer device, characterized in that: include: processor and memory; The above-mentioned processor is connected to a memory; the memory is used to store program code, and the processor is used to call the program code to execute the method according to any one of claims 1-15.

18. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 1 to 15.

19. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 15 are implemented.

Citation Information

Patent Citations

  • Information recommendation method, device and equipment and computer readable storage medium

    CN111680221A