Training Method, Device, Equipment and Medium of Feedback Information Estimation Model

By using multimedia samples of different data types in the feedback information prediction model training, the problem of low accuracy in the prior art is solved, and more accurate multimedia content recommendation is achieved.

CN114492750BActive Publication Date: 2025-07-04TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210080869.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-24
Publication Date
2025-07-04
Estimated Expiration
2042-01-24

AI Technical Summary

Technical Problem

In the prior art, the feedback information prediction model uses only one data type multimedia sample, resulting in low training accuracy, which in turn affects the accuracy of multimedia content recommendation.

Method used

Multimedia samples containing different data types are trained, and the accuracy of the model is improved through feature extraction and parameter adjustment, and the content feature correlation between multimedia samples is used to alleviate the problems of sparsity and delay.

Benefits of technology

The accuracy of the feedback information prediction model is improved, the accuracy of multimedia content recommendation is enhanced, and the content of interest to the target object can be recommended more accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114492750B_ABST
    Figure CN114492750B_ABST
Patent Text Reader

Abstract

This application relates to the field of computer technologies, and in particular, to a method, apparatus, device, and medium for training a feedback information prediction model, which can be applied to various scenarios such as cloud technologies, artificial intelligence, intelligent transportation, and vehicle networking to quickly and flexibly detect objects. The method includes: obtaining a training data sample set, where each training data sample includes: object features corresponding to the sample object, first and second multimedia samples with different data types and their respective content features, and a first actual feedback information generated by the sample object for the first multimedia sample and a second actual feedback information generated for the second multimedia sample; based on the training data sample set, iteratively training the feedback information prediction model to be trained until the trained target feedback information prediction model is output. Since the second multimedia sample with non-sparse and low-latency feedback information is added in this application, the accuracy of the feedback information prediction model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the rapid development of the Internet, a multimedia recommendation system can display multimedia content of different types or different contents on a multimedia content browsing page to achieve personalized display of multimedia content.

[0003] Specifically, a multimedia recommendation system generally sorts according to the estimated results of the feedback information of each multimedia content for a target object; after sorting the estimated results of the feedback information of each multimedia content, the multimedia recommendation system will recommend interesting multimedia content for the target object according to the obtained sorting results.

[0004] In the related art, a multimedia sample of one data type is generally used to train a feedback information estimation model, and based on the trained feedback information estimation model, the estimated results of the feedback information of each multimedia for a target object are obtained. Since the feedback information estimation model in the related art only uses a multimedia sample corresponding to one data type during training, and the feedback information generated by the sample object for the multimedia sample corresponding to this data type has high sparsity and high latency, the accuracy of the trained feedback information estimation model is not high, which in turn leads to low accuracy of the obtained estimated results of the feedback information, thus affecting the accuracy of the result of recommending interesting multimedia content for the target object. Summary of the Invention

[0005] Embodiments of the present application provide a training method, device, equipment and medium for a feedback information estimation model, so as to improve the accuracy of the feedback information estimation model and the accuracy of recommending interesting multimedia content for a target object.

[0006] A training method for a feedback information estimation model provided by an embodiment of the present application includes:

[0007] Obtain a training data sample set, where each training data sample includes: object features corresponding to a sample object, a first multimedia sample and a second multimedia sample with different data types and their respective content features, and a first actual feedback information generated by the sample object for the first multimedia sample and a second actual feedback information generated for the second multimedia sample;

[0008] Based on the training data sample set, perform iterative training on the feedback information estimation model to be trained until the trained target feedback information estimation model is output. During one iteration process, perform the following operations:

[0009] Input the extracted training data sample into the feedback information estimation model to be trained, and output a first estimated feedback information of the corresponding first multimedia sample and a second estimated feedback information of the corresponding second multimedia sample;

[0010] Adjust the parameters in the feedback information prediction model to be trained according to the difference between the first predicted feedback information and the corresponding first actual feedback information, and the difference between the second predicted feedback information and the corresponding second actual feedback information.

[0011] An apparatus for training a feedback information prediction model provided by an embodiment of the present application includes:

[0012] An acquisition unit, configured to acquire a training data sample set, where each training data sample includes: object features corresponding to a sample object, a first multimedia sample and a second multimedia sample with different data types and their respective content features, and a first actual feedback information generated by the sample object for the first multimedia sample and a second actual feedback information generated for the second multimedia sample;

[0013] A training unit, configured to perform iterative training on the feedback information prediction model to be trained based on the training data sample set until a trained target feedback information prediction model is output. In one iteration process, the following operations are performed: input the extracted training data sample into the feedback information prediction model to be trained, and output the first predicted feedback information of the corresponding first multimedia sample and the second predicted feedback information of the corresponding second multimedia sample; adjust the parameters in the feedback information prediction model to be trained according to the difference between the first predicted feedback information and the corresponding first actual feedback information, and the difference between the second predicted feedback information and the corresponding second actual feedback information.

[0014] Optionally, the training unit is specifically configured to:

[0015] Based on the feedback information prediction model to be trained, perform feature extraction on the extracted training data sample to obtain a first feature vector corresponding to the content features of the first multimedia sample, a second feature vector corresponding to the content features of the second multimedia sample, and a corresponding comprehensive feature vector; where the comprehensive feature vector is obtained by performing feature fusion according to the content features of the first multimedia sample, the content features of the second multimedia sample, and the object features of the sample object;

[0016] Perform feature analysis on the first feature vector, the second feature vector, and the comprehensive feature vector respectively to obtain a first predicted feedback information matrix corresponding to the first multimedia sample and a second predicted feedback information matrix corresponding to the second multimedia sample;

[0017] Based on the first estimated feedback information matrix, the second estimated feedback information matrix, and a preset normalization function, respectively determine the first estimated feedback information corresponding to the first multimedia sample and the second estimated feedback information corresponding to the second multimedia sample.

[0018] Optionally, the training unit is specifically configured to:

[0019] Concatenate the first feature vector, the second feature vector, and the comprehensive feature vector to obtain a comprehensive concatenated feature matrix;

[0020] Based on the comprehensive concatenated feature matrix, respectively perform at least one round of joint feature extraction operations using each preset joint feature extraction rule to obtain corresponding joint feature matrices;

[0021] Perform feature weighting on each obtained joint feature matrix to obtain the first weighted feature matrix corresponding to the first multimedia sample and the second weighted feature matrix corresponding to the second multimedia sample;

[0022] Concatenate the multiple joint feature matrices with the first weighted feature matrix and the second weighted feature matrix respectively to obtain the first concatenated feature matrix corresponding to the first multimedia sample and the second concatenated feature matrix corresponding to the second multimedia sample;

[0023] Based on the first concatenated feature matrix and the second concatenated feature matrix, respectively perform at least one round of comprehensive feature extraction operations to obtain the first estimated feedback information matrix corresponding to the first multimedia sample and the second estimated feedback information matrix corresponding to the second multimedia sample.

[0024] Optionally, the training unit is specifically configured to:

[0025] For each round of comprehensive feature extraction operation, respectively execute the following process:

[0026] Based on the result matrix corresponding to the first multimedia sample in the previous round after the comprehensive feature extraction operation and the first network parameter matrix saved in advance, determine the result matrix corresponding to the first multimedia sample in the current round;

[0027] Wherein, the result matrix corresponding to the first multimedia sample in the first round is the first concatenated feature matrix; the result matrix corresponding to the first multimedia sample in the previous round is positively correlated with the result matrix corresponding to the first multimedia sample in the current round.

[0028] Optionally, the training unit is specifically configured to:

[0029] For each round of comprehensive feature extraction operation, respectively execute the following process:

[0030] Determine the result matrix corresponding to the current round of the second multimedia sample based on the result matrix corresponding to the previous round of the second multimedia sample output after performing the comprehensive feature extraction operation and the pre - saved second network parameter matrix;

[0031] Among them, the result matrix corresponding to the first round of the second multimedia sample is the second splicing feature matrix, and the result matrix corresponding to the previous round of the second multimedia sample is positively correlated with the result matrix corresponding to the current round of the second multimedia sample.

[0032] Optionally, the training unit is specifically configured to:

[0033] Perform a linear transformation on the splicing result corresponding to the spliced obtained joint feature matrices to determine the first multi - dimensional matrix corresponding to the first multimedia sample and the second multi - dimensional matrix corresponding to the second multimedia sample;

[0034] Based on a preset normalization function, perform normalization processing on the first multi - dimensional matrix and the second multi - dimensional matrix respectively to obtain the first weight corresponding to the first multimedia sample for each joint feature matrix, and the second weight corresponding to the second multimedia sample for each joint feature matrix respectively;

[0035] Determine the first weighted feature matrix corresponding to the first multimedia sample according to the respective joint feature matrices and the corresponding first weights, and determine the second weighted feature matrix corresponding to the second multimedia sample according to the respective joint feature matrices and the corresponding second weights.

[0036] Optionally, the training unit is specifically configured to:

[0037] Determine the loss value corresponding to the corresponding first multimedia sample based on the difference between the first estimated feedback information and the corresponding first actual feedback information;

[0038] Determine the loss value corresponding to the second multimedia sample based on the difference between the second estimated feedback information and the corresponding second actual feedback information;

[0039] Determine the sum of weights based on the loss value corresponding to the first multimedia sample and the corresponding first weights, the loss value corresponding to the corresponding second multimedia sample and the corresponding second weights, and determine the sum of weights as the target loss value;

[0040] Adjust the parameters in the feedback information prediction model to be trained based on the target loss value.

[0041] An electronic device provided by an embodiment of the present application includes a processor and a memory. Among them, the memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of the above-mentioned method for a feedback information prediction model.

[0042] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of any of the above-mentioned methods for a feedback information prediction model.

[0043] An embodiment of the present application provides a computer-readable storage medium, which includes program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps of the above-mentioned method for a feedback information prediction model.

[0044] The beneficial effects of the present application are as follows:

[0045] In the training method, device, equipment and medium of the feedback information prediction model provided by the embodiment of the present application, since in the process of training the feedback information prediction model in the embodiment of the present application, each training data sample includes a first multimedia sample and a second multimedia sample of different data types. Further, based on the content features corresponding to the first multimedia sample and the second multimedia sample respectively, the feedback information prediction model is trained. The second multimedia sample with non-sparse and low-latency feedback information is added, which alleviates the problem that the accuracy of the feedback information prediction model is not high caused by training the feedback information prediction model only based on the feedback information of the first multimedia sample with high sparsity and high latency. In addition, since the content features of the second multimedia sample are correlated with the content features of the first multimedia sample, it is more conducive for the feedback information prediction model to predict the first predicted feedback information corresponding to the first multimedia sample. Finally, after the feedback information prediction model is trained, based on the trained feedback information prediction model, the result of more accurately recommending multimedia content of interest to the target object can be obtained. By training the feedback information prediction model in the above manner, the problem that the accuracy of the feedback information prediction model is not high when only adding multimedia samples of one data type is solved, and the accuracy of the result of recommending multimedia content of interest to the target object is improved.

[0046] Other features and advantages of the present application will be described in the following specification, and part of them will become obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be realized and obtained by the structures specifically pointed out in the written specification, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The drawings described herein are provided to further understand the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0048] Figure 1 is a schematic diagram of an application scenario of the feedback information prediction model method in an embodiment of the present application;

[0049] Figure 2 is a flowchart of an implementation of a method for training a feedback information prediction model provided by an embodiment of the present application;

[0050] Figure 3 is a schematic diagram of the process for determining the first predicted feedback information and the second predicted feedback information in an embodiment of the present application;

[0051] Figure 4 is a schematic diagram of the process for converting a high-dimensional feature vector into a low-dimensional feature vector based on an embedding layer provided by an embodiment of the present application;

[0052] Figure 5 is a schematic diagram of the process for determining the first predicted feedback information corresponding to the first multimedia sample and the second predicted feedback information corresponding to the second multimedia sample provided by an embodiment of the present application;

[0053] Figure 6a is a specific flowchart of joint feature extraction based on an MLP layer provided by an embodiment of the present application;

[0054] Figure 6b is a specific flowchart of joint feature extraction based on an MLP layer and an expert network provided by an embodiment of the present application;

[0055] Figure 7 is a specific flowchart of comprehensive feature extraction based on an MLP layer provided by an embodiment of the present application;

[0056] Figure 8 is a schematic diagram of the process for determining the first weighted feature matrix and the second weighted feature matrix provided by an embodiment of the present application;

[0057] Figure 9 is a schematic diagram of the structure for determining the weight corresponding to the joint feature matrix based on a gating network provided by an embodiment of the present application;

[0058] Figure 10 is a schematic diagram of the structure of a feedback information prediction model provided by an embodiment of the present application;

[0059] Figure 11A schematic flowchart of multimedia content recommendation provided by an embodiment of the present application;

[0060] Figure 12 A schematic structural diagram of a feedback information prediction model training device provided by an embodiment of the present application;

[0061] Figure 13 A schematic composition diagram of an electronic device in an embodiment of the present application;

[0062] Figure 14 A schematic composition diagram of another electronic device applying an embodiment of the present application. Detailed implementation manners

[0063] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the technical solutions of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments recorded in this application document without creative efforts shall fall within the scope of protection of the technical solutions of the present application.

[0064] To facilitate better understanding of the technical solutions of the present application by those skilled in the art, the terms related to the present application are introduced below.

[0065] First multimedia sample: A human-computer interactive information communication and dissemination media sample combining two or more media, such as advertisements, information, short videos, news, mini-programs, etc.

[0066] Second multimedia sample: A human-computer interactive information communication and dissemination media sample combining two or more media, such as advertisements, information, short videos, news, mini-programs, etc., and the data types corresponding to the first multimedia sample and the second multimedia sample are different. For example, if the first multimedia sample is an advertisement sample, the second multimedia sample can be a sample of other data types except the advertisement sample, such as an information sample, a news sample, etc.

[0067] Actual feedback information: Information representing the actual feedback behavior of a sample object for a multimedia sample. Among them, the feedback behavior includes clicking, browsing, liking, sharing, collecting, deep conversion, shallow conversion, saving, etc. Correspondingly, the actual feedback information includes but is not limited to: actual click-through rate, actual browsing rate, actual like rate, actual sharing rate, actual collection rate, actual deep conversion rate, actual shallow conversion rate, actual saving rate.

[0068] Estimated feedback information: Information about the feedback behavior estimated by the sample object for the multimedia sample, where the feedback behavior includes clicking, browsing, liking, sharing, collecting, deep conversion, shallow conversion, saving, etc. Correspondingly, the estimated feedback information includes, but is not limited to: estimated click-through rate, estimated browsing rate, estimated liking rate, estimated sharing rate, estimated collection rate, estimated deep conversion rate, estimated shallow conversion rate, estimated saving rate.

[0069] Click-through rate: It refers to the ratio from multimedia exposure to multimedia click, where multimedia click represents that the object clicks on the item or link contained in the multimedia content.

[0070] Deep conversion rate: It refers to the ratio from multimedia click to multimedia deep conversion, where multimedia deep conversion represents that after the object clicks on the item or link contained in the multimedia content, it jumps to the corresponding page and performs corresponding operations, and the operations include: paying, downloading the application program and achieving the next-day retention of the application program, etc.

[0071] Shallow conversion rate: It refers to the ratio from multimedia click to multimedia shallow conversion, where multimedia shallow conversion represents that after the object clicks on the item or link contained in the multimedia content, it jumps to the corresponding page and performs corresponding operations, and the operations include: application program download, application program activation, and form registration, etc.

[0072] Feedback information estimation model: A neural network model used to input the object features of the target object and the content features of the first multimedia content, and output the estimated feedback information of the target object for the first multimedia content, so as to facilitate subsequent ranking of each first multimedia sample according to the estimated feedback information of the target object for each first multimedia content, and recommend the first multimedia content that the target object is interested in.

[0073] Expert network: A fully connected neural network using the relu activation function, which learns from different angles. In the embodiments of the present application, based on the preset joint feature extraction rules corresponding to different expert networks, joint feature extraction operations can be performed on the feature matrix to obtain the corresponding joint feature matrix.

[0074] Gating network: It is used to control how much of the input features need to be retained and how much need to be discarded to achieve feature isolation. In the embodiments of the present application, based on the gating network, a corresponding weight can be assigned to the joint feature vector output by each expert network, and the weighted feature matrix determined according to the joint feature vector output by each expert network and the corresponding weight can be output.

[0075] The following briefly introduces the design concept of the embodiments of the present application:

[0076] With the rapid development of the Internet, a multimedia recommendation system can display multimedia content of different types or different contents on a multimedia content browsing page to achieve personalized display of multimedia content. Specifically, a general multimedia recommendation system usually sorts according to the estimated results of the feedback information of the target object for each multimedia content. After sorting the estimated results of the feedback information of each multimedia content, the multimedia recommendation system will recommend interesting multimedia content for the target object according to the obtained sorting results.

[0077] If only one data type of multimedia sample is used to train the feedback information estimation model, and based on the trained feedback information estimation model, when obtaining the estimated results of the feedback information of the target object for each multimedia, due to the particularly small amount of feedback information generated by the sample object for the multimedia sample corresponding to this data type, and the feedback information needs to be sent back by the object owner after a period of time, resulting in low accuracy of the trained feedback information estimation model, and further resulting in low accuracy of the obtained estimated results of the feedback information, which will affect the accuracy of the result of recommending interesting multimedia content for the target object.

[0078] In view of this, the training method, device, equipment and medium of the feedback information estimation model provided by the embodiments of the present application obtain a training data sample set, where each training data sample includes: the object features corresponding to a sample object, a first multimedia sample and a second multimedia sample with different data types and their respective content features, and a first actual feedback information generated by the sample object for the first multimedia sample and a second actual feedback information generated for the second multimedia sample. Based on the training data sample set, the feedback information estimation model to be trained is iteratively trained until the trained target feedback information estimation model is output. Among them, in one iteration process, the following operations are performed:

[0079] Input the extracted training data sample into the feedback information estimation model to be trained, and output the first estimated feedback information of the corresponding first multimedia sample and the second estimated feedback information of the corresponding second multimedia sample. Adjust the parameters in the feedback information estimation model to be trained according to the difference between the first estimated feedback information and the corresponding first actual feedback information, and the difference between the second estimated feedback information and the corresponding second actual feedback information.

[0080] In the process of training the feedback information prediction model in the embodiments of the present application, each training data sample includes a first multimedia sample and a second multimedia sample of different data types. Further, based on the content features corresponding to the first multimedia sample and the second multimedia sample respectively, the feedback information prediction model is trained. By adding the second multimedia sample with non-sparse and low-latency feedback information, the problem of low accuracy of the feedback information prediction model caused by training the feedback information prediction model only based on the feedback information of the first multimedia sample with high sparsity and high latency is alleviated.

[0081] In addition, since the content features of the second multimedia sample may be correlated with the content features of the first multimedia sample, it is more conducive for the feedback information prediction model to predict the first predicted feedback information corresponding to the first multimedia sample.

[0082] Finally, after training the feedback information prediction model, based on the trained feedback information prediction model, the result of recommending interesting multimedia content to the target object can be more accurate. By training the feedback information prediction model in the above manner, the problem of low accuracy of the feedback information prediction model caused by adding only one type of multimedia sample is solved, and the accuracy of the result of recommending interesting multimedia content to the target object is improved.

[0083] The following describes the preferred embodiments of the present application with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.

[0084] As Figure 1 shown, it is a schematic diagram of the application scenario of the embodiments of the present application. The application scenario diagram includes two terminal devices 110 and a server 120.

[0085] In the embodiments of the present application, the terminal device 110 includes, but is not limited to, devices such as mobile phones, tablet computers, laptop computers, desktop computers, e - book readers, intelligent voice interaction devices, intelligent household appliances, vehicle terminals, etc.; a client for multimedia content recommendation can be installed on the terminal device, and the client for multimedia content recommendation can be software (such as a browser, multimedia content recommendation software, etc.), or a web page, a small program, etc. The server 120 is the background server corresponding to the software, web page, small program, etc., or a server dedicated to multimedia content recommendation, and the present application does not make specific limitations. The server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0086] It should be noted that the feedback information prediction model training method in the embodiments of the present application can be executed by an electronic device, and the electronic device can be the server 120 or the terminal device 110, that is, the method can be executed independently by the server 120 or the terminal device 110, or jointly executed by the server 120 and the terminal device 110. For example, when executed independently by the server 120, a large number of training data samples can be stored in the server 120. Each training data sample contains object features corresponding to a sample object, a first multimedia sample and a second multimedia sample with different data types and their respective content features, and a first actual feedback information generated by the sample object for the first multimedia sample and a second actual feedback information generated for the second multimedia sample, which are used to train the feedback information prediction model. In the embodiments of the present application, after the feedback information prediction model is trained based on the training method in the embodiments of the present application, the trained feedback information prediction model can be directly deployed on the server 120 or the terminal device 110. Generally, the feedback information prediction model is directly deployed on the server 120. In the embodiments of the present application, the feedback information prediction model is often used to sort multimedia content, and then recommend interesting multimedia content to users.

[0087] In an alternative embodiment, the terminal device 110 and the server 120 can communicate through a communication network.

[0088] In an alternative embodiment, the communication network is a wired network or a wireless network.

[0089] It should be noted that Figure 1The above is only an example, and actually the number of terminal devices and servers is not limited, and no specific limitation is made in the embodiments of the present application.

[0090] In the embodiments of the present application, when the number of servers is multiple, the multiple servers can form a blockchain, and the server is a node on the blockchain; for example, in the feedback information prediction model training method disclosed in the embodiments of the present application, the training data samples involved can be stored on the blockchain.

[0091] In addition, the embodiments of the present application can be applied to various scenarios, including not only the multimedia content recommendation scenario, but also other scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving.

[0092] Next, in combination with the above-described application scenarios, the feedback information prediction model training method provided by the exemplary embodiments of the present application will be described with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard.

[0093] Refer to Figure 2 As shown, it is a flowchart of the implementation of a method for training a feedback information prediction model provided by the embodiments of the present application. Here, mainly taking the server as the execution subject as an example for illustration, the specific implementation process of the method is as follows:

[0094] S21: The server obtains a training data sample set, where each training data sample includes: object features corresponding to a sample object, a first multimedia sample and a second multimedia sample with different data types and their respective content features, and a first actual feedback information generated by the sample object for the first multimedia sample and a second actual feedback information generated for the second multimedia sample;

[0095] Among them, the object features corresponding to the sample object include but are not limited to any one of the following: age, gender, occupation, hobbies, and the multimedia sample (the first multimedia sample or the second multimedia sample) includes but is not limited to any one of the following: advertisement sample, information sample, video sample, applet sample, news sample.

[0096] It can be understood that in the specific implementation of the present application, when it comes to data related to object features corresponding to the sample object, etc., when the above embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions.

[0097] In the embodiments of the present application, some or all of the information in the first actual feedback information corresponding to the first multimedia sample has the problems of high sparsity and high latency. For example, the first multimedia sample can be an advertisement sample, and the first actual feedback information can include the deep conversion rate and the shallow conversion rate of the advertisement corresponding to the advertisement sample.

[0098] In the embodiments of the present application, in order to alleviate the problems of high sparsity and high latency of the first actual feedback information of the first multimedia sample, and in order to improve the accuracy of the feedback information prediction model, in each training data sample in the training data sample set, in addition to including the first multimedia sample, the content features of the first multimedia sample, and the first actual feedback information of the first multimedia sample, it also includes the second actual feedback information of the second multimedia sample that does not have high sparsity and high latency, the second multimedia sample, and the content features of the second multimedia sample. Wherein, the data types corresponding to the first multimedia sample and the second multimedia are different.

[0099] Optionally, taking the first multimedia sample as an advertisement sample and the second multimedia sample as an information sample as an example, the first actual feedback information of the first multimedia sample can include the actual click-through rate of clicking on the advertisement sample, the actual shallow conversion rate of the advertisement sample for shallow conversion, and the actual deep conversion rate of the advertisement sample for deep conversion. The second actual feedback information corresponding to the second multimedia sample can be the actual view rate of the information sample, the actual like rate of the information sample, the actual share rate of the information sample, and so on. Among them, the actual shallow conversion rate and the actual deep conversion rate in the first actual feedback information have the problems of high sparsity and high latency, while the actual click-through rate in the first actual feedback information, the actual view rate, the actual like rate, and the actual share rate in the second actual feedback information do not have the problems of high sparsity and high latency.

[0100] In the embodiments of the present application, for the convenience of description, the present application will be described below by taking the first multimedia sample as an advertisement sample and the second multimedia sample as an information sample as an example.

[0101] In the embodiments of the present application, the content features corresponding to the multimedia sample (the first multimedia sample and the second multimedia sample) include, but are not limited to, any one of the following: the identification information corresponding to the multimedia sample, the category features corresponding to the multimedia sample, the semantic features corresponding to the text included in the multimedia sample, and the image features corresponding to the images included in the multimedia sample.

[0102] S22: The server iteratively trains the feedback information prediction model to be trained based on the training data sample set until the trained target feedback information prediction model is output. Wherein, in one iteration process, the following operations are performed:

[0103] S221: Input the extracted training data samples into the feedback information prediction model to be trained, and output the first predicted feedback information of the corresponding first multimedia sample and the second predicted feedback information of the corresponding second multimedia sample;

[0104] S222: Adjust the parameters in the feedback information prediction model to be trained according to the differences between the first predicted feedback information and the corresponding first actual feedback information, and between the second predicted feedback information and the corresponding second actual feedback information.

[0105] For example, training data sample A is extracted. Training data sample A includes: object features corresponding to sample object B, advertisement sample C and information sample D and their respective content features, and the actual click-through rate a1 of sample object B for advertisement sample C, the actual deep conversion rate b1 for advertisement sample C, and the actual shallow conversion rate c1 for advertisement sample C, the actual browsing rate d1 of sample object B for information sample D, the actual like rate e1 for information sample D, and the actual sharing rate f1 for information sample D. Input this training data sample A into the feedback information prediction model to be trained, and output the predicted click-through rate a2 for advertisement sample C, the predicted deep conversion rate b2 for advertisement sample C, and the predicted shallow conversion rate c2 for advertisement sample C, and the predicted browsing rate d2 for information sample D, the predicted like rate e2 for information sample D, and the predicted sharing rate f2 for information sample D.

[0106] Adjust the parameters in the feedback information prediction model to be trained according to the differences between the predicted click-through rate a1 and the actual click-through rate a2, between the predicted deep conversion rate b1 and the actual deep conversion rate b2, between the predicted shallow conversion rate c1 and the actual shallow conversion rate c2, between the predicted browsing rate d1 and the actual browsing rate d2, between the predicted like rate e1 and the actual like rate e2, and between the predicted sharing rate f1 and the actual sharing rate f2.

[0107] In an optional implementation manner, S221 can be implemented according to the Figure 3 flowchart shown, including the following steps:

[0108] S301: Based on the feedback information prediction model to be trained, perform feature extraction on the extracted training data samples to obtain the first feature vector corresponding to the content features of the first multimedia sample, the second feature vector corresponding to the content features of the second multimedia sample, and the corresponding comprehensive feature vector; wherein, the comprehensive feature vector is obtained by performing feature fusion according to the content features of the first multimedia sample, the content features of the second multimedia sample, and the object features of the sample object;

[0109] In the embodiment of the present application, in order to enable the content features of the first multimedia sample, the content features of the second multimedia sample, and the object features of the object sample in the extracted training data sample to be recognized by the feedback information prediction model to be trained, the extracted training data sample is first subjected to feature extraction based on the feedback information prediction model to be trained, and a first feature vector corresponding to the content features of the first multimedia sample, a second feature vector corresponding to the second multimedia sample, and a corresponding comprehensive feature vector are obtained. Among them, the corresponding comprehensive feature vector is obtained by fusing features according to the content features of the first multimedia sample, the content features of the second multimedia sample, and the object features corresponding to the sample object. Among them, the first feature vector, the second feature vector, and the corresponding comprehensive feature vector have the same dimension.

[0110] For example, if the first feature vector corresponding to the content features of the first multimedia sample is (a1, b1), and the second feature vector corresponding to the content features of the second multimedia sample is (a2, b2), for the convenience of description, the vector corresponding to the object features of the sample object is called the third feature vector, and the third feature vector is (a3, b3), then the corresponding comprehensive feature vector can be (a1 + a2 + a3, b1 + b2 + b3).

[0111] Since there may be high-dimensional feature vectors in the obtained first feature vector, second feature vector, and corresponding comprehensive feature vector, especially for the content features such as the identification information corresponding to the first multimedia sample and the identification information corresponding to the second multimedia sample, due to the large number of the first multimedia sample and the second multimedia sample, the dimensions of the first feature vector corresponding to the identification information of the first multimedia sample and the second feature vector corresponding to the identification information of the second multimedia sample are very high, and the high-dimensional feature vectors are not convenient for subsequent processing. Therefore, the high-dimensional feature vectors in the first feature vector, the second feature vector, and the corresponding comprehensive feature vector can be converted into low-dimensional feature vectors.

[0112] In a possible implementation, the embedding layer in the prediction model can be estimated based on the feedback information, and the first feature vector, the second feature vector, and the corresponding comprehensive feature vector can be converted from high-dimensional feature vectors to low-dimensional feature vectors. Specifically, taking the first feature vector as an example, the first feature vector is input into a hash function to obtain the result output by the hash function. This result is used as the key, and the low-dimensional feature vector output by the embedding layer is used as the value, which is stored in a query table for facilitating the subsequent direct conversion of high-dimensional feature vectors to low-dimensional feature vectors. Before training the feedback information prediction model to be trained, all the stored feature vectors in the query table are initialized first. Subsequently, during the model training process, the feature vectors in the query table are updated based on the reverse gradient. Among them, the more times the model is trained, the higher the accuracy of converting high-dimensional feature vectors to low-dimensional feature vectors.

[0113] Refer to Figure 4 As shown, it is a schematic flowchart of a process for converting high-dimensional feature vectors to low-dimensional feature vectors based on an embedding layer provided by an embodiment of the present application.

[0114] The first feature vector corresponding to the first multimedia sample is respectively input into the embedding layer of the first multimedia sample, the second feature vector corresponding to the second multimedia sample is input into the embedding layer of the second multimedia sample, and the first feature vector corresponding to the first multimedia sample, the second feature vector corresponding to the second multimedia sample, and the corresponding comprehensive feature vector are input into the shared embedding layer, and the first target feature vector after dimensionality reduction, the second target feature vector after dimensionality reduction, and the target comprehensive feature vector after dimensionality reduction are respectively output. Then, the first target feature vector is used to replace the initial first feature vector to obtain a new first feature vector, the second target feature vector is used to replace the initial second feature vector to obtain a new second feature vector, and the target comprehensive feature vector is used to replace the initial comprehensive feature vector to obtain a new comprehensive feature vector.

[0115] In the present application, since the first feature vector corresponding to the first multimedia sample, the second feature vector corresponding to the second multimedia sample, and the corresponding comprehensive feature vector are input into the shared embedding layer, this sharing mechanism can alleviate the high sparsity problem of the first actual feedback information of the first multimedia sample.

[0116] S302: Perform feature analysis on the first feature vector, the second feature vector, and the comprehensive feature vector respectively to obtain the first predicted feedback information matrix corresponding to the first multimedia sample and the second predicted feedback information matrix corresponding to the second multimedia sample;

[0117] Wherein, the first estimated feedback information matrix is the matrix expression form corresponding to the first estimated feedback information, and the second estimated feedback information matrix is the matrix expression form corresponding to the second estimated feedback information.

[0118] S303: Based on the first estimated feedback information matrix, the second estimated feedback information matrix, and a preset normalization function, respectively determine the first estimated feedback information corresponding to the first multimedia sample and the second estimated feedback information corresponding to the second multimedia sample.

[0119] To facilitate viewing the first estimated feedback information corresponding to the first multimedia sample and the second estimated feedback information corresponding to the second multimedia sample, in the embodiments of the present application, normalization processing can be performed according to the first estimated feedback information matrix, the second estimated feedback information matrix, and the preset normalization function, and based on the normalization result, determine the first estimated feedback information corresponding to the first multimedia sample and the second estimated feedback information corresponding to the second multimedia sample.

[0120] Optionally, the preset normalization function can be an activation function (softmax), an activation function (tanh), an activation function (sigmoid), etc.

[0121] In a possible implementation manner, in order to determine the first estimated feedback information corresponding to the first multimedia sample and the second estimated feedback information corresponding to the second multimedia sample based on the feedback information estimation model to be trained. Refer to Figure 5 As shown, it is a schematic diagram of the process for determining the first estimated feedback information corresponding to the first multimedia sample and the second estimated feedback information corresponding to the second multimedia sample provided by the embodiments of the present application.

[0122] S501: Concatenate the first feature vector, the second feature vector, and the comprehensive feature vector to obtain a comprehensive concatenated feature matrix;

[0123] For example, if the first feature vector is (a1, b1), the second feature vector is (a2, b2), and the third feature vector is (a3, b3), then the comprehensive concatenated feature matrix can be (a1, b1, a2, b2, a3, b3).

[0124] S502: Based on the comprehensive concatenated feature matrix, respectively perform at least one round of joint feature extraction operations using each preset joint feature extraction rule to obtain corresponding joint feature matrices;

[0125] In the embodiments of the present application, in order for the feedback information prediction model to be trained to achieve learning from different angles and with different focuses, various joint feature extraction rules are preset. For this comprehensive splicing feature matrix, at least one round of joint feature extraction operations are respectively performed using the preset various joint feature extraction rules to obtain corresponding joint feature matrices.

[0126] Among them, different joint feature extraction rules have different corresponding learning angles and learning focuses, and the joint feature matrices obtained based on different joint feature extraction rules are also different.

[0127] For example, three joint feature extraction rules are preset, namely joint feature extraction rule A, joint feature extraction rule B, and joint feature extraction rule C. Among them, the learning angles and focuses of the three preset joint feature extraction rules are different. For example, joint feature extraction rule A can focus on joint feature extraction of semantic features corresponding to the text contained in the advertisement sample, joint feature extraction rule B can focus on joint feature extraction of semantic features corresponding to the text contained in the information sample, and joint feature extraction rule C can focus on joint feature extraction of image features corresponding to the image contained in the advertisement sample, etc. Specifically, each joint feature extraction rule can be set according to requirements.

[0128] Based on the comprehensive splicing feature matrix, after performing at least one round of joint feature extraction operation using joint feature extraction rule A, joint feature matrix a is obtained. Based on the comprehensive splicing feature matrix, after performing at least one round of joint feature extraction operation using joint feature extraction rule B, joint feature matrix b is obtained. Based on the comprehensive splicing feature matrix, after performing at least one round of joint feature extraction operation using joint feature extraction rule C, joint feature matrix expert c is obtained. Among them, joint feature matrix a, joint feature matrix b, and joint feature matrix c are different.

[0129] In a possible implementation manner, joint feature extraction can be performed based on a multi-layer perceptron (MLP). Among them, one MLP layer corresponds to performing one round of joint feature extraction, and at least one round of joint feature extraction operation is performed based on at least one MLP layer. Among them, different MLP layers have the same network structure but different network parameters, and different joint extraction rules correspond to different MLP layers.

[0130] The process of performing one round of joint feature extraction operation based on one MLP layer can be expressed as:

[0131] Z l+1 =W l+1 a l +b l+1 =W l+1 σ(zl ) + b l+1 ;

[0132] wherein, Z l+1 the output of the MLP layer, z l represents the input of the output of the MLP layer, σ() represents the sigmoid function, W l+1 and b l+1 are respectively the network parameter matrices corresponding to the MLP layer.

[0133] In the embodiment of the present application, when performing the first round of joint feature extraction, the input of the MLP layer corresponding to the first round is the comprehensive splicing feature matrix, and when performing the last round of joint feature extraction, the output of the MLP layer corresponding to the last round is the joint feature matrix.

[0134] Refer to Figure 6a , which is a schematic diagram of a specific process of joint feature extraction based on the MLP layer provided by the embodiment of the present application.

[0135] Taking three joint feature extraction rules as an example, when performing at least one round of joint feature extraction operation using joint feature extraction rule A, feature extraction operation is performed based on all the MLPs included in MLP set 1 to obtain joint feature matrix a; when performing at least one round of joint feature extraction operation using joint feature extraction rule B, feature extraction operation is performed based on all the MLPs included in MLP set 2 to obtain joint feature matrix b; when performing at least one round of joint feature extraction operation using joint feature extraction rule C, feature extraction operation is performed based on all the MLPs included in MLP set 3 to obtain joint feature matrix c, wherein the structures corresponding to different MLP layers are the same, and the corresponding network parameters are different.

[0136] In another possible implementation manner, joint feature extraction can be performed based on at least one MLP layer and multiple expert networks, wherein one expert network is equivalent to two MLP layers.

[0137] Refer to Figure 6b , which is a schematic diagram of a specific process of joint feature extraction based on the MLP layer and expert networks provided by the embodiment of the present application.

[0138] Taking three jointly feature extraction rules as an example, when performing at least one round of jointly feature extraction operation using jointly feature extraction rule A, a feature extraction operation is performed based on the shared MLP layer and expert network 1 to obtain a jointly feature matrix a; when performing at least one round of jointly feature extraction operation using jointly feature extraction rule B, a feature extraction operation is performed based on the shared MLP layer and expert network 2 to obtain a jointly feature matrix b; when performing at least one round of jointly feature extraction operation using jointly feature extraction rule C, a feature extraction operation is performed based on the shared MLP layer and expert network 3 to obtain a jointly feature matrix c, where the structures of different expert networks are the same and the corresponding network parameters are different.

[0139] S503: Feature weighting is respectively performed on the obtained jointly feature matrices to obtain a first weighted feature matrix corresponding to the first multimedia sample and a second weighted feature matrix corresponding to the second multimedia sample;

[0140] In the embodiments of the present application, during the process of determining the first estimated feedback information of the first multimedia sample or during the process of determining the second estimated feedback information of the second multimedia sample, the importance degrees of the obtained jointly feature matrices for determining the first estimated feedback information or the second estimated feedback information are different. For example, the jointly feature matrix a is very important for determining the first estimated information of the first multimedia sample, the jointly feature matrix b is of general importance for determining the first estimated information of the first multimedia sample, and the jointly feature matrix c is not very important for determining the first estimated information of the second multimedia sample. The jointly feature matrix a is not very important for determining the first estimated information of the second multimedia sample, the jointly feature matrix b is very important for determining the first estimated information of the first multimedia sample, and the jointly feature matrix c is of general importance for determining the first estimated information of the second multimedia sample.

[0141] In order to determine the first estimated feedback information and the second estimated feedback information, in the embodiments of the present application, feature weighting is respectively performed on the obtained jointly feature matrices to obtain a first weighted feature matrix corresponding to the first multimedia sample and a second weighted feature matrix corresponding to the second multimedia sample. For example: for the first multimedia sample, the weight corresponding to the pre-stored jointly feature matrix a is 0.6, the weight corresponding to the pre-stored jointly feature matrix b is 0.3, and the weight corresponding to the pre-stored jointly feature matrix c is 0.1. Then, the first weighted feature matrix corresponding to the first multimedia sample is: jointly feature matrix a * 0.6 + jointly feature matrix b * 0.3 + jointly feature matrix c * 0.1.

[0142] S504: The multiple jointly feature matrices are respectively concatenated with the first weighted feature matrix and the second weighted feature matrix to obtain a first concatenated feature matrix corresponding to the first multimedia sample and a second concatenated feature matrix corresponding to the second multimedia sample;

[0143] S505: Based on the first splicing feature matrix and the second splicing feature matrix, perform at least one round of comprehensive feature extraction operations respectively to obtain the first estimated feedback information matrix corresponding to the first multimedia sample and the second estimated feedback information matrix corresponding to the second multimedia sample.

[0144] In the embodiments of the present application, in order to obtain the first estimated feedback information matrix corresponding to the first multimedia sample and the second estimated feedback information matrix corresponding to the second multimedia sample, comprehensive feature extraction operations are respectively performed on the first splicing matrix and the second splicing feature matrix to obtain the first estimated feedback information matrix corresponding to the first multimedia sample and the second estimated feedback information matrix corresponding to the second multimedia sample.

[0145] In a possible implementation manner, comprehensive feature extraction can be performed based on the MLP layer in the feedback information estimation model to be trained. Among them, one MLP layer performs one round of comprehensive feature extraction, and at least one round of comprehensive feature extraction operations are performed based on at least one MLP layer. Among them, the network structures corresponding to different MLP layers are the same, and the network parameters are different.

[0146] Refer to Figure 7 , which is a schematic diagram of a specific process for performing comprehensive feature extraction based on the MLP layer provided by the embodiments of the present application.

[0147] Input the first splicing feature matrix into the first MLP layer, and input the second splicing feature matrix into the second MLP layer, and perform one round of comprehensive feature extraction operations respectively to obtain the first estimated feedback information matrix corresponding to the first multimedia sample and the second estimated feedback information matrix corresponding to the second multimedia sample.

[0148] In a possible implementation manner, the first estimated feedback information matrix corresponding to the first multimedia sample is obtained through the following method:

[0149] Based on the result matrix corresponding to the first multimedia sample in the previous round after performing the comprehensive feature extraction operation and the first network parameter matrix pre-stored, determine the result matrix corresponding to the first multimedia sample in the current round;

[0150] Among them, the result matrix corresponding to the first multimedia sample in the first round is the first splicing feature matrix; the result matrix corresponding to the first multimedia sample in the previous round is positively correlated with the result matrix corresponding to the first multimedia sample in the current round.

[0151] In a possible implementation, the first product of the result matrix corresponding to the previous round of the first multimedia sample and the first network parameter matrix saved in advance can be determined as the result matrix corresponding to the current round of the first multimedia sample. For example, if the result matrix corresponding to the previous round of the first multimedia sample is matrix A and the first network parameter matrix saved in advance is matrix B, then the result matrix corresponding to the first round of the first multimedia sample can be determined as A·B.

[0152] In another possible embodiment, the first product of the result matrix corresponding to the previous round of the first multimedia sample and the first network parameter matrix saved in advance can be determined first, and then the sum value of the first product and the third network parameter matrix saved in advance can be determined as the result matrix corresponding to the current round of the first multimedia sample. For example, if the result matrix corresponding to the previous round of the first multimedia sample is matrix A, the first network parameter matrix saved in advance is matrix B, and the third network parameter matrix saved in advance is matrix C, then the result matrix corresponding to the first round of the first multimedia sample can be determined as A·B + C.

[0153] In a possible implementation, the second estimated feedback information matrix corresponding to the second multimedia sample is obtained by the following method:

[0154] Based on the result matrix corresponding to the previous round of the second multimedia sample output after the comprehensive feature extraction operation and the second network parameter matrix saved in advance, determine the result matrix corresponding to the current round of the second multimedia sample; among them, the result matrix corresponding to the first round of the second multimedia sample is the second splicing feature matrix, and the result matrix corresponding to the previous round of the second multimedia sample is positively correlated with the result matrix corresponding to the current round of the second multimedia sample.

[0155] In a possible embodiment, the second product of the result matrix corresponding to the previous round of the second multimedia sample and the second network parameter matrix saved in advance can be determined as the result matrix corresponding to the current round of the second multimedia sample. For example, if the result matrix corresponding to the previous round of the second multimedia sample is matrix C and the second network parameter matrix saved in advance is matrix D, then the result matrix corresponding to the first round of the second multimedia sample can be determined as C·D.

[0156] In another possible embodiment, the second product of the result matrix corresponding to the previous round of the second multimedia sample and the second network parameter matrix saved in advance can be determined first, and then the sum value of the second product and the third network parameter matrix saved in advance can be determined as the result matrix corresponding to the current round of the first multimedia sample. For example, if the result matrix corresponding to the previous round of the first multimedia sample is matrix C, the first network parameter matrix saved in advance is matrix D, and the third network parameter matrix saved in advance is matrix E, then the result matrix corresponding to the first round of the first multimedia sample can be determined as C·D + E.

[0157] In the embodiments of the present application, the first network parameter matrix and the second network parameter matrix corresponding to different comprehensive feature extraction operations are different.

[0158] In a possible implementation manner, in order to determine the first weighted feature matrix corresponding to the first multimedia sample and the second weighted feature matrix corresponding to the second multimedia sample. Refer to Figure 8 As shown, it is a schematic diagram of the process for determining the first weighted feature matrix and the second weighted feature matrix provided by the embodiments of the present application.

[0159] S801: Perform a linear transformation on the concatenation result corresponding to the concatenation of the obtained joint feature matrices to determine the first multi-dimensional matrix corresponding to the first multimedia sample and the second multi-dimensional matrix corresponding to the second multimedia sample;

[0160] S802: Based on a preset normalization function, perform normalization processing on the first multi-dimensional matrix and the second multi-dimensional matrix respectively to obtain the first weight corresponding to the first multimedia sample for each joint feature matrix and the second weight corresponding to the second multimedia sample for each joint feature matrix respectively;

[0161] In the embodiments of the present application, in order to determine the first weight corresponding to the first multimedia sample for each joint feature matrix, the first multi-dimensional matrix is normalized to obtain the first weight corresponding to the first multimedia sample for each joint feature matrix. Among them, the greater the first weight corresponding to the joint matrix, the greater the influence degree of the joint weight on the determination of the first estimated feedback information of the first multimedia sample.

[0162] In order to determine the second weight corresponding to the second multimedia sample for each joint feature matrix, the second multi-dimensional matrix is normalized to obtain the second weight corresponding to the second multimedia sample for each joint feature matrix. Among them, the greater the second weight corresponding to the joint matrix, the greater the influence degree of the joint weight on the determination of the second estimated feedback information of the second multimedia sample.

[0163] Optionally, the preset normalization function may be a softmax function, a tanh function, a sigmoid function, etc.

[0164] In a possible implementation manner, the first weight corresponding to the first multimedia sample for each joint feature matrix and the second weight corresponding to the second multimedia sample for each joint feature matrix may be determined based on a gating network.

[0165] The process of determining the weights (the first weight and the second weight) corresponding to each joint feature matrix for the multimedia samples (the first multimedia sample and the second multimedia sample) based on the gating network can be expressed as

[0166] g k (x) = softmax(W gk x);

[0167] where g k (x) is the weight determined by the k-th gating network, x represents the input of the k-th gating network, and W gk is the network parameter matrix corresponding to the k-th gating network.

[0168] In the embodiments of the present application, the gating network for determining the first weight is different from the gating network for determining the second weight.

[0169] Refer to Figure 9 , which is a schematic structural diagram of a method for determining the weight corresponding to the joint feature matrix based on the gating network provided by the embodiments of the present application.

[0170] Taking the output of four joint feature matrices as an example, namely joint feature matrix A, joint feature matrix B, joint feature matrix C, and joint feature matrix D, the first multimedia sample corresponds to gating network 1, and the second multimedia sample corresponds to gating network 2. After splicing the four joint feature matrices into the first multi-dimensional matrix and the second multi-dimensional matrix respectively, the first multi-dimensional matrix is input into gating network 1 for normalization processing, and the first weights corresponding to the joint feature matrix A, joint feature matrix B, joint feature matrix C, and joint feature matrix D for the first multimedia sample are obtained as w1, w2, w3, and w4 respectively. The second multi-dimensional matrix is input into gating network 2 for normalization processing, and the first weights corresponding to the joint feature matrix A, joint feature matrix B, joint feature matrix C, and joint feature matrix D for the second multimedia sample are obtained as p1, p2, p3, and p4 respectively.

[0171] S803: Determine the first weighted feature matrix corresponding to the first multimedia sample according to each joint feature matrix and the corresponding first weight, and determine the second weighted feature matrix corresponding to the second multimedia sample according to each joint feature matrix and the corresponding second weight.

[0172] In the embodiments of the present application, the first weighted feature matrix can be expressed as:

[0173]

[0174] where x1 represents the first multimedia sample, f(x1) is the first weighted feature matrix corresponding to the first multimedia sample, and f m(x1) is the m-th combined feature matrix corresponding to the first multimedia sample, w m (x1) is the first weight corresponding to the m-th combined feature matrix corresponding to the first multimedia sample, and N is the total number of combined feature matrices.

[0175] In the embodiments of the present application, the second weighted feature matrix can be expressed as:

[0176]

[0177] where x2 represents the second multimedia sample, f(x2) is the second weighted feature matrix corresponding to the second multimedia sample, f m (x2) is the m-th combined feature matrix corresponding to the second multimedia sample, w m (x2) is the second weight corresponding to the m-th combined feature matrix corresponding to the second multimedia sample, and N is the total number of combined feature matrices. The m-th combined feature matrix corresponding to the first multimedia sample is the same as the m-th combined feature matrix corresponding to the second multimedia sample, and the first weight corresponding to the m-th combined feature matrix corresponding to the first multimedia sample is different from the second weight corresponding to the m-th combined feature matrix corresponding to the second multimedia sample.

[0178] For example, if there are two combined feature matrices, namely combined feature matrix A and combined feature matrix B, where combined feature matrix A is Combined feature matrix B is The first weight corresponding to the combined feature matrix A is 0.8, and the first weight corresponding to the combined matrix B is 0.2, then the first weighted feature matrix is

[0179] In the embodiments of the present application, after determining the weights (the first weight and the second weight) corresponding to each combined feature matrix for the multimedia samples (the first multimedia sample and the second multimedia sample) based on the gating network, the first weighted feature matrix corresponding to the first multimedia sample is determined according to each combined feature matrix and the corresponding first weight and output from the gating network, and the second weighted feature matrix corresponding to the second multimedia sample is determined according to each combined feature matrix and the corresponding second weight and output from the gating network.

[0180] In the embodiments of the present application, the gating network for outputting the first weighted feature matrix is different from the gating network for outputting the first weighted feature matrix, and the gating network for determining the first weight outputs the first weighted feature matrix, and the gating network for determining the second weight outputs the second weighted feature matrix.

[0181] In a possible implementation manner, the parameters in the feedback information prediction model to be trained are adjusted by the following method:

[0182] First, determine the loss value corresponding to the corresponding first multimedia sample based on the difference between the first estimated feedback information and the corresponding first actual feedback information; determine the loss value corresponding to the second multimedia sample based on the difference between the second estimated feedback information and the corresponding second actual feedback information; then determine the sum of weights based on the loss value corresponding to the first multimedia sample and the corresponding first weight, and the loss value corresponding to the corresponding second multimedia sample and the corresponding second weight, and determine the sum of weights as the target loss value; finally, adjust the parameters in the feedback information prediction model to be trained based on the target loss value.

[0183] In the embodiments of the present application, taking the first multimedia sample as an advertisement sample, the first estimated feedback information includes the actual click-through rate, the actual deep conversion rate, and the actual shallow conversion rate, and the second multimedia sample as an information sample, and the second estimated feedback information includes the browsing rate, the like rate, and the sharing rate as an example for illustration. To determine the target loss value, first determine the loss value corresponding to the first multimedia sample according to the difference between the first estimated feedback information and the corresponding first actual feedback information respectively. Among them, the difference between the first estimated feedback information and the corresponding first actual feedback information includes: the difference between the actual click-through rate and the estimated click-through rate, the difference between the actual deep conversion rate and the estimated deep conversion rate, and the difference between the actual shallow conversion rate and the estimated shallow conversion rate.

[0184] Then, determine the loss value corresponding to the second multimedia sample according to the difference between the second estimated feedback information and the corresponding second actual feedback information. Among them, the difference between the second estimated feedback information and the corresponding second actual feedback information includes: the difference between the actual browsing rate and the estimated browsing rate, the difference between the actual like rate and the estimated like rate, and the difference between the actual sharing rate and the estimated sharing rate.

[0185] In the embodiments of the present application, the difference between the actual click-through rate and the estimated click-through rate can be expressed as:

[0186]

[0187] where x i is the i-th data training sample in the extracted training data samples, f(x i ; θ CTR ) is the estimated click-through rate of the i-th data training sample, is the actual click-through rate of the i-th data training sample, L(θ CTR ) is the difference between the actual click-through rate of the i-th data training sample and the estimated click-through rate of the i-th data training sample, N is the total number of the extracted training data samples, and θ CTR is the network parameter corresponding to the click-through rate pre-stored.

[0188] In the embodiments of the present application, the difference between the actual shallow conversion rate and the estimated shallow conversion rate can be expressed as:

[0189]

[0190] where x i is the i-th data training sample in the extracted training data samples, and f(x i ; θ CTR_1 ) is the estimated shallow conversion rate, is the actual shallow conversion rate of the i-th data training sample, L(θ CVR_1 ) is the difference between the actual deep conversion rate and the estimated deep conversion rate of the i-th data training sample, N is the total number of the extracted training data samples, and θ CTR_1 is the network parameter corresponding to the pre-stored shallow conversion rate.

[0191] In the embodiments of the present application, the difference between the actual deep conversion rate and the estimated deep conversion rate can be expressed as:

[0192]

[0193] where x i is the i-th data training sample in the extracted training data samples, and f(x i ; θ CVR_2 ) is the estimated deep conversion rate, is the actual deep conversion rate of the i-th data training sample, L(θ CVR_2 ) is the difference between the actual deep conversion rate and the estimated deep conversion rate of the i-th data training sample, N is the total number of the extracted training data samples, and θ CVR_2 is the network parameter corresponding to the pre-stored deep conversion rate.

[0194] In the embodiments of the present application, the loss value corresponding to the first multimedia sample can be expressed as:

[0195]

[0196] where Loss1 is the loss value corresponding to the first multimedia sample, where L(θ CTR ) is the difference between the actual click-through rate and the estimated click-through rate, is the difference between the actual shallow conversion rate and the estimated shallow conversion rate, The difference between the actual deep conversion rate and the predicted deep conversion rate, w1 is the weight corresponding to the difference between the actual click-through rate and the predicted click-through rate, w2 is the weight corresponding to the difference between the actual shallow conversion rate and the predicted shallow conversion rate, and w3 is the weight corresponding to the difference between the actual deep conversion rate and the predicted deep conversion rate.

[0197] In the embodiments of the present application, the difference between the actual view rate and the predicted view rate can be expressed as:

[0198]

[0199] where x i is the i-th data training sample in the extracted training data samples, f(x i ; θ view ) is the predicted view rate, is the actual view rate of the i-th data training sample, and L(θ view ) is the difference between the actual view rate and the predicted view rate of the i-th data training sample. N is the total number of the extracted training data samples, and θ view is the network parameter corresponding to the pre-stored view rate.

[0200] In the embodiments of the present application, the difference between the actual like rate and the predicted like rate can be expressed as:

[0201]

[0202] where x i is the i-th data training sample in the extracted training data samples, f(x i ; θ like ) is the predicted like rate of the i-th data training sample, is the actual like rate of the i-th data training sample, and L(θ like ) is the difference between the actual like rate and the predicted like rate of the i-th data training sample. N is the total number of the extracted training data samples, and θ like is the network parameter corresponding to the pre-stored like rate.

[0203] In the embodiments of the present application, the difference between the actual share rate and the predicted share rate can be expressed as:

[0204]

[0205] where x i is the i-th data training sample in the extracted training data samples, f(x i ; θ share ) is the predicted share rate, is the actual sharing rate of the i-th data training sample, and L(θ share ) is the difference between the actual sharing rate of the i-th data training sample and the predicted sharing rate of the i-th data training sample. N is the total number of extracted training data samples, and θ share is the network parameter corresponding to the pre-saved sharing rate.

[0206] In the embodiment of the present application, the loss value corresponding to the second multimedia sample can be expressed as:

[0207] Loss2 = w4 * L(θ view ) + w5 * L(θ like ) + w6 * L(θ share );

[0208] where Loss2 is the loss value corresponding to the second multimedia sample. Among them, L(θ view ) is the difference between the actual viewing rate and the predicted viewing rate, L(θ like ) is the difference between the actual like rate and the predicted like rate, L(θ share ) is the difference between the actual sharing rate and the predicted sharing rate. w4 is the weight corresponding to the difference between the actual viewing rate and the predicted viewing rate, w5 is the weight corresponding to the difference between the actual like rate and the predicted like rate, and w6 is the weight corresponding to the difference between the actual sharing rate and the predicted sharing rate.

[0209] In the present application, the target loss value can be expressed as:

[0210] Loss = w7 * Loss1 + w8 * Loss2;

[0211] where Loss is the target loss value, Loss1 is the loss value corresponding to the first multimedia sample, Loss2 is the loss value corresponding to the second multimedia sample, w7 is the first weight corresponding to the loss value of the first multimedia sample, and w8 is the second weight corresponding to the loss value of the second multimedia sample.

[0212] For the convenience of description, the weight corresponding to the difference between the actual click-through rate and the predicted click-through rate is called the first target weight, the weight corresponding to the difference between the actual shallow conversion rate and the predicted shallow conversion rate is called the second target weight, the weight corresponding to the difference between the actual deep conversion rate and the predicted deep conversion rate is called the third target weight, the weight corresponding to the difference between the actual viewing rate and the predicted viewing rate is called the fourth target weight, the weight corresponding to the difference between the actual like rate and the predicted like rate is called the fifth target weight, and the weight corresponding to the difference between the actual sharing rate and the predicted sharing rate is called the sixth target weight.

[0213] For example, the predicted click-through rate, predicted deep conversion rate, and predicted shallow conversion rate in the first predicted feedback information are 0.8, 0.6, and 0.4 respectively. The predicted view rate, predicted like rate, and predicted share rate in the second predicted feedback information are 0.7, 0.5, and 0.3 respectively. The actual click-through rate, actual deep conversion rate, and actual shallow conversion rate in the first actual feedback information are 0.85, 0.75, and 0.5. The actual view rate, actual like rate, and actual share rate in the second actual feedback information are 0.8, 0.6, and 0.2 respectively. The first weight is 0.6, the second weight is 0.4, the first target weight is 0.5, the second target weight is 0.3, the third target weight is 0.2, the fourth target weight is 0.2, the fifth target weight is 0.7, and the sixth target weight is 0.1. Then the loss value corresponding to the first multimedia sample is 0.5 * (0.85 - 0.8) + 0.3 * (0.75 - 0.6) + 0.2 * (0.5 - 0.4) = 0.06. The loss value corresponding to the second multimedia sample is 0.2 * (0.8 - 0.7) + 0.7 * (0.6 - 0.5) + 0.1 * (0.2 - 0.3) = 0.08. The target loss value is 0.6 * 0.06 + 0.4 * 0.08 = 0.068.

[0214] Refer to Figure 10 , which is a schematic structural diagram of a feedback information prediction model provided by an embodiment of the present application.

[0215] First, input the training data sample into the feedback information prediction model to be trained. Based on the first multimedia sample embedding layer in the feedback information prediction model to be trained, obtain the first feature vector corresponding to the content feature of the first multimedia sample. Based on the second multimedia embedding layer, obtain the second feature vector corresponding to the content feature of the second multimedia sample. Based on the shared embedding layer, obtain the comprehensive feature vector obtained by fusing the content features of the first multimedia sample, the content features of the second multimedia sample, and the object features of the sample object.

[0216] Concatenate the first feature vector, the second feature vector, and the comprehensive feature vector to obtain a comprehensive concatenated feature matrix. Based on the comprehensive concatenated feature matrix and the shared MLP layer, perform a round of joint feature extraction operations. Input the results output from the shared MLP layer into expert network 1, expert network 2, expert network 3, and expert network 4 respectively, and perform learning with different focuses and different angles. That is, perform comprehensive feature extraction operations based on the preset joint feature extraction rules, and output joint feature matrix a, joint feature matrix b, joint feature matrix c, and joint feature matrix d from expert network 1, expert network 2, expert network 3, and expert network 4 respectively.

[0217] Based on the gating network 1, determine that the first weights of the joint feature matrices a, b, c, and d for the first multimedia sample are w1, w2, w3, and w4 respectively, and output the first weighted feature matrix from the gating network 1. This first weighted feature matrix is a*w1 + b*w2 + c*w3 + d*w4. Based on the gating network 2, determine that the first weights of the joint feature matrices a, b, c, and d for the first multimedia sample are p1, p2, p3, and p4 respectively, and output the second weighted feature matrix from the gating network 2. This second weighted feature matrix is a*p1 + b*p2 + c*p3 + d*p4.

[0218] Input the result of concatenating the first weighted feature matrix, the joint feature matrices a, b, c, and d into the MLP1 layer for linear transformation, and output the first estimated feedback information matrix corresponding to the first multimedia sample. Input the second weighted feature matrix into the MLP2 for linear transformation, and output the second estimated feedback information matrix corresponding to the second multimedia sample.

[0219] Take the example where the first estimated feedback information contains three types of first estimated feedback information and the second estimated feedback information contains three types of second estimated feedback information. Input the first estimated feedback information matrix into the MLP layers corresponding to various types of first estimated feedback information, and output the first estimated feedback information that meets the type of estimated feedback information. Input the second estimated feedback information matrix into the MLP layers corresponding to various types of second estimated feedback information, and output the second estimated feedback information that meets the type of second estimated feedback information. Specifically, input the first estimated feedback information matrix into the MLP layers 3, 4, and 5 corresponding to various types of first estimated feedback information, and output the first estimated feedback information of the corresponding type. Input the second estimated feedback information matrix into the MLP layers 6, 7, and 8 corresponding to various types of second estimated feedback information, and output the second estimated feedback information of the corresponding type.

[0220] It should be noted that Figure 10 The above is only an example. In fact, the structure included inside the model is not limited and is not specifically defined in the embodiments of the present application.

[0221] In the embodiments of the present application, after training the feedback information prediction model, it is possible to predict only the feedback information corresponding to the first multimedia, or only the feedback information corresponding to the second multimedia, or simultaneously predict the feedback information corresponding to the first multimedia and the second multimedia. Taking the first multimedia as an advertisement and the predicted feedback information corresponding to the advertisement being the predicted click-through rate, the predicted deep conversion rate, and the predicted shallow conversion rate as an example for illustration.

[0222] Input the feature vector corresponding to the object feature of the target object and the feature vector corresponding to the content feature of the advertisement into the trained feedback information prediction model, and output the predicted click-through rate, the predicted deep conversion rate, and the predicted shallow conversion rate of the target object for the advertisement.

[0223] After determining the predicted click-through rate, the predicted deep conversion rate, and the predicted shallow conversion rate of the target object for each advertisement, determine the score of the target object's interest in the advertisement according to the predicted click-through rate, the predicted deep conversion rate, and the predicted shallow conversion rate of each advertisement, and sort each advertisement according to the score.

[0224] In the embodiments of the present application, advertisements are divided into ordinary advertisements and multi-target advertisements. Among them, an ordinary advertisement refers to an advertisement for which a user can only perform clicks and shallow conversions, and a multi-target advertisement refers to an advertisement for which a user can not only perform clicks and shallow conversions, but also perform deep conversions.

[0225] In the embodiments of the present application, if the advertisement is an ordinary advertisement, the score of the target object's interest in the advertisement can be determined according to the predicted click-through rate and the predicted shallow conversion rate, which can be expressed as:

[0226] ecpm = ectr * ecvr1 * bid

[0227] Wherein, ectr is the predicted click-through rate of the advertisement, ecvr1 is the predicted shallow conversion rate of the advertisement, and bid is the shallow bid of the advertisement pre-saved.

[0228] In the embodiments of the present application, if the advertisement is a multi-target advertisement, the score of the target object's interest in the advertisement can be determined according to the predicted click-through rate, the predicted shallow conversion rate, and the predicted deep conversion rate, which can be expressed as:

[0229] ecpm = ectr * ecvr1 * bid

[0230] Wherein, ectr is the predicted click-through rate of the advertisement, ecvr1 is the predicted shallow conversion rate of the advertisement, and bid is the shallow bid of the advertisement pre-saved.

[0231] When ecvr1 > ecvr2:

[0232] ecpm = ectr * ecvr1 * bid1

[0233] Otherwise:

[0234] ecpm = ectr * ecvr2 * bid2

[0235] Wherein, ectr is the estimated click-through rate of the advertisement, ecvr1 is the estimated shallow conversion rate of the advertisement, ecvr2 is the estimated deep conversion rate of the advertisement, bid1 is the shallow bid price of the advertisement saved in advance, and id2 is the deep bid price of the advertisement saved in advance.

[0236] Refer to Figure 11 , which is a schematic flowchart of a multimedia content recommendation provided by an embodiment of the present application.

[0237] Obtain a training data sample set, train a feedback information prediction model based on the training data sample set, and obtain a trained feedback information prediction model. Based on the trained feedback information prediction model, as well as the feature vectors corresponding to the object features of the target object and the feature vectors corresponding to the content features of each advertisement, determine the estimated click-through rate, estimated deep conversion rate, and estimated shallow conversion rate of the target object for each advertisement. Based on the estimated click-through rate, estimated deep conversion rate, and estimated shallow conversion rate corresponding to each advertisement, determine the score of the target object's interest in each advertisement, and recommend the advertisement with the highest score to the target object.

[0238] Based on the same inventive concept, an embodiment of the present application also provides an object detection device. As Figure 12 shown, it is a schematic structural diagram of a feedback information prediction model training device 1200 provided by an embodiment of the present application. The device may include:

[0239] An acquisition unit 1201, configured to acquire a training data sample set, wherein each training data sample includes: object features corresponding to a sample object, a first multimedia sample and a second multimedia sample with different data types and their respective content features, and a first actual feedback information generated by the sample object for the first multimedia sample and a second actual feedback information generated for the second multimedia sample;

[0240] A training unit 1202 for iteratively training a feedback information prediction model to be trained based on a training data sample set until a trained target feedback information prediction model is output. During one iteration process, the following operations are performed: inputting the extracted training data samples into the feedback information prediction model to be trained, and outputting first predicted feedback information of a corresponding first multimedia sample and second predicted feedback information of a corresponding second multimedia sample; and adjusting parameters in the feedback information prediction model to be trained according to the difference between the first predicted feedback information and the corresponding first actual feedback information and the difference between the second predicted feedback information and the corresponding second actual feedback information.

[0241] Optionally, the training unit 1202 is specifically configured to:

[0242] Based on the feedback information prediction model to be trained, perform feature extraction on the extracted training data samples to obtain a first feature vector corresponding to the content features of the first multimedia sample, a second feature vector corresponding to the content features of the second multimedia sample, and a corresponding comprehensive feature vector; wherein, the comprehensive feature vector is obtained by performing feature fusion according to the content features of the first multimedia sample, the content features of the second multimedia sample, and the object features of the sample object.

[0243] Perform feature analysis on the first feature vector, the second feature vector, and the comprehensive feature vector respectively to obtain a first predicted feedback information matrix corresponding to the first multimedia sample and a second predicted feedback information matrix corresponding to the second multimedia sample.

[0244] Based on the first predicted feedback information matrix, the second predicted feedback information matrix, and a preset normalization function, determine the first predicted feedback information corresponding to the first multimedia sample and the second predicted feedback information corresponding to the second multimedia sample respectively.

[0245] Optionally, the training unit 1202 is specifically configured to:

[0246] Concatenate the first feature vector, the second feature vector, and the comprehensive feature vector to obtain a comprehensive concatenated feature matrix.

[0247] Based on the comprehensive concatenated feature matrix, perform at least one round of joint feature extraction operations respectively using preset joint feature extraction rules to obtain corresponding joint feature matrices.

[0248] Perform feature weighting on the obtained joint feature matrices respectively to obtain a first weighted feature matrix corresponding to the first multimedia sample and a second weighted feature matrix corresponding to the second multimedia sample.

[0249] Concatenate multiple combined feature matrices with the first weighted feature matrix and the second weighted feature matrix respectively to obtain the first concatenated feature matrix corresponding to the first multimedia sample and the second concatenated feature matrix corresponding to the second multimedia sample;

[0250] Based on the first concatenated feature matrix and the second concatenated feature matrix, perform at least one round of comprehensive feature extraction operations respectively to obtain the first estimated feedback information matrix corresponding to the first multimedia sample and the second estimated feedback information matrix corresponding to the second multimedia sample.

[0251] Optionally, the training unit 1202 is specifically configured to:

[0252] For each round of comprehensive feature extraction operation, perform the following processes respectively:

[0253] Based on the result matrix corresponding to the previous round of the first multimedia sample output after performing the comprehensive feature extraction operation and the pre-saved first network parameter matrix, determine the result matrix corresponding to the current round of the first multimedia sample;

[0254] Among them, the result matrix corresponding to the first round of the first multimedia sample is the first concatenated feature matrix; the result matrix corresponding to the previous round of the first multimedia sample is positively correlated with the result matrix corresponding to the current round of the first multimedia sample.

[0255] Optionally, the training unit 1202 is specifically configured to:

[0256] For each round of comprehensive feature extraction operation, perform the following processes respectively:

[0257] Based on the result matrix corresponding to the previous round of the second multimedia sample output after performing the comprehensive feature extraction operation and the pre-saved second network parameter matrix, determine the result matrix corresponding to the current round of the second multimedia sample;

[0258] Among them, the result matrix corresponding to the first round of the second multimedia sample is the second concatenated feature matrix, and the result matrix corresponding to the previous round of the second multimedia sample is positively correlated with the result matrix corresponding to the current round of the second multimedia sample.

[0259] Optionally, the training unit 1202 is specifically configured to:

[0260] Perform a linear transformation on the concatenated result corresponding to the concatenated combined feature matrices obtained to determine the first multi-dimensional matrix corresponding to the first multimedia sample and the second multi-dimensional matrix corresponding to the second multimedia sample;

[0261] Based on a preset normalization function, perform normalization processing on the first multi-dimensional matrix and the second multi-dimensional matrix respectively to obtain the first weights corresponding to each joint feature matrix for the first multimedia sample, and the second weights corresponding to each joint feature matrix for the second multimedia sample;

[0262] According to each joint feature matrix and the corresponding first weights, determine the first weighted feature matrix corresponding to the first multimedia sample, and according to each joint feature matrix and the corresponding second weights, determine the second weighted feature matrix corresponding to the second multimedia sample.

[0263] Optionally, the training unit 1202 is specifically configured to:

[0264] Based on the difference between the first estimated feedback information and the corresponding first actual feedback information, determine the loss value corresponding to the corresponding first multimedia sample;

[0265] Based on the difference between the second estimated feedback information and the corresponding second actual feedback information, determine the loss value corresponding to the second multimedia sample;

[0266] Based on the loss value corresponding to the first multimedia sample and the corresponding first weights, the loss value corresponding to the corresponding second multimedia sample and the corresponding second weights, determine the weight sum, and determine the weight sum as the target loss value;

[0267] Based on the target loss value, adjust the parameters in the feedback information estimation model to be trained.

[0268] For the convenience of description, the above parts are divided into various modules (or units) according to functions and described separately. Of course, when implementing this application, the functions of each module (or unit) can be implemented in the same or multiple software or hardware.

[0269] After introducing the feedback information estimation model training method and device of the exemplary embodiment of the present application, next, introduce a device for training a feedback information estimation model according to another exemplary embodiment of the present application.

[0270] Those skilled in the art of the present technology can understand that various aspects of the present application can be implemented as a system, method, or program product. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.

[0271] Based on the same inventive concept as the above method embodiment, an electronic device is further provided in the embodiment of the present application. In one embodiment, the electronic device can be a server, such as Figure 1The server 120 shown. In this embodiment, the structure of the electronic device may be as Figure 13 shown, including a memory 1301, a communication module 1303, and one or more processors 1302.

[0272] The memory 1301 is used to store computer programs executed by the processor 1302. The memory 1301 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, programs required to run the instant messaging function, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.

[0273] The memory 1301 may be a volatile memory, such as a random-access memory (RAM); the memory 1301 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 1301 is any other medium that can be used to carry or store a desired computer program in the form of instructions or data structures and can be accessed by a computer, but is not limited to this. The memory 1301 may be a combination of the above memories.

[0274] The processor 1302 may include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 1302 is used to implement the above feedback information prediction model training method when calling the computer program stored in the memory 1301.

[0275] The communication module 1303 is used to communicate with the terminal device and other servers.

[0276] In the embodiments of the present application, the specific connection medium between the above memory 1301, communication module 1303, and processor 1302 is not limited. In the embodiments of the present application Figure 13 it is described that the memory 1301 and the processor 1302 are connected through a bus 1304, and the bus 1304 is described in thick lines in Figure 13 The connection methods between other components are only for illustrative purposes and are not to be construed as limiting. The bus 1304 may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of description, Figure 13 it is only described by a thick line in

[0277] The computer storage medium is stored in the memory 1301. The computer executable instructions are stored in the computer storage medium, and the computer executable instructions are used to implement the feedback information prediction model training method of the embodiments of the present application. The processor 1302 is used to execute the above-mentioned feedback information prediction model training method, such as Figure 2 shown.

[0278] In another embodiment, the electronic device may also be other electronic devices, such as Figure 1 shown terminal device 110. In this embodiment, the structure of the electronic device may be as Figure 14 shown, including: communication component 1410, memory 1420, display unit 1430, camera 1440, sensor 1450, audio circuit 1460, Bluetooth module 1470, processor 1480 and other components.

[0279] The communication component 1410 is used to communicate with the server. In some embodiments, it may include a Wireless Fidelity (WiFi) module. The WiFi module belongs to short-range wireless transmission technology, and the electronic device can help users send and receive information through the WiFi module.

[0280] The memory 1420 can be used to store software programs and data. The processor 1480 executes various functions and data processing of the terminal device 110 by running the software programs or data stored in the memory 1420. The memory 1420 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices. The memory 1420 stores an operating system that enables the terminal device 110 to run. In the present application, the memory 1420 can store the operating system and various application programs, and can also store the code for executing the feedback information prediction model training method of the embodiments of the present application.

[0281] The display unit 1430 can also be used to display the information input by the user or the information provided to the user and the graphical user interface (GUI) of various menus of the terminal device 110. Specifically, the display unit 1430 may include a display screen 1432 disposed on the front of the terminal device 110. Among them, the display screen 1432 may be configured in the form of a liquid crystal display, a light-emitting diode, etc. The display unit 1430 can be used to display the interface of multimedia content recommendation in the embodiments of the present application.

[0282] The display unit 1430 can also be used to receive input digital or character information and generate signal inputs related to the user settings and function control of the terminal device 110. Specifically, the display unit 1430 can include a touch screen 1431 disposed on the front of the terminal device 110, which can collect touch operations of the user thereon or nearby, such as clicking buttons, dragging scroll boxes, etc.

[0283] Among them, the touch screen 1431 can cover the display screen 1432, or the touch screen 1431 and the display screen 1432 can be integrated to implement the input and output functions of the terminal device 110. After integration, it can be simply referred to as a touch display screen. In this application, the display unit 1430 can display application programs and corresponding operation steps.

[0284] The camera 1440 can be used to capture static images, and the user can post comments on the images captured by the camera 1440 through an application. The camera 1440 can be one or multiple. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the processor 1480 to convert it into a digital image signal.

[0285] The terminal device may further include at least one sensor 1450, such as an acceleration sensor 1451, a distance sensor 1452, a fingerprint sensor 1453, and a temperature sensor 1454. The terminal device may also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, and a motion sensor.

[0286] The audio circuit 1460, the speaker 1461, and the microphone 1462 can provide an audio interface between the user and the terminal device 110. The audio circuit 1460 can transmit the electrical signal converted from the received audio data to the speaker 1461, and the speaker 1461 converts it into a sound signal for output. The terminal device 110 can also be configured with volume buttons for adjusting the volume of the sound signal. On the other hand, the microphone 1462 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1460 and then converted into audio data, and then the audio data is output to the communication component 1410 to be sent to, for example, another terminal device 110, or the audio data is output to the memory 1420 for further processing.

[0287] The Bluetooth module 1470 is used to interact with other Bluetooth devices having Bluetooth modules through the Bluetooth protocol. For example, the terminal device can establish a Bluetooth connection with a wearable electronic device (such as a smart watch) also equipped with a Bluetooth module through the Bluetooth module 1470, so as to perform data interaction.

[0288] The processor 1480 is the control center of the terminal device, connecting various parts of the entire terminal using various interfaces and circuits. By running or executing software programs stored in the memory 1420, and by calling data stored in the memory 1420, it executes various functions of the terminal device and processes data. In some embodiments, the processor 1480 may include one or more processing units; the processor 1480 may also integrate an application processor and a baseband processor, where the application processor mainly processes the operating system, user interface, and application programs, etc., and the baseband processor mainly processes wireless communication. It can be understood that the above baseband processor may not be integrated into the processor 1480. In this application, the processor 1480 can run the operating system, application programs, user interface display and touch response, as well as the feedback information prediction model training method of the embodiments of this application. In addition, the processor 1480 is coupled to the display unit 1430.

[0289] In some possible implementation manners, various aspects of the feedback information prediction model training method provided in this application can also be implemented in the form of a program product, which includes a computer program. When the program product runs on a computer device, the computer program is used to cause the computer device to execute the steps in the feedback information prediction model training method according to various exemplary embodiments of this application described above in this specification. For example, the computer device can execute the steps as shown in Figure 2 shown.

[0290] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0291] The program product of the embodiments of the present application may be a portable compact disc read-only memory (CD-ROM) and include a computer program, and may be run on a computing device. However, the program product of the present application is not limited thereto. In this document, a readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with a command execution system, apparatus, or device.

[0292] A readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a readable computer program is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with a command execution system, apparatus, or device.

[0293] The computer program contained on a readable medium may be transmitted with any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0294] The computer program for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The computer program may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any kind of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0295] It should be noted that although several units or subunits of the apparatus are mentioned in the foregoing detailed description, such a division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-mentioned units may be embodied in one unit. Conversely, the features and functions of one unit described above may be further divided and embodied by multiple units.

[0296] In addition, although the operations of the method of the present application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all of the shown operations must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step and performed, and / or one step may be decomposed into multiple steps and performed.

[0297] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable computer programs.

[0298] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.

[0299] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these changes and modifications.

Claims

1. A method for training a feedback information prediction model, characterized in that, The method includes: Obtaining a training data sample set, where each training data sample includes: object features corresponding to a sample object, a first multimedia sample and a second multimedia sample with different data types and their respective content features, and first actual feedback information generated by the sample object for the first multimedia sample and second actual feedback information generated for the second multimedia sample; Based on the training data sample set, iteratively training a feedback information prediction model to be trained until a trained target feedback information prediction model is output, where, in one iteration process, the following operations are performed: Based on the feedback information prediction model to be trained, performing feature extraction on the extracted training data sample to obtain a first feature vector corresponding to the content features of the first multimedia sample in the training data sample, a second feature vector corresponding to the content features of the second multimedia sample, and a comprehensive feature vector; wherein, the comprehensive feature vector is obtained by performing feature fusion according to the content features of the first multimedia sample, the content features of the second multimedia sample, and the object features of the sample object in the training data sample; Performing feature analysis on the first feature vector, the second feature vector, and the comprehensive feature vector respectively to obtain a first predicted feedback information matrix corresponding to the corresponding first multimedia sample and a second predicted feedback information matrix corresponding to the corresponding second multimedia sample; Based on the first predicted feedback information matrix, the second predicted feedback information matrix, and a preset normalization function, respectively determining the first predicted feedback information corresponding to the first multimedia sample and the second predicted feedback information corresponding to the second multimedia sample; Adjusting the parameters in the feedback information prediction model to be trained according to the difference between the first predicted feedback information and the corresponding first actual feedback information, the second predicted feedback information, and the difference between the corresponding second actual feedback information.

2. The method according to claim 1, characterized in that, The performing feature analysis on the first feature vector, the second feature vector, and the comprehensive feature vector to obtain a first predicted feedback information matrix corresponding to the corresponding first multimedia sample and a second predicted feedback information matrix corresponding to the corresponding second multimedia sample includes: Concatenating the first feature vector, the second feature vector, and the comprehensive feature vector to obtain a comprehensive concatenated feature matrix; Based on the comprehensive concatenated feature matrix, respectively performing at least one round of joint feature extraction operations using preset respective joint feature extraction rules to obtain corresponding joint feature matrices; Performing feature weighting on the obtained respective joint feature matrices to obtain a first weighted feature matrix corresponding to the first multimedia sample and a second weighted feature matrix corresponding to the second multimedia sample; Concatenating multiple joint feature matrices with the first weighted feature matrix and the second weighted feature matrix respectively to obtain a first concatenated feature matrix corresponding to the first multimedia sample and a second concatenated feature matrix corresponding to the second multimedia sample; Based on the first splicing feature matrix and the second splicing feature matrix, perform at least one round of comprehensive feature extraction operations respectively to obtain the first estimated feedback information matrix corresponding to the first multimedia sample and the second estimated feedback information matrix corresponding to the second multimedia sample.

3. The method according to claim 2, characterized in that, Based on the first splicing feature matrix, perform at least one round of comprehensive feature extraction operations to obtain the first estimated feedback information matrix corresponding to the first multimedia sample, including: For each round of comprehensive feature extraction operations, perform the following processes respectively: Based on the result matrix corresponding to the previous round of the first multimedia sample output after performing the comprehensive feature extraction operation and the pre - saved first network parameter matrix, determine the result matrix corresponding to the current round of the first multimedia sample; Among them, the result matrix corresponding to the first round of the first multimedia sample is the first splicing feature matrix; the result matrix corresponding to the previous round of the first multimedia sample is positively correlated with the result matrix corresponding to the current round of the first multimedia sample.

4. The method according to claim 2, wherein Based on the second splicing feature matrix, perform at least one round of comprehensive feature extraction operations to obtain the second estimated feedback information matrix corresponding to the second multimedia sample, including: For each round of comprehensive feature extraction operations, perform the following processes respectively: Based on the result matrix corresponding to the previous round of the second multimedia sample output after performing the comprehensive feature extraction operation and the pre - saved second network parameter matrix, determine the result matrix corresponding to the current round of the second multimedia sample; Among them, the result matrix corresponding to the first round of the second multimedia sample is the second splicing feature matrix, and the result matrix corresponding to the previous round of the second multimedia sample is positively correlated with the result matrix corresponding to the current round of the second multimedia sample.

5. The method according to claim 2, characterized in that, The step of performing feature weighting on each obtained joint feature matrix respectively to obtain the first weighted feature matrix corresponding to the first multimedia sample and the second weighted feature matrix corresponding to the second multimedia sample, includes: Perform a linear transformation on the splicing result corresponding to the concatenation of each obtained joint feature matrix to determine the first multi - dimensional matrix corresponding to the first multimedia sample and the second multi - dimensional matrix corresponding to the second multimedia sample; Based on a preset normalization function, perform normalization processing on the first multi - dimensional matrix and the second multi - dimensional matrix respectively to obtain the first weight corresponding to the first multimedia sample for each joint feature matrix and the second weight corresponding to the second multimedia sample for each joint feature matrix respectively; According to each joint feature matrix and the corresponding first weight, determine the first weighted feature matrix corresponding to the first multimedia sample, and according to each joint feature matrix and the corresponding second weight, determine the second weighted feature matrix corresponding to the second multimedia sample.

6. The method according to any one of claims 1 to 5, characterized in that, The step of adjusting the parameters in the feedback information estimation model to be trained according to the difference between the first estimated feedback information and the corresponding first actual feedback information, the difference between the second estimated feedback information and the corresponding second actual feedback information, includes: Determine the loss value corresponding to the corresponding first multimedia sample based on the difference between the first estimated feedback information and the corresponding first actual feedback information; Determine the loss value corresponding to the second multimedia sample based on the difference between the second estimated feedback information and the corresponding second actual feedback information; Determine the sum of weights based on the loss value corresponding to the first multimedia sample and the corresponding first weight, the loss value corresponding to the corresponding second multimedia sample and the corresponding second weight, and determine the sum of weights as the target loss value; Adjust the parameters in the feedback information estimation model to be trained based on the target loss value.

7. A feedback information prediction model training device, characterized in that Includes: An acquisition unit for acquiring a training data sample set, where each training data sample includes: object features corresponding to a sample object, a first multimedia sample and a second multimedia sample with different data types and their respective content features, and a first actual feedback information generated by the sample object for the first multimedia sample and a second actual feedback information generated for the second multimedia sample; A training unit for iteratively training the feedback information estimation model to be trained based on the training data sample set until the trained target feedback information estimation model is output. During one iteration process, the following operations are performed: Based on the feedback information estimation model to be trained, perform feature extraction on the extracted training data samples to obtain a first feature vector corresponding to the content features of the first multimedia sample in the training data samples, a second feature vector corresponding to the content features of the second multimedia sample, and a comprehensive feature vector; wherein, the comprehensive feature vector is obtained by performing feature fusion based on the content features of the first multimedia sample, the content features of the second multimedia sample, and the object features of the sample object in the training data samples; Perform feature analysis on the first feature vector, the second feature vector, and the comprehensive feature vector respectively to obtain a first estimated feedback information matrix corresponding to the corresponding first multimedia sample and a second estimated feedback information matrix corresponding to the corresponding second multimedia sample; Based on the first estimated feedback information matrix, the second estimated feedback information matrix, and a preset normalization function, determine the first estimated feedback information corresponding to the first multimedia sample and the second estimated feedback information corresponding to the second multimedia sample respectively; Adjust the parameters in the feedback information estimation model to be trained according to the difference between the first estimated feedback information and the corresponding first actual feedback information, the second estimated feedback information, and the difference between the corresponding second actual feedback information.

8. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores program code, and when the program code is executed by the processor, the processor is caused to execute the steps of any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, It includes program code, and when the storage medium runs on an electronic device, the program code is used to cause the electronic device to execute the steps of any one of claims 1 to 6.

10. A computer program product, characterized in that, Comprising computer instructions, characterized in that when the computer instructions are executed by a processor, the steps of any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Information processing method, recommendation method and related equipment

    CN110851713A