Model training method and device, electronic equipment and storage medium
By determining the intersection sample objects in the target application and the associated application and performing operation continuity analysis, the media content recommendation model is trained, and the problem of inaccurate analysis of sample data multiplexing effect in the prior art is solved, and the accuracy and effect of the recommended model is improved.
Patent Information
- Application Number
- CN202410003293.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-02
- Publication Date
- 2025-07-04
AI Technical Summary
The existing method of analysis of sample data reuse effect is not accurate enough, resulting in poor recommendation effects of media content recommendation models, and the overall evaluation indicators cannot reflect the differences in the model in different content types.
By obtaining the object operation data of the target application and the associated application, determining the intersection sample objects, and performing operation continuity analysis, obtaining the target operation indicator data of each preset content type, and training the media content recommendation model with the object operation data of the target content type.
It improves the recommendation accuracy of the media content recommendation model, avoids the reduction in recommendation accuracy caused by using data that does not belong to the target content type, and improves the recommendation effect.
Smart Images

Figure CN120256649A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a model training method, apparatus, electronic device, and storage medium. Background Art
[0002] In the currently commonly used scenario of sample data reuse in the industry, mainly by introducing sample data of associated applications (such as applications with similar click-through rates), the sample data of the target application is supplemented, so that the model "sees more and knows more". For example, for the multimedia resource display scenario, the sample object operation data of the associated application can be added to the training sample set of the target application, that is, the sample object operation data of the target application and the sample object operation data of the associated application are combined to train the media content recommendation model of the target application.
[0003] Furthermore, for the sample data of the associated application reused in the above training process, it is necessary to analyze its reuse effect. Currently, the way to analyze the reuse effect of sample data is usually to look at the AUC (Area Under the Curve) index of the model, or the general multimedia resource display effect indexes in A / B tests, such as the improvement of click-through rate, conversion rate, revenue, etc. However, the indexes in the above analysis methods of sample data reuse effect are more from the overall perspective to analyze the effect of the model, which is a general and overall evaluation index. Therefore, the analysis results obtained by the above analysis methods of sample data reuse effect are inaccurate, which in turn leads to a poor recommendation effect of the trained media content recommendation model. Summary of the Invention
[0004] In view of the above existing technical problems, the present disclosure provides a model training method, apparatus, electronic device, and storage medium.
[0005] According to an aspect of an embodiment of the present disclosure, there is provided a model training method, including:
[0006] Obtain first object operation data of a first sample object in a target application and second object operation data of a second sample object in an associated application of the target application, where the first object operation data is operation data of the first sample object performing a preset service operation on a corresponding first historical media content, and the second object operation data is operation data of the second sample object performing the preset service operation on a corresponding second historical media content;
[0007] According to the content type of the second historical media content, determine at least one intersection sample object corresponding to each preset content type from the first sample object and the second sample object;
[0008] Based on the third object operation data and the fourth object operation data, perform operation continuity analysis on each preset content type to obtain the target operation index data corresponding to each preset content type; the third object operation data is the object operation data of the intersection sample objects corresponding to each preset content type in the first object operation data, and the fourth object operation data is the object operation data of the intersection sample objects corresponding to each preset content type in the second object operation data;
[0009] Based on the object operation data corresponding to at least one target content type in the first object operation data and the second object operation data, train the media content recommendation model corresponding to the target application to obtain the target recommendation model corresponding to the target application, where the at least one target content type is the preset content type in the at least one preset content type whose corresponding target operation index data is not lower than the preset index data.
[0010] According to another aspect of the embodiments of the present disclosure, there is provided a model training device, including:
[0011] An operation data acquisition module, configured to acquire the first object operation data of the first sample object in the target application and the second object operation data of the second sample object in the associated application of the target application, where the first object operation data is the operation data of the first sample object performing a preset service operation on the corresponding first historical media content, and the second object operation data is the operation data of the second sample object performing the preset service operation on the corresponding second historical media content;
[0012] An intersection object acquisition module, configured to determine the intersection sample objects corresponding to at least one preset content type respectively from the first sample object and the second sample object according to the content type of the second historical media content;
[0013] A continuity analysis module, configured to perform operation continuity analysis on each preset content type based on the third object operation data and the fourth object operation data to obtain the target operation index data corresponding to each preset content type; the third object operation data is the object operation data of the intersection sample objects corresponding to each preset content type in the first object operation data, and the fourth object operation data is the object operation data of the intersection sample objects corresponding to each preset content type in the second object operation data;
[0014] A training module, configured to train a media content recommendation model corresponding to the target application based on object operation data corresponding to at least one target content type in the first object operation data and the second object operation data, to obtain a target recommendation model corresponding to the target application, where the at least one target content type is a preset content type in the at least one preset content type for which the corresponding target operation index data is not lower than the preset index data.
[0015] Optionally, the continuity analysis module includes:
[0016] A first quantity acquisition module, configured to determine the number of intersection sample objects corresponding to each preset content type based on the third object operation data;
[0017] A reused object determination module, configured to determine, based on the third object operation data and the fourth object operation data, a reused sample object corresponding to each preset content type from the intersection sample objects corresponding to each preset content type; the reused sample object is a sample object that, when performing the preset service operation on the media content corresponding to each preset content type in the associated application, also performs the preset service operation on the media content corresponding to each preset content type in the target application;
[0018] A second quantity acquisition module, configured to determine the number of reused objects of the reused sample objects corresponding to each preset content type;
[0019] A first ratio analysis module, configured to perform a continuous operation ratio analysis on each preset content type based on the number of intersection sample objects corresponding to each preset content type and the number of reused objects corresponding to each preset content type, to obtain target operation index data corresponding to each preset content type.
[0020] Optionally, the number of intersection sample objects corresponding to each preset content type includes the number of first sample objects corresponding to multiple preset sub-ranges in a first preset time range, the number of reused objects corresponding to each preset content type includes the number of second sample objects corresponding to the multiple preset sub-ranges respectively, and the target operation index data corresponding to each preset content type includes first operation index data and second operation index data; the first ratio analysis module includes:
[0021] The second proportion analysis module is used to perform object proportion analysis on each preset sub - range based on the number of the first sample objects corresponding to each preset sub - range in the number of intersection sample objects corresponding to each preset content type, and the number of the second sample objects corresponding to each preset sub - range in the number of reused objects corresponding to each preset content type, so as to obtain the third operation index data corresponding to each preset content type within each preset sub - range;
[0022] The third proportion analysis module is used to perform time - range proportion analysis on each preset content type based on the third operation index data corresponding to each preset content type within each preset sub - range, so as to obtain the first operation index data corresponding to each preset content type and the second operation index data corresponding to each preset content type.
[0023] Optionally, the third proportion analysis module includes:
[0024] The mean processing module is used to perform mean processing on the third operation index data of each of the multiple preset sub - ranges corresponding to each preset content type, so as to obtain the first operation index data corresponding to each preset content type;
[0025] The nearest sub - range determination module is used to determine the nearest preset sub - range corresponding to each preset content type from the multiple preset sub - ranges corresponding to each preset content type;
[0026] The index data generation module is used to generate the second operation index data corresponding to each preset content type based on the third operation index data of the nearest preset sub - range corresponding to each preset content type.
[0027] Optionally, the intersection object acquisition module includes:
[0028] The third object determination module is used to determine the third sample object corresponding to each preset content type from the second sample objects according to the content type of the second historical media content and the second object operation data; the third sample object corresponding to each preset content type is the object in the second sample objects that performs the preset service operation on the media content corresponding to each preset content type in the second historical media content;
[0029] The object intersection processing module is used to perform object intersection processing on the first sample objects and the third sample objects corresponding to each preset content type to obtain the intersection sample objects corresponding to each preset content type.
[0030] Optionally, the first sample object includes fourth sample objects corresponding to respective preset sub - ranges within a first preset time range, the third sample object corresponding to each preset content type includes a fifth sample object corresponding to a second preset time range, and the intersection sample object corresponding to each preset content type includes sixth sample objects corresponding to respective ones of the plurality of preset sub - ranges; the object intersection processing module includes:
[0031] An intersection processing module, configured to perform object intersection processing on the fifth sample object and the fourth sample objects corresponding to each preset sub - range to obtain sixth sample objects corresponding to each preset sub - range.
[0032] Optionally, the apparatus further includes:
[0033] A time interval acquisition module, configured to acquire a sample collection time interval corresponding to the target application and a model push time interval corresponding to the target application, where the sample collection time interval represents the time range corresponding to sample data used for each training of the media content recommendation model corresponding to the target application, and the model push time interval represents the time interval between two adjacent pushes of the media content recommendation model corresponding to the target application;
[0034] A second range determination module, configured to determine the second preset time range based on the minimum time interval among the sample collection time interval and the model push time interval.
[0035] Optionally, the apparatus further includes:
[0036] An operation mean analysis module, configured to perform operation mean analysis on the target application based on first object operation data of the first sample object in the target application to obtain fourth operation index data, where the fourth operation index data represents the number of sample object operations per preset duration on average among a plurality of preset durations; the number of sample object operations is the number of operations performed by the first sample object on the first historical media content for the preset service operation;
[0037] An average interval determination module, configured to determine an operation average interval duration corresponding to the target application based on the fourth operation index data;
[0038] A third range determination module, configured to determine a third preset time range based on the operation average interval duration and the second preset time range; the third preset time range is an integer multiple of the second preset time range, and the duration corresponding to the third preset time range is greater than the operation average interval duration;
[0039] A first range determination module, configured to use the third preset time range as the first preset time range.
[0040] Optionally, the target operation index data corresponding to each preset content type includes first operation index data and second operation index data, and the apparatus further includes:
[0041] A first comparison module, configured to compare the first operation index data corresponding to each preset content type with first preset index data, to obtain a first comparison result corresponding to each preset content type;
[0042] A second comparison module, configured to compare the second operation index data corresponding to each preset content type with second preset index data, to obtain a second comparison result corresponding to each preset content type;
[0043] A target type determination module, configured to use the preset content types that meet a preset index condition among the at least one preset content type as the at least one target content type; the preset index condition is that the corresponding first comparison result indicates that the first operation index data is not lower than the first preset index data, or the corresponding second comparison result indicates that the second operation index data is not lower than the second preset index data.
[0044] Optionally, the training module includes:
[0045] A current data determination module, configured to determine current sample operation data and label operation data corresponding to the current sample operation data from the object operation data corresponding to at least one target content type in the first object operation data and the second object operation data;
[0046] An operation prediction processing module, configured to input the current sample operation data into the media content recommendation model for operation prediction processing, to obtain operation prediction information;
[0047] A loss determination module, configured to determine target loss information based on the operation prediction information and the label operation data;
[0048] An update module, configured to update the media content recommendation model based on the target loss information, and based on the updated media content recommendation model, repeat the steps of determining current sample operation data and label operation data corresponding to the current sample operation data from the object operation data corresponding to at least one target content type in the first object operation data and the second object operation data, to the step of updating the media content recommendation model based on the target loss information, until a preset convergence condition is met;
[0049] A recommendation model generation module, configured to use the media content recommendation model when the preset convergence condition is reached as the target recommendation model.
[0050] According to another aspect of the embodiments of the present disclosure, an electronic device is provided, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the above-mentioned model training method.
[0051] According to another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device can execute the above-mentioned model training method.
[0052] According to another aspect of the embodiments of the present disclosure, a computer program product containing instructions is provided. When it runs on a computer, the computer is enabled to execute the above-mentioned model training method.
[0053] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:
[0054] By obtaining the first object operation data of the first sample object in the target application and the second object operation data of the second sample object in the associated application of the target application, and determining at least one intersection sample object corresponding to each preset content type from the first sample object and the second sample object according to the content type of the second historical media content, it is possible to obtain the intersection sample objects of each preset content type in the target application and the associated application. Then, by combining the third object operation data and the fourth object operation data, the operation continuity analysis is performed on each preset content type to obtain the target operation index data corresponding to each preset content type, and the accurate prediction of the reuse effect of the operation data corresponding to each preset content type in the second object operation data can be realized. Then, by combining the object operation data corresponding to at least one target content type in the first object operation data and the second object operation data, the media content recommendation model corresponding to the target application is trained to obtain the target recommendation model corresponding to the target application. At least one target content type is a preset content type in which the corresponding target operation index data is not lower than the preset index data, which can avoid reducing the recommendation accuracy of the model caused by training the model with the second object operation data that does not belong to the target content type, improve the recommendation prediction accuracy of the target recommendation model, and further improve the recommendation effect of the target recommendation model.
[0055] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation of the present disclosure.
[0057] Figure 1 is a schematic diagram of an application system shown according to an exemplary embodiment;
[0058] Figure 2 is a flowchart of a model training method shown according to an exemplary embodiment;
[0059] Figure 3 is a schematic flow diagram of a model training method shown according to an exemplary embodiment;
[0060] Figure 4 is a block diagram of a model training apparatus shown according to an exemplary embodiment;
[0061] Figure 5 is a block diagram of an electronic device for training a target recommendation model shown according to an exemplary embodiment;
[0062] Figure 6 is a block diagram of another electronic device for training a target recommendation model shown according to an exemplary embodiment. Detailed implementation manners
[0063] Various exemplary embodiments, features, and aspects of the present application will be described in detail below with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0064] The special term "exemplary" herein means "serving as an example, embodiment, or illustration". Any embodiment described as "exemplary" herein need not be construed as superior to or better than other embodiments.
[0065] In addition, for a better description of the present application, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present application can be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present application.
[0066] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system for perceiving the environment, acquiring knowledge, and using knowledge to obtain the best results. Artificial intelligence software technology mainly includes several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0067] In recent years, with the research and progress of artificial intelligence technology, artificial intelligence technology has been widely applied in multiple fields. The solution provided in the embodiments of this application involves technologies such as machine learning / deep learning, and will be specifically described through the following embodiments:
[0068] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an application system shown according to an exemplary embodiment. The application system can be used for the model training method of this application. As Figure 1 shown, the application system can at least include a server 01 and a terminal 02.
[0069] In the embodiments of this application, the server 01 can be used to train a target recommendation model. Specifically, the above-mentioned server 01 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0070] In the embodiments of this application, the terminal 02 can be used to generate object operation data. The above-mentioned terminal 02 can include entity devices of various types such as smart phones, desktop computers, tablet computers, laptop computers, smart speakers, vehicle-mounted terminals, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices, or can also include software running on the entity devices, such as application programs, etc. In the embodiments of this application, the operating systems running on the above-mentioned terminal 02 can include but are not limited to Android systems, IOS systems, linux, windows, etc.
[0071] In addition, it should be noted that Figure 1 what is shown is only an application environment provided by this disclosure. In actual applications, other application environments may also be included. For example, for the training process of the target recommendation model, it can also be implemented on the terminal 02.
[0072] In the embodiments of this specification, the above-mentioned terminal 02 and the server 01 can be directly or indirectly connected through wired or wireless communication methods, and this application does not make any limitations in this regard.
[0073] It should be noted that the step sequence shown in the following figures is a possible one, and in fact, it is not necessary to strictly follow this order. Some steps can be executed in parallel without depending on each other.
[0074] Specifically, Figure 2It is a flowchart of a model training method shown according to an exemplary embodiment. As Figure 2 shown, this model training method can be used in electronic devices such as terminals or servers, and specifically may include the following steps:
[0075] S201: Obtain the first object operation data of the first sample object in the target application and the second object operation data of the second sample object in the associated application of the target application.
[0076] In a specific embodiment, the target application may refer to an application that needs to perform media content recommendation prediction based on a media content recommendation model and display the corresponding media content based on the media content prediction result. The target application may include video applications or information applications, etc. Exemplarily, the video application may include an online video platform application.
[0077] In a specific embodiment, the associated application may refer to an application whose corresponding media content includes the target media content and has a need for media content recommendation; wherein, the content type corresponding to the target media content may be the same as the content type of some media content in the target application. Specifically, the associated application may include at least one application. Further, the associated application may include an application whose difference degree between the corresponding click-through rate and the click-through rate of the target application is lower than a preset difference degree. Exemplarily, when the target application is an online video platform application, the associated application may be an information display application.
[0078] In a specific embodiment, the first sample object may refer to an object that performs a preset service operation on the first historical media content. The first sample object may include multiple sample objects corresponding to the target application. Among them, the sample object may include a user account. The first historical media content may refer to the media content displayed in the target application before the current moment. Specifically, the media content type included in the first historical media content may include at least one preset content type. Among them, the preset content type may refer to the content type to be analyzed; specifically, at least one preset content type may be set according to actual application needs, and the present disclosure does not limit it.
[0079] In a specific embodiment, the first object operation data may refer to the operation data of the first sample object performing a preset service operation on the corresponding first historical media content. The first object operation data may include operation data corresponding to multiple preset data dimensions. The multiple preset data dimensions may include an operation object dimension, an operation time dimension, an operation content type dimension, an operation type dimension, and an operation application dimension. Among them, the preset service operation may be used to request a corresponding service from the corresponding background server by execution. Specifically, the preset service operation may include a click operation, a conversion operation, or an uninterested operation, etc. Exemplarily, any object operation data in the first object operation data may include the operation data "user1" corresponding to the operation object dimension, the operation data "2023.11.16; 16:49" corresponding to the operation time dimension, the operation data "clothing type" corresponding to the operation content type dimension, the operation data "click operation" corresponding to the operation type dimension, and the operation data "Application A" corresponding to the operation application dimension.
[0080] In a specific embodiment, the second sample object may refer to an object that performs a preset service operation on the second historical media content. The second sample object may include multiple sample objects corresponding to the associated application. Among them, the second historical media content may refer to the media content displayed in the associated application before the current moment. Specifically, the media content types included in the second historical media content may include at least one preset content type. Further, in the case where the associated application is one of multiple applications associated with the target application, the second historical media content may be the media content displayed in any one of the associated applications before the current moment.
[0081] In a specific embodiment, the second object operation data may refer to the operation data of the second sample object performing a preset service operation on the corresponding second historical media content. The second object operation data may include operation data corresponding to multiple preset data dimensions.
[0082] In a specific embodiment, the first object operation data may be obtained from the background server corresponding to the target application by sending an operation data acquisition instruction to the background server corresponding to the target application. Correspondingly, the second object operation data may be obtained from the background server corresponding to the associated application by sending an operation data acquisition instruction to the background server corresponding to the associated application.
[0083] S203: According to the content type of the second historical media content, determine the intersection sample objects corresponding to at least one preset content type from the first sample object and the second sample object.
[0084] In a specific embodiment, the intersection sample object corresponding to any preset content type may refer to the sample object obtained by performing a preset service operation on the first historical media content when a preset service operation is performed on the media content corresponding to any preset content type in the second historical media content. It can be understood that the intersection sample object corresponding to any preset content type may belong to the first sample object or the second sample object.
[0085] In a specific embodiment, step S203 may include:
[0086] Determine the third sample object corresponding to each preset content type from the second sample objects according to the content type of the second historical media content and the second object operation data;
[0087] Perform object intersection processing on the first sample object and the third sample object corresponding to each preset content type to obtain the intersection sample object corresponding to each preset content type.
[0088] In a specific embodiment, the third sample object corresponding to any preset content type may refer to the object in the second sample objects that performs a preset service operation on the media content corresponding to any preset content type in the second historical media content. The third sample object may include multiple sample objects.
[0089] In a specific embodiment, according to the content type of the second historical media content, the fifth object operation data corresponding to any preset content type may be obtained from the second object operation data; then, by performing deduplication processing on the sample objects corresponding to the object operation data in the fifth object operation data corresponding to any preset content type, the third sample object corresponding to any preset content type may be obtained.
[0090] In a specific embodiment, the intersection sample object corresponding to any preset content type can be obtained by performing an object intersection process on the first sample object and the third sample object corresponding to any preset content type. Specifically, the intersection sample object corresponding to any preset content type can be obtained by performing an object intersection process on the first sample object and the third sample object corresponding to any preset content type. Further, the same sample objects (objects that belong to both the first sample object and the third sample object corresponding to any preset content type) can be selected from the first sample object and the third sample object corresponding to any preset content type as the intersection sample object corresponding to any preset content type. Exemplarily, assume that the first sample object includes "sample object A", "sample object B", and "sample object C", and the third sample object corresponding to "preset content type 1" includes "sample object B", "sample object C", and "sample object D". By performing an object intersection process on the first sample object and the third sample object corresponding to "preset content type 1", the intersection sample object corresponding to "preset content type 1" can be obtained, including "sample object B" and "sample object C".
[0091] In a specific embodiment, the above first sample object may include fourth sample objects corresponding to multiple preset sub - ranges within the first preset time range. The first preset time range may include multiple different preset sub - ranges. Specifically, the durations corresponding to the above multiple different preset sub - ranges may be the same. Exemplarily, the multiple preset sub - ranges may include [t, t + T a ), [t + T a , t + 2T a )... [t + T b - T a , t + T b ).
[0092] In a specific embodiment, according to the operation time corresponding to each object operation data, the fifth object operation data within each preset sub - range can be obtained from the first object operation data; by performing a duplicate - removal process on the sample objects included in the fifth object operation data within each preset sub - range, the fourth sample objects within each preset sub - range can be obtained.
[0093] In a specific embodiment, the third sample object corresponding to each preset content type may include fifth sample objects corresponding to the second preset time range. The duration corresponding to the second preset time range is the same as the duration corresponding to each preset sub - range. The second preset time range may be the time range before the first preset time range. Specifically, the end time of the second preset time range and the start time of the first preset time range may be the same. Exemplarily, the first preset time range may be [t, t + Tb ; The second preset time range can be [t - T a , t], where T a is less than T b . Specifically, the second preset time range and the first preset time range can be set according to actual application requirements.
[0094] In a specific embodiment, according to the operation time corresponding to each object operation data and the content type of the second historical media content, the sixth object operation data corresponding to any preset content type within the second preset time range can be obtained from the second object operation data; then, by de-duplicating the sample objects included in the sixth object operation data corresponding to any preset content type within the second preset time range, the fifth sample object corresponding to each preset content type can be obtained.
[0095] In a specific embodiment, the intersection sample object corresponding to each preset content type may include the sixth sample objects corresponding to multiple preset sub-ranges respectively.
[0096] In a specific embodiment, the above-mentioned object intersection processing of the first sample object and the third sample object corresponding to each preset content type to obtain the intersection sample object corresponding to each preset content type may include:
[0097] Performing object intersection processing on the fifth sample object and the fourth sample object corresponding to each preset sub-range to obtain the sixth sample object corresponding to each preset sub-range.
[0098] In a specific embodiment, performing object intersection processing on the fifth sample object corresponding to any preset content type within the second preset time range and the fourth sample object within any preset sub-range can obtain the sixth sample object corresponding to any preset content type within the above-mentioned any preset sub-range.
[0099] In a specific embodiment, the above-mentioned second preset time range can be obtained in the following manner:
[0100] Obtain the sample collection time interval corresponding to the target application and the model push time interval corresponding to the target application;
[0101] Based on the minimum time interval among the sample collection time interval and the model push time interval, determine the second preset time range.
[0102] In a specific embodiment, the sample collection time interval may represent the time range corresponding to the sample data required for training the media content recommendation model corresponding to each training target application. Specifically, in the process of training the media content recommendation model corresponding to the target application, object operation data within the sample collection time interval before the current time may be selected from the first object operation data and the second object operation data as the training data for training the media content recommendation model. Furthermore, the current media content recommendation model may be trained based on the above training data to obtain a trained media content recommendation model, and the trained media content recommendation model may be pushed to the online application.
[0103] In a specific embodiment, the model push time interval may represent the time interval between two adjacent pushes of the media content recommendation model corresponding to the target application. Specifically, the trained media content recommendation model may be pushed to the online application at a time corresponding to the model push time interval after the time of the previous push of the media content recommendation model.
[0104] In a specific embodiment, preset model configuration information may be obtained from the background server corresponding to the target application; correspondingly, the sample collection time interval corresponding to the target application and the model push time interval corresponding to the target application may be obtained from the preset model configuration information.
[0105] In a specific embodiment, the minimum time interval among the sample collection time interval and the model push time interval may be used as the duration corresponding to the second preset time range to generate the second preset time range.
[0106] In a specific embodiment, the above first preset time range may be obtained in the following manner:
[0107] Based on the first object operation data of the first sample object in the target application, operation mean analysis is performed on the target application to obtain the fourth operation index data;
[0108] Based on the fourth operation index data, the operation average interval duration corresponding to the target application is determined;
[0109] Based on the operation average interval duration and the second preset time range, the third preset time range is determined;
[0110] The third preset time range is used as the first preset time range.
[0111] In a specific embodiment, the fourth operation metric data may characterize the number of sample object operations corresponding to each of multiple preset time periods on average. The number of sample object operations may refer to the number of operations performed by the first sample object on the first historical media content within each preset time period for a preset service operation. Optionally, the above preset time period may be one day. Exemplarily, when the preset time period is one day, the fourth operation metric data may characterize the average number of sample object operations per day.
[0112] In a specific embodiment, based on the first object operation data of the first sample object in the target application, the number of sample object operations corresponding to the target application within multiple preset time periods may be determined; the number of sample object operations corresponding to the target application within the above multiple preset time periods may be averaged to obtain the average number of sample object operations per preset time period; correspondingly, the average number of sample object operations per preset time period may be used as the fourth operation metric data.
[0113] In a specific embodiment, the average operation interval duration corresponding to the target application may characterize the time interval between any two executions of the preset service operation by the sample object on average.
[0114] In a specific embodiment, the average operation interval duration corresponding to the target application may be determined based on the preset time period and the fourth operation metric data. Specifically, the above average operation interval duration may be obtained through the following formula:
[0115] D ave = D1 / N4
[0116] where D ave is the average operation interval duration; D1 is the preset time period; N4 is the fourth operation metric data. Exemplarily, when the preset time period is one day, D1 may be 1440 min.
[0117] In a specific embodiment, the third preset time range may be an integer multiple of the second preset time range, and the duration corresponding to the third preset time range is greater than the average operation interval duration.
[0118] In a specific embodiment, the duration corresponding to the third preset time range may be obtained through the following formula:
[0119] D3 = CEILING(D ave / D2) * D2 + D2
[0120] where D3 is the third preset time range; D ave is the average operation interval duration; D2 is the duration corresponding to the second preset time range.
[0121] In a specific embodiment, the end time of the second preset time range and the duration corresponding to the third preset time range can be combined to generate the third preset time range. Specifically, the end time of the second preset time range can be used as the start time of the third preset time range, and by combining the duration corresponding to the above-mentioned third preset time range, the end time of the third preset time range can be determined. Correspondingly, the third preset time range can be used as the first preset time range.
[0122] S205: Based on the third object operation data and the fourth object operation data, perform operation continuity analysis on each preset content type to obtain the target operation index data corresponding to each preset content type.
[0123] In a specific embodiment, the target operation index data can represent the probability that the corresponding intersection sample object performs a preset service operation on the first preset media content when performing a preset service operation on the second preset media content. Among them, the first preset media content can refer to the media content corresponding to each preset content type in the target application. The second preset media content can refer to the media content corresponding to each preset content type in the associated application.
[0124] In a specific embodiment, the third object operation data can refer to the object operation data of the intersection sample object corresponding to each preset content type in the first object operation data.
[0125] In a specific embodiment, the fourth object operation data can refer to the object operation data of the intersection sample object corresponding to each preset content type in the second object operation data.
[0126] In a specific embodiment, the above step S205 may include:
[0127] Based on the third object operation data, determine the number of intersection sample objects corresponding to each preset content type;
[0128] Based on the third object operation data and the fourth object operation data, determine the reused sample objects corresponding to each preset content type from the intersection sample objects corresponding to each preset content type;
[0129] Determine the number of reused objects of the reused sample objects corresponding to each preset content type;
[0130] Based on the number of intersection sample objects corresponding to each preset content type and the number of reused objects corresponding to each preset content type, perform continuous operation ratio analysis on each preset content type to obtain the target operation index data corresponding to each preset content type.
[0131] In a specific embodiment, the number of intersection sample objects corresponding to any preset content type may refer to the number of sample objects included in the intersection sample objects corresponding to any preset content type.
[0132] In a specific embodiment, the reusable sample objects corresponding to each preset content type may refer to the sample objects in the intersection sample objects that, when performing a preset service operation on the media content corresponding to each preset content type in the associated application, perform the preset service operation on the media content corresponding to each preset content type in the target application.
[0133] In a specific embodiment, based on the third object operation data of the intersection sample objects corresponding to any preset content type and the fourth object operation data of the intersection sample objects corresponding to any preset content type, the sample objects in the intersection sample objects that, when performing a preset service operation on the media content corresponding to any preset content type in the associated application, perform the preset service operation on the media content corresponding to any preset content type in the target application can be used as the reusable sample objects corresponding to any preset content type.
[0134] In a specific embodiment, the number of reusable objects of the reusable sample objects corresponding to each preset content type may refer to the number of sample objects included in the reusable sample objects corresponding to each preset content type.
[0135] In a specific embodiment, the number of intersection sample objects corresponding to each preset content type may include the first sample object numbers corresponding to multiple preset sub - ranges in the first preset time range. Among them, the first sample object number corresponding to any preset sub - range may refer to the number of sample objects included in the intersection sample objects corresponding to each preset content type within any preset sub - range. Specifically, the first sample object number corresponding to any preset content type within any preset sub - range may be determined based on the sixth sample object corresponding to any preset content type within any preset sub - range.
[0136] In a specific embodiment, the number of reusable objects corresponding to each preset content type may include the number of second sample objects corresponding to multiple preset sub-ranges respectively. Among them, the number of second sample objects corresponding to any preset content type within any preset sub-range may represent the number of sample objects that perform a preset service operation on the media content corresponding to the above-mentioned any preset content type in the target application within the above-mentioned any preset sub-range when performing a preset service operation on the media content corresponding to the above-mentioned any preset content type in the associated application within the second preset time range. Specifically, based on the operation time corresponding to the object operation data, in combination with the third object operation data of the intersection sample objects corresponding to any preset content type and the fourth object operation data of the intersection sample objects corresponding to the above-mentioned any preset content type, the number of second sample objects corresponding to any preset content type within any preset sub-range can be determined.
[0137] In a specific embodiment, the target operation index data corresponding to each preset content type may include first operation index data and second operation index data. Among them, the first operation index data corresponding to each preset content type may represent the central tendency of the ratio between the number of reusable objects corresponding to each preset content type and the number of corresponding intersection sample objects within the first preset time range. The second operation index data corresponding to each preset content type may represent the ratio between the number of reusable objects corresponding to each preset content type and the number of corresponding intersection sample objects within the most recent preset sub-range.
[0138] In a specific embodiment, the first preset time range may be a time range after the second preset time range, and the duration corresponding to the second preset time range may be the same as the duration corresponding to each preset sub-range.
[0139] In a specific embodiment, based on the number of intersection sample objects corresponding to each preset content type and the number of reusable objects corresponding to each preset content type, performing a continuous operation ratio analysis on each preset content type to obtain the target operation index data corresponding to each preset content type may include:
[0140] Based on the number of first sample objects corresponding to each preset sub-range in the number of intersection sample objects corresponding to each preset content type, and the number of second sample objects corresponding to each preset sub-range in the number of reusable objects corresponding to each preset content type, performing an object ratio analysis on each preset sub-range to obtain the third operation index data corresponding to each preset content type within each preset sub-range;
[0141] Based on the third operation index data corresponding to each preset content type within each preset sub - range, perform a time - range proportion analysis on each preset content type to obtain the first operation index data corresponding to each preset content type and the second operation index data corresponding to each preset content type.
[0142] In a specific embodiment, the third operation index data corresponding to any preset content type within any preset sub - range can represent the proportion between the second sample object quantity corresponding to any preset sub - range in the number of reuse objects corresponding to any of the above - mentioned preset content types and the first sample object quantity corresponding to any preset sub - range in the intersection sample object quantity corresponding to any of the above - mentioned preset content types. Specifically, the third operation index data corresponding to any preset content type within any preset sub - range can be obtained through the following formula:
[0143]
[0144] where retio i,j is the third operation index data corresponding to the j - th preset content type within the i - th preset sub - range; N 2i,j is the second sample object quantity corresponding to the j - th preset content type within the i - th preset sub - range; N 1i,j is the first sample object quantity corresponding to the j - th preset content type within the i - th preset sub - range.
[0145] In a specific embodiment, based on the third operation index data corresponding to each preset content type within each preset sub - range, performing a time - range proportion analysis on each preset content type to obtain the first operation index data corresponding to each preset content type and the second operation index data corresponding to each preset content type may include:
[0146] Perform a mean - value processing on the third operation index data of each of the multiple preset sub - ranges corresponding to each preset content type to obtain the first operation index data corresponding to each preset content type;
[0147] Determine the nearest preset sub - range corresponding to each preset content type from the multiple preset sub - ranges corresponding to each preset content type;
[0148] Generate the second operation index data corresponding to each preset content type based on the third operation index data of the nearest preset sub - range corresponding to each preset content type.
[0149] In a specific embodiment, the third operation index data of each of the multiple preset sub-ranges corresponding to any preset content type is averaged to obtain the first operation index data corresponding to any of the above preset content types. Specifically, the first operation index data corresponding to any of the above preset content types can be obtained through the following formula:
[0150]
[0151] where retio all,j is the first operation index data corresponding to the j-th preset content type; retio i,j is the third operation index data corresponding to the j-th preset content type within the i-th preset sub-range; and n is the number of multiple preset sub-ranges.
[0152] In a specific embodiment, the most recent preset sub-range may refer to the preset sub-range that is closest to the current moment among the multiple preset sub-ranges. Specifically, the current moment can be obtained first. Correspondingly, the preset sub-range that is closest to the current moment can be selected from the multiple preset sub-ranges corresponding to each preset content type as the most recent preset sub-range corresponding to each preset content type.
[0153] In a specific embodiment, the third operation index data of the most recent preset sub-range corresponding to any preset content type can be used as the second operation index data corresponding to any of the above preset content types.
[0154] In the above embodiments, by analyzing the object proportion of each preset sub - range based on the first sample object quantity corresponding to each preset sub - range in the intersection sample object quantity corresponding to each preset content type and the second sample object quantity corresponding to each preset sub - range in the reused object quantity corresponding to each preset content type, the third operation index data corresponding to each preset content type in each preset sub - range is obtained. Based on the third operation index data corresponding to each preset content type in each preset sub - range, the time - range proportion analysis of each preset content type is carried out to obtain the first operation index data corresponding to each preset content type and the second operation index data corresponding to each preset content type. It can realize the detection of abnormal continuity operation data reuse of media content of any preset content type from the associated application to the target application within the cycle, and can timely obtain the abnormal situation of the object operation data reuse of the associated application on the continuity effect of media content recommendation of the target application. Furthermore, it can avoid reducing the recommendation accuracy of the model caused by training the model with the second object operation data that does not belong to the target content type, and can improve the recommendation prediction accuracy of the target recommendation model. It can be understood that the first operation index data corresponding to any preset content type can be used to indicate the continuity operation situation of the media content of any preset content type from the associated application to the target application within the first preset time range; the second operation index data corresponding to any preset content type can be used to indicate the continuity operation situation of the media content of any preset content type from the associated application to the target application within the time of the most recent preset sub - range.
[0155] S207: Train the media content recommendation model corresponding to the target application based on the object operation data corresponding to at least one target content type in the first object operation data and the second object operation data, to obtain the target recommendation model corresponding to the target application.
[0156] In a specific embodiment, at least one target content type may be a preset content type in at least one preset content type whose corresponding target operation index data is not lower than the preset index data.
[0157] In a specific embodiment, the target recommendation model may refer to the trained media content recommendation model, that is, the media content recommendation model to be pushed.
[0158] In a specific embodiment, at least one target content type can be obtained in the following way:
[0159] Compare the first operation index data corresponding to each preset content type with the first preset index data to obtain the first comparison result corresponding to each preset content type;
[0160] Compare the second operation indicator data corresponding to each preset content type with the second preset indicator data to obtain a second comparison result corresponding to each preset content type;
[0161] A preset content type that meets a preset indicator condition among at least one preset content type is used as at least one target content type.
[0162] In a specific embodiment, the first comparison result corresponding to any preset content type can be used to indicate the size relationship between the first operation indicator data corresponding to any preset content type and the first preset indicator data.
[0163] In a specific embodiment, the second comparison result corresponding to any preset content type can be used to indicate the size relationship between the second operation indicator data corresponding to any preset content type and the second preset indicator data.
[0164] In a specific embodiment, the preset index condition may be that the corresponding first comparison result indicates that the first operation index data is not lower than the first preset index data, or the corresponding second comparison result indicates that the second operation index data is not lower than the second preset index data. Specifically, the first preset index data and the second preset index data can be set according to actual application needs. Optionally, the value range of the first preset index data and the second preset index data can be 0.4 to 0.6; illustratively, the first preset index data can be 0.5, and the second preset index data can be 0.4.
[0165] In a specific embodiment, each preset content type can be matched with the preset index condition to obtain a target matching result. Accordingly, at least one target content type can be determined from at least one preset content type based on the target matching result. Specifically, the preset content type that satisfies the preset index condition indicated by the target matching result can be used as the target content type.
[0166] In a specific embodiment, when at least one preset content type does not meet the above-mentioned preset indicator conditions, the above-mentioned at least one target content type can be empty, that is, the media content recommendation model corresponding to the target application can be trained based on the first object operation data to obtain the target recommendation model corresponding to the target application.
[0167] In a specific embodiment, in the case where any of the preset content types does not meet the above preset index conditions, it can be determined that there is an anomaly in the object operation data corresponding to any of the preset content types in the second object operation data. By avoiding using the second object operation data corresponding to any of the preset content types to train the model, the influence of the object operation data with anomalies in the media content recommendation continuity effect on the model can be reduced, so as to improve the recommendation prediction accuracy of the target recommendation model, and further improve the recommendation effect of the target recommendation model.
[0168] In a specific embodiment, the above step S207 may include:
[0169] Determine the current sample operation data and the label operation data corresponding to the current sample operation data from the object operation data corresponding to at least one target content type in the first object operation data and the second object operation data;
[0170] Input the current sample operation data into the media content recommendation model for operation prediction processing to obtain operation prediction information;
[0171] Based on the operation prediction information and the label operation data, determine the target loss information;
[0172] Based on the target loss information, update the media content recommendation model, and based on the updated media content recommendation model, repeat the steps of determining the current sample operation data and the label operation data corresponding to the current sample operation data from the object operation data corresponding to at least one target content type in the first object operation data and the second object operation data, until based on the target loss information, update the media content recommendation model in the model training step until the preset convergence condition is met;
[0173] Use the media content recommendation model when the preset convergence condition is reached as the target recommendation model.
[0174] In a specific embodiment, the current sample operation data may refer to the sample data used in the current model training round. The current sample operation data may include at least one object operation data in the object operation data corresponding to at least one target content type in the first object operation data and the second object operation data.
[0175] In a specific embodiment, the label operation data corresponding to the current sample operation data can be used to provide a reference for the training of the media content recommendation model. The label operation data may include at least one object operation data in the object operation data corresponding to at least one target content type in the first object operation data and the second object operation data. Specifically, the operation time corresponding to the label operation data may be later than the operation time corresponding to the current sample operation data.
[0176] In a specific embodiment, the operation prediction information may be used to characterize the prediction probability of the sample object performing a preset service operation on the media content corresponding to each preset content type in the target application. The operation prediction information may include the prediction probability corresponding to each preset content type.
[0177] In a specific embodiment, the target loss information may be used to characterize the degree of difference between the operation prediction result corresponding to the operation prediction information and the labeled operation data. Specifically, the labeled prediction probability corresponding to the labeled operation data may be determined from the operation prediction information; then, based on the above-mentioned labeled prediction probability and the cross-entropy loss function, the target loss information may be determined. The labeled prediction probability may refer to the prediction probability corresponding to the labeled operation data in the operation prediction information.
[0178] In a specific embodiment, the preset convergence condition may include that the current iteration number meets the preset iteration number, or the target loss information is less than the preset loss information. Specifically, the preset iteration number and the preset loss information may be set according to actual application needs, and the present disclosure does not make any limitation.
[0179] In a specific embodiment, the update gradient may be determined based on the target loss information; then, based on the above-mentioned update gradient, the model parameters in the media content recommendation model may be updated to obtain an updated media content recommendation model. Correspondingly, based on the updated media content recommendation model, the model training steps may be repeated until the preset convergence condition is met.
[0180] In the above embodiments, by obtaining the first object operation data of the first sample object in the target application and the second object operation data of the second sample object in the associated application of the target application, and determining at least one intersection sample object corresponding to each preset content type from the first sample object and the second sample object according to the content type of the second historical media content, it is possible to obtain the intersection sample object of each preset content type in the target application and the associated application. Then, by combining the third object operation data and the fourth object operation data, the operation continuity analysis is performed on each preset content type to obtain the target operation index data corresponding to each preset content type, and it is possible to accurately predict the reuse effect of the operation data corresponding to each preset content type in the second object operation data. Next, by combining the object operation data corresponding to at least one target content type in the first object operation data and the second object operation data, the media content recommendation model corresponding to the target application is trained to obtain the target recommendation model corresponding to the target application. At least one target content type is a preset content type in at least one preset content type whose corresponding target operation index data is not lower than the preset index data, which can avoid reducing the recommendation accuracy of the model caused by using the second object operation data that does not belong to the target content type to train the model, can improve the recommendation prediction accuracy of the target recommendation model, and further improve the recommendation effect of the target recommendation model.
[0181] Figure 3 is a schematic flowchart of a model training method shown according to an exemplary embodiment. Specifically, as Figure 3 shown, it is possible to obtain the sample collection time interval corresponding to the target application and the model push time interval corresponding to the target application; based on the smallest time interval among the sample collection time interval and the model push time interval, the second preset time range can be determined; based on the fourth operation index data, the average operation interval duration corresponding to the target application can be determined; based on the average operation interval duration and the second preset time range, the third preset time range can be determined; correspondingly, the third preset time range can be used as the first preset time range.
[0182] Next, according to the content type of the second historical media content and the operation time corresponding to the object operation data, from the second sample objects, the fifth sample objects corresponding to each preset content type within the second preset time range can be determined; according to the operation time corresponding to the object operation data, from the first sample objects, the fourth sample objects corresponding to each preset sub-range within the first preset time range can be determined; by performing an object intersection process on the fourth sample objects corresponding to each preset sub-range within the first preset time range and the fifth sample objects corresponding to the second preset time range, the sixth sample objects corresponding to each preset sub-range in the intersection sample objects corresponding to each preset content type can be obtained; based on the third object operation data and the fourth object operation data, from the intersection sample objects corresponding to each preset content type within each preset sub-range, the reused sample objects corresponding to each preset content type within each preset sub-range can be determined; the number of the first sample objects corresponding to each of the multiple preset sub-ranges in the number of intersection sample objects and the number of the second sample objects corresponding to each of the multiple preset sub-ranges in the number of reused objects can be determined; based on the number of the first sample objects corresponding to each preset sub-range in the number of intersection sample objects corresponding to each preset content type and the number of the second sample objects corresponding to each preset sub-range in the number of reused objects corresponding to each preset content type, an object ratio analysis is performed on each preset sub-range, and the third operation index data corresponding to each preset content type within each preset sub-range can be obtained; based on the third operation index data corresponding to each preset content type within each preset sub-range, a time range ratio analysis is performed on each preset content type, and the first operation index data corresponding to each preset content type and the second operation index data corresponding to each preset content type in the target operation index data can be obtained.
[0183] Then, at least one preset content type that meets the preset index condition among at least one preset content type can be used as at least one target content type; the object operation data corresponding to at least one target content type in the first object operation data and the second object operation data can be used as training data; based on the above training data, the media content recommendation model corresponding to the target application can be trained to obtain the target recommendation model corresponding to the target application.
[0184] Figure 4 is a block diagram of a model training device shown according to an exemplary embodiment. As Figure 4 shown, the device may include:
[0185] The operation data acquisition module 410 can be used to acquire the first object operation data of the first sample object in the target application and the second object operation data of the second sample object in the associated application of the target application. The first object operation data is the operation data of the first sample object performing a preset service operation on the corresponding first historical media content, and the second object operation data is the operation data of the second sample object performing a preset service operation on the corresponding second historical media content;
[0186] The intersection object acquisition module 420 can be used to determine at least one intersection sample object corresponding to each preset content type from the first sample object and the second sample object according to the content type of the second historical media content;
[0187] The continuity analysis module 430 can be used to perform operation continuity analysis on each preset content type based on the third object operation data and the fourth object operation data to obtain the target operation index data corresponding to each preset content type. The third object operation data is the object operation data of the intersection sample object corresponding to each preset content type in the first object operation data, and the fourth object operation data is the object operation data of the intersection sample object corresponding to each preset content type in the second object operation data;
[0188] The training module 440 can be used to train the media content recommendation model corresponding to the target application based on the object operation data corresponding to at least one target content type in the first object operation data and the second object operation data to obtain the target recommendation model corresponding to the target application. At least one target content type is a preset content type in at least one preset content type whose corresponding target operation index data is not lower than the preset index data.
[0189] In a specific embodiment, the above continuity analysis module 430 may include:
[0190] The first quantity acquisition module can be used to determine the quantity of intersection sample objects corresponding to each preset content type based on the third object operation data;
[0191] The reusable object determination module can be used to determine the reusable sample objects corresponding to each preset content type from the intersection sample objects corresponding to each preset content type based on the third object operation data and the fourth object operation data. The reusable sample objects are the sample objects that perform the preset service operation on the media content corresponding to each preset content type in the target application when performing the preset service operation on the media content corresponding to each preset content type in the associated application;
[0192] The second quantity acquisition module can be used to determine the quantity of reusable objects of the reusable sample objects corresponding to each preset content type;
[0193] The first proportion analysis module can be used to perform continuous operation proportion analysis on each preset content type based on the number of intersection sample objects corresponding to each preset content type and the number of reused objects corresponding to each preset content type, so as to obtain the target operation index data corresponding to each preset content type.
[0194] In a specific embodiment, the above-mentioned first proportion analysis module may include:
[0195] The second proportion analysis module can be used to perform object proportion analysis on each preset sub-range based on the number of first sample objects corresponding to each preset sub-range in the number of intersection sample objects corresponding to each preset content type and the number of second sample objects corresponding to each preset sub-range in the number of reused objects corresponding to each preset content type, so as to obtain the third operation index data corresponding to each preset content type within each preset sub-range;
[0196] The third proportion analysis module can be used to perform time range proportion analysis on each preset content type based on the third operation index data corresponding to each preset content type within each preset sub-range, so as to obtain the first operation index data corresponding to each preset content type and the second operation index data corresponding to each preset content type.
[0197] In a specific embodiment, the above-mentioned third proportion analysis module may include:
[0198] The mean processing module can be used to perform mean processing on the third operation index data of each of the multiple preset sub-ranges corresponding to each preset content type to obtain the first operation index data corresponding to each preset content type;
[0199] The nearest sub-range determination module can be used to determine the nearest preset sub-range corresponding to each preset content type from the multiple preset sub-ranges corresponding to each preset content type;
[0200] The index data generation module can be used to generate the second operation index data corresponding to each preset content type based on the third operation index data of the nearest preset sub-range corresponding to each preset content type.
[0201] In a specific embodiment, the above-mentioned intersection object acquisition module 420 may include:
[0202] The third object determination module can be used to determine the third sample object corresponding to each preset content type from the second sample objects according to the content type of the second historical media content and the second object operation data; the third sample object corresponding to each preset content type is the object in the second sample objects that performs a preset service operation on the media content corresponding to each preset content type in the second historical media content;
[0203] The object intersection processing module can be used to perform object intersection processing on the first sample object and the third sample object corresponding to each preset content type, so as to obtain the intersection sample object corresponding to each preset content type.
[0204] In a specific embodiment, the above-mentioned first sample object includes the fourth sample object corresponding to each of multiple preset sub-ranges in the first preset time range, the third sample object corresponding to each preset content type includes the fifth sample object corresponding to the second preset time range, and the intersection sample object corresponding to each preset content type includes the sixth sample object corresponding to each of multiple preset sub-ranges; the above-mentioned object intersection processing module may include:
[0205] The intersection processing module can be used to perform object intersection processing on the fifth sample object and the fourth sample object corresponding to each preset sub-range, so as to obtain the sixth sample object corresponding to each preset sub-range.
[0206] In a specific embodiment, the above-mentioned device may further include:
[0207] The time interval acquisition module can be used to acquire the sample collection time interval corresponding to the target application and the model push time interval corresponding to the target application. The sample collection time interval represents the time range corresponding to the sample data used for each training of the media content recommendation model corresponding to the target application, and the model push time interval represents the time interval between two adjacent pushes of the media content recommendation model corresponding to the target application;
[0208] The second range determination module can be used to determine the second preset time range based on the minimum time interval among the sample collection time interval and the model push time interval.
[0209] In a specific embodiment, the above-mentioned device may further include:
[0210] The operation mean analysis module can be used to perform operation mean analysis on the target application based on the first object operation data of the first sample object in the target application, so as to obtain the fourth operation index data. The fourth operation index data represents the number of sample object operations per preset duration on average among multiple preset durations; the number of sample object operations is the number of operations performed by the first sample object on the first historical media content for the preset service operation;
[0211] The average interval determination module can be used to determine the operation average interval duration corresponding to the target application based on the fourth operation index data;
[0212] The third range determination module can be used to determine the third preset time range based on the operation average interval duration and the second preset time range; the third preset time range is an integer multiple of the second preset time range, and the duration corresponding to the third preset time range is greater than the operation average interval duration;
[0213] The first range determination module can be used to use the third preset time range as the first preset time range.
[0214] In a specific embodiment, the target operation index data corresponding to each preset content type includes first operation index data and second operation index data, and the above device may further include:
[0215] The first comparison module can be used to compare the first operation index data corresponding to each preset content type with the first preset index data to obtain the first comparison result corresponding to each preset content type;
[0216] The second comparison module can be used to compare the second operation index data corresponding to each preset content type with the second preset index data to obtain the second comparison result corresponding to each preset content type;
[0217] The target type determination module can be used to use the preset content types that meet the preset index conditions in at least one preset content type as at least one target content type; the preset index condition is that the corresponding first comparison result indicates that the first operation index data is not lower than the first preset index data, or the corresponding second comparison result indicates that the second operation index data is not lower than the second preset index data.
[0218] In a specific embodiment, the above training module 440 may include:
[0219] The current data determination module can be used to determine the current sample operation data and the label operation data corresponding to the current sample operation data from the object operation data corresponding to at least one target content type in the first object operation data and the second object operation data;
[0220] The operation prediction processing module can be used to input the current sample operation data into the media content recommendation model for operation prediction processing to obtain operation prediction information;
[0221] The loss determination module can be used to determine the target loss information based on the operation prediction information and the label operation data;
[0222] The update module can be used to update the media content recommendation model based on the target loss information, and based on the updated media content recommendation model, repeat the steps of determining the current sample operation data and the label operation data corresponding to the current sample operation data from the object operation data corresponding to at least one target content type in the first object operation data and the second object operation data, to the step of updating the media content recommendation model based on the target loss information, until the preset convergence condition is met;
[0223] A recommended model generation module can be used to use the media content recommendation model when the preset convergence condition is reached as the target recommendation model.
[0224] Regarding the device in the above embodiments, the specific manners in which each module and unit perform operations have been described in detail in the embodiments related to the method, and will not be elaborated here.
[0225] Figure 5 is a block diagram of an electronic device for training a target recommendation model shown according to an exemplary embodiment. The electronic device can be a server, and its internal structure diagram can be as Figure 5 shown. The electronic device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a model training method.
[0226] Figure 6 is a block diagram of another electronic device for training a target recommendation model shown according to an exemplary embodiment. The electronic device can be a terminal, and its internal structure diagram can be as Figure 6 shown. The electronic device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a model training method. The display screen of the electronic device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the electronic device can be a touch layer covered on the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the electronic device, or an external keyboard, a touchpad, or a mouse, etc.
[0227] Those skilled in the art can understand, Figure 5 or Figure 6The structure shown is only a block diagram of some structures related to the present disclosure, and does not constitute a limitation on the electronic device to which the present disclosure is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0228] In an exemplary embodiment, an electronic device is further provided, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the model training method in the embodiments of the present disclosure.
[0229] In an exemplary embodiment, a computer-readable storage medium is further provided. When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute the model training method in the embodiments of the present disclosure.
[0230] In an exemplary embodiment, a computer program product containing instructions is further provided. When it runs on a computer, the computer is enabled to execute the model training method in the embodiments of the present disclosure.
[0231] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it may include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application may include non-volatile and / or volatile memories. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0232] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the function of that module or unit.
[0233] It can be understood that in the specific embodiments of the present application, when it comes to relevant data such as user information, when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0234] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.
[0235] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A model training method, characterized in that, The method includes: Obtaining first object operation data of a first sample object in a target application and second object operation data of a second sample object in an associated application of the target application, where the first object operation data is operation data of the first sample object performing a preset service operation on corresponding first historical media content, and the second object operation data is operation data of the second sample object performing the preset service operation on corresponding second historical media content; Determining at least one intersection sample object corresponding to each preset content type from the first sample object and the second sample object according to the content type of the second historical media content; Based on third object operation data and fourth object operation data, performing operation continuity analysis on each preset content type to obtain target operation index data corresponding to each preset content type; the third object operation data is object operation data of the intersection sample object corresponding to each preset content type in the first object operation data, and the fourth object operation data is object operation data of the intersection sample object corresponding to each preset content type in the second object operation data; Training a media content recommendation model corresponding to the target application based on object operation data corresponding to at least one target content type in the first object operation data and the second object operation data to obtain a target recommendation model corresponding to the target application, where the at least one target content type is a preset content type in the at least one preset content type whose corresponding target operation index data is not lower than a preset index data.
2. The method according to claim 1, characterized in that, The performing operation continuity analysis on each preset content type based on the third object operation data and the fourth object operation data to obtain target operation index data corresponding to each preset content type includes: Determining the number of intersection sample objects corresponding to each preset content type based on the third object operation data; Based on the third object operation data and the fourth object operation data, determining a reused sample object corresponding to each preset content type from the intersection sample objects corresponding to each preset content type; the reused sample object is a sample object that, when performing the preset service operation on media content corresponding to each preset content type in the associated application, performs the preset service operation on media content corresponding to each preset content type in the target application; Determining the number of reused objects of the reused sample object corresponding to each preset content type; Based on the number of intersection sample objects corresponding to each preset content type and the number of reused objects corresponding to each preset content type, performing continuous operation ratio analysis on each preset content type to obtain target operation index data corresponding to each preset content type.
3. The method according to claim 2, characterized in that, The number of intersection sample objects corresponding to each preset content type includes the number of first sample objects corresponding to each of a plurality of preset sub - ranges in the first preset time range. The number of reused objects corresponding to each preset content type includes the number of second sample objects corresponding to each of the plurality of preset sub - ranges. The target operation index data corresponding to each preset content type includes first operation index data and second operation index data. The continuous operation ratio analysis of each preset content type based on the number of intersection sample objects corresponding to each preset content type and the number of reused objects corresponding to each preset content type to obtain the target operation index data corresponding to each preset content type includes: Based on the number of first sample objects corresponding to each preset sub - range in the number of intersection sample objects corresponding to each preset content type, and the number of second sample objects corresponding to each preset sub - range in the number of reused objects corresponding to each preset content type, perform an object ratio analysis on each preset sub - range to obtain third operation index data corresponding to each preset content type within each preset sub - range. Based on the third operation index data corresponding to each preset content type within each preset sub - range, perform a time - range ratio analysis on each preset content type to obtain the first operation index data corresponding to each preset content type and the second operation index data corresponding to each preset content type.
4. The method according to claim 3, wherein The performing a time - range ratio analysis on each preset content type based on the third operation index data corresponding to each preset content type within each preset sub - range to obtain the first operation index data corresponding to each preset content type and the second operation index data corresponding to each preset content type includes: Perform a mean processing on the third operation index data of each of the plurality of preset sub - ranges corresponding to each preset content type to obtain the first operation index data corresponding to each preset content type. Determine the nearest preset sub - range corresponding to each preset content type from the plurality of preset sub - ranges corresponding to each preset content type. Generate the second operation index data corresponding to each preset content type based on the third operation index data of the nearest preset sub - range corresponding to each preset content type.
5. The method according to claim 1, characterized in that The determining at least one intersection sample object corresponding to each preset content type from the first sample objects and the second sample objects according to the content type of the second historical media content includes: According to the content type of the second historical media content and the second object operation data, determine the third sample object corresponding to each preset content type from the second sample objects. The third sample object corresponding to each preset content type is the object in the second sample objects that performs the preset service operation on the media content corresponding to each preset content type in the second historical media content. Perform an object intersection process on the first sample object and the third sample object corresponding to each preset content type to obtain the intersection sample object corresponding to each preset content type.
6. The method according to claim 5, wherein The first sample object includes the fourth sample object corresponding to each of a plurality of preset sub - ranges within a first preset time range. The third sample object corresponding to each preset content type includes the fifth sample object corresponding to a second preset time range. The intersection sample object corresponding to each preset content type includes the sixth sample object corresponding to each of the plurality of preset sub - ranges. The performing an object intersection process on the first sample object and the third sample object corresponding to each preset content type to obtain the intersection sample object corresponding to each preset content type includes: Perform an object intersection process on the fifth sample object and the fourth sample object corresponding to each preset sub - range to obtain the sixth sample object corresponding to each preset sub - range.
7. The method according to claim 6, wherein The second preset time range is obtained in the following manner: Obtain the sample collection time interval corresponding to the target application and the model push time interval corresponding to the target application. The sample collection time interval represents the time range corresponding to the sample data used for each training of the media content recommendation model corresponding to the target application. The model push time interval represents the time interval between two adjacent pushes of the media content recommendation model corresponding to the target application. Based on the minimum time interval among the sample collection time interval and the model push time interval, determine the second preset time range.
8. The method according to claim 7, characterized in that The first preset time range is obtained in the following manner: Based on the first object operation data of the first sample object in the target application, perform an operation mean analysis on the target application to obtain the fourth operation index data. The fourth operation index data represents the number of sample object operations per preset duration on average among a plurality of preset durations. The number of sample object operations is the number of operations performed by the first sample object on the first historical media content for the preset business operation. Based on the fourth operation index data, determine the operation average interval duration corresponding to the target application. Based on the operation average interval duration and the second preset time range, determine the third preset time range. The third preset time range is an integer multiple of the second preset time range, and the duration corresponding to the third preset time range is greater than the operation average interval duration. Use the third preset time range as the first preset time range.
9. The method according to any one of claims 1-8, characterized in that, The target operation index data corresponding to each preset content type includes the first operation index data and the second operation index data. The at least one target content type is obtained in the following manner: Compare the first operation index data corresponding to each preset content type with the first preset index data to obtain the first comparison result corresponding to each preset content type. Compare the second operation index data corresponding to each preset content type with the second preset index data to obtain the second comparison result corresponding to each preset content type. Use the preset content types that meet the preset index conditions in the at least one preset content type as the at least one target content type; The preset index condition is that the corresponding first comparison result indicates that the first operation index data is not lower than the first preset index data, or the corresponding second comparison result indicates that the second operation index data is not lower than the second preset index data.
10. The method according to claim 1, wherein Training the media content recommendation model corresponding to the target application based on the object operation data corresponding to at least one target content type in the first object operation data and the second object operation data to obtain the target recommendation model corresponding to the target application includes: Determine the current sample operation data and the label operation data corresponding to the current sample operation data from the object operation data corresponding to at least one target content type in the first object operation data and the second object operation data; Input the current sample operation data into the media content recommendation model for operation prediction processing to obtain operation prediction information; Determine the target loss information based on the operation prediction information and the label operation data; Based on the target loss information, update the media content recommendation model, and based on the updated media content recommendation model, repeat the steps of determining the current sample operation data and the label operation data corresponding to the current sample operation data from the object operation data corresponding to at least one target content type in the first object operation data and the second object operation data to the step of updating the media content recommendation model based on the target loss information until the preset convergence condition is met; Use the media content recommendation model when the preset convergence condition is reached as the target recommendation model.
11. A model training device, characterized in that, The device includes: An operation data acquisition module, configured to acquire the first object operation data of the first sample object in the target application and the second object operation data of the second sample object in the associated application of the target application, where the first object operation data is the operation data of the first sample object performing a preset service operation on the corresponding first historical media content, and the second object operation data is the operation data of the second sample object performing the preset service operation on the corresponding second historical media content; An intersection object acquisition module, configured to determine at least one intersection sample object corresponding to each preset content type from the first sample object and the second sample object according to the content type of the second historical media content; A continuity analysis module, configured to perform operation continuity analysis on each preset content type based on the third object operation data and the fourth object operation data to obtain the target operation index data corresponding to each preset content type; the third object operation data is the object operation data of the intersection sample object corresponding to each preset content type in the first object operation data, and the fourth object operation data is the object operation data of the intersection sample object corresponding to each preset content type in the second object operation data; A training module, configured to train a media content recommendation model corresponding to the target application based on object operation data corresponding to at least one target content type in the first object operation data and the second object operation data, to obtain a target recommendation model corresponding to the target application, where the at least one target content type is a preset content type in the at least one preset content type for which the corresponding target operation index data is not lower than the preset index data.
12. An electronic device, characterized in that, Comprising: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to execute the executable instructions to implement the model training method according to any one of claims 1 to 10.
13. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the model training method according to any one of claims 1 to 10 is implemented.