Multi-domain data-based model training method and device, electronic equipment and storage medium
By dividing and processing the sample data of candidate scenarios and target scenarios in multi-domain data model training, we ensure that the model is not over-influenced under weight adjustment, and solving the problem of low model learning and low accuracy in the prior art, achieving higher model accuracy.
Patent Information
- Application Number
- CN202510204437.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-10
AI Technical Summary
The existing technology cannot realize the migration of data and information from other scenarios to target scenarios with relatively insufficient data while ensuring that the model is not biased, resulting in low model accuracy in target scenarios.
By obtaining multiple sample data in candidate scenarios and target scenarios, each sample data is divided according to the preset feature type, and the features corresponding to each sample data and each feature type are obtained. Then, the training model is trained through multiple sample data to obtain the target model. The model includes a first feature domain trained by the first unique feature and a second feature domain trained by the second unique feature, with the weight of the first feature domain higher than the weight of the second feature domain.
It realizes that data from other scenarios are migrated to the target scenario without causing model learning bias, thereby improving the model accuracy in the target scenario.
Smart Images

Figure CN120123908A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a multi-domain data model training method, apparatus, electronic device, and storage medium. Background Art
[0002] With the development of the affiliate advertising business, more and more affiliate traffic is accessed, and the accumulation of new traffic data is small. How to make full use of in-site advertising data and data from other traffic sources is crucial for improving the prediction accuracy of affiliate traffic. Currently, multi-domain joint modeling in the industry is generally achieved in the following ways: 1. Sample migration, using a multi-scenario and multi-task model structure to fuse and learn data samples; 2. Feature migration, training a separate model for samples in other scenarios to learn user embedding information and using it as the input information for the target domain as a feature.
[0003] The advantage of sample migration is that it can input complete information into the target domain model, but it will cause the target domain to over-learn other information, resulting in the model being biased. The advantage of feature migration is that it can reduce the learning cost of the target domain model, but the debugging cost of splitting the two models is relatively high, and a model cannot learn complete information end-to-end.
[0004] It can be seen that in the related art, it is impossible to migrate data and information from other scenarios to a target scenario with relatively insufficient data without causing the model to be biased, thereby resulting in relatively low accuracy of the model in the target scenario. Summary of the Invention
[0005] This application provides a multi-domain data model training method, apparatus, electronic device, and storage medium, so as to at least solve the problem in the related art that it is impossible to migrate data and information from other scenarios to a target scenario with relatively insufficient data without causing the model to be biased, thereby resulting in relatively low accuracy of the model in the target scenario.
[0006] According to one aspect of the embodiments of the present application, a multi-domain data model training method is provided, including:
[0007] Obtaining multiple sample data in a candidate scenario and / or a target scenario, where the sample data in the candidate scenario is used to supplement the sample data in the target scenario;
[0008] Dividing each sample data according to a preset feature type to obtain features corresponding to each sample data and each feature type, where the feature type includes: a first exclusive feature in the target scenario, a second exclusive feature in the candidate scenario, and an exclusive feature is a feature that only exists in the corresponding scenario;
[0009] Training the model to be trained with the multiple pieces of sample data according to the features corresponding to each piece of sample data and each feature type to obtain a target model, where the model to be trained is a model applied to prediction in a target scenario, and the model to be trained includes: a first feature sub-domain trained with the first exclusive feature, and a second feature sub-domain trained with the second exclusive feature, and the weight of the first feature sub-domain is higher than the weight of the second feature sub-domain.
[0010] Optionally, in the method as described above, the dividing each piece of sample data according to a preset feature type to obtain the features corresponding to each piece of sample data and each feature type includes:
[0011] Dividing the sample data according to a preset feature type to obtain scenario features corresponding to a scenario type, general features corresponding to a general type, first exclusive features corresponding to a target scenario type, and second exclusive features corresponding to a candidate scenario type.
[0012] Optionally, in the method as described above, the training the model to be trained with the multiple pieces of sample data according to the features corresponding to each piece of sample data and each feature type includes:
[0013] Training the scenario feature sub-domain in the model to be trained with the scenario features corresponding to the sample data;
[0014] Training the general embedding sub-domain in the model to be trained with the general features corresponding to the sample data;
[0015] Training the exclusive feature sub-domain in the model to be trained with the exclusive features.
[0016] Optionally, in the method as described above, the exclusive feature sub-domain includes:
[0017] A first exclusive embedding layer for inputting the first exclusive feature and a second exclusive embedding layer for inputting the second exclusive feature;
[0018] A routing layer for inputting the embedding vector output by the first exclusive embedding layer and / or the embedding vector output by the second exclusive embedding layer into an expert network set, where the expert network set includes multiple expert networks.
[0019] Optionally, in the method as described above, the training the scenario feature sub-domain in the model to be trained with the scenario features corresponding to the sample data includes:
[0020] Training the gating in the scene feature sub-domains through the scene features to obtain a weight vector, where the weight vector is used to control the relative importance of different feature domains, and the feature domains include: a general embedding sub-domain, a first exclusive feature sub-domain corresponding to the first exclusive feature in the exclusive feature sub-domain, and a second exclusive feature sub-domain corresponding to the second exclusive feature in the exclusive feature sub-domain.
[0021] Optionally, as in the foregoing method, the method further includes:
[0022] Inputting the behavior data of the target user into the target model to obtain the predicted advertisement click-through rate result and / or the predicted advertisement conversion rate result of the target user in the target scenario predicted by the target model, where the target scenario is the scenario of placing advertisements on the alliance platform, and the candidate scenario is the scenario of placing advertisements on the main website platform.
[0023] Optionally, as in the foregoing method, training the model to be trained with the multiple sample data according to the features corresponding to each sample data and each feature type to obtain a target model includes:
[0024] Training the model to be trained with the multiple sample data according to the features corresponding to each sample data and each feature type to obtain a trained model;
[0025] In the case where it is determined that the prediction accuracy of the trained model in the target scenario is lower than the preset requirement, increasing the weight of the first feature domain and training the trained model again, and repeating this cycle until the target model that meets the preset requirement is obtained.
[0026] According to another aspect of the embodiments of the present application, there is also provided a multi-domain data model training device, including:
[0027] An acquisition module, configured to acquire multiple sample data in the candidate scenario and / or the target scenario, where the sample data in the candidate scenario is used to supplement the sample data for the target scenario;
[0028] A partitioning module, configured to partition each sample data according to a preset feature type to obtain the features corresponding to each sample data and each feature type, where the feature types include: a first exclusive feature in the target scenario, a second exclusive feature in the candidate scenario, and the exclusive feature is a feature that only exists in the corresponding scenario;
[0029] A training module, configured to train a model to be trained with the multiple pieces of sample data according to the features corresponding to each piece of sample data and each feature type, so as to obtain a target model, where the model to be trained is a model applied to prediction in a target scenario, and the model to be trained includes: a first feature sub-domain trained with the first exclusive feature, and a second feature sub-domain trained with the second exclusive feature, and the weight of the first feature sub-domain is higher than the weight of the second feature sub-domain.
[0030] According to another aspect of the embodiments of the present application, an electronic device is further provided, including a processor, a communication interface, a memory, and a communication bus. The processor, the communication interface, and the memory complete communication with each other through the communication bus. The memory is used to store a computer program, and the processor is configured to execute the method steps in any of the above embodiments by running the computer program stored on the memory.
[0031] According to another aspect of the embodiments of the present application, a computer-readable storage medium is further provided. The storage medium stores a computer program, where the computer program is configured to execute the method steps in any of the above embodiments when running.
[0032] In the embodiments of the present application, a method of setting weights for different feature domains is adopted. By obtaining multiple sample data in a candidate scenario and / or a target scenario, where the sample data in the candidate scenario is used to supplement the sample data in the target scenario; each sample data is divided according to a preset feature type to obtain the features corresponding to each sample data and each feature type, where the feature types include: a first exclusive feature in the target scenario, a second exclusive feature in the candidate scenario, and an exclusive feature is a feature that only exists in the corresponding scenario; according to the features corresponding to each sample data and each feature type, the multiple sample data is used to train a model to be trained to obtain a target model, where the model to be trained is a model applied to prediction in the target scenario, and the model to be trained includes: a first feature sub-domain trained by the first exclusive feature, a second feature sub-domain trained by the second exclusive feature, and the weight of the first feature sub-domain is higher than the weight of the second feature sub-domain. Since the weight of the first feature sub-domain is higher than the weight of the second feature sub-domain, it can be realized that even if the second exclusive feature in the second feature sub-domain is more than the first exclusive feature in the first feature sub-domain, the trained target model can be adjusted by the weight and will not be overly affected by the second exclusive feature and learn wrongly, achieving the technical effect of improving the model accuracy, and further solving the problem in the related art that it is impossible to transfer data and information in other scenarios to the target scenario with relatively insufficient data under the premise of ensuring that the model does not learn wrongly, resulting in low accuracy of the model in the target scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.
[0034] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0035] Figure 1 It is a schematic diagram of the hardware environment of an optional multi-domain data model training method according to an embodiment of the present application;
[0036] Figure 2 It is a schematic flowchart of an optional multi-domain data model training method according to an embodiment of the present application;
[0037] Figure 3It is a schematic structural diagram of an optional model for implementing a multi - domain data model training method according to an embodiment of the present application;
[0038] Figure 4 It is a schematic flow diagram of another optional multi - domain data model training method according to an embodiment of the present application;
[0039] Figure 5 It is a schematic block diagram of an optional multi - domain data model training device according to an embodiment of the present application;
[0040] Figure 6 It is a schematic block diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners
[0041] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0042] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above - mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non - exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0043] First, some nouns or terms that appear during the description of the embodiments of the present application are applicable to the following explanations:
[0044] 1. Affiliate advertising business, also known as affiliate marketing or performance - based marketing, is a performance - based online marketing model.
[0045] 2. Affiliate traffic refers to the traffic introduced through the affiliate advertising business, which has the characteristics of diversification, strong targeting, and pay - per - performance.
[0046] 3. Feature Domain Partitioning (or Feature Splitting) refers to the process of classifying the features in the input data according to different sources, natures, or uses of the features in machine learning and deep learning models, and processing each category independently. This technique helps to better capture the unique information of different types of features, thereby improving the performance and generalization ability of the model.
[0047] According to one aspect of the embodiments of the present application, a method for training a multi-domain data model is provided. As an optional implementation, in this embodiment, the above-mentioned method for training a multi-domain data model can be applied to a hardware environment composed of a terminal 1402 and a server 1404 as shown in Figure 1 the figure. As shown in Figure 1 the figure, the server 1404 is connected to the terminal 1402 through a network, and can be used to provide services (such as game services, application services, etc.) for the terminal or the client installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for the server 1404.
[0048] The above network can include but is not limited to at least one of the following: wired network, wireless network. The above wired network can include but is not limited to at least one of the following: wide area network, metropolitan area network, local area network. The above wireless network can include but is not limited to at least one of the following: WIFI (Wireless Fidelity), Bluetooth. The terminal is not limited to a PC, mobile phone, tablet computer, etc.
[0049] The method for training a multi-domain data model according to the embodiments of the present application can be executed by the server, or can be executed by the terminal, or can be jointly executed by the server and the terminal. Among them, when the terminal executes the method for training a multi-domain data model according to the embodiments of the present application, it can also be executed by the client installed on it.
[0050] Taking the execution of the method for training a multi-domain data model in this embodiment by the server as an example, Figure 1 a method for training a multi-domain data model provided by the embodiments of the present application includes the following steps:
[0051] Step S202, obtaining multiple pieces of sample data in a candidate scenario and / or a target scenario, where the sample data in the candidate scenario is used to supplement the sample data in the target scenario.
[0052] The method for training a multi-domain data model in this embodiment can be applied to the scenario of training a model in a scenario with few training samples.
[0053] The candidate scenario can be a scenario with a large number of training samples. For example, the training samples of the main station. The main station is the advertising placement platform owned by the model training party itself. The target scenario can be a scenario with fewer training samples. For example, the training samples of the alliance. The alliance is a platform that can be used by the main station to place advertisements for the advertisers of the main station.
[0054] Among the multiple sample data, there is sample data under the candidate scenario and / or the target scenario. Specifically, the multiple sample data can include sample data under the candidate scenario, can also include sample data under the target scenario, and can further include sample data that involves both the candidate scenario and the target scenario.
[0055] Step S204: Divide each piece of sample data according to the preset feature types to obtain the features corresponding to each piece of sample data and each feature type. Among them, the feature types include: the first exclusive feature under the target scenario, and the second exclusive feature under the candidate scenario. The exclusive feature is a feature that only exists in the corresponding scenario.
[0056] Specifically, after obtaining each piece of sample data, since each piece of sample data can include multiple features, therefore, all the features included in each piece of sample data can be classified according to the feature types to obtain the features corresponding to each feature type. And in this embodiment, the feature types at least include the first exclusive feature under the target scenario, and the first exclusive feature is a feature that only exists under the target scenario. For example, when the target scenario is the scenario of an audiobook platform, the first exclusive feature can be the click behavior on the audio; the second exclusive feature under the candidate scenario, and the second exclusive feature is a feature that only exists under the candidate scenario. For example, when the candidate scenario is the scenario of a video platform, the second exclusive feature can be the click behavior on the video.
[0057] Step S206: Train the model to be trained with the multiple sample data according to the features corresponding to each piece of sample data and each feature type to obtain the target model. Among them, the model to be trained is a model applied to the prediction under the target scenario. The model to be trained includes: the first feature sub-domain trained with the first exclusive feature, and the second feature sub-domain trained with the second exclusive feature. The weight of the first feature sub-domain is higher than the weight of the second feature sub-domain.
[0058] Specifically, after determining the feature types corresponding to the features in each sample data, the model to be trained can be trained with multiple sample data. Moreover, for different feature domains, during the training phase, the feature types used for training in each feature domain are different from each other. The first feature domain is trained with the first exclusive feature, and the second feature domain is trained with the second exclusive feature. Therefore, the first feature domain is used to capture the unique information of the first exclusive feature in the target scenario, while the second feature domain is used to capture the unique information of the second exclusive feature in the candidate scenario. In this embodiment, to increase the weight of the target scenario, the weight of the first feature domain is set higher than the weight of the second feature domain, so that the model can pay more attention to the features in the target scenario. Thus, after determining the first embedding vector of the first exclusive feature in the first feature domain, the first embedding vector can be multiplied by the weight of the first feature domain to generate the feature representation output by the first feature domain. Similarly, the second feature domain can also output the corresponding feature representation in a similar manner. Furthermore, after the above training, the trained target model can be obtained, and the target model can achieve a better prediction effect in the target scenario.
[0059] In the embodiment of the present application, by setting the weights of different feature domains, multiple sample data in the candidate scenario and / or the target scenario are obtained, where the sample data in the candidate scenario is used to supplement the sample data in the target scenario; each sample data is divided according to the preset feature types to obtain the features corresponding to each sample data and each feature type, where the feature types include: the first exclusive feature in the target scenario, the second exclusive feature in the candidate scenario, and the exclusive feature is a feature that only exists in the corresponding scenario; according to the features corresponding to each sample data and each feature type, the model to be trained is trained with multiple sample data to obtain the target model, where the model to be trained is a model applied to the prediction in the target scenario, and the model to be trained includes: the first feature domain trained with the first exclusive feature, the second feature domain trained with the second exclusive feature, and the weight of the first feature domain is higher than the weight of the second feature domain. Since the weight of the first feature domain is higher than the weight of the second feature domain, it can be realized that even if the second exclusive feature in the second feature domain is more than the first exclusive feature in the first feature domain, the trained target model can be adjusted by the weight and will not be overly influenced by the second exclusive feature and learn wrongly, achieving the technical effect of improving the model accuracy. Furthermore, it solves the problem in the related art that it is impossible to transfer the data and information in other scenarios to the target scenario with relatively insufficient data under the premise of ensuring that the model does not learn wrongly, resulting in a low accuracy of the model in the target scenario.
[0060] As an optional implementation manner, like the method described above, the foregoing step S202 can be implemented through the following steps. Each piece of sample data is divided according to a preset feature type to obtain the feature corresponding to each piece of sample data and each feature type:
[0061] The sample data is divided according to a preset feature type to obtain a scene feature corresponding to a scene type, a general feature corresponding to a general type, a first exclusive feature corresponding to a target scene type, and a second exclusive feature corresponding to a candidate scene type.
[0062] Optionally, the scene type feature can be used to represent information about a specific scene, such as the position, time, device type, etc. of an advertisement display. The general type feature can be used to represent basic features common to all scenes or tasks, such as user ID, advertisement ID, user age, user gender, etc. The first exclusive feature can be used to represent features specific to the target scene type, such as the click behavior and browsing behavior of a user in the target scene. The second exclusive feature can be used to represent features specific to the candidate scene type, such as the click behavior and browsing behavior of a user in the candidate scene. Furthermore, when each piece of sample data includes the following fields: user ID, advertisement ID, user age, user gender, advertisement display position, advertisement display time, device type, user click behavior in the target scene, and user click behavior in the candidate scene, the sample data can be divided according to a preset feature type to obtain a scene feature (advertisement display position, advertisement display time, device type) corresponding to the scene type, a general feature (user ID, advertisement ID, user age, user gender) corresponding to the general type, a first exclusive feature (user click behavior in the target scene) corresponding to the target scene type, and a second exclusive feature (user click behavior in the candidate scene) corresponding to the candidate scene type.
[0063] By dividing the sample data according to a preset feature type, features corresponding to different scene types can be obtained, including scene features, general features, the first exclusive feature of the target scene, and the second exclusive feature of the candidate scene. This division helps the model better capture the unique information of different types of features, thereby improving the prediction accuracy.
[0064] As an optional implementation manner, like the method described above, the foregoing step S206 can be implemented through the following steps. According to the feature corresponding to each piece of sample data and each feature type, the model to be trained is trained with multiple pieces of sample data:
[0065] The scene feature domain in the model to be trained is trained with the scene feature corresponding to the sample data;
[0066] Train the general embedding sub-domains in the model to be trained through the general features corresponding to the sample data;
[0067] Train the exclusive feature sub-domains in the model to be trained through the exclusive features.
[0068] Specifically, the exclusive features include the aforementioned first exclusive feature and second exclusive feature, and the exclusive feature sub-domains include a first exclusive feature sub-domain corresponding to the target scenario type and a second exclusive feature sub-domain corresponding to the candidate standard scenario type.
[0069] That is to say, after dividing the sample data, the scenario feature sub-domains in the model to be trained can be trained through the scenario features in the sample data; and the general embedding sub-domains in the model to be trained can be trained through the general features in the sample data; the first exclusive feature sub-domains in the model to be trained can be trained through the first exclusive features in the sample data, and the second exclusive feature sub-domains in the model to be trained can be trained through the second exclusive features in the sample data.
[0070] Optionally, the model structure of the above-mentioned feature sub-domains can adopt the EPNET structure.
[0071] As an optional implementation manner, as in the aforementioned method, the exclusive feature sub-domains include:
[0072] A first exclusive embedding layer for inputting the first exclusive feature and a second exclusive embedding layer for inputting the second exclusive feature;
[0073] A routing layer for inputting the embedding vectors output by the first exclusive embedding layer and / or the embedding vectors output by the second exclusive embedding layer into an expert network set, where the expert network set includes multiple expert networks.
[0074] As Figure 3 shown, in the exclusive feature sub-domains, there is a first exclusive embedding layer corresponding to the first exclusive feature sub-domain (i.e., Figure 3 shown, exclusive embedding (coalition)) and a second exclusive embedding layer corresponding to the second exclusive feature sub-domain (i.e., Figure 3 shown, exclusive embedding (master station)). Furthermore, after the first exclusive embedding layer inputs the first exclusive feature, it can convert it into a low-dimensional dense vector; after the second exclusive embedding layer inputs the second exclusive feature, it can convert it into a low-dimensional dense vector.
[0075] In this embodiment, an expert network set (such as MMOE, Multi-gate Mixture-of-Experts) architecture is used for prediction. The exclusive feature domain can include multiple exclusive Embedding layers, and these Embedding vectors are input into a routing layer connected to the expert network set (i.e., Figure 3 the Router shown in Figure 3 , a gating network) for further processing. Expert Networks: A set of shared neural network modules, and each expert network is responsible for extracting specific types of features. Gating Networks: Each task has an independent gating network for selecting appropriate expert networks from the expert network set.
[0076] As an alternative implementation, as described in the foregoing method, the training of the scenario feature domain in the to-be-trained model by the scenario features corresponding to the sample data in the foregoing steps can be achieved through the following steps:
[0077] Train the gating in the scenario feature domain through the scenario features to obtain a weight vector, where the weight vector is used to control the relative importance of different feature domains, and the feature domains include: a general embedding domain, a first exclusive feature domain corresponding to the first exclusive feature in the exclusive feature domain, and a second exclusive feature domain corresponding to the second exclusive feature in the exclusive feature domain.
[0078] Specifically, as Figure 3 shown, the gating is Gate NU. This model will use the scenario features to dynamically adjust the weights of different feature domains (general Embedding domain, first exclusive feature domain, and second exclusive feature domain), so as to better capture the patterns in the data. Train the gating through the scenario features to obtain a weight vector for controlling the relative importance of different feature domains.
[0079] As an alternative implementation, as described in the foregoing method, the method further includes:
[0080] Input the behavior data of the target user into the target model to obtain the predicted results of the ad click-through rate and / or ad conversion rate of the target user in the target scenario predicted by the target model, where the target scenario is the scenario of placing ads on the alliance platform, and the candidate scenario is the scenario of placing ads on the main site platform.
[0081] That is to say, after training the above target model, when the behavior data of the target user is obtained, the data of the target user can be input into the target model, so that the target model can make predictions based on the behavior data to obtain the predicted results of the ad click-through rate and / or ad conversion rate of the target user in the target scenario.
[0082] As an alternative embodiment, for the method described above, the foregoing step S206 may be implemented through the following steps: According to the features corresponding to each sample data and each feature type, train the model to be trained with multiple sample data to obtain a target model:
[0083] Step S402: Train the model to be trained with multiple sample data according to the scenario features corresponding to the sample data to obtain a trained model.
[0084] That is to say, after training the model to be trained with multiple sample data according to the scenario features corresponding to the sample data, a trained model obtained by training the model to be trained can be obtained.
[0085] Step S404: In the case where it is determined that the prediction accuracy of the trained model in the target scenario is lower than the preset requirement, increase the weight of the first feature sub-domain, and train the trained model again. Repeat this cycle until a target model that meets the preset requirement is obtained.
[0086] That is to say, after each training, an evaluation should be carried out to evaluate the prediction accuracy of the trained model. The preset requirement may be the lower limit of the prediction accuracy in the target scenario (for example, 80%, 90%), or it may be that the prediction result does not deviate from the candidate scenario, etc. When it is determined that the prediction accuracy of the trained model in the target scenario is lower than the preset requirement, increase the weight of the first feature sub-domain, and then train the trained model again. Repeat this cycle until a target model that meets the preset requirement is obtained.
[0087] Through the method of this embodiment, when the trained model does not meet the preset requirement, training can be continued so that the finally obtained target model can meet the preset requirement, thereby ensuring the prediction accuracy of the target model in the target scenario.
[0088] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0089] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM (Read-Only Memory), RAM (Random Access Memory), magnetic disk, optical disk), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.
[0090] According to another aspect of the embodiments of the present application, there is also provided a multi-domain data model training device for implementing the above multi-domain data model training method. Figure 5 FIG. is a structural block diagram of an optional multi-domain data model training device according to an embodiment of the present application, as Figure 5 shown. The device may include:
[0091] An acquisition module 51, configured to acquire multiple pieces of sample data in a candidate scenario and / or a target scenario, where the sample data in the candidate scenario is used to supplement the sample data in the target scenario;
[0092] A division module 52, configured to divide each piece of sample data according to a preset feature type to obtain features corresponding to each piece of sample data and each feature type, where the feature type includes: a first exclusive feature in the target scenario, a second exclusive feature in the candidate scenario, and the exclusive feature is a feature that only exists in the corresponding scenario;
[0093] A training module 53, configured to train a model to be trained with multiple pieces of sample data according to the features corresponding to each piece of sample data and each feature type to obtain a target model, where the model to be trained is a model applied to prediction in the target scenario, and the model to be trained includes: a first feature sub-domain trained with the first exclusive feature, and a second feature sub-domain trained with the second exclusive feature, and the weight of the first feature sub-domain is higher than the weight of the second feature sub-domain.
[0094] It should be noted that the acquisition module 51 in this embodiment may be used to execute the above step S202, the division module 52 in this embodiment may be used to execute the above step S204, and the training module 53 in this embodiment may be used to execute the above step S206.
[0095] Through the above-mentioned module, by making the weight of the first feature sub-domain higher than that of the second feature sub-domain, it can be achieved that even if the second exclusive feature in the second feature sub-domain is more than the first exclusive feature in the first feature sub-domain, the trained target model can, under the adjustment of the weight, not be overly affected by the second exclusive feature and learn wrongly, achieving the technical effect of improving the model accuracy, and further solving the problem in the related technology that it is impossible to transfer data and information in other scenarios to the target scenario with relatively insufficient data under the premise of ensuring that the model does not learn wrongly, thereby resulting in a low accuracy of the model in the target scenario.
[0096] The device in this embodiment, in addition to including the above-mentioned module, may further include a module for executing any method in any of the foregoing embodiments of the multi-domain data model training method.
[0097] It should be noted here that the examples and application scenarios implemented by the above-mentioned module and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above-mentioned module, as a part of the device, can run in the hardware environment as shown in Figure 1 and can be implemented by software or by hardware, where the hardware environment includes a network environment.
[0098] According to another aspect of the embodiments of the present application, there is also provided an electronic device for implementing the above-mentioned multi-domain data model training method, and the electronic device may be a server, a terminal, or a combination thereof.
[0099] According to another embodiment of the present application, there is also provided an electronic device, including: as shown in Figure 6 the electronic device may include: a processor 1501, a communication interface 1502, a memory 1503, and a communication bus 1504, where the processor 1501, the communication interface 1502, and the memory 1503 communicate with each other through the communication bus 1504.
[0100] The memory 1503 is used to store a computer program;
[0101] The processor 1501, when executing the program stored on the memory 1503, implements the following steps:
[0102] Step S202, obtaining a plurality of sample data in the candidate scenario and / or the target scenario, where the sample data in the candidate scenario is used to supplement the sample data in the target scenario.
[0103] Step S204: Divide each piece of sample data according to the preset feature types to obtain the features corresponding to each piece of sample data and each feature type. Among them, the feature types include: the first exclusive feature in the target scenario, the second exclusive feature in the candidate scenario, and the exclusive feature is a feature that only exists in the corresponding scenario.
[0104] Step S206: Train the model to be trained with multiple pieces of sample data according to the features corresponding to each piece of sample data and each feature type to obtain the target model. Among them, the model to be trained is a model applied to the prediction in the target scenario. The model to be trained includes: the first feature sub-domain trained with the first exclusive feature, the second feature sub-domain trained with the second exclusive feature, and the weight of the first feature sub-domain is higher than the weight of the second feature sub-domain.
[0105] Optionally, in this embodiment, the above communication bus may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus. The communication interface is used for communication between the above electronic device and other devices.
[0106] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0107] As an example, the above memory 1503 may but is not limited to include the acquisition module 51, the division module 52, and the training module 53 in the above multi-domain data model training device. In addition, it may also include but is not limited to other module units in the above multi-domain data model training device, which will not be elaborated in this example.
[0108] The above-mentioned processor can be a general-purpose processor, including but not limited to: CPU (Central Processing Unit, central processing unit), NP (Network Processor, network processor), etc.; it can also be a DSP (Digital Signal Processor, digital signal processor), ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), FPGA (Field-Programmable Gate Array, field-programmable gate array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0109] The embodiments of the present application also provide a computer-readable storage medium, where the storage medium includes a stored program, and when the program runs, it executes the method steps of the above-mentioned method embodiments.
[0110] Optionally, in this embodiment, the above-mentioned storage medium may include but not be limited to: various media that can store program codes such as USB flash drives, ROMs, RAMs, mobile hard disks, magnetic disks or optical discs.
[0111] The serial numbers of the above-mentioned embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0112] If the integrated unit in the above-mentioned embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in the above-mentioned computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in the storage medium and includes several instructions for causing one or more computer devices (which may be personal computers, servers or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application.
[0113] In the above-mentioned embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0114] In several embodiments provided by the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in electrical or other forms.
[0115] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution provided in this embodiment.
[0116] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0117] The above is only the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A multi-domain data model training method, characterized in that: include: Acquire multiple pieces of sample data in candidate scenes and / or target scenes, wherein the sample data in the candidate scenes are used to supplement sample data in the target scenes; Divide each piece of sample data according to a preset feature type to obtain a feature corresponding to each piece of sample data and each feature type, wherein the feature type includes: a first exclusive feature in a target scene, a second exclusive feature in a candidate scene, and an exclusive feature is a feature that only exists in the corresponding scene; According to the features corresponding to each sample data and each feature type, the model to be trained is trained using the multiple sample data to obtain a target model, wherein the model to be trained is a model applied to prediction in a target scenario, and the model to be trained includes: a first feature domain trained using the first exclusive feature, and a second feature domain trained using the second exclusive feature, wherein a weight of the first feature domain is higher than a weight of the second feature domain.
2. The method according to claim 1, characterized in that The step of dividing each piece of sample data according to a preset feature type to obtain a feature corresponding to each piece of sample data and each feature type includes: The sample data is divided according to preset feature types to obtain scene features corresponding to the scene type, general features corresponding to the general type, first unique features corresponding to the target scene type, and second unique features corresponding to the candidate scene type.
3. The method according to claim 2, characterized in that The training of the to-be-trained model by using the plurality of sample data according to the features corresponding to each sample data and each feature type includes: Training the scene feature domains in the to-be-trained model using the scene features corresponding to the sample data; Training the universal embedding sub-domain in the model to be trained by using the universal features corresponding to the sample data; The exclusive feature domains in the model to be trained are trained using the exclusive features.
4. The method according to claim 3, characterized in that: The exclusive feature domains include: A first exclusive embedding layer for inputting the first exclusive feature and a second exclusive embedding layer for inputting the second exclusive feature; Used to input the embedding vector output by the first exclusive embedding layer and / or the embedding vector output by the second exclusive embedding layer to the routing layer of the expert network set, wherein the expert network set includes multiple expert networks.
5. The method according to claim 3, characterized in that: The training of the scene feature domains in the to-be-trained model by using the scene features corresponding to the sample data includes: The gating in the scene feature domain is trained by the scene feature to obtain a weight vector, wherein the weight vector is used to control the relative importance of different feature domains, and the feature domains include: a common embedding domain, a first exclusive feature domain in the exclusive feature domain corresponding to the first exclusive feature, and a second exclusive feature domain in the exclusive feature domain corresponding to the second exclusive feature.
6. The method according to claim 1, characterized in that The method further comprises: The behavior data of the target user is input into the target model to obtain the advertising click-through rate prediction result and / or the advertising conversion rate prediction result of the target user in the target scenario predicted by the target model, wherein the target scenario is a scenario of advertising on the alliance platform, and the candidate scenario is a scenario of advertising on the main site platform.
7. The method according to any one of claims 1 to 6, characterized in that The step of training the to-be-trained model by using the plurality of sample data according to the features corresponding to each sample data and each feature type to obtain a target model includes: According to the features corresponding to each piece of sample data and each feature type, the to-be-trained model is trained using the plurality of sample data to obtain a trained model; When it is determined that the prediction accuracy of the trained model in the target scenario is lower than the preset requirement, the weight of the first feature domain is increased, and the trained model is trained again, and this cycle is repeated until the target model that meets the preset requirement is obtained through training.
8. A multi-domain data model training device, characterized in that: include: An acquisition module, used to acquire multiple pieces of sample data in a candidate scene and / or a target scene, wherein the sample data in the candidate scene is used to supplement the sample data in the target scene; A division module, used to divide each piece of sample data according to a preset feature type, and obtain a feature corresponding to each piece of sample data and each feature type, wherein the feature type includes: a first exclusive feature in a target scene, a second exclusive feature in a candidate scene, and an exclusive feature is a feature that only exists in the corresponding scene; A training module is used to train a model to be trained using the multiple sample data according to the features corresponding to each sample data and each feature type, to obtain a target model, wherein the model to be trained is a prediction model applied to a target scenario, and the model to be trained includes: a first feature sub-domain trained using the first exclusive feature, and a second feature sub-domain trained using the second exclusive feature, wherein a weight of the first feature sub-domain is higher than a weight of the second feature sub-domain.
9. An electronic device comprising a processor, a communication interface, a memory and a communication bus, wherein: The processor, the communication interface and the memory communicate with each other via the communication bus, wherein: The memory is used to store computer programs; The processor is configured to execute the method according to any one of claims 1 to 7 by running the computer program stored in the memory.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 7 when executed.