Model construction method and electronic device

CN115471279BActive Publication Date: 2026-09-25VIVO SOFTWARE TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211251827.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-13
Publication Date
2026-09-25
Estimated Expiration
2042-10-13

AI Technical Summary

Technical Problem

[0004]本申请实施例的目的是提供一种模型构建方法和电子设备,能够解决模型构建方法往往难以同时满足模型结构简单,且能够缓解负迁移现象的需求的问题

Benefits of technology

[0019]在本申请实施例中,电子设备能够根据包括多个归因节点的关系树,划分多个归因节点,得到N个归因节点组,每个归因节点组包括至少M个归因节点;构建模型的共享向量层,共享向量层包括与N个归因节点组一一对应的N个专家网络,各专家网络在输入样本的情况下,输出各专家网络对应的归因节点组的共享向量;构建模型的向量融合层,向量融合层包括与多个归因节点一一对应的多个融合模块,各融合模块连接至少一个第一专家网络,第一专家网络对应的归因节点组中包括与第一专家网络连接的融合模块对应的归因节点,各融合模块在输入至少一个第一专家网络输出的共享向量的情况下,输出融合向量;构建模型的输出层,输出层包括多个目标塔,多个目标塔与多个融合模块一一对应连接,各目标塔在输入融合模块输出的融合向量的情况下,输出融合模块对应的归因节点的预测值。这样,可以基于多个归因节点的关系树来构建专家网络,从而使各归因节点对应的融合模块与至少一个第一专家网络之间的连接关系是以该关系树为依据,换而言之,每个专家网络输出的共享向量能够输入至具有相关性的一些归因节点对应的融合模块中,从而可以有效缓解负迁移现象。同时,无需每个专家网络与所有归因节点对应的融合模块连接,模型结构更简单。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115471279B_ABST
    Figure CN115471279B_ABST
Patent Text Reader

Abstract

The application discloses a model construction method and an electronic device. The method comprises the following steps: obtaining N attribution node groups according to a relationship tree of a plurality of attribution nodes, each attribution node group comprising at least M attribution nodes; constructing a shared vector layer of a model, the shared vector layer comprising N expert networks, each expert network outputting a shared vector of an attribution node group corresponding to the expert network under the condition that an input sample is input; constructing a vector fusion layer of the model, the vector fusion layer comprising a plurality of fusion modules corresponding to the plurality of attribution nodes one by one, each fusion module being connected to at least one first expert network, and each fusion module outputting a fusion vector under the condition that the shared vectors output by the at least one first expert network are input; and constructing an output layer of the model, the output layer comprising a plurality of target towers, the plurality of target towers being connected to the plurality of fusion modules one by one, and each target tower outputting a predicted value of an attribution node corresponding to the fusion module under the condition that the fusion vector output by the fusion module is input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, specifically relating to a model building method and an electronic device. Background Technology

[0002] In advertising recommendation scenarios, it is often necessary to learn multiple attribution nodes through models. Currently, the construction method for models that learn multiple attribution nodes can adopt a heterogeneous multi-expert network parameter sharing approach. Heterogeneous multi-expert network parameter sharing models typically include multi-gate mixture-of-experts (MMOE) models and customized gate control (CGC) / progressive layered extraction (PLE) models.

[0003] In the MMOE model, each expert network is connected to all target towers in the upper layer, and each target tower needs to use a gate network to learn the weights of different shared vectors connected to it. This easily leads to a rapid increase in the number of parameters in the model, making the model structure difficult to train and converge, and resulting in a complex model structure. The CGC / PLE model uses a shared expert network combined with multiple other expert networks. Besides the shared expert network being connected to all target towers, each of the other expert networks is only randomly and uniformly connected to a subset of target towers. Compared to the MMOE model, the CGC / PLE model has a simpler structure, but the random and uniform connection of the shared vectors output by the expert networks to the target towers has the drawback of negative transfer due to irrelevant attribution nodes. It is evident that the model construction methods in related technologies often fail to simultaneously satisfy the requirements of a simple model structure and the ability to mitigate negative transfer. Summary of the Invention

[0004] The purpose of this application is to provide a model building method and an electronic device that can solve the problem that model building methods often cannot simultaneously meet the requirements of simple model structure and the ability to alleviate negative transfer phenomenon.

[0005] In a first aspect, embodiments of this application provide a model construction method, the method comprising:

[0006] Based on the relation tree containing multiple attribution nodes, divide the attribution nodes into multiple groups to obtain N groups of attribution nodes. Each group of attribution nodes includes at least M attribution nodes, where N and M are integers greater than 1.

[0007] The model is constructed with a shared vector layer, which includes N expert networks that correspond one-to-one with the N attribution node groups. Each expert network outputs the shared vector of the attribution node group corresponding to its input sample.

[0008] The model is constructed with a vector fusion layer, which includes multiple fusion modules that correspond one-to-one with multiple attribution nodes. Each fusion module is connected to at least one first expert network. The attribution node group corresponding to the first expert network includes attribution nodes corresponding to the fusion modules connected to the first expert network. Each fusion module outputs a fusion vector when it receives a shared vector output by at least one first expert network.

[0009] The model's output layer is constructed, which includes multiple target towers. Each target tower is connected to a fusion module in a one-to-one correspondence. Each target tower outputs the predicted value of the attribution node corresponding to the fusion module, given the fusion vector output by the fusion module.

[0010] Secondly, embodiments of this application provide a model building apparatus, the apparatus comprising:

[0011] The partitioning module is used to partition multiple attribution nodes according to the relation tree containing multiple attribution nodes, resulting in N attribution node groups. Each attribution node group includes at least M attribution nodes, where N and M are integers greater than 1.

[0012] The first building module is used to build the shared vector layer of the model. The shared vector layer includes N expert networks that correspond one-to-one with the N attribution node groups. Each expert network outputs the shared vector of the attribution node group corresponding to the input sample.

[0013] The second building module is used to build the vector fusion layer of the model. The vector fusion layer includes multiple fusion modules that correspond one-to-one with multiple attribution nodes. Each fusion module is connected to at least one first expert network. The attribution node group corresponding to the first expert network includes attribution nodes corresponding to the fusion modules connected to the first expert network. Each fusion module outputs a fusion vector when it is input with a shared vector output by at least one first expert network.

[0014] The third building module is used to build the output layer of the model. The output layer includes multiple target towers, which are connected one-to-one with multiple fusion modules. Each target tower outputs the predicted value of the attribution node corresponding to the fusion module, given the fusion vector output by the fusion module.

[0015] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions implementing the steps of the method as described in the first aspect when executed by the processor.

[0016] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, and when the program or instructions are executed by a processor, they implement the steps of the method as described in the first aspect.

[0017] Fifthly, embodiments of this application provide a chip, which includes a processor and a communication interface. The communication interface and the processor are coupled, and the processor is used to run programs or instructions to implement the method as described in the first aspect.

[0018] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method as described in the first aspect.

[0019] In this embodiment, the electronic device can divide multiple attribution nodes into N attribution node groups based on a relation tree including multiple attribution nodes, each attribution node group including at least M attribution nodes; construct a shared vector layer of the model, the shared vector layer including N expert networks corresponding one-to-one with the N attribution node groups, each expert network outputting a shared vector of its corresponding attribution node group when given input samples; construct a vector fusion layer of the model, the vector fusion layer including multiple fusion modules corresponding one-to-one with the multiple attribution nodes, each fusion module being connected to at least one first expert network, the attribution node group corresponding to the first expert network including attribution nodes corresponding to the fusion modules connected to the first expert network, each fusion module outputting a fusion vector when given a shared vector output by at least one first expert network; and construct an output layer of the model, the output layer including multiple target towers, each target tower being connected one-to-one with the multiple fusion modules, each target tower outputting the predicted value of the attribution node corresponding to the fusion module when given a fusion vector output by the fusion module. In this way, an expert network can be constructed based on the relationship tree of multiple attribution nodes. This ensures that the connection between the fusion module corresponding to each attribution node and at least one first expert network is based on this relationship tree. In other words, the shared vector output by each expert network can be input into the fusion modules corresponding to some relevant attribution nodes, thus effectively mitigating negative transfer. Furthermore, it eliminates the need for each expert network to be connected to the fusion modules corresponding to all attribution nodes, resulting in a simpler model structure. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating the model building method provided in an embodiment of this application;

[0021] Figure 2 This is a schematic diagram of the relationship tree in the model construction method provided in the embodiments of this application;

[0022] Figure 3 This is a schematic diagram of a model in the model construction method provided in the embodiments of this application;

[0023] Figure 4 This is a schematic flowchart of a scenario embodiment of the model building method provided in this application;

[0024] Figure 5 This is a schematic diagram of the structure of the model building apparatus provided in the embodiments of this application;

[0025] Figure 6 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application;

[0026] Figure 7 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0028] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0029] The model construction method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0030] Figure 1 This is a flowchart illustrating the model building method provided in an embodiment of this application. The model building method may include:

[0031] Step 101: Based on the relation tree containing multiple attribution nodes, divide the attribution nodes into multiple groups to obtain N groups of attribution nodes. Each group of attribution nodes includes at least M attribution nodes, where N and M are integers greater than 1.

[0032] In step 101, a relational tree including multiple attribution nodes can be pre-constructed. For example, taking ad placement as an example, such as... Figure 2As shown, multiple attribution nodes can include attribution node 1 to attribution node 13, which can correspond to "click", "download", "game registration", "game payment", "custom activation", "second-day retention activation", "custom registration", "custom payment", "custom activation", "form submission", "web purchase", "button click" and "game reservation" respectively.

[0033] A relationship tree can be constructed based on the business category and order of these attribution nodes. For example, "click" (i.e., attribution node 1) can be used as the attribution node corresponding to level "0" in the relationship tree. Then, based on the business category, attribution nodes can be divided into application download category, general webpage category, and game reservation category. Finally, the level of the attribution nodes under each category is determined by combining the business order.

[0034] For example, the attribution nodes corresponding to level "1" under the app download category can include "download" (attribution node 2) and "custom activation" (attribution node 9). Then, depending on the application (APP), the attribution nodes corresponding to levels "3" and "4" under the game APP can be "game registration" (attribution node 3) and "game payment" (attribution node 4), respectively. The attribution nodes corresponding to level "3" under the app can include "custom activation" (attribution node 5), "custom registration" (attribution node 7), and "custom payment" (attribution node 8), where the "custom activation" link can also include "second-day retention activation" (attribution node 6) corresponding to level "4".

[0035] The attribution nodes corresponding to level "1" under the general webpage category can include "form submission" (i.e., attribution node 10), "webpage purchase" (i.e., attribution node 11), and "button click" (i.e., attribution node 12). The attribution nodes corresponding to level "1" under the game reservation category can include "game reservation" (i.e., attribution node 13).

[0036] It is understood that the above attribution nodes are only an example for understanding the technical solutions provided in the embodiments of this application, and are not intended to limit the specific embodiments of this application.

[0037] After constructing the relationship tree, multiple attribution nodes can be divided based on this relationship tree, resulting in N attribution node groups. Each attribution node group includes at least M attribution nodes. The value of M can be set based on empirical values ​​or determined based on the number of levels in the relationship tree; for example, M can be 4.

[0038] Based on business classification, all attribution nodes under the same category in the relationship tree can be identified as an attribution node group. Alternatively, based on business classification and business order, attribution nodes on the same attribution chain can be identified as an attribution node group. For example, attribution nodes 1 to 4 can be considered as one attribution node group, and attribution nodes 1, 2, 5, and 6 can also be considered as one attribution node group.

[0039] In some examples, all attribution nodes, i.e., attribution nodes 1 to 13, can also be treated as one group of attribution nodes.

[0040] Step 102: Construct the shared vector layer of the model. The shared vector layer includes N expert networks that correspond one-to-one with the N attribution node groups. Each expert network outputs the shared vector of the attribution node group corresponding to its input sample.

[0041] In step 102, as Figure 3 As shown, the number of expert networks can be determined based on the number of attribution node groups, thereby constructing the shared vector layer of the model. The shared vector layer can include N expert networks corresponding one-to-one with the N attribution node groups, and each expert network can consist of a single-layer fully connected neural network.

[0042] The input to an expert network can be samples, and the output can be a shared vector of fixed dimensions. The dimension of the shared vector can be set according to the specific situation; for example, the dimension of the shared vector can be 1024, and the shared vector M... i The expression for can be shown in formula (1):

[0043]

[0044] The sample can be obtained by combining information such as users, materials, and contextual features, and then processing these features. Taking advertising as an example, users can refer to those who watch the advertisement, materials can refer to contextual information such as the advertising platform and the format of the advertisement, and contextual features can refer to feature information such as the frequency and weight of the advertisement.

[0045] like Figure 3 As shown, the attribution node group corresponding to expert network 1 can include attribution nodes 1 to 4. Therefore, expert network 1 can output the shared vector of attribution nodes 1 to 4 when given input samples.

[0046] In some examples, the N expert networks may also include a global expert network, whose attribution node group includes all attribution nodes. For example... Figure 3 As shown, the N attribution node groups may also include attribution node group 6 (attribution nodes 1 to 13), with the global expert network corresponding to attribution node 6.

[0047] In this way, by adding a global expert network whose output shared vectors include all attribution nodes, the overfitting phenomenon of other shared vectors during training can be further improved.

[0048] Step 103: Construct the vector fusion layer of the model. The vector fusion layer includes multiple fusion modules that correspond one-to-one with multiple attribution nodes. Each fusion module is connected to at least one first expert network. The attribution node group corresponding to the first expert network includes attribution nodes corresponding to the fusion modules connected to the first expert network. Each fusion module outputs a fusion vector when it receives a shared vector output by at least one first expert network.

[0049] In step 103, each attribution node can correspond to a fusion module, and the vector fusion layer includes multiple fusion modules that correspond one-to-one with multiple attribution nodes. The connection relationship between the expert network and the fusion modules can be determined based on the attribution node group corresponding to the expert network. For example, such as... Figure 3 As shown, each fusion module is connected to at least one first expert network, and the attribution node group corresponding to the first expert network includes the attribution node corresponding to the fusion module connected to the first expert network.

[0050] For example, if fusion module 2 corresponds to attribution node 2, then at least one first expert network can be all expert networks in the corresponding attribution node group that include attribution node 2, i.e., expert network 1, expert network 2, expert network 3, expert network 5, and the global expert network. Similarly, if fusion module 6 corresponds to attribution node 6, then at least one first expert network can be all expert networks in the corresponding attribution node group that include attribution node 6, i.e., expert network 2 and the global expert network. In other words, the attribution node group corresponding to expert network 1 can include attribution nodes 1 to 4. Therefore, expert network 1 can be integrated with fusion module 1 (corresponding to attribution node 1), fusion module 2 (corresponding to attribution node 2), fusion module 3 (corresponding to attribution node 3), and fusion module 4 (corresponding to attribution node 4). This further ensures that the shared vector output by each expert network can be input into the fusion modules corresponding to some relevant attribution nodes, thereby further mitigating the negative transfer phenomenon.

[0051] The fusion module can take shared vectors from at least one first expert network connected to it as input, and output a fusion vector. The fusion vector can be calculated by summing the products of the weights of each shared vector and the weights of the shared vectors. The weights of each shared vector can be obtained by matching the corresponding attribution node group, or calculated based on feature information such as the hierarchy of different attribution nodes in the attribution node group.

[0052] Step 104: Construct the output layer of the model. The output layer includes multiple target towers, which are connected one-to-one with multiple fusion modules. Each target tower outputs the predicted value of the attribution node corresponding to the fusion module, given the fusion vector output by the fusion module.

[0053] In step 104, as Figure 3 As shown, an output layer can also be constructed for the model. A target tower can be built for each attribution node. The output layer can include multiple constructed target towers, which can be connected to the fusion module. After the fusion vector output by the fusion module is input into the target tower, the target tower can output the predicted value of the corresponding attribution node from the fusion module. This predicted value can refer to the trigger rate of the attribution node. For example, if the attribution node is "download", the predicted value could be the probability of clicking "download".

[0054] In this embodiment, the model building method can divide multiple attribution nodes into N attribution node groups based on a relation tree including multiple attribution nodes, each attribution node group including at least M attribution nodes; construct a shared vector layer of the model, which includes N expert networks corresponding one-to-one with the N attribution node groups, each expert network outputting a shared vector of its corresponding attribution node group when given input samples; construct a vector fusion layer of the model, which includes multiple fusion modules corresponding one-to-one with the multiple attribution nodes, each fusion module being connected to at least one first expert network, the attribution node group corresponding to the first expert network including attribution nodes corresponding to the fusion modules connected to the first expert network, each fusion module outputting a fusion vector when given a shared vector output by at least one first expert network; and construct an output layer of the model, which includes multiple target towers, each target tower being connected one-to-one with the multiple fusion modules, each target tower outputting the predicted value of the attribution node corresponding to the fusion module when given a fusion vector output by the fusion module. In this way, an expert network can be constructed based on the relationship tree of multiple attribution nodes. This ensures that the connection between the fusion module corresponding to each attribution node and at least one first expert network is based on this relationship tree. In other words, the shared vector output by each expert network can be input into the fusion modules corresponding to some relevant attribution nodes, thus effectively mitigating negative transfer. Furthermore, it eliminates the need for each expert network to be connected to the fusion modules corresponding to all attribution nodes, resulting in a simpler model structure.

[0055] Optionally, in some embodiments, before step 101 above, the model building method may further include:

[0056] Obtain the business classification and business order of multiple attribution nodes;

[0057] Based on business classification and business order, a relation tree with multiple attribution nodes is constructed. The business classification indicates the horizontal category relationship of the relation tree, and the business order indicates the vertical hierarchical relationship of the relation tree.

[0058] In this embodiment, as Figure 2 As shown, the business categories and business order of multiple attribution nodes can be obtained, and a relationship tree can be constructed based on the business categories and business order. For example, the horizontal category relationships of the relationship tree can be based on business categories, and the vertical hierarchical relationships of the relationship tree can be based on business order.

[0059] As mentioned above, please refer to Figure 2 The attribution node for level "0" in this relationship tree can be "click". The attribution node for level "1" under the app download category can include "download" and "custom activation". Then, the attribution nodes for levels "3" and "4" under the game app category can be "game registration" and "game payment" respectively. The attribution node for level "3" under the app category can include "custom activation", "custom registration", and "custom payment", with "custom activation" further including "second-day retention activation" at level "4". The attribution node for level "1" under the general webpage category can include "form submission", "webpage purchase", and "button click". The attribution node for level "1" under the game reservation category can include "game reservation".

[0060] In other words, taking an advertisement for a game app as an example, its attribution path might be "click - download - game registration - game payment"; taking an advertisement for an application app as an example, its attribution path might be "click - download - custom activation - second-day retention activation".

[0061] In this embodiment, a relationship tree including multiple attribution nodes is constructed based on the business classification and business order of multiple attribution nodes. This can intuitively reflect the correlation between each attribution node and serve as a reference for the connection relationship between the expert network and the fusion module. This can alleviate the negative migration phenomenon caused by irrelevant attribution nodes.

[0062] Optionally, in some embodiments, step 101 above may include:

[0063] Attribution nodes on the same attribution path in the relationship tree are identified as a first group;

[0064] The first group with a number of attribution nodes greater than or equal to M is identified as an attribution node group;

[0065] Merge the second group of the same category in the business classification to obtain the third group. The second group is the first group with less than M attribution nodes.

[0066] The third group with a number of attribution nodes greater than or equal to M is identified as an attribution node group;

[0067] The fourth combination at the same level in the business sequence is combined to obtain the fifth combination. The fourth combination is the third combination with less than M attribution nodes.

[0068] The fifth group with a number of attribution nodes greater than or equal to M is identified as an attribution node group;

[0069] The sixth combination is supplemented by attribution nodes at the same level in the business sequence. The sixth combination is the fifth combination with less than M attribution nodes.

[0070] If the number of attribution nodes in the supplemented sixth group is greater than or equal to M, the supplemented sixth group will be determined as an attribution node group.

[0071] In this embodiment, please refer to Figure 2 First, the attribution links in the relationship tree can be used to determine the first group, thus ensuring that attribution nodes on the same attribution link are grouped in the same shared vector layer. For example, attribution nodes on the same attribution link can be determined as a first group, which may include:

[0072] Group 1 (attribution nodes 1 to 4), Group 2 (attribution nodes 1, 2, 5 and 6), Group 3 (attribution nodes 1, 2 and 7), Group 4 (1, 2 and 8), Group 5 (attribution nodes 1 and 9), Group 6 (attribution nodes 1 and 10), Group 7 (attribution nodes 1 and 11), Group 8 (attribution nodes 1 and 12), and Group 9 (attribution nodes 1 and 13).

[0073] Taking M as an example, the first combination with 4 or more attribution nodes can be identified as an attribution node group. That is, the first combination 1 and the first combination 2 can be identified as attribution node groups respectively.

[0074] The first group with fewer than 4 attribution nodes can be considered the second group. That is, groups 3 through 9 are all second groups. Second groups within the same business category can be merged to obtain the third group. For example, since groups 3 and 4 belong to the "App" business category, they can be merged to obtain the third group 1 (attribution nodes 1, 2, 7, and 8). Since groups 6, 7, and 8 belong to the "Regular URL" business category, they can be merged to obtain the third group 2 (attribution nodes 1, 10, 11, and 12). Since groups 5 and 9 do not have other second groups in the same category, their merging objects can be considered empty, meaning groups 5 and 9 can also be directly considered as the third group.

[0075] A third combination with a number of attribution nodes greater than or equal to M can be identified as an attribution node group. That is, third combination 1 and third combination 2 can be identified as attribution node groups respectively.

[0076] The third combination, where the number of attribution nodes is less than M, can be considered the fourth combination. That is, the first combination 5 and the first combination 9 can be the fourth combination. Then, the fourth combinations at the same level in the business sequence can be concatenated to obtain the fifth combination. Since the first combination 5 and the first combination 9 belong to the same level, the first combination 5 and the first combination 9 can be concatenated to obtain the fifth combination 1 (attribution nodes 1, 9, and 13).

[0077] A fifth combination with a number of attribution nodes greater than or equal to M can be identified as an attribution node group. Conversely, a fifth combination with a number of attribution nodes less than M can be considered a sixth combination. Since the number of attribution nodes in fifth combination 1 is less than 4, fifth combination 1 can be considered the sixth combination. In this case, the sixth combination can be supplemented using attribution nodes at the same level in the business sequence. For example, the target attribution node for supplementing the sixth combination can be determined based on the correlation between the attribution nodes at the same level and the attribution nodes in the sixth combination to be concatenated. For instance, attribution node 2 can be identified as the target attribution node to supplement fifth combination 1. The supplemented sixth combination includes attribution nodes 1, 2, 9, and 13. Since the number of attribution nodes is greater than or equal to M, the supplemented sixth combination can be identified as an attribution node group.

[0078] In the example above, the N attribution node groups can include attribution node group 1 (attribution nodes 1 to 4), attribution node group 2 (attribution nodes 1, 2, 5, and 6), attribution node group 3 (attribution nodes 1, 2, 7, and 8), attribution node group 4 (attribution nodes 1, 10, 11, and 12), and attribution node group 5 (attribution nodes 1, 2, 9, and 13). For example... Figure 3As shown, expert networks 1 to 5 can be constructed that correspond one-to-one with attribution node groups 1 to 5.

[0079] In this embodiment, by merging, splicing, and supplementing, the number of attribution nodes in each attribution node group is ensured, which effectively improves the phenomenon that the shared vector layer is prone to overfitting during training due to the insufficient number of attribution nodes.

[0080] Optionally, in some embodiments, the above-described fusion modules, when given a shared vector output by at least one first expert network, output a fusion vector, which may include:

[0081] Each fusion module performs a first operation to obtain a fusion vector by taking at least one shared vector output by the first expert network as input.

[0082] Output the fused vector;

[0083] The first operation may include:

[0084] Obtain the weight values ​​of each first expert network in at least one first expert network;

[0085] The fusion vector is obtained based on the weight values ​​and the shared vector output by at least one first expert network.

[0086] In this embodiment, the shared vector output by at least one first expert network can be sent to the fusion module. In this case, the fusion module can replace the gate network mechanism to obtain the weight values ​​of each first expert network. It is understood that obtaining the weight values ​​of each first expert network can be achieved by directly matching the most suitable preset weight value as the weight value of each first expert network based on the correspondence between preset weight values ​​and preset attribution node groups, or it can be calculated based on some feature parameters of the attribution node group corresponding to each first expert network. For example, the weight values ​​can be calculated based on the hierarchy of the attribution nodes included in the attribution node group.

[0087] A fusion vector can be obtained and output based on the weight values ​​and a shared vector output by at least one first expert network. The fusion vector M is then output. 融合 The calculation method can be shown in formula (2):

[0088]

[0089] Where W(n,M) i ) represents the shared vector M i The weight value when connected to the fusion module n This represents the set of shared vectors for the input fusion module n.

[0090] For example, taking fusion module 6 as an example, fusion module 6 connects expert network 2 and global expert network, therefore Combining formulas (1) and (2), the expression for the fusion vector output by fusion module 6 is shown in formula (3):

[0091]

[0092] In this way, the fusion module can directly output the fusion vector based on the weight values ​​and shared vectors, without having to learn through complex gate networks, effectively reducing the number of model parameters and making the model structure simpler.

[0093] Optionally, in some embodiments, obtaining the weight values ​​of each first expert network in at least one first expert network may include:

[0094] Based on the first attribution node corresponding to the target fusion module and the second attribution node group corresponding to each target expert network in at least one target expert network, determine the first distance between the target fusion module and each target expert network.

[0095] The weight values ​​of each target expert network are calculated based on the first distance between the target fusion module and each target expert network.

[0096] The target fusion module is any one of multiple fusion modules, and at least one target expert network is connected to the target fusion module.

[0097] In this embodiment, the following description will take the target fusion module as fusion module 6, and at least one target expert network as expert network 2 and global expert network as examples.

[0098] The first distance between the fusion module 6 and the expert network 2 can be determined based on the attribution node group 2 (attribution nodes 1, 2, 5, and 6) corresponding to attribution node 6 and expert network 2. The first distance between the fusion module 6 and the global expert network can also be determined based on the attribution node group 6 (attribution nodes 1 to 13) corresponding to attribution node 6 and global expert network. The formula for calculating the first distance is shown in formula (4):

[0099]

[0100] Where D(n,M) i ) represents the fusion module n and the shared vector M i The first distance, d(n,m), represents the distance between the first attribution node and any attribution node in the second attribution node group. This distance can be determined based on the hierarchy of different attribution nodes.

[0101] The weight values ​​of each target expert network can be calculated based on the first distance between the target fusion module and each target expert network. The formula for calculating the weight values ​​is shown in formula (5):

[0102]

[0103] Where W(n,M) i ) represents the shared vector M i The weight value when connected to the fusion module n D(n,M) represents the set of shared vectors of the input fusion module n. i ) represents the fusion module n and the shared vector M i The first distance.

[0104] In this way, the weight values ​​of each target expert network can be directly calculated based on the first distance between the target fusion module and each target expert network, without the need for learning through complex gate networks, which further simplifies the model.

[0105] Optionally, in some embodiments, determining the first distance between the target fusion module and each target expert network based on the first attribution node corresponding to the target fusion module and the second attribution node group corresponding to each target expert network in at least one target expert network may include:

[0106] The second distance between the first attribution node and each second attribution node is determined based on the level of the first attribution node, the level of each second attribution node in the second attribution node group, and the level of the least common ancestor of the first attribution node and each second attribution node.

[0107] The sum of the second distances between the first attribution node and each of the second attribution nodes is calculated to obtain the first distance between the target fusion module and each target expert network.

[0108] In this embodiment, the second distance between the first attribution node and each of the second attribution nodes can be determined based on the level of the first attribution node, the level of each of the second attribution nodes in the second attribution node group, and the level of the least common ancestor of the first attribution node and each of the second attribution nodes. The formula for calculating the second distance is shown in formula (6):

[0109] d(i,j)=d(root,i)+d(root,j)-2d(root,lca) (6)

[0110] Where d(i,j) represents the second distance between attribution node i and attribution node j, root represents the level, and lca represents the least common ancestor of attribution node i and attribution node j.

[0111] Taking the calculation of the second distance between attribution node 3 and attribution node 5 as an example, the least common ancestor of attribution node 3 and attribution node 5 is attribution node 2, therefore:

[0112] d(3,5)=d(root,3)+d(root,5)-2d(root,2)=2+2-2*1=2

[0113] We can see that the second distance between attribution node 3 and attribution node 5 is 2.

[0114] The sum of the second distances between the first attribution node and each of the second attribution nodes is calculated to obtain the first distance between the target fusion module and each target expert network. The formula for calculating the first distance is shown in formula (4), which will not be elaborated here.

[0115] Based on this, by combining formulas (4), (5) and (6), the weight values ​​of the expert network 2 and the global expert network fusion module 6 can be calculated.

[0116] The first distance between the shared vector M2 output by expert network 2 and fusion module 6 is:

[0117] D(6,M2)=d(6,1)+d(6,2)+d(6,5)+d(6,6)=6

[0118] The shared vector M output by the global expert network share The first distance between the fusion module 6 and the fusion module 6 is:

[0119] D(6,M share )=d(6,1)+d(6,2)+…+d(6,12)+d(6,13)=39

[0120] The weight values ​​when the shared vector M2 is connected to the fusion module 6 are:

[0121]

[0122] Shared vector M share The weight value when connected to fusion module 6 is:

[0123]

[0124] In this embodiment, the second distance between attribution nodes can be calculated based on the hierarchy, and then the first distance between the target fusion module and each target expert network can be calculated based on the second distance. In this way, the weight value of each target expert network can be calculated, which eliminates the need to learn through complex gate networks and further simplifies the model.

[0125] Optionally, in some embodiments, after constructing the output layer as described above, the model construction method may further include:

[0126] The model's loss layer is constructed, which includes multiple loss functions that correspond one-to-one with multiple target towers. Each loss function outputs the loss value of the target tower when given the predicted value of the target tower's output and the true value of the attribution node corresponding to the target tower in the sample.

[0127] Based on the loss value, adjust the neuron parameters on the connection links of the target tower.

[0128] In this embodiment, a loss layer of the model can also be constructed. The loss layer may include multiple loss functions that correspond one-to-one with multiple target towers, wherein the loss function may be a cross-entropy function.

[0129] The predicted value output by the target tower and the true value of the attribution node corresponding to the target tower in the sample can both be input into the loss function. Here, the true value can refer to the conversion rate of the attribution node. For example, if the attribution node is "download", the true value can be the probability that "download" has been successfully completed.

[0130] In some examples, the product of the predicted value and the true value can be calculated to obtain the loss value. If the loss value meets a preset threshold, the neuron parameters inside the model can be considered relatively accurate. If the loss value does not meet the preset threshold, backpropagation can be performed based on the loss value to continuously adjust and optimize the neuron parameters on the connection link of the target tower until the loss value meets the condition.

[0131] In this way, the neuron parameters in the model can be continuously optimized by using the loss value of the loss layer, which effectively improves the accuracy of the model.

[0132] To facilitate understanding of the model building method provided in the above embodiments, the following describes the model building method using a specific scenario embodiment. Figure 4 This is a schematic diagram of a scenario embodiment of the model building method provided in this application.

[0133] This scenario embodiment may include the following steps:

[0134] Step 401, Begin.

[0135] Step 402: Construct a relation tree that includes multiple attribution nodes. For example, the relation tree can be constructed based on the business categories and business order of the multiple attribution nodes.

[0136] Step 403: Divide the multiple attribution nodes to obtain attribution node groups. For example, multiple attribution nodes can be grouped based on the same attribution link, and then processed through combination merging, combination splicing, and combination supplementation to obtain attribution node groups with a sufficient number of attribution nodes.

[0137] Step 404: Construct the shared vector layer of the model. For example, determine the number of expert networks based on the number of attribution node groups, and construct the expert networks accordingly to form a shared vector layer. The input of each expert network is a sample, and the output is a shared vector of fixed dimensions.

[0138] Step 405: Construct the vector fusion layer of the model. For example, a fusion module is constructed for each attribution node to form a vector fusion layer. The connection relationship between the fusion module and the expert network is determined according to the relationship tree. The input of each fusion module is the shared vector output by the expert network connected to it. The weight values ​​of different shared vectors are calculated during fusion, and the fused vector is output after weighted merging.

[0139] Step 406: Construct the output layer of the model. Each attribution node constructs a target tower, which is a fully connected neural network. The target towers are connected to the corresponding fusion modules. The input of each target tower is the fusion vector output by the fusion module connected to it, and the output is the predicted value of that attribution node.

[0140] Step 407: Construct the loss layer of the model. A loss function is constructed for each target tower. The input to the loss function is the predicted value output by the target tower and the corresponding true value in the sample. The output is the loss value, which is used for backpropagation to optimize the neuron parameters in the model.

[0141] In this way, an expert network can be constructed based on a relational tree containing multiple attribution nodes. This ensures that the connection between the fusion module corresponding to each attribution node and at least one first expert network is based on this relational tree. In other words, the shared vector output by each expert network can be input into the fusion modules corresponding to some relevant attribution nodes, effectively mitigating negative transfer. Furthermore, it eliminates the need for each expert network to be connected to the fusion modules corresponding to all attribution nodes, resulting in a simpler model structure.

[0142] The model building method provided in this application can be executed by a model building device. This application uses the model building device executing the model building method as an example to illustrate the model building device provided in this application.

[0143] like Figure 5 As shown in the figure, this application embodiment provides a model building apparatus 500, which may include:

[0144] The partitioning module 501 is used to partition multiple attribution nodes according to the relation tree containing multiple attribution nodes, to obtain N attribution node groups, each attribution node group including at least M attribution nodes, where N and M are integers greater than 1;

[0145] The first construction module 502 is used to construct the shared vector layer of the model. The shared vector layer includes N expert networks that correspond one-to-one with the N attribution node groups. Each expert network outputs the shared vector of the attribution node group corresponding to the input sample.

[0146] The second building module 503 is used to build the vector fusion layer of the model. The vector fusion layer includes multiple fusion modules that correspond one-to-one with multiple attribution nodes. Each fusion module is connected to at least one first expert network. The attribution node group corresponding to the first expert network includes attribution nodes corresponding to the fusion modules connected to the first expert network. Each fusion module outputs a fusion vector when it is input with a shared vector output by at least one first expert network.

[0147] The third building module 504 is used to build the output layer of the model. The output layer includes multiple target towers, which are connected one-to-one with multiple fusion modules. Each target tower outputs the predicted value of the attribution node corresponding to the fusion module when it receives the fusion vector output by the fusion module.

[0148] In this way, an expert network can be constructed based on the relationship tree of multiple attribution nodes. This ensures that the connection between the fusion module corresponding to each attribution node and at least one first expert network is based on this relationship tree. In other words, the shared vector output by each expert network can be input into the fusion modules corresponding to some relevant attribution nodes, thus effectively mitigating negative transfer. Furthermore, it eliminates the need for each expert network to be connected to the fusion modules corresponding to all attribution nodes, resulting in a simpler model structure.

[0149] Optionally, in some embodiments, the model building apparatus 504 may further include:

[0150] The fourth building module is used to build the loss layer of the model. The loss layer includes multiple loss functions that correspond one-to-one with multiple target towers. Each loss function outputs the loss value of the target tower when given the predicted value of the target tower output and the true value of the attribution node corresponding to the target tower in the sample.

[0151] Based on the loss value, adjust the neuron parameters on the connection links of the target tower.

[0152] In this way, the neuron parameters in the model can be continuously optimized by using the loss value of the loss layer, which effectively improves the accuracy of the model.

[0153] Optionally, in some embodiments, the second building module 503 described above can also be used for:

[0154] Each fusion module performs a first operation to obtain a fusion vector by taking at least one shared vector output by the first expert network as input.

[0155] Output the fused vector;

[0156] The first operation may include:

[0157] Obtain the weight values ​​of each first expert network in at least one first expert network;

[0158] The fusion vector is obtained based on the weight values ​​and the shared vector output by at least one first expert network.

[0159] In this way, the fusion module can directly output the fusion vector based on the weight values ​​and shared vectors, without having to learn through complex gate networks, effectively reducing the number of model parameters and making the model structure simpler.

[0160] Optionally, in some embodiments, the second building module 503 described above can also be used for:

[0161] Based on the first attribution node corresponding to the target fusion module and the second attribution node group corresponding to each target expert network in at least one target expert network, determine the first distance between the target fusion module and each target expert network.

[0162] The weight values ​​of each target expert network are calculated based on the first distance between the target fusion module and each target expert network.

[0163] The target fusion module is any one of multiple fusion modules, and at least one target expert network is connected to the target fusion module.

[0164] In this way, the weight values ​​of each target expert network can be directly calculated based on the first distance between the target fusion module and each target expert network, without the need for learning through complex gate networks, which further simplifies the model.

[0165] Optionally, in some embodiments, the second building module 503 described above can also be used for:

[0166] The second distance between the first attribution node and each second attribution node is determined based on the level of the first attribution node, the level of each second attribution node in the second attribution node group, and the level of the least common ancestor of the first attribution node and each second attribution node.

[0167] The sum of the second distances between the first attribution node and each of the second attribution nodes is calculated to obtain the first distance between the target fusion module and each target expert network.

[0168] In this embodiment, the second distance between attribution nodes can be calculated based on the hierarchy, and then the first distance between the target fusion module and each target expert network can be calculated based on the second distance. In this way, the weight value of each target expert network can be calculated, which eliminates the need to learn through complex gate networks and further simplifies the model.

[0169] Optionally, in some embodiments, the model building apparatus 500 may further include:

[0170] The acquisition module is used to acquire the business classification and business order of multiple attribution nodes;

[0171] The relationship tree building module is used to construct a relationship tree with multiple attribution nodes based on business categories and business order. The business category indicates the horizontal category relationship of the relationship tree, and the business order indicates the vertical hierarchical relationship of the relationship tree.

[0172] In this embodiment, a relationship tree including multiple attribution nodes is constructed based on the business classification and business order of multiple attribution nodes. This can intuitively reflect the correlation between each attribution node and serve as a reference for the connection relationship between the expert network and the fusion module. This can alleviate the negative migration phenomenon caused by irrelevant attribution nodes.

[0173] Optionally, in some embodiments, the partitioning module 501 can also be used for:

[0174] Attribution nodes on the same attribution path in the relationship tree are identified as a first group;

[0175] The first group with a number of attribution nodes greater than or equal to M is identified as an attribution node group;

[0176] Merge the second group of the same category in the business classification to obtain the third group. The second group is the first group with less than M attribution nodes.

[0177] The third group with a number of attribution nodes greater than or equal to M is identified as an attribution node group;

[0178] The fourth combination at the same level in the business sequence is combined to obtain the fifth combination. The fourth combination is the third combination with less than M attribution nodes.

[0179] The fifth group with a number of attribution nodes greater than or equal to M is identified as an attribution node group;

[0180] The sixth combination is supplemented by attribution nodes at the same level in the business sequence. The sixth combination is the fifth combination with less than M attribution nodes.

[0181] If the number of attribution nodes in the supplemented sixth group is greater than or equal to M, the supplemented sixth group is determined as an attribution node group.

[0182] In this embodiment, by merging, splicing, and supplementing, the number of attribution nodes in each attribution node group is ensured, which effectively improves the phenomenon that the shared vector layer is prone to overfitting during training due to the insufficient number of attribution nodes.

[0183] The model building device in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television set (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the specific device.

[0184] The model building apparatus in this application embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit it.

[0185] The model building apparatus provided in this application embodiment can achieve... Figures 1 to 4 The various processes implemented in the method implementation examples will not be described again here to avoid repetition.

[0186] Optionally, such as Figure 6 As shown, this application embodiment also provides an electronic device 600, including a processor 601 and a memory 602. The memory 602 stores a program or instructions that can run on the processor 601. When the program or instructions are executed by the processor 601, they implement the various steps of the above-described model construction method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0187] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0188] Figure 7A schematic diagram of the hardware structure of an electronic device to implement an embodiment of this application.

[0189] The electronic device 700 includes, but is not limited to, components such as: radio frequency unit 701, network module 702, audio output unit 703, input unit 704, sensor 705, display unit 706, user input unit 707, interface unit 708, memory 709, and processor 710.

[0190] Those skilled in the art will understand that the electronic device 700 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 710 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 7 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0191] The processor 710 can be used for:

[0192] Based on the relation tree containing multiple attribution nodes, divide the attribution nodes into multiple groups to obtain N groups of attribution nodes. Each group of attribution nodes includes at least M attribution nodes, where N and M are integers greater than 1.

[0193] A shared vector layer is constructed, which includes N expert networks that correspond one-to-one with N attribution node groups. Each expert network outputs the shared vector of its corresponding attribution node group when given input samples.

[0194] The model is constructed with a vector fusion layer, which includes multiple fusion modules that correspond one-to-one with multiple attribution nodes. Each fusion module is connected to at least one first expert network. The attribution node group corresponding to the first expert network includes attribution nodes corresponding to the fusion modules connected to the first expert network. Each fusion module outputs a fusion vector when it receives a shared vector output by at least one first expert network.

[0195] The model's output layer is constructed, which includes multiple target towers. Each target tower is connected to a fusion module in a one-to-one correspondence. Each target tower outputs the predicted value of the attribution node corresponding to the fusion module, given the fusion vector output by the fusion module.

[0196] In this way, an expert network can be constructed based on the relationship tree of multiple attribution nodes. This ensures that the connection between the fusion module corresponding to each attribution node and at least one first expert network is based on this relationship tree. In other words, the shared vector output by each expert network can be input into the fusion modules corresponding to some relevant attribution nodes, thus effectively mitigating negative transfer. Furthermore, it eliminates the need for each expert network to be connected to the fusion modules corresponding to all attribution nodes, resulting in a simpler model structure.

[0197] Alternatively, in some embodiments, the processor 710 may be used for:

[0198] The model's loss layer is constructed, which includes multiple loss functions that correspond one-to-one with multiple target towers. Each loss function outputs the loss value of the target tower when given the predicted value of the target tower's output and the true value of the attribution node corresponding to the target tower in the sample.

[0199] Based on the loss value, adjust the neuron parameters on the connection links of the target tower.

[0200] In this way, the neuron parameters in the model can be continuously optimized by using the loss value of the loss layer, which effectively improves the accuracy of the model.

[0201] Alternatively, in some embodiments, the processor 710 may be used for:

[0202] Each fusion module performs a first operation to obtain a fusion vector by taking at least one shared vector output by the first expert network as input.

[0203] Output the fused vector;

[0204] The first operation may include:

[0205] Obtain the weight values ​​of each first expert network in at least one first expert network;

[0206] The fusion vector is obtained based on the weight values ​​and the shared vector output by at least one first expert network.

[0207] In this way, the fusion module can directly output the fusion vector based on the weight values ​​and shared vectors, without having to learn through complex gate networks, effectively reducing the number of model parameters and making the model structure simpler.

[0208] Alternatively, in some embodiments, the processor 710 may be used for:

[0209] Based on the first attribution node corresponding to the target fusion module and the second attribution node group corresponding to each target expert network in at least one target expert network, determine the first distance between the target fusion module and each target expert network.

[0210] The weight values ​​of each target expert network are calculated based on the first distance between the target fusion module and each target expert network.

[0211] The target fusion module is any one of multiple fusion modules, and at least one target expert network is connected to the target fusion module.

[0212] In this way, the weight values ​​of each target expert network can be directly calculated based on the first distance between the target fusion module and each target expert network, without the need for learning through complex gate networks, which further simplifies the model.

[0213] Alternatively, in some embodiments, the processor 710 may be used for:

[0214] The second distance between the first attribution node and each second attribution node is determined based on the level of the first attribution node, the level of each second attribution node in the second attribution node group, and the level of the least common ancestor of the first attribution node and each second attribution node.

[0215] The sum of the second distances between the first attribution node and each of the second attribution nodes is calculated to obtain the first distance between the target fusion module and each target expert network.

[0216] In this embodiment, the second distance between attribution nodes can be calculated based on the hierarchy, and then the first distance between the target fusion module and each target expert network can be calculated based on the second distance. In this way, the weight value of each target expert network can be calculated, which eliminates the need to learn through complex gate networks and further simplifies the model.

[0217] Optionally, in some embodiments, the radio frequency unit 701 can be used to: acquire the service classification and service order of multiple attribution nodes;

[0218] The processor 710 can be used to: construct a relation tree including multiple attribution nodes based on business classification and business order, wherein business classification indicates the horizontal category relationship of the relation tree and business order indicates the vertical hierarchical relationship of the relation tree.

[0219] In this embodiment, a relationship tree including multiple attribution nodes is constructed based on the business classification and business order of multiple attribution nodes. This can intuitively reflect the correlation between each attribution node and serve as a reference for the connection relationship between the expert network and the fusion module. This can alleviate the negative migration phenomenon caused by irrelevant attribution nodes.

[0220] Alternatively, in some embodiments, the processor 710 may be used for:

[0221] Attribution nodes on the same attribution path in the relationship tree are identified as a first group;

[0222] The first group with a number of attribution nodes greater than or equal to M is identified as an attribution node group;

[0223] Merge the second group of the same category in the business classification to obtain the third group. The second group is the first group with less than M attribution nodes.

[0224] The third group with a number of attribution nodes greater than or equal to M is identified as an attribution node group;

[0225] The fourth combination at the same level in the business sequence is combined to obtain the fifth combination. The fourth combination is the third combination with less than M attribution nodes.

[0226] The fifth group with a number of attribution nodes greater than or equal to M is identified as an attribution node group;

[0227] The sixth combination is supplemented by attribution nodes at the same level in the business sequence. The sixth combination is the fifth combination with less than M attribution nodes.

[0228] If the number of attribution nodes in the supplemented sixth group is greater than or equal to M, the supplemented sixth group is determined as an attribution node group.

[0229] In this embodiment, by merging, splicing and supplementing, the number of attribution nodes in each attribution node group is guaranteed, which effectively improves the phenomenon that the shared vector layer is prone to overfitting during training due to the insufficient number of attribution nodes.

[0230] It should be understood that, in this embodiment, the input unit 704 may include a graphics processing unit (GPU) 7041 and a microphone 7042. The GPU 7041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 706 may include a display panel 7061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 707 includes at least one of a touch panel 7071 and other input devices 7072. The touch panel 7071 is also called a touch screen. The touch panel 7071 may include a touch detection device and a touch controller. Other input devices 7072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0231] The memory 709 can be used to store software programs and various data. The memory 709 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 709 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 709 in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0232] Processor 710 may include one or more processing units; optionally, processor 710 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 710.

[0233] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described model construction method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0234] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0235] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described model construction method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0236] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0237] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the model building method embodiment described above, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0238] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0239] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0240] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for building a model for advertising delivery, characterized in that, include: Based on the relation tree including multiple attribution nodes, the multiple attribution nodes are divided to obtain N attribution node groups, each of the attribution node groups including at least M attribution nodes, where N and M are integers greater than 1; the relation tree is constructed based on the business classification and business order of the multiple attribution nodes, the business classification including attribution nodes for application download category, attribution nodes for ordinary web page category, and attribution nodes for game reservation category; A shared vector layer is constructed for the model. The shared vector layer includes N expert networks that correspond one-to-one with the N attribution node groups. Each expert network outputs a shared vector corresponding to its attribution node group when given input samples. The samples are obtained by feature processing after concatenating user, material, and context features. The user refers to the user who watches the advertisement. The material includes information on the advertising platform and the form of advertising. The context features include information on advertising frequency and advertising weight. A vector fusion layer for constructing the model is provided. The vector fusion layer includes multiple fusion modules that correspond one-to-one with the multiple attribution nodes. Each fusion module is connected to at least one first expert network. The attribution node group corresponding to the first expert network includes attribution nodes corresponding to the fusion modules connected to the first expert network. Each fusion module outputs a fusion vector when it receives a shared vector output by the at least one first expert network. The model's output layer is constructed, which includes multiple target towers. Each target tower is connected to a corresponding fusion module. Each target tower outputs the predicted value of the attribution node corresponding to the fusion module when it receives the fusion vector output by the fusion module.

2. The method according to claim 1, characterized in that, After constructing the output layer of the model, the method further includes: The model's loss layer is constructed, which includes multiple loss functions that correspond one-to-one with the multiple target towers. Each loss function outputs the loss value of the target tower when given the predicted value output by the target tower and the true value of the attribution node corresponding to the target tower in the sample. Based on the loss value, the neuron parameters on the connection links of the target tower are adjusted.

3. The method according to claim 1, characterized in that, Each of the fusion modules outputs a fusion vector upon receiving a shared vector output by at least one first expert network, including: Each of the fusion modules performs a first operation to obtain a fusion vector when it receives a shared vector output by at least one first expert network. Output the fusion vector; The first operation includes: Obtain the weight values ​​of each of the at least one first expert networks; The fusion vector is obtained based on the weight values ​​and the shared vector output by the at least one first expert network.

4. The method according to claim 3, characterized in that, Obtaining the weight values ​​of each of the at least one first expert networks includes: Based on the first attribution node corresponding to the target fusion module and the second attribution node group corresponding to each of the target expert networks in the at least one target expert network, a first distance between the target fusion module and each of the target expert networks is determined. The weight values ​​of each target expert network are calculated based on the first distance between the target fusion module and each target expert network. The target fusion module is any one of the plurality of fusion modules, and the at least one target expert network is connected to the target fusion module.

5. The method according to claim 4, characterized in that, The step of determining the first distance between the target fusion module and each of the target expert networks based on the first attribution node corresponding to the target fusion module and the second attribution node group corresponding to each of the at least one target expert network includes: The second distance between the first attribution node and each of the second attribution nodes is determined based on the level of the first attribution node, the level of each of the second attribution nodes in the second attribution node group, and the level of the least common ancestor of the first attribution node and each of the second attribution nodes. The sum of the second distances between the first attribution node and each of the second attribution nodes is calculated to obtain the first distance between the target fusion module and each of the target expert networks.

6. The method according to claim 1, characterized in that, Before dividing the multiple attribution nodes into N groups based on the relation tree including multiple attribution nodes, the method further includes: Obtain the business classification and business order of the multiple attribution nodes; Based on the business classification and the business order, a relation tree including the multiple attribution nodes is constructed, wherein the business classification indicates the horizontal category relationship of the relation tree, and the business order indicates the vertical hierarchical relationship of the relation tree.

7. The method according to claim 6, characterized in that, The process involves dividing the multiple attribution nodes into N attribution node groups based on a relational tree comprising multiple attribution nodes. Each attribution node group includes at least M attribution nodes, including: The attribution nodes on the same attribution link in the relationship tree are identified as a first combination; The first group with a number of attribution nodes greater than or equal to M is defined as an attribution node group. Merge the second group of the same category in the business classification to obtain the third group, where the second group is the first group whose number of attribution nodes is less than M; The third combination with a number of attribution nodes greater than or equal to M is determined as an attribution node group; By concatenating the fourth combination at the same level in the business sequence, a fifth combination is obtained, wherein the fourth combination is the third combination in which the number of attribution nodes is less than M; The fifth combination with a number of attribution nodes greater than or equal to M is defined as an attribution node group. The sixth combination is supplemented by attribution nodes at the same level in the business sequence, and the sixth combination is the fifth combination with less than M attribution nodes; If the number of attribution nodes in the supplemented sixth group is greater than or equal to M, the supplemented sixth group is determined as an attribution node group.

8. A model building device for advertising placement, characterized in that, include: The partitioning module is used to partition the multiple attribution nodes according to the relation tree including multiple attribution nodes, to obtain N attribution node groups, each attribution node group including at least M attribution nodes, where N and M are integers greater than 1; the relation tree is constructed based on the business classification and business order of the multiple attribution nodes, and the business classification includes attribution nodes for application download category, attribution nodes for ordinary web page category, and attribution nodes for game reservation category; The first construction module is used to construct the shared vector layer of the model. The shared vector layer includes N expert networks that correspond one-to-one with the N attribution node groups. Each expert network outputs the shared vector of the attribution node group corresponding to the input sample. The sample is obtained by feature processing after splicing user, material, and context features. The user is the user who watches the advertisement. The material includes information on the advertising platform and the form of advertising. The context features include information on advertising frequency and advertising weight. The second construction module is used to construct the vector fusion layer of the model. The vector fusion layer includes multiple fusion modules that correspond one-to-one with the multiple attribution nodes. Each fusion module is connected to at least one first expert network. The attribution node group corresponding to the first expert network includes attribution nodes corresponding to the fusion modules connected to the first expert network. Each fusion module outputs a fusion vector when it is input with the shared vector output by the at least one first expert network. The third construction module is used to construct the output layer of the model. The output layer includes multiple target towers, which are connected one-to-one with the multiple fusion modules. Each target tower outputs the predicted value of the attribution node corresponding to the fusion module when it is input with the fusion vector output by the fusion module.

9. The apparatus according to claim 8, characterized in that, Also includes: The fourth construction module is used to construct the loss layer of the model. The loss layer includes multiple loss functions that correspond one-to-one with the multiple target towers. Each loss function outputs the loss value of the target tower when it is given the predicted value output by the target tower and the true value of the attribution node corresponding to the target tower in the sample. Based on the loss value, the neuron parameters on the connection links of the target tower are adjusted.

10. The apparatus according to claim 8, characterized in that, The second building module is also used for: Each of the fusion modules performs a first operation to obtain a fusion vector when it receives a shared vector output by at least one first expert network. Output the fusion vector; The first operation includes: Obtain the weight values ​​of each of the at least one first expert networks; The fusion vector is obtained based on the weight values ​​and the shared vector output by the at least one first expert network.

11. The apparatus according to claim 10, characterized in that, The second building module is also used for: Based on the first attribution node corresponding to the target fusion module and the second attribution node group corresponding to each of the target expert networks in the at least one target expert network, a first distance between the target fusion module and each of the target expert networks is determined. The weight values ​​of each target expert network are calculated based on the first distance between the target fusion module and each target expert network. The target fusion module is any one of the plurality of fusion modules, and the at least one target expert network is connected to the target fusion module.

12. The apparatus according to claim 11, characterized in that, The second building module is also used for: The second distance between the first attribution node and each of the second attribution nodes is determined based on the level of the first attribution node, the level of each of the second attribution nodes in the second attribution node group, and the level of the least common ancestor of the first attribution node and each of the second attribution nodes. The sum of the second distances between the first attribution node and each of the second attribution nodes is calculated to obtain the first distance between the target fusion module and each of the target expert networks.

13. The apparatus according to claim 8, characterized in that, Also includes: The acquisition module is used to acquire the business classification and business order of the multiple attribution nodes; A relationship tree construction module is used to construct a relationship tree including the multiple attribution nodes according to the business classification and the business order, wherein the business classification indicates the horizontal category relationship of the relationship tree, and the business order indicates the vertical hierarchical relationship of the relationship tree.

14. The apparatus according to claim 13, characterized in that, The partitioning module is also used for: The attribution nodes on the same attribution link in the relationship tree are identified as a first combination; The first group with a number of attribution nodes greater than or equal to M is defined as an attribution node group. Merge the second group of the same category in the business classification to obtain the third group, where the second group is the first group whose number of attribution nodes is less than M; The third combination with a number of attribution nodes greater than or equal to M is determined as an attribution node group; By concatenating the fourth combination at the same level in the business sequence, a fifth combination is obtained, wherein the fourth combination is the third combination in which the number of attribution nodes is less than M; The fifth combination with a number of attribution nodes greater than or equal to M is defined as an attribution node group. The sixth combination is supplemented by attribution nodes at the same level in the business sequence, and the sixth combination is the fifth combination with less than M attribution nodes; If the number of attribution nodes in the supplemented sixth group is greater than or equal to M, the supplemented sixth group is determined as an attribution node group.

15. An electronic device, characterized in that, It includes a processor and a memory, the memory storing a program or instructions that can run on the processor, the program or instructions being executed by the processor to implement the steps of the method as described in any one of claims 1-7.

16. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-task processing and model training method and device, medium and equipment

    CN114282681A

  • Recommendation method, training method and device

    CN114997412A