Data processing method and device, equipment, medium and program product

By introducing threshold gating networks into the multitasking model, filtering and dynamically selecting target processing subnets, the problem of poor computing performance and processing effects in Internet business processing is solved, and more efficient and accurate data processing is achieved.

CN120047190APending Publication Date: 2025-05-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311604648.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-27
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When handling Internet services, traditional multitasking models have problems such as large computing volume, poor performance and poor processing effects, making it difficult to take into account both computing performance and processing effects.

Method used

Using a data processing method based on a multi-task model, by predicting the weight distribution of multiple processing subnets, the target processing subnet for performing tasks is selected, and a second weight distribution is generated based on its weight value, and finally these processing subnets are called to perform tasks to predict the processing results.

Benefits of technology

It improves the data processing performance and effectiveness of Internet services, ensures the accuracy of task processing results, and avoids the problems of computing redundancy and low performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047190A_ABST
    Figure CN120047190A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a data processing method and device, equipment, a medium and a program product. The method comprises the following steps: predicting first weight distribution of N processing sub-networks according to task data of a to-be-processed task; based on the data volume of the task data, screening out M target processing sub-networks for executing tasks from the N processing sub-networks; generating second weight distribution of the M target processing sub-networks according to the first weight value corresponding to each target processing sub-network; and calling the M target processing sub-networks to execute the task respectively, and predicting a processing result of the task based on the second weight distribution and task execution results obtained by the M target processing sub-networks respectively. By adopting the embodiment of the invention, the calculation performance and the processing effect of the task can be considered; the embodiment of the invention can be applied to internet business scenes such as advertisements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, in particular to the field of artificial intelligence, and specifically to a data processing method based on a multi-task model, a data processing device, a computer equipment, a computer-readable storage medium, and a computer program product. Background Art

[0002] With the rapid development of Internet technology, the data processing requirements of various Internet services are becoming more and more complex.

[0003] At present, it is supported to use multi-task models to process data for Internet services to meet business needs. However, when traditional multi-task models are used to process data for Internet services, there are not only problems of large computational complexity and poor performance, but also problems such as poor processing effects and negative impacts on processing results. Therefore, how to improve the processing effect and performance for Internet services has become a research hotspot. Summary of the invention

[0004] The embodiments of the present application provide a data processing method, apparatus, device, medium and program product based on a multi-task model, which can take into account both computing performance and processing effect when processing data for tasks in Internet services.

[0005] On the one hand, an embodiment of the present application provides a data processing method based on a multi-task model, wherein the multi-task model includes N processing sub-networks, where N is an integer greater than 1; the method includes:

[0006] Predicting first weight distributions of N processing subnetworks according to task data of the task to be processed; the first weight distributions include first weight values ​​corresponding to each processing subnetwork, and the first weight values ​​are used to indicate the degree of influence of a task execution result obtained by executing the task by the corresponding processing subnetwork on a processing result of the task;

[0007] Based on the amount of task data, M target processing subnetworks for executing the task are selected from N processing subnetworks; the first weight value corresponding to each target processing subnetwork is greater than the target threshold gating; M is a positive integer, and M is less than or equal to N;

[0008] Generate a second weight distribution of M target processing subnetworks according to the first weight value corresponding to each target processing subnetwork; the second weight distribution includes the second weight value corresponding to each target processing subnetwork;

[0009] The M target processing sub-networks are called to respectively execute tasks, and based on the second weight distribution and the task execution results respectively obtained by the M target processing sub-networks, the processing results of the tasks are predicted.

[0010] On the other hand, an embodiment of the present application provides a data processing device based on a multi-task model, wherein the multi-task model includes N processing sub-networks, where N is an integer greater than 1; the device includes:

[0011] A prediction unit, used to predict a first weight distribution of N processing subnetworks according to task data of the task to be processed; the first weight distribution includes a first weight value corresponding to each processing subnetwork, and the first weight value is used to indicate the degree of influence of the task execution result obtained by the corresponding processing subnetwork executing the task on the processing result of the task;

[0012] A processing unit, configured to select M target processing subnetworks for executing the task from N processing subnetworks based on the amount of task data; the first weight value corresponding to each target processing subnetwork is greater than the target threshold gating; M is a positive integer, and M is less than or equal to N;

[0013] The processing unit is further used to generate a second weight distribution of the M target processing subnetworks according to the first weight value corresponding to each target processing subnetwork; the second weight distribution includes the second weight value corresponding to each target processing subnetwork;

[0014] The processing unit is further used to call the M target processing sub-networks to respectively execute tasks, and predict the processing results of the tasks based on the second weight distribution and the task execution results respectively obtained by the M target processing sub-networks.

[0015] In one implementation, the processing unit is used to select M target processing subnetworks for executing the task from N processing subnetworks based on the data volume of the task data, specifically to:

[0016] The task data is dimensionally transformed to obtain a target threshold gating that matches the data volume of the task data; the target threshold gating is used to indicate: the minimum degree of influence of the task execution result obtained by the processing sub-network executing the task on the task processing result;

[0017] The target threshold gating is used to perform threshold truncation processing on the first weight distribution, and M target processing subnetworks for performing the task are screened out from the N processing subnetworks.

[0018] In one implementation, the processing unit is used to perform dimension transformation processing on the task data to obtain a target threshold gating that matches the data volume of the task data, specifically for:

[0019] Obtaining a distribution characteristic of the first weight distribution; the distribution characteristic is used to indicate the distribution of the N first weight values ​​in the weight range;

[0020] According to the data volume of the task data and the distribution characteristics of the first weight distribution, the task data is subjected to a first dimension transformation process to obtain an initial threshold gating; the dimension of the initial threshold gating is one-dimensional;

[0021] The initial threshold gating is subjected to threshold conversion processing to obtain a target threshold gating that matches the data volume of the task data.

[0022] In one implementation, the processing unit is used to perform threshold conversion processing on the initial threshold gating to obtain a target threshold gating that matches the data volume of the task data, specifically for:

[0023] The target conversion function is used to perform threshold compression processing on the initial threshold gating to obtain the intermediate threshold gating; the target conversion function is obtained by fine-tuning the normalization function;

[0024] Screening out a maximum first weight value from the first weight distribution, and comparing the maximum first weight value with an intermediate threshold gating to obtain a comparison result;

[0025] According to the comparison result, a target threshold gating matching the data volume of the task data is determined;

[0026] Among them, when the comparison result indicates that the maximum first weight value is less than the intermediate threshold gating, the target threshold gating is the maximum first weight value; when the comparison result indicates that the maximum first weight value is greater than the intermediate threshold gating, the target threshold gating is the intermediate threshold gating.

[0027] In one implementation, the processing unit is used to perform threshold truncation processing on the first weight distribution by using target threshold gating, and when M target processing subnetworks for performing the task are screened out from N processing subnetworks, specifically for:

[0028] Subtract each first weight value in the first weight distribution from the target threshold gate to obtain N subtraction results;

[0029] Filter out the subtraction results with positive values ​​from the N subtraction results;

[0030] Select a processing subnetwork corresponding to a subtraction result with a positive value from the N processing subnetworks as a target processing subnetwork for executing the task;

[0031] The number M of networks of the target processing subnetwork is the same as the number of subtraction results that are positive numbers.

[0032] In one implementation, the processing unit is used to generate the second weight distribution of the M target processing subnetworks according to the first weight value corresponding to each target processing subnetwork, specifically to:

[0033] Obtaining difference information corresponding to each target processing subnetwork; the difference information is obtained by subtracting the first weight value corresponding to the corresponding target processing subnetwork from the target threshold gating;

[0034] Based on the difference information corresponding to each target processing subnetwork, a second weight distribution of the M target processing subnetworks is generated.

[0035] In one implementation, the processing unit is used to generate the second weight distribution of the M target processing subnetworks based on the difference information corresponding to each target processing subnetwork, specifically to:

[0036] The difference information corresponding to the M target processing sub-networks is normalized to generate a second weight distribution of the M target processing sub-networks.

[0037] In one implementation, the processing unit is used to predict the processing result of the task based on the second weight distribution and the task execution results respectively obtained by the M target processing subnetworks, specifically to:

[0038] M target processing sub-networks are used to process the task data of the task respectively, and the task execution result of each target processing sub-network for the task is obtained;

[0039] Performing weighted operations on the task execution results of each target processing subnetwork for the task and the corresponding second weight values ​​in the second weight distribution, respectively, to obtain weighted results corresponding to each target processing subnetwork;

[0040] The weighted results corresponding to each target processing sub-network are summed to obtain the processing result of the task.

[0041] In one implementation, each of the N processing subnetworks has a different processing capability for different tasks; the processing unit is used to predict the first weight distribution of the N processing subnetworks according to the task data of the task to be processed, specifically to:

[0042] According to the matching relationship between the task and the processing capability of each of the N processing subnetworks, the task data is transformed in the second dimension to obtain a gating vector; the dimension of the gating vector is N, and each element in the gating vector corresponds to a processing subnetwork;

[0043] The gating vector is normalized to obtain the first weight distribution of the N processing sub-networks.

[0044] In one implementation, the task is a task generated when an Internet service has a data processing requirement; the Internet service includes a multimedia data recommendation service and a search service;

[0045] Wherein, when the Internet service is a multimedia data recommendation service, the task corresponding to the multimedia data recommendation service includes a task of predicting data indicators for multimedia data; the data indicators include at least one of the following: a click rate indicator, a conversion rate indicator, and a resource value estimation indicator;

[0046] When the task is to predict a click rate indicator for multimedia data, the processing result of the task includes the click rate prediction result corresponding to the multimedia data; or, when the task is to predict a conversion rate indicator for multimedia data, the processing result of the task includes the conversion rate prediction result corresponding to the multimedia data; or, when the task is to predict a resource value indicator for multimedia data, the processing result of the task includes the resource value prediction result corresponding to the multimedia data.

[0047] On the other hand, an embodiment of the present application provides a computer device, the device comprising:

[0048] A processor for loading and executing computer programs;

[0049] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the data processing method is implemented.

[0050] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and the computer program is suitable for being loaded by a processor and executing the above-mentioned data processing method.

[0051] On the other hand, an embodiment of the present application provides a computer program product, which includes computer instructions, and when the computer instructions are executed by a processor, the above-mentioned data processing method is implemented.

[0052] In the embodiment of the present application, after obtaining the task data of the task to be processed, the first weight distribution of N processing subnetworks in the multi-task model can be predicted according to the task data, and the first weight distribution includes the first weight value corresponding to each processing subnetwork. It is also possible to screen M target processing subnetworks for performing the task from N processing subnetworks based on the data volume of the task data, and the first weight value corresponding to each target processing subnetwork is greater than the target threshold gating. In this way, the second weight distribution of M target processing subnetworks can be generated based on the first weight value corresponding to each target processing subnetwork, and the second weight value corresponding to each target processing subnetwork distribution in the second weight distribution. Finally, calling the M target processing subnetworks screened for the task and their second weight values ​​can realize the prediction for the task and obtain the processing result of the task. It can be seen from the above scheme that the embodiment of the present application can flexibly screen the target processing subnetwork in the multi-task model for the task based on the data volume of the task data of the task. On the one hand, the number of selected target processing subnetworks matches the amount of task data, which enables the selection of an appropriate number of target processing subnetworks to perform the task, avoiding problems such as computational redundancy, poor computational performance, and high computational overhead caused by too many selected processing subnetworks. On the other hand, the first weight value of the selected target processing subnetwork is greater than the target threshold gating, which enables the use of a target processing subnetwork that has a greater impact on the task to perform the task, thereby ensuring the accuracy of the task processing result, that is, ensuring the task processing effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0054] Figure 1a is a schematic diagram of a hybrid expert network provided by an exemplary embodiment of the present application;

[0055] Figure 1b is a schematic diagram of another hybrid expert network provided by an exemplary embodiment of the present application;

[0056] Figure 2a is a schematic diagram of an existing N-Gate solution;

[0057] Figure 2b It is a schematic diagram of an existing TopK-Gate solution;

[0058] Figure 2cis a schematic diagram of a Threshold-Gate solution provided by an exemplary embodiment of the present application;

[0059] Figure 3 is a schematic diagram of an Internet service scenario provided by an exemplary embodiment of the present application;

[0060] Figure 4 It is a flowchart of a data processing method provided by an exemplary embodiment of the present application;

[0061] Figure 5 is a schematic diagram of a composition of task data provided by an exemplary embodiment of the present application;

[0062] Figure 6 is a schematic diagram of a weight distribution provided by an exemplary embodiment of the present application;

[0063] Figure 7 is a schematic diagram of the calculation logic of a Threshold-Gate solution provided by an exemplary embodiment of the present application;

[0064] Figure 8 is a structural schematic diagram of a data processing device provided by an exemplary embodiment of the present application;

[0065] Fig. 9 It is a structural diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0066] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0067] In the embodiment of the present application, a data processing scheme based on a multi-task model is proposed; the multi-task model belongs to the field of artificial intelligence (AI). Among them, artificial intelligence is the theory, method, technology and application system of simulating, extending and expanding human intelligence, perceiving the environment, acquiring knowledge and using knowledge to obtain the best results using digital computers or machines controlled by digital computers. In other words, artificial intelligence is a comprehensive technology of computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning and decision-making. Artificial intelligence technology is a comprehensive discipline involving a wide range of fields, including both hardware-level technology and software-level technology. Basic artificial intelligence technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction system, mechatronics, etc. Among them, pre-trained models are also called large models and basic models, which can be widely used in downstream tasks in various major directions of artificial intelligence after fine-tuning. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0068] Further, the multi-task model involved in the embodiment of the present application belongs to the machine learning / deep learning direction in the field of artificial intelligence. Among them, machine learning (Machine Learning, ML) is a multi-field cross-disciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. Specializes in how computers simulate or implement human learning behavior to acquire new knowledge or skills, and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all fields of artificial intelligence. Machine learning and deep learning generally include technologies such as neural networks (such as artificial neural networks), belief networks, reinforcement learning, transfer learning, inductive learning, and learning by teaching.

[0069] Furthermore, the multi-task model provided in the embodiment of the present application is an ensemble learning technology developed in the field of neural networks (NN). In detail, the multi-task model can be called a mixture of experts (Mixture of Expert, MoE), which is a deep learning model controlled by sparse gates; the deep learning model can be specifically understood as a modeling method that can solve multiple task problems, that is, the mixed expert network can simultaneously realize predictive processing for multiple tasks and improve the processing efficiency of multiple tasks. For the sake of ease of explanation, no distinction will be made between multi-task models and mixed expert networks in the following.

[0070] The hybrid expert network can be expressed as MoE[1,2], which is mainly composed of a group of processing subnetworks and at least one gated network. Among them: ① The processing subnetwork (Expert) can be called an expert model or an expert network, etc.; the processing subnetwork is generally a multilayer perceptron (Multilayer Perceptron, MLP) structure, and the MLP structure is a feedforward artificial neural network model, which mainly maps multiple input data sets to a single output data set; in the hybrid expert network, the processing subnetwork is mainly used to perform task processing on the input data. In addition, the processing subnetwork can be shared (such as used in a multi-task learning network (Multi-gate Mixture-of-Experts, MMOE) structure to realize multiple tasks sharing an expert network), or it can be task-specific (such as used in a personal network (Personal Learning Environment, PLE) structure); and the embodiment of the present application does not limit the number of processing subnetworks included in a group of processing subnetworks. ② The gated network (Gate) can be called a gated model, which acts as a router and selects a set of experts (Experts) for task processing for each task's task data (or input data); specifically, it supports receiving the input data of the task and then outputs the weight distribution; the weight distribution represents the contribution of the processing sub-network selected by the gated network to the processing of the input data. For example, if the hybrid expert network has two processing sub-networks, and the weight values ​​predicted by the gated network for each processing sub-network are distributed as 0.7 and 0.3, then it means that one of the two processing sub-networks contributes 70% to the processing of the input data, while the other processing sub-network contributes 30% to the processing of the input data.

[0071] For example, a schematic diagram of a hybrid expert network including one gating network and multiple gating networks (such as two gating networks) can be seen in Figure 1a and Figure 1b .like Figure 1aAs shown, assuming that the hybrid expert network includes three expert networks and one gating network, and the hybrid expert network is used to realize the parallel processing of two tasks (such as task A and task B); then Figure 1a The structure shown can be described from bottom to top as follows: the hybrid expert network obtains input data, which includes task data of task A and task data of task B (that is, the task data of the two tasks can be combined as input data, such as the data volume of task data of task A in the input data accounts for 80%, while the data volume of task data of task B accounts for 20%). Then, the role of the MoE layer in the hybrid expert network (including a gate network Gate and three expert networks Expert) is to set different weight values ​​for different tasks and perform weighted sum operations; specifically: the gate network in the MoE layer can generate weight values ​​for different expert networks according to the different processing capabilities (or expertise) of different expert networks for different tasks, and each expert network can process the data of task A and task B in the input data respectively to obtain the task execution results of each expert network for each task; then each expert network performs a weighted sum operation on the task execution results of each task and the corresponding weight values. Finally, the results of the weighted sum operation can be input into the task modules corresponding to different tasks (which can be referred to as towers in the embodiments of the present application) according to the task type. The structure of the tower is mainly a multi-layer perceptron structure, including independent parameters of the corresponding tasks; in this way, the output of the task results of the tasks can be realized by the towers corresponding to different tasks.

[0072] Similarly, if Figure 1b As shown, assuming that the hybrid expert network includes 3 expert networks and 2 gating networks, and the hybrid expert network is used to realize the parallel processing of two tasks (such as task A and task B); then Figure 1b The structure shown can be described from bottom to top as follows: the hybrid expert network obtains input data, which includes task data of task A and task data of task B. Then, the gated network A of the two gated networks included in the MoE layer in the hybrid expert network can generate weight values ​​for the processing of task A by the three expert networks, and the gated network B can generate weight values ​​for the processing of task B by the three experts; and each expert network can process the data of task A and task B in the input data respectively to obtain the task execution results of each expert network for each task respectively; then each expert network performs a weighted sum operation on the task execution result of task A and the corresponding weight value output by the gated network A, and each expert network performs a weighted sum operation on the task execution result of task B and the corresponding weight value output by the gated network B. Finally, the result of the weighted sum operation for task A can be input to the task module A corresponding to task A, and similarly, the result of the weighted sum operation for task B can be input to the task module A corresponding to task B, so that the task modules corresponding to different tasks can realize the output of the task results of the tasks.

[0073] It should be understood that: ① The embodiments of the present application do not limit the number of expert networks and gated networks in the hybrid expert network. For ease of explanation, the hybrid expert network includes N processing sub-networks, N is an integer greater than 1, and includes a gated network as an example for introduction; the data processing solution provided in the embodiments of the present application can be easily expanded from one gated network to multiple gated networks, that is, the embodiments of the present application are also applicable to hybrid expert networks with multiple gated networks. ② The embodiments of the present application do not limit the neural networks included in each module (such as a gated network, an expert network or a tower) in the hybrid expert network; such as the number of hidden layers, the number of neurons, the depth of the neural network (increasing or decreasing the number of hidden layers) and the width of the neural network (increasing or decreasing the number of neurons in each hidden layer), as well as the type of neural network, are not limited.

[0074] In practical applications, when using a multi-task model to process tasks in Internet services, there are problems such as poor computing performance (i.e., high computing overhead, or many resources required for computing), or poor processing effect (such as inaccurate processing results for tasks). In order to make the multi-task model take into account both computing performance and processing effect when processing tasks, the data processing solution based on the multi-task model proposed in the embodiment of the present application introduces a threshold gated network Threshold-Gate; by introducing Threshold-Gate in the multi-task model and through end-to-end learning, the multi-task model can adaptively calculate the number of processing sub-networks required for different tasks, that is, automatically and dynamically select some target processing sub-networks for different tasks to perform tasks according to different tasks, so as to achieve the purpose of improving the model calculation efficiency while maintaining the processing effect of the task. In other words, by introducing Threshold-Gate in the multi-task model, a reasonable number of Experts can be selected for each task, and the number of Experts selected for different tasks can be different.

[0075] For the convenience of explanation, the following takes the multi-task model for task processing for a single task as an example to introduce the data processing scheme based on the multi-task model after the threshold gating network Threshold-Gate is introduced. The general process of the data processing scheme may include: first, after obtaining the task data of the task to be processed, the threshold gating network in the multi-task model can be used to predict the first weight distribution of the N processing subnetworks in the multi-task model for the task; at this time, the first weight distribution includes the first weight value corresponding to each processing subnetwork, and any first weight value is used to indicate the degree of influence of the task execution result obtained by the processing subnetwork corresponding to any first weight value when performing the task on the processing result of the task, specifically, when N processing subnetworks are used to perform the task, the degree of influence of the task execution result of the corresponding processing subnetwork on the processing result of the task. Then, based on the data volume of the task data of the task, M target processing subnetworks for performing the task can be screened from the N processing subnetworks, and the first weight value corresponding to each target processing subnetwork in the M target processing subnetworks is greater than the target threshold gating. Then, according to the first weight value corresponding to each target processing subnetwork, a second weight distribution of M target processing subnetworks can be generated, and the second weight distribution includes the second weight corresponding to each target processing subnetwork in the M target processing subnetworks; any second weight value is used to indicate the degree of influence of the task execution result obtained by the processing subnetwork corresponding to any second weight value in executing the task on the processing result of the task, specifically, when M target processing subnetworks are used to execute the task, the degree of influence of the task execution result of the corresponding target processing subnetwork in executing the task on the processing result of the task. Finally, the M target processing subnetworks are called to execute the task respectively, and the processing result of the task is predicted based on the second weight distribution and the task execution results obtained by the M target processing subnetworks respectively.

[0076] It has been found in practice that the data processing solution provided by the embodiment of the present application has obvious advantages in predicting tasks. The following is an example of comparing the solution of the present application with the existing mainstream multi-task model solution to illustrate the advantages of the embodiment of the present application, where:

[0077] The existing mainstream solutions based on multi-task models can be divided into two types of access, including: N-Gate solution and TopK-Gate solution. Among them, when the N-Gate solution uses the multi-task model for task prediction, it supports selecting all processing sub-networks in the multi-task model to perform tasks, such as Figure 2aThe N processing subnetworks included in the multi-task model shown are all used to perform tasks; at this time, the gating network needs to calculate the corresponding weight value for each expert, which makes the amount of calculation from all processing subnetworks and gating networks very large, with a large computational overhead, reducing computing performance. When the TopK-Gate solution uses a multi-task model for task prediction, it supports selecting K processing subnetworks from the N processing subnetworks included in the multi-task model to perform tasks, where K is a fixed value; if K = 2, Figure 2b Two processing subnetworks (e.g., processing subnetwork 2 and processing subnetwork N are selected) out of the N processing subnetworks in the multi-task model shown are selected to perform the task. However, for task data with a large sample size (or a large amount of data), a larger K value is often required, that is, more parameters are required to fit well, while for task data with a small sample size, only less than K processing subnetworks are required; this makes the fixed K value unable to adapt to task data with different data amounts, such as the fixed K value is too redundant for task data with a small amount of data, and insufficient for task data with a large amount of data, thereby causing a negative impact on the processing effect of the task and greatly affecting the accuracy of task prediction.

[0078] However, the embodiment of the present application introduces a threshold gating network Threshold-Gate in the multi-task model. The threshold gating network can automatically select a reasonable number of target processing subnetworks for different tasks, and the first weight values ​​corresponding to the selected target processing subnetworks are all greater than the target threshold gating. On the one hand, an appropriate number of target processing subnetworks are selected to perform tasks (that is, a certain number of target processing subnetworks that match the data volume of the task data of the task are used to perform the task) to avoid problems such as computational redundancy, poor computational performance, and high computational overhead caused by too many selected processing subnetworks. On the other hand, using a target processing subnetwork that has a greater impact on the task to perform the task can ensure the accuracy of the task processing results, that is, ensure the task processing effect. Figure 2c As shown, assuming that the data volume of task data x1 of task A is smaller than the data volume of task data x2 of task B, then when selecting target processing subnetworks for task A and task B respectively, the number of target processing subnetworks selected for task A may be smaller than the number of target processing subnetworks selected for task B; for example, processing subnetwork 2 and processing subnetwork N are selected for task A, while processing subnetwork 1, processing subnetwork 2, and processing subnetwork 3 are selected for task B.

[0079] The data processing scheme provided in the embodiment of the present application can be implemented based on an Internet business model, and the Internet business model can be deployed in a computer device used to execute Internet business; the Internet business model here can be a model that combines the multi-task model described above with a specific Internet business, and the comparison is not limited. When an Internet business has a data processing demand, a task corresponding to the data processing demand can be generated. At this time, the computer device can call the Internet business model deployed with the data processing scheme to automatically implement business processing for the business. For the sake of ease of explanation, the data processing scheme involved in the embodiment of the present application will be introduced later with the computer device as the execution subject.

[0080] Among them, computer devices may include but are not limited to: smart phones (such as Android phones, iOS phones, etc.), tablets, personal computers, portable personal computers, mobile Internet devices (Mobile Internet Devices, MID for short), smart TVs, car-mounted devices, head-mounted devices and other smart devices that can touch the screen. Applications that provide Internet services (such as advertising applications, audio and video applications or car-mounted applications, etc.) can be run in computer devices; applications may include but are not limited to: clients installed and running in computer devices; installation-free applications, that is, applications that can be used without downloading and installing, such applications are also commonly known as mini-programs, which are usually run in clients as sub-programs; or web applications opened through browsers; etc. Computer devices may also include servers for providing background computing and application service support for Internet services; servers may be independent physical servers, or server clusters or distributed systems composed of multiple physical servers, or cloud servers that provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal and the server may be connected directly or indirectly via wired or wireless communication, which is not limited in this application.

[0081] In detail, when the above-mentioned Internet services have data processing needs, corresponding tasks can be generated, that is, the tasks involved in the embodiments of the present application are tasks generated when the Internet services have data processing needs. The embodiments of the present application do not limit the business types of the above-mentioned Internet services, and the task types of tasks in the Internet services. Exemplarily, Internet services may include multimedia data recommendation services and search services, etc.; among which, multimedia data recommendation services can be simply understood as services that actively or passively recommend multimedia data to users; search services can be simply understood as services that recommend multiple outputs related to the input to users based on the user's input, such as services that answer user questions during artificial intelligence question-and-answer sessions, and services that search for knowledge through search engines, etc.

[0082] The following takes the multimedia data recommendation service as an example to briefly introduce the Internet service and the tasks corresponding to the Internet service. Multimedia data may include but is not limited to: advertisements, videos, audio (such as music), text, and images, etc.; accordingly, multimedia data recommendation services may include but are not limited to: advertisement recommendation services, audio and video recommendation services (such as short video recommendations), text recommendation services (such as novel recommendations), and image recommendation services, etc. When the Internet service is a multimedia data recommendation service, the tasks corresponding to the multimedia data recommendation service may include but are not limited to: tasks for predicting data indicators for multimedia data; the data indicators here include at least one of the following: click-through rate indicators, conversion rate indicators, and resource value estimation indicators, etc. For example: when the task is to predict the click rate index for multimedia data, the processing result of the task using the Internet business model provided in the embodiment of the present application includes the click rate prediction result corresponding to the multimedia data; or, when the task is to predict the conversion rate index for multimedia data, the processing result of the task using the Internet business model provided in the embodiment of the present application includes the conversion rate prediction result corresponding to the multimedia data; or, when the task is to predict the resource value index for multimedia data, the processing result of the task using the Internet business model provided in the embodiment of the present application includes the resource value prediction result corresponding to the multimedia data.

[0083] Taking the multimedia data recommendation service as an advertisement recommendation service as an example, the advertisement recommendation service can be applied to other application scenarios, such as implementing advertisement recommendation in film and television drama scenarios, and implementing advertisement recommendation in game scenarios, etc.; the tasks corresponding to the advertisement recommendation service can include tasks of estimating click-through rate for advertisements, or tasks of estimating conversion rate for advertisements, etc. Moreover, when the multimedia data recommendation service is an advertisement recommendation service, the Internet business model can be an advertisement ranking model; according to the tasks corresponding to the advertisement recommendation service, the advertisement ranking model here can include a recall model, a rough ranking model, and a fine ranking model, etc.; the fine ranking model can include a pCTR (Click-Through Rate Prediction, click-through rate prediction) model, a pCVR (Predict Conversion Rate, conversion rate prediction) model, a pDCVR (Predict Deep-Conversion Rate, deep conversion rate prediction) model, and a pLTV (Predict Life-Time Value, long-term value prediction) model, etc.

[0084] An exemplary Internet service is a multimedia data recommendation service. For example, a system for an advertisement recommendation service can be found in Figure 3 .like Figure 3 In the advertising recommendation system shown (such as recommending advertisements in a web page), it is assumed that the advertising recommendation model (specifically, it can be a conversion rate prediction model, i.e., a pCVR model) is deployed in the server 301; when the server 301 needs to recommend advertisements to users based on the advertisement conversion rate, it determines that the task to be processed is the task of predicting the conversion rate index of the advertisement. Then, the object characteristics of the user (such as age, gender, province, and interests) and the advertisement characteristics of the advertisement to be recommended (such as advertisement ID and advertisement category) can be obtained. The object characteristics and advertisement characteristics here together constitute the task data of the task; then, the server 301 uses the threshold gating scheme (i.e., data processing scheme) provided by the embodiment of the present application to deploy the advertising recommendation model to perform task processing on the task data, and the processing result of the task obtained is the conversion rate of the user for the advertisement, such as the possibility of the user downloading, registering, or paying for the advertisement. In this way, after the server 301 estimates the conversion rates between multiple advertisements and users respectively, it can push multiple advertisements with higher estimated conversion rates to the terminal held by the user, thereby recommending more preferred advertisements to the user. It can be seen that when the data processing solution provided in the embodiment of the present application is applied to the advertising recommendation business, its functional characteristics are to improve the accuracy of advertising recommendations; its performance characteristics are to improve the training and estimation efficiency of advertising recommendation models (such as conversion rate estimation models), so as to achieve the purpose of balancing task processing effects and computing performance.

[0085] It should be noted that the above Figure 3This is only an exemplary system of the advertisement recommendation system; in actual applications, the advertisement recommendation system may also include other servers or terminals, and the embodiments of the present application do not limit the number and types of terminals and servers.

[0086] It should also be noted that the collection and processing of relevant data in the embodiments of this application should be strictly in accordance with the requirements of relevant laws and regulations. The acquisition of personal information requires the knowledge or consent of the individual subject (or the legal basis for obtaining the information), and the subsequent use and processing of data shall be carried out within the scope of authorization of laws and regulations and the subject of personal information. For example, when the embodiments of this application are applied to specific products or technologies, such as in the advertising recommendation scenario, if the user's object characteristics are to be obtained, then the user's permission or consent is required, and the collection, use and processing of relevant data (such as the collection and release of the barrage posted by the object, etc.) need to comply with the relevant laws, regulations and standards of the relevant regions.

[0087] Based on the data processing scheme described above, the embodiment of the present application proposes a more detailed data processing method. The data processing method proposed in the embodiment of the present application will be described in detail below in conjunction with the accompanying drawings. Figure 4 A flow chart of a data processing method provided by an exemplary embodiment of the present application is shown; the data processing method may be executed by the aforementioned computer device; the data processing method may include but is not limited to steps S401-S404:

[0088] S401: Predicting first weight distributions of N processing sub-networks according to task data of a task to be processed.

[0089] The task data of the task to be processed is some data required to perform the task; the task data of the task may vary depending on the Internet business to which the task to be processed belongs, or the task type of different tasks under the same Internet business. For example, if the task to be processed is a business of estimating conversion rate under the advertising recommendation business, then the task data of the task may be obtained based on the object characteristics of the user (such as some characteristics of the user's identity) and the advertising characteristics of the advertisement; for another example, if the task to be processed is a search business, then the task data of the task may be obtained based on the content characteristics of the user input and the content characteristics of the content to be searched; and so on.

[0090] Take the task to be processed as an example of an advertising recommendation business, specifically, a conversion rate estimation task generated when the advertising recommendation business has a conversion rate estimation requirement; then considering that the conversion rate is the probability of evaluating the user user performing a conversion behavior after clicking an ad (advertisement) (such as downloading the application rendered by the ad after clicking the ad, or paying for the ad after clicking the ad, or registering after clicking the ad, etc.), therefore, the task data of this task can be obtained based on the user's object features and the advertisement features of the advertisement. Taking the pCVR model as an example of the multi-task model corresponding to this task, the pCVR model can be expressed as:

[0091] pCVR=f(x=<user,ad> )

[0092] Among them, f(x=<user,ad> ) is the estimated conversion rate, x =<user,ad> It represents input data (i.e., task data), user represents the object features of the user, and ad represents the advertising features of the advertisement.

[0093] It can be seen that the task data of the conversion rate estimation task can be composed of two parts of features, which can include object features and advertising features. For the sake of convenience, the object features of the user include: age, gender, province and interest, and the advertising features of the advertisement include: advertising identification (or ID) and advertising category. For example, the object features and advertising features can be expressed in natural language; for example, the object features are expressed as: the user's age is 25 years old, the gender is male, the province is Province A, and the interest is games, and the advertising features are expressed as: the advertising identification is 123 and the advertising category is stand-alone games. Then, the user's object features and the advertising features of the advertisement are respectively represented by vectors, specifically, each value of these features (such as the value of the user's age is 25 years old, the value of the gender is male) corresponds to a learnable vector parameter; then, by representing the values ​​of these features by vectors, the feature vector corresponding to each feature (specifically the value of the feature) can be obtained; the feature vector corresponding to any feature can be used to represent the semantics or meaning of any feature, etc. Among them, the vector representation can be implemented by using but not limited to the word embedding model, etc. The word embedding is a general term for a set of language modeling and feature learning technologies in natural language processing (NLP), which can be used to map characters or strings expressed in natural language to vectors of real numbers. Finally, in order to better express the feature vectors of each feature, the feature vectors of each feature are directly converted into vectors of fixed size, and then the feature vectors of the same length of each feature are spliced, etc., to obtain the task data of the task. In short, the task data of the task is obtained by vectorizing each feature and performing conversion and splicing on the feature vectors after the vector representation; the task data of the task can be represented as a multi-dimensional (or simply multi-dimensional) vector.

[0094] A schematic diagram of an exemplary task data for constructing a task can be found in Figure 5 ;like Figure 5 As shown, the user's age feature (specifically, the value of the age feature "25 years old"), gender feature (specifically, the value of the gender feature "male"), province feature (specifically, the value of the province feature "Province A"), and interest feature (specifically, the value of the interest feature "game") are respectively represented by vectors to obtain four feature vectors of user-side features; similarly, the advertisement identification feature (specifically, the value of the advertisement identification feature "123") and advertisement type feature (specifically, the value of the advertisement type feature "stand-alone game") are respectively represented by vectors to obtain two feature vectors of advertisement-side features. Then, the 6 feature vectors obtained by the above vector representation are spliced ​​to obtain the task data of the task.

[0095] It should be understood that the above examples only use age, gender, province and interest as object features, and only use ad logo and ad category as main ad features; in actual application scenarios, there are many object features, ad features and even context features used to train models, and it is impossible to cite them one by one. However, for new features not described, the method of inputting them into the multi-task model can be promoted and used with reference to the embodiments of this application.

[0096] Further, after the multi-task model is trained, each of the N processing subnetworks included in it has different processing capabilities for different tasks, specifically, each processing subnetwork can be trained to learn the processing capabilities of different task data during the model training stage. For example, when the processing subnetwork has a strong processing capability for the task data of a certain type of task, it means that the processing subnetwork is better at predicting this type of task, so in the model application stage, the processing subnetwork has a greater degree of influence on the processing results of this type of task. Based on this, after the embodiment of the present application obtains the task data of the task to be processed based on the above-mentioned related description, the threshold gated network in the multi-task model can be used to predict the first weight distribution of the N processing subnetworks for the task through the task data. Among them, the first weight distribution includes the first weight value corresponding to each processing subnetwork in the multi-task model, and the first weight value can be used to indicate: the corresponding processing subnetwork performs the task The degree of influence of the task execution result obtained by the task on the processing result of the task; or, the first weight value can be used to indicate that if N processing subnetworks are used to perform the task, then the corresponding processing subnetwork in the N processing subnetworks contributes to the task. The greater the degree of influence of the task execution result obtained by the processing sub-network on the processing result of the task, the greater the processing sub-network has on the task, and the greater the first weight value corresponding to the processing sub-network. It can be seen that by learning the processing capabilities of each processing sub-network for different tasks in the multi-task model, it can ensure that the same multi-task model can have good processing capabilities for different types of tasks, thereby improving the scalability and model performance of the multi-task model.

[0097] In a specific implementation, after obtaining the task data of the task to be processed, the task data can be transformed in the second dimension according to the matching relationship between the task and the processing capabilities of the N processing subnetworks in the multi-task model to obtain a gating vector. Among them: ① The matching relationship between the task and the processing capabilities of the processing subnetwork mentioned above is learned by the processing subnetwork in the model training stage. Specifically, by training the processing subnetwork in the multi-task model, the processing subnetwork can have a better processing capability for the task data of a certain type of task, that is, the processing subnetwork is better at processing the task data of the task; therefore, in the model application stage, if the task data of the task to be processed is similar to the data characteristics or structural type of the task data of a certain type of task in the model training stage, then it is determined that the processing subnetwork in the multi-task model has a better processing capability for the task, which is the processing subnetwork that is better at processing the task data of the certain type of task. ② The dimension of the gating vector obtained after the second dimensional transformation is N, and each dimension / element in the gating vector corresponds to a processing subnetwork. That is to say, the second dimensional transformation can map the multi-dimensional task data into an N-dimensional gating vector that matches the processing capability of each processing sub-network in the multi-task model according to the processing capability of the N processing sub-networks for the task. The distribution characteristics of each element in the N-dimensional gating vector reflect the proficiency of each processing sub-network in the multi-task model for the task; that is, the influence of the task processing results of each processing sub-network in the multi-task model on the processing results of the task can be analyzed through the N-dimensional gating vector. Among them, when the processing sub-network is better at processing the task, the task execution results of the processing sub-network performing the task have a greater influence on the processing results of the task. Conversely, when the processing sub-network is less good at processing the task, the task execution results of the processing sub-network performing the task have a smaller influence on the processing results of the task.

[0098] In order to better characterize the proficiency of each processing sub-network in the multi-task model in the task, the embodiment of the present application, after performing a second-dimensional transformation on the task data of the task to obtain an N-dimensional gating vector, also supports normalizing the gating vector to obtain the first weight distribution of the N processing sub-networks; specifically, each dimensional element in the gating vector is mapped to the range of [0,1], so that the degree of influence of the task execution results obtained by each processing sub-network on the processing results of the task can be characterized in the form of weight values. Among them, the normalization process can be simply understood as converting N-dimensional data into a number between [0,1], and its purpose is to cancel the order of magnitude difference between the data of each dimension, and avoid problems such as excessive network prediction errors caused by large differences in the order of magnitude of input and output data.

[0099] A schematic diagram of the processing capabilities of the processing subnetwork for different tasks in an exemplary multi-task model can be found in Figure 6 .like Figure 6 As shown, it is assumed that the multi-task model includes processing sub-network 1, processing sub-network 2 and processing sub-network 3, and the multi-task model can process task A and task B, and the task data of the task to be processed include task data of task A and task data of task B respectively. If the multi-task model is used to predict the task data of the task to be processed, the first weight distribution of the three processing sub-networks for task A is (0.6, 0.1, 0.3), and the first weight distribution of the three processing sub-networks for task B is (0.1, 0.7, 0.2), indicating that: processing sub-network 1 is better at processing task A, processing sub-network 2 is better at processing task B, and processing sub-network 3 is better at processing task A. Of course, the processing ability of each processing sub-network in the multi-task model for different tasks is obtained by training the multi-task model with rich sample tasks so that each processing sub-network can learn.

[0100] S402: Based on the data volume of the task data, M target processing sub-networks for executing the task are selected from the N processing sub-networks.

[0101] In order to enable the multi-task model to have both the processing effect and computing performance for the task when executing the task, the embodiment of the present application supports analyzing the data volume of the task data of the task and the first weight distribution of the N processing subnetworks for the task, and selecting a suitable number of target processing subnetworks with higher weight values ​​for the task; wherein, the data volume of the task data can be understood as the storage capacity occupied by storing the task data, and the unit of the data volume can be bytes or bits, etc. In this way, the computing overhead caused by the execution of the task by all processing subnetworks is avoided, and the task can be executed based on the partial processing subnetworks with higher weight values ​​to ensure the processing effect of the task.

[0102] In a specific implementation, the general process of selecting a target processing subnetwork for a task based on the data volume of the task data in the embodiment of the present application is as follows: first, a target threshold gating is determined for the task based on the data volume of the task data and the distribution characteristics of the first weight value; then, based on the target threshold gating, a suitable number M of target processing subnetworks are selected for the task from the N processing subnetworks. For example, N=3, and the first weight distribution of the N processing subnetworks is (0.6, 0.1, 0.3). If the target threshold gating is 0.2, it means that the minimum weight value selected from the three processing subnetworks for executing the task is 0.2, then the two processing subnetworks corresponding to 0.6 and 0.3 can be selected from the first weight distribution as the processing subnetworks for executing the task; that is, the two processing subnetworks corresponding to 0.6 and 0.3 are the target processing subnetworks that match the data volume of the task data.

[0103] The specific implementation process of the above-mentioned screening target processing sub-network may include but is not limited to the following steps (1)-(2), wherein:

[0104] (1) Performing a dimension transformation process on the task data to obtain a target threshold gating that matches the data volume of the task data; the purpose of the dimension transformation process here is to determine a threshold gating for the task based on the data volume of the task data and the first weight distribution, so that the threshold gating can be used to indicate: the task execution result obtained by the processing subnetwork executing the task, the minimum degree of influence on the processing result of the task, that is, the target threshold gating indicates the minimum weight value that should be selected from the N processing subnetworks to execute the task.

[0105] Specifically, the specific process of the dimensional transformation processing may include: obtaining the distribution characteristics of the first weight distribution of the N processing subnetworks for the task in the multi-task model, and the distribution characteristics are used to indicate the distribution of the N first weight values ​​in the weight range (such as 0 to 1). For example, the first weight distribution is (0.9, 0.09, 0.01), and the distribution characteristics of the first weight distribution can be used to indicate that the distribution of the three first weight values ​​between (0, 1) is that one first weight value is close to 1, and the remaining two first weight values ​​are close to 0; of course, the above is only an exemplary description of the distribution situation, and the distribution situation can also be described as a large difference between one first weight value and the remaining two weight values ​​in the three first weight values. Then, according to the data volume of the task data and the distribution characteristics of the first weight distribution, the task data is subjected to a first dimensional transformation process to obtain an initial threshold gating, and the dimension of the initial threshold gating is one-dimensional. In other words, the embodiment of the present application supports converting multi-dimensional task data into a one-dimensional initial threshold gating through linear or nonlinear processing. Finally, in order to compare the initial threshold gating with the first weight value in the first weight distribution, it is also necessary to perform threshold conversion processing on the initial threshold gating to obtain a target threshold gating that matches the data volume of the task data.

[0106] The specific process of threshold conversion processing may include: using the target conversion function to perform threshold compression processing on the initial threshold gating to obtain the intermediate threshold gating; the target conversion function is obtained by fine-tuning the normalization function, which can make the intermediate threshold gating compressed based on the target conversion function fall into a smaller weight range, thereby avoiding the problem that the target processing sub-network cannot be selected or the number of selected target processing sub-networks is insufficient due to the large value of the intermediate threshold gating. The normalization function may include linear normalization, nonlinear normalization and standardized normalization, etc.; specifically including Log normalization function, Sigmoid normalization function, σ normalization function and Softmax normalization function, etc. Then, in order to ensure that at least one processing subnetwork can be selected from the N processing subnetworks as the target processing subnetwork to perform the task, the embodiment of the present application also needs to ensure that the value of the intermediate threshold gating must be less than or equal to the maximum first weight value in the first weight distribution, so that at least one target processing subnetwork can be selected for the task; based on this, after obtaining the intermediate threshold gating, the embodiment of the present application needs to filter out the maximum first weight value from the first weight distribution, and compare the maximum first weight value with the intermediate threshold gating to obtain a comparison result. Finally, determine the target threshold gating that matches the data volume of the task data based on the comparison result.

[0107] Among them, when the comparison result indicates that the largest first weight value in the first weight distribution is less than the intermediate threshold gating, it indicates that: when the intermediate threshold gating is directly used to select the target processing subnetwork, the processing subnetwork whose first weight value is greater than the intermediate threshold gating cannot be selected from the N processing subnetworks, that is, there is no processing subnetwork that can be used to perform the task in the N processing subnetworks; at this time, the intermediate threshold gating is abandoned as the target threshold gating to select the target processing subnetwork, but the target threshold gating is set to the largest first weight value. In this way, it can be ensured that at least one target processing subnetwork can be selected from the N processing subnetworks to perform the task, ensuring the smooth execution of the task. Conversely, when the comparison result indicates that the largest first weight value in the first weight distribution is greater than the intermediate threshold gating, it indicates that: when the intermediate threshold gating is directly used to select the target processing subnetwork, the processing subnetwork whose first weight value is greater than the intermediate threshold gating can be first selected from the N processing subnetworks, and the target threshold gating is directly set to the intermediate threshold gating.

[0108] (2) The target threshold gating is used to perform threshold truncation processing on the first weight distribution, and M target processing subnetworks for performing the task are selected from the N processing subnetworks. Among them, the threshold truncation processing mainly selects one or more processing subnetworks whose first weight value is greater than the target threshold gating from the N processing subnetworks as the target processing subnetworks for performing the task; for different tasks, the target threshold gating and the first weight distribution are different, so the number M of target processing subnetworks selected from the N processing subnetworks is different, such as the number M of target processing subnetworks is the same as the number of subtraction results whose values ​​are positive numbers. In other words, the embodiment of the present application supports the flexible selection of an appropriate number of target processing subnetworks for a task based on the task data of the task, avoiding problems such as large performance or insufficient number of target processing subnetworks for performing tasks, improving the processing effect of the task, and ensuring that the computing performance is not too large.

[0109] Specifically, the embodiment of the present application supports subtracting each first weight value in the first weight distribution from the target threshold gate, and obtaining N subtraction results; then, the subtraction results with positive values ​​are screened out from the N subtraction results, and the processing subnetwork corresponding to the subtraction result with positive values ​​is screened out from the N processing subnetworks as the target processing subnetwork for executing the task. In other words, when any first weight value in the first weight distribution is subtracted from the target threshold gate, if the subtraction result is a positive number, it means that any first weight value is greater than the target threshold gate, that is, the processing subnetwork corresponding to any first weight value has a higher processing capability for the task, then the processing subnetwork is selected as the target processing subnetwork for executing the task; on the contrary, if the subtraction result is a negative value, it means that any first weight value is less than the target threshold gate, that is, the processing subnetwork corresponding to any first weight value has a lower processing capability for the task, then the subtraction result can be set to zero, and the selection of the processing subnetwork to execute the task is abandoned. It is worth noting that, considering that the subtraction result of the first weight value and the target threshold gate may be zero, and when the subtraction result is subsequently normalized, the zero value may be used as the denominator; therefore, the embodiment of the present application supports subtracting each first weight value from the target threshold gate respectively, and then adding a very small positive value to the subtraction result to avoid the denominator taking a zero value during the normalization process.

[0110] It can be seen that the embodiment of the present application innovatively improves the traditional gating method in N processing subnetworks into a threshold gating method (i.e., a method for dynamically / flexibly screening target processing subnetworks based on target threshold gating), while maintaining a sufficient number of processing subnetworks to perform tasks, while reducing the calculation of non-important and repeated information in the N processing subnetworks (i.e., reducing the computational overhead caused by the execution of the task by the processing subnetworks with weak / low processing capabilities), so that this solution has both processing effect and computing performance when processing tasks. In addition, considering that the advertising ranking model under the advertising recommendation business is a large-scale storage and sparse computing characteristic of trillion-dimensional parameters, the threshold gating method provided by the embodiment of the present application can better meet the sparse computing requirements, improve the fitting accuracy and computing efficiency of the N processing subnetworks, and improve the efficiency and quality of advertising recommendations.

[0111] S403: Generate a second weight distribution of M target processing sub-networks according to the first weight value corresponding to each target processing sub-network.

[0112] Considering that the sum of the M first weight values ​​corresponding to the selected M target processing subnetworks is less than 1, in order to obtain the contribution of each target processing subnetwork to the task when the M target processing subnetworks are used to perform the task; the embodiment of the present application also supports generating the second weight distribution of the M target processing subnetworks based on the first weight values ​​corresponding to the M target processing subnetworks. Among them, the second weight distribution includes the second weight value corresponding to each target processing subnetwork, and the second weight value is used to indicate the degree of influence of the task execution result obtained by the corresponding target processing subnetwork performing the task on the processing result of the task, specifically indicating the contribution of the corresponding target processing subnetwork to the task if the M target processing subnetworks perform the task.

[0113] In the specific implementation, it supports obtaining the difference information corresponding to each target processing subnetwork in the M target processing subnetworks, and the difference information refers to the subtraction result mentioned above, that is, the difference information corresponding to any target processing subnetwork is obtained by subtracting the first weight value corresponding to the any target processing subnetwork and the target threshold gating. As described above, in order to avoid the situation where the denominator is zero when generating the second weight distribution, the difference information corresponding to any target processing subnetwork in the embodiment of the present application has the result of subtracting the first weight value corresponding to the any target processing subnetwork and the target threshold gating plus a very small positive value. Then, based on the difference information corresponding to each target processing subnetwork, the second weight distribution of M target processing subnetworks is generated; specifically, the difference information corresponding to the N target processing subnetworks is normalized to generate the second weight distribution of M target processing subnetworks.

[0114] S404: Calling the M target processing sub-networks to respectively execute tasks, and predicting the processing results of the tasks based on the second weight distribution and the task execution results respectively obtained by the M target processing sub-networks.

[0115] After screening M target processing subnetworks and the second weight distribution for the task based on the aforementioned steps S401-S403, the embodiment of the present application can adopt the M target processing subnetworks to realize the task processing for the task, and obtain the processing result of the task. Specifically, the task data of the task is input into the multi-task model, and the M target processing subnetworks selected for the task in the multi-task model are used to perform task processing on the task data of the task respectively, and the task execution result of each target processing subnetwork for the task is obtained. Then, the task execution result (i.e., the task execution result of each target processing subnetwork output) is weighted with the corresponding second weight value in the second weight distribution, respectively, to obtain the weighted result corresponding to each target processing subnetwork. Finally, the weighted result corresponding to each target processing subnetwork is summed to obtain the processing result of the task.

[0116] To facilitate a better understanding of the process of the complete data processing method shown in the above steps S401-S404, the following formula is used to illustrate the calculation logic of the threshold gating method Threshold-Gate introduced in the embodiment of the present application to execute the task. Specifically, assuming that the task data of the task is represented by x, the calculation logic may at least include:

[0117] Step 1: Given the task data x of the task and the hyperparameters β and γ of the multi-task model, flexibly select M target processing subnetworks for the task and calculate the second weight distribution of the M target processing subnetworks. Specifically, it includes:

[0118] ①Use the model parameter 1 of the threshold gating network to perform the second dimension transformation on the task data x. The formula is as follows:

[0119] H(x)=x·W g

[0120] Where H(x) is the gate vector of N processing sub-networks; W g is the model parameter 1 of the threshold gating network, and the specific value of the model parameter 1 is obtained through model training.

[0121] Take the task data x as Figure 7 As an example, assuming that the number of processing subnetworks in the multi-task model is N=3, then the model parameter W g It should be a matrix with 4 rows and 3 columns. In this way, the task data x and model parameters W gAfter the product operation, a vector with 1 row and 3 columns can be obtained. This vector is the gate vector, and the three elements arranged in sequence in this vector correspond to the three processing sub-networks in the multi-task model. Figure 7 As shown, assuming that the gating vector is (5,4,1), the element 5 in the gating vector corresponds to the processing sub-network 1 in the multi-task model, the element 4 corresponds to the processing sub-network 2, and the element 1 corresponds to the processing sub-network 3.

[0122] ② Normalize the gate vector to obtain the first weight distribution of N processing sub-networks. The normalization formula is as follows:

[0123] G 1 (x) = Softmax(H(x))

[0124] In the formula, G 1 (x) is the first weight distribution; Softmax(H(x)) is the normalization function; H(x) is the gate vector. Among them, Softmax(H(x)) can be expressed as:

[0125]

[0126] Among them, z i Represents the element of the i-th processing subnetwork in N processing subnetworks in H(x), i∈[1,2,…,N]. This normalization function is mainly used to convert the gating vector H(x) into a probability distribution in the range [0,1] and with a sum of 1 (i.e., the first weight distribution).

[0127] Continue to see Figure 7 , it supports the use of a normalization function to normalize the gating vector H(x) to map each element to the range of [0,1]; assuming that the first weight distribution obtained by the normalization process is approximately (0.72, 0.27, 0.01), it is determined that: element 0.72 corresponds to the contribution of processing subnetwork 1 in the multi-task model to the task, element 0.27 corresponds to the contribution of processing subnetwork 2 in the multi-task model to the task, and element 0.01 corresponds to the contribution of processing subnetwork 3 in the multi-task model to the task.

[0128] ③Use the model parameter 2 of the threshold gating network to transform the task data in the first dimension to obtain the initial threshold gating. The formula is as follows:

[0129] T(x)=x·W t

[0130] Where T(x) is the initial threshold gating; W t is the model parameter 2 of the threshold gating network, and the specific value of the model parameter 2 is obtained through model training.

[0131] Continue to see Figure 7 , when the task data x is a vector of 1 row and 4 columns, the model parameter W t It should be a matrix with 4 rows and 1 column. In this way, the task data x and model parameters W t After the product operation, we can get a vector with 1 row and 1 column, that is, the vector is a real number T(x).

[0132] ④ Use the target conversion function to perform threshold compression processing on the initial threshold gating T(x) to obtain the intermediate threshold gating. The formula is as follows:

[0133] t=σ(T(x),β,γ)

[0134] Where t is the intermediate threshold gate; σ(T(x), β, γ) is the target conversion function, which is obtained by fine-tuning the normalization function. The normalization function here is the σ normalization function. The target conversion function can be expressed as follows:

[0135]

[0136] Where z represents the initial threshold gating T(x) corresponding to any task.

[0137] It is worth noting that by comparing the target conversion function with the normalization function, it can be seen that this solution introduces hyperparameters β and γ in the denominator of the normalization function. The target conversion function after introducing the hyperparameters can compress or limit the maximum value of the intermediate threshold gating to a smaller value in the range of (0,1). The specific value of the smaller value is determined according to the value of the hyperparameter; in this way, it can be avoided that the intermediate threshold gating value is large, which makes it difficult to obtain a suitable number of target processing subnetworks from N processing subnetworks. For example, if the intermediate threshold gating calculated by the traditional normalization function is 0.8, if the first weight distribution is (0.5, 0.4, 0.1), the target processing subnetwork cannot be selected for the task. On the contrary, this scheme introduces hyperparameters in the normalized access, such as hyperparameters β=16, γ=4.7. Then, when the data volume of z is close to positive infinity, that is, the data volume of the task data is large, the intermediate threshold gating calculated by the target conversion function with the introduced hyperparameters can be as large as about 0.0625; that is, the maximum value that the intermediate threshold gating can take is compressed from 1 to 0.0625, which is conducive to the selection of the target processing sub-network based on the intermediate threshold gating.

[0138] ⑤ Based on the intermediate threshold gating and the first weight distribution, determine the target threshold gating for the task. The formula is as follows:

[0139] t′=min(max(G 1 (X),t))

[0140] Where t′ represents the target threshold gating; min(max(G 1 (X),t) is the minimum function; max(G 1 (X),t) is the maximum value function.

[0141] It should be noted that if the intermediate threshold gating t is used directly to truncate the first weight distribution G 1 (x), the extremely small probability due to the first weight distribution G 1 The maximum first weight value in (x) is greater than the intermediate threshold gating t, resulting in the inability to screen the target processing subnetwork for the task. Based on this, in order to avoid the intermediate threshold gating t exceeding the first weight distribution G 1 (x) results in the selected network set (i.e., the Exoert set) being empty, i.e., the target processing subnetwork cannot be selected for the task; the embodiment of the present application also supports limiting the intermediate threshold gating t to be further limited to the target threshold gating t′, ensuring that the maximum value of the target threshold gating t′ is the first weight distribution G 1 The largest first weight value in (x) ensures that at least one target processing subnetwork can be selected for the task to ensure that the task will be executed.

[0142] ⑥ Use target threshold gating to perform threshold truncation processing on the first weight distribution and select M target processing subnetworks for the task. The formula is as follows:

[0143] G 2 (x) = relu(G 1 (x)-t+∈)

[0144] In the formula, G 2 (x) is the subtraction result between each first weight value in the first weight distribution and the target threshold gate t, plus the difference information obtained by the minimum positive value ∈, and the comparison result with the zero value. The value of the minimum positive value ∈ can be (1e-10). relu(G 1 (x)-t+∈) is the maximum value function, and the relu function can be expressed as:

[0145] relu=max(ω,0)

[0146] In the formula, ω=G 1 (x)-t+∈ represents the subtraction result between the first weight value and the target threshold gate t, plus the difference information obtained by adding the minimum positive value ∈.

[0147] It can be seen that by calculating the difference information between the first weight value and the target threshold gating, the difference information corresponding to the first weight value in the first weight distribution when it is less than zero (i.e., the difference information is a negative number) can be set to zero, and only the processing subnetwork with a positive difference information is retained as the target processing subnetwork for performing the task; at this time, the network set (or subscript set) selected for the task can be represented as ThresholdxSet. The network set includes the network identifier of each target processing subnetwork in the M target processing subnetworks screened for the task (used to uniquely identify the target processing subnetwork, such as the network number of the target processing subnetwork, etc.).

[0148] like Figure 7 As shown, if the first weight value 0.5 in the first weight distribution is greater than the target threshold gate 0.3, the difference information between the first weight value 0.5 and the target threshold gate can be expressed as 0.2+ε; if the difference information 0.2+ε is greater than zero, it is retained; if the first weight value 0.4 in the first weight distribution is greater than the target threshold gate 0.3, the difference information between the first weight value 0.4 and the target threshold gate can be expressed as 0.1+∈; if the difference information 0.1+∈ is greater than zero, it is retained; if the first weight value 0.1 in the first weight distribution is less than the target threshold gate 0.3, the difference information between the first weight value 0.1 and the target threshold gate can be expressed as 10.2+∈; if the difference information -0.2+ε is less than zero, the value of the difference information is set to 0; in this implementation, G 2 (x) = (0.2 + ε, 0.1 + ε, 0), and the network set selected for the task can be expressed as [1, 2], that is, processing subnetwork 1 and processing subnetwork 2 in the multi-task model are selected as target processing subnetworks to perform the task.

[0149] ⑦ Generate the second weight distribution of M target processing sub-networks. The formula is as follows:

[0150]

[0151] Where i represents the i-th processing subnetwork among N processing subnetworks, and the value of i is i∈{1,2,…,N}; G 2 (x) i Represents the difference information corresponding to the i-th processing sub-network; G(x) i represents the second weight value of the i-th processing sub-network; Represents the sum of the difference information of N processing sub-networks.

[0152] like Figure 7As shown in FIG. 1 , considering that the difference information corresponding to the unselected processing subnetwork among the three processing subnetworks is 0, its second weight value is 0, then the second weight distribution of the selected two target processing subnetworks (i.e., processing subnetwork 1 and processing subnetwork 2) can be expressed as G(x)=(G(x) 1 ,G(x) 2 ).

[0153] Step 2: Use the selected M target processing subnetworks to perform task processing on the task data x of the task, and obtain the output result corresponding to each target processing subnetwork, which is the task execution result of the task. In the embodiment of the present application, the task execution result corresponding to the kth target processing subnetwork among the M target processing subnetworks can be expressed as E k (x), the value of k is k∈ThresholdxSet.

[0154] Step 3: Perform a weighted operation using the output result of each target processing subnetwork and the corresponding second weight value to obtain the first intermediate processing result of each target processing subnetwork for the task. For example, the task execution result E corresponding to the kth target processing subnetwork among the M target processing subnetworks is k (x), and the second weight value G(x) of the kth target processing subnetwork k Perform weighted operations to obtain the first intermediate processing result E of the kth target processing subnetwork for the task k (x)·G(x) k . Further, the M target processing sub-networks are summed for the first intermediate processing results of the task, that is, the second intermediate processing result of the task is obtained. The formulas of the weighted operation and summation operation (referred to as weighted sum operation) described above are as follows:

[0155]

[0156] Where y is the result of the weighted sum operation. Figure 7 As shown, the second intermediate processing result of the task is represented by E 1 (x)·G(x) 1 +E 2 (x)·G(x) 2 Furthermore, the output result of the MoE layer in the multi-task model (ie, the second intermediate processing result) is input into the task module corresponding to the task to realize personalized processing of the task, thereby obtaining the processing result of the task.

[0157] In summary, the network number of the target processing subnetwork screened by the embodiment of the present application is matched with the data volume of the task data of the task, which makes it possible to select the target processing subnetwork of the appropriate number to perform the task, avoiding the problems of computational redundancy, poor computational performance and large computational overhead caused by the selected processing subnetwork too much. In addition, the first weight value of the target processing subnetwork screened for the task is greater than the target threshold gating, which makes it possible to use the target processing subnetwork with a greater degree of influence on the task to perform the task, and the accuracy of the processing result of the task can be ensured, that is, the processing effect of the task can be ensured. In other words, the data processing scheme provided by the embodiment of the present application is applied by applying a learnable parameter (such as the target threshold gating mentioned above, or the second model parameter, etc.), so that the multi-task model learns and adjusts the gate threshold end-to-end in model training, solves the drawbacks of the traditional gating method by artificial prior adjustment of the hyperparameter N or K, that is, selects the target processing subnetwork of the appropriate number for the task, avoids the problems of large computational overhead or poor processing effect caused by too many or too few networks, and can take into account both computing performance and processing effect.

[0158] The method of the embodiment of the present application is described in detail above. In order to facilitate the above-mentioned scheme of the embodiment of the present application to be better implemented, the device of the embodiment of the present application is provided below accordingly. In the embodiment of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the module or unit function.

[0159] Figure 8 A schematic diagram of the structure of a data processing device provided by an exemplary embodiment of the present application is shown; the data processing device can be used to execute Figure 4 Some or all of the steps in the method embodiment shown.

[0160] See also Figure 8 , the device comprises the following units:

[0161] The prediction unit 801 is used to predict the first weight distribution of N processing subnetworks according to the task data of the task to be processed; the first weight distribution includes a first weight value corresponding to each processing subnetwork, and the first weight value is used to indicate the influence degree of the task execution result obtained by the corresponding processing subnetwork executing the task on the processing result of the task;

[0162] The processing unit 802 is used to select M target processing subnetworks for executing the task from the N processing subnetworks based on the data volume of the task data; the first weight value corresponding to each target processing subnetwork is greater than the target threshold gating; M is a positive integer, and M is less than or equal to N;

[0163] The processing unit 802 is further configured to generate a second weight distribution of the M target processing subnetworks according to the first weight value corresponding to each target processing subnetwork; the second weight distribution includes the second weight value corresponding to each target processing subnetwork;

[0164] The processing unit 802 is further used to call the M target processing sub-networks to respectively execute tasks, and predict the processing results of the tasks based on the second weight distribution and the task execution results respectively obtained by the M target processing sub-networks.

[0165] In one implementation, the processing unit 802 is used to select M target processing subnetworks for executing the task from N processing subnetworks based on the data volume of the task data, specifically to:

[0166] The task data is dimensionally transformed to obtain a target threshold gating that matches the data volume of the task data; the target threshold gating is used to indicate: the minimum degree of influence of the task execution result obtained by the processing sub-network executing the task on the task processing result;

[0167] The target threshold gating is used to perform threshold truncation processing on the first weight distribution, and M target processing subnetworks for performing the task are screened out from the N processing subnetworks.

[0168] In one implementation, the processing unit 802 is used to perform dimension transformation processing on the task data to obtain a target threshold gating that matches the data volume of the task data, specifically to:

[0169] Obtaining a distribution characteristic of the first weight distribution; the distribution characteristic is used to indicate the distribution of the N first weight values ​​in the weight range;

[0170] According to the data volume of the task data and the distribution characteristics of the first weight distribution, the task data is subjected to a first dimension transformation process to obtain an initial threshold gating; the dimension of the initial threshold gating is one-dimensional;

[0171] The initial threshold gating is subjected to threshold conversion processing to obtain a target threshold gating that matches the data volume of the task data.

[0172] In one implementation, the processing unit 802 is used to perform threshold conversion processing on the initial threshold gating to obtain a target threshold gating that matches the data volume of the task data, specifically to:

[0173] The target conversion function is used to perform threshold compression processing on the initial threshold gating to obtain the intermediate threshold gating; the target conversion function is obtained by fine-tuning the normalization function;

[0174] Screening out a maximum first weight value from the first weight distribution, and comparing the maximum first weight value with an intermediate threshold gating to obtain a comparison result;

[0175] According to the comparison result, a target threshold gating matching the data volume of the task data is determined;

[0176] Among them, when the comparison result indicates that the maximum first weight value is less than the intermediate threshold gating, the target threshold gating is the maximum first weight value; when the comparison result indicates that the maximum first weight value is greater than the intermediate threshold gating, the target threshold gating is the intermediate threshold gating.

[0177] In one implementation, the processing unit 802 is used to perform threshold truncation processing on the first weight distribution by using target threshold gating, and when M target processing subnetworks for performing the task are screened out from N processing subnetworks, specifically for:

[0178] Subtract each first weight value in the first weight distribution from the target threshold gate to obtain N subtraction results;

[0179] Filter out the subtraction results with positive values ​​from the N subtraction results;

[0180] Select a processing subnetwork corresponding to a subtraction result with a positive value from the N processing subnetworks as a target processing subnetwork for executing the task;

[0181] The number M of networks of the target processing subnetwork is the same as the number of subtraction results that are positive numbers.

[0182] In one implementation, the processing unit 802 is configured to generate the second weight distribution of the M target processing subnetworks according to the first weight value corresponding to each target processing subnetwork, specifically to:

[0183] Obtaining difference information corresponding to each target processing subnetwork; the difference information is obtained by subtracting the first weight value corresponding to the corresponding target processing subnetwork from the target threshold gating;

[0184] Based on the difference information corresponding to each target processing sub-network, a second weight distribution of the M target processing sub-networks is generated.

[0185] In one implementation, the processing unit 802 is configured to generate the second weight distribution of the M target processing subnetworks based on the difference information corresponding to each target processing subnetwork, specifically to:

[0186] The difference information corresponding to the M target processing sub-networks is normalized to generate a second weight distribution of the M target processing sub-networks.

[0187] In one implementation, the processing unit 802 is used to predict the processing result of the task based on the second weight distribution and the task execution results respectively obtained by the M target processing subnetworks, specifically to:

[0188] M target processing sub-networks are used to process the task data of the task respectively, and the task execution result of each target processing sub-network for the task is obtained;

[0189] Performing weighted operations on the task execution results of each target processing subnetwork for the task and the corresponding second weight values ​​in the second weight distribution, respectively, to obtain weighted results corresponding to each target processing subnetwork;

[0190] The weighted results corresponding to each target processing sub-network are summed to obtain the processing result of the task.

[0191] In one implementation, each of the N processing subnetworks has different processing capabilities for different tasks; the processing unit 802 is used to predict the first weight distribution of the N processing subnetworks according to the task data of the task to be processed, specifically for:

[0192] According to the matching relationship between the task and the processing capability of each of the N processing subnetworks, the task data is transformed in the second dimension to obtain a gating vector; the dimension of the gating vector is N, and each element in the gating vector corresponds to a processing subnetwork;

[0193] The gating vector is normalized to obtain the first weight distribution of the N processing sub-networks.

[0194] In one implementation, the task is a task generated when an Internet service has a data processing requirement; the Internet service includes a multimedia data recommendation service and a search service;

[0195] Wherein, when the Internet service is a multimedia data recommendation service, the task corresponding to the multimedia data recommendation service includes a task of predicting data indicators for multimedia data; the data indicators include at least one of the following: a click rate indicator, a conversion rate indicator, and a resource value estimation indicator;

[0196] When the task is to predict a click rate indicator for multimedia data, the processing result of the task includes the click rate prediction result corresponding to the multimedia data; or, when the task is to predict a conversion rate indicator for multimedia data, the processing result of the task includes the conversion rate prediction result corresponding to the multimedia data; or, when the task is to predict a resource value indicator for multimedia data, the processing result of the task includes the resource value prediction result corresponding to the multimedia data.

[0197] According to one embodiment of the present application, Figure 8 The various units in the data processing device shown can be individually or completely combined into one or several other units to form, or one (some) of the units can be further divided into multiple functionally smaller units to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In practical applications, the function of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the data processing device may also include other units. In practical applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units. According to another embodiment of the present application, it can be executed by running on a general computing device such as a computer including processing elements and storage elements such as a central processing unit (CPU), a memory access storage medium (RAM), and a read-only storage medium (ROM). Figure 4 A computer program (including program code) for each step involved in the corresponding method shown in FIG. Figure 8 The data processing device shown in the embodiment of the present application is implemented by the data processing device. The computer program can be recorded on a computer-readable recording medium, for example, and loaded into the above-mentioned computing device through the computer-readable recording medium and run therein.

[0198] In the embodiment of the present application, after obtaining the task data of the task to be processed, the first weight distribution of N processing subnetworks in the multi-task model can be predicted according to the task data, and the first weight distribution includes the first weight value corresponding to each processing subnetwork. It is also possible to screen M target processing subnetworks for performing the task from N processing subnetworks based on the data volume of the task data, and the first weight value corresponding to each target processing subnetwork is greater than the target threshold gating. In this way, the second weight distribution of M target processing subnetworks can be generated based on the first weight value corresponding to each target processing subnetwork, and the second weight value corresponding to each target processing subnetwork distribution in the second weight distribution. Finally, calling the M target processing subnetworks screened for the task and their second weight values ​​can realize the prediction for the task and obtain the processing result of the task. It can be seen from the above scheme that the embodiment of the present application can flexibly screen the target processing subnetwork in the multi-task model for the task based on the data volume of the task data of the task. On the one hand, the number of selected target processing subnetworks matches the amount of task data, which enables the selection of an appropriate number of target processing subnetworks to perform the task, avoiding problems such as computational redundancy, poor computational performance, and high computational overhead caused by too many selected processing subnetworks. On the other hand, the first weight value of the selected target processing subnetwork is greater than the target threshold gating, which enables the use of a target processing subnetwork that has a greater impact on the task to perform the task, thereby ensuring the accuracy of the task processing result, that is, ensuring the task processing effect.

[0199] Fig. 9 FIG. 1 shows a schematic diagram of a computer device provided by an exemplary embodiment of the present application. Fig. 9 , the computer device includes a processor 901, a communication interface 902 and a computer-readable storage medium 903. The processor 901, the communication interface 902 and the computer-readable storage medium 903 can be connected via a bus or other means. The communication interface 902 is used to receive and send data. The computer-readable storage medium 903 can be stored in the memory of the computer device, the computer-readable storage medium 903 is used to store a computer program, the computer program includes program instructions, and the processor 901 is used to execute the program instructions stored in the computer-readable storage medium 903. The processor 901 (or CPU (Central Processing Unit)) is the computing core and control core of the computer device, which is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.

[0200] The embodiment of the present application also provides a computer-readable storage medium (Memory), which is a memory device in a computer device for storing programs and data. It is understandable that the computer-readable storage medium here can include both built-in storage media in the computer device and, of course, extended storage media supported by the computer device. The computer-readable storage medium provides a storage space that stores the processing system of the computer device. In addition, one or more instructions suitable for being loaded and executed by the processor 901 are also stored in the storage space, and these instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage; optionally, it can also be at least one computer-readable storage medium located away from the aforementioned processor.

[0201] In one embodiment, the computer-readable storage medium stores one or more instructions; the processor 901 loads and executes the one or more instructions stored in the computer-readable storage medium to implement the corresponding steps in the above data processing method embodiment; in a specific implementation, the one or more instructions in the computer-readable storage medium are loaded by the processor 901 and the following steps are executed:

[0202] Predicting first weight distributions of N processing subnetworks according to task data of the task to be processed; the first weight distributions include first weight values ​​corresponding to each processing subnetwork, and the first weight values ​​are used to indicate the degree of influence of a task execution result obtained by executing the task by the corresponding processing subnetwork on a processing result of the task;

[0203] Based on the amount of task data, M target processing subnetworks for executing the task are selected from N processing subnetworks; the first weight value corresponding to each target processing subnetwork is greater than the target threshold gating; M is a positive integer, and M is less than or equal to N;

[0204] Generate a second weight distribution of M target processing subnetworks according to the first weight value corresponding to each target processing subnetwork; the second weight distribution includes the second weight value corresponding to each target processing subnetwork;

[0205] The M target processing sub-networks are called to respectively execute tasks, and based on the second weight distribution and the task execution results respectively obtained by the M target processing sub-networks, the processing results of the tasks are predicted.

[0206] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 901 and when the processor 901 selects M target processing subnetworks for executing the task from N processing subnetworks based on the data amount of the task data, the following steps are specifically performed:

[0207] The task data is dimensionally transformed to obtain a target threshold gating that matches the data volume of the task data; the target threshold gating is used to indicate: the minimum degree of influence of the task execution result obtained by the processing sub-network executing the task on the task processing result;

[0208] The target threshold gating is used to perform threshold truncation processing on the first weight distribution, and M target processing subnetworks for performing the task are screened out from the N processing subnetworks.

[0209] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 901 and when performing dimension transformation processing on the task data to obtain a target threshold gating that matches the data volume of the task data, the following steps are specifically performed:

[0210] Obtaining a distribution characteristic of the first weight distribution; the distribution characteristic is used to indicate the distribution of the N first weight values ​​in the weight range;

[0211] According to the data volume of the task data and the distribution characteristics of the first weight distribution, the task data is subjected to a first dimension transformation process to obtain an initial threshold gating; the dimension of the initial threshold gating is one-dimensional;

[0212] The initial threshold gating is subjected to threshold conversion processing to obtain a target threshold gating that matches the data volume of the task data.

[0213] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 901 and when performing threshold conversion processing on the initial threshold gating to obtain a target threshold gating that matches the data volume of the task data, the following steps are specifically performed:

[0214] The target conversion function is used to perform threshold compression processing on the initial threshold gating to obtain the intermediate threshold gating; the target conversion function is obtained by fine-tuning the normalization function;

[0215] Screening out a maximum first weight value from the first weight distribution, and comparing the maximum first weight value with an intermediate threshold gating to obtain a comparison result;

[0216] According to the comparison result, a target threshold gating matching the data volume of the task data is determined;

[0217] Among them, when the comparison result indicates that the maximum first weight value is less than the intermediate threshold gating, the target threshold gating is the maximum first weight value; when the comparison result indicates that the maximum first weight value is greater than the intermediate threshold gating, the target threshold gating is the intermediate threshold gating.

[0218] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 901 and when performing threshold truncation processing on the first weight distribution using target threshold gating to select M target processing subnetworks for performing the task from N processing subnetworks, the following steps are specifically performed:

[0219] Subtract each first weight value in the first weight distribution from the target threshold gate to obtain N subtraction results;

[0220] Filter out the subtraction results with positive values ​​from the N subtraction results;

[0221] Select a processing subnetwork corresponding to a subtraction result with a positive value from the N processing subnetworks as a target processing subnetwork for executing the task;

[0222] The number M of networks of the target processing subnetwork is the same as the number of subtraction results that are positive numbers.

[0223] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 901 and when executing to generate a second weight distribution of M target processing subnetworks according to the first weight value corresponding to each target processing subnetwork, the following steps are specifically performed:

[0224] Obtaining difference information corresponding to each target processing subnetwork; the difference information is obtained by subtracting the first weight value corresponding to the corresponding target processing subnetwork from the target threshold gating;

[0225] Based on the difference information corresponding to each target processing sub-network, a second weight distribution of the M target processing sub-networks is generated.

[0226] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 901 and when executing to generate a second weight distribution of M target processing subnetworks based on the difference information corresponding to each target processing subnetwork, the following steps are specifically performed:

[0227] The difference information corresponding to the M target processing sub-networks is normalized to generate a second weight distribution of the M target processing sub-networks.

[0228] In one implementation, one or more instructions in the computer-readable storage medium are loaded by the processor 901 and when executing the task execution results obtained based on the second weight distribution and the M target processing subnetworks respectively, predicting the processing result of the task, specifically perform the following steps:

[0229] M target processing sub-networks are used to process the task data of the task respectively, and the task execution result of each target processing sub-network for the task is obtained;

[0230] Performing weighted operations on the task execution results of each target processing subnetwork for the task and the corresponding second weight values ​​in the second weight distribution, respectively, to obtain weighted results corresponding to each target processing subnetwork;

[0231] The weighted results corresponding to each target processing sub-network are summed to obtain the processing result of the task.

[0232] In one implementation, each of the N processing subnetworks has different processing capabilities for different tasks; one or more instructions in the computer-readable storage medium are loaded by the processor 901 and when executing the task data of the task to be processed, predicting the first weight distribution of the N processing subnetworks, specifically performing the following steps:

[0233] According to the matching relationship between the task and the processing capability of each of the N processing subnetworks, the task data is transformed in the second dimension to obtain a gating vector; the dimension of the gating vector is N, and each element in the gating vector corresponds to a processing subnetwork;

[0234] The gating vector is normalized to obtain the first weight distribution of the N processing sub-networks.

[0235] In one implementation, the task is a task generated when an Internet service has a data processing requirement; the Internet service includes a multimedia data recommendation service and a search service;

[0236] Wherein, when the Internet service is a multimedia data recommendation service, the task corresponding to the multimedia data recommendation service includes a task of predicting data indicators for multimedia data; the data indicators include at least one of the following: a click rate indicator, a conversion rate indicator, and a resource value estimation indicator;

[0237] When the task is to predict a click rate indicator for multimedia data, the processing result of the task includes the click rate prediction result corresponding to the multimedia data; or, when the task is to predict a conversion rate indicator for multimedia data, the processing result of the task includes the conversion rate prediction result corresponding to the multimedia data; or, when the task is to predict a resource value indicator for multimedia data, the processing result of the task includes the resource value prediction result corresponding to the multimedia data.

[0238] The embodiment of the present application can flexibly screen the target processing subnetwork in the multi-task model for the task based on the data volume of the task data of the task. On the one hand, the network number of the selected target processing subnetwork matches the data volume of the task data, which makes it possible to select a suitable number of target processing subnetworks to perform the task, avoiding the problems of computational redundancy, poor computational performance and large computational overhead caused by too many selected processing subnetworks. On the other hand, the first weight value of the selected target processing subnetwork is greater than the target threshold gating, which makes it possible to use a target processing subnetwork with a greater degree of influence on the task to perform the task, which can ensure the accuracy of the processing result of the task, that is, ensure the processing effect of the task.

[0239] The embodiment of the present application also provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above-mentioned model training method.

[0240] A person skilled in the art can appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0241] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When loading and executing computer program instructions on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable device. The computer instruction can be stored in a computer-readable storage medium or transmitted by a computer-readable storage medium. The computer instruction can be transmitted from a website site, a computer, a server or a data center to another website site, a computer, a server or a data center by wired (for example, coaxial cable, optical fiber, digital line (DSL)) or wireless (for example, infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can access or a data processing device such as a server, a data center, etc. that contains one or more available media integration. Available media can be magnetic media (for example, floppy disk, hard disk, tape), optical media (for example, DVD), or semiconductor media (for example, solid-state drive (Solid State Disk, SSD)) and the like.

[0242] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technical object familiar with the technical field within the technical scope disclosed in the present application can be easily thought of for changes or replacements, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A data processing method based on a multi-task model, It is characterized in that The multi-task model includes N processing sub-networks, where N is an integer greater than 1; the method includes: Predicting a first weight distribution of N processing subnetworks according to task data of the task to be processed; the first weight distribution includes a first weight value corresponding to each of the processing subnetworks, and the first weight value is used to indicate the degree of influence of a task execution result obtained by the corresponding processing subnetwork executing the task on a processing result of the task; Based on the data volume of the task data, M target processing subnetworks for executing the task are selected from the N processing subnetworks; the first weight value corresponding to each of the target processing subnetworks is greater than the target threshold gating; M is a positive integer, and M is less than or equal to N; According to the first weight value corresponding to each of the target processing subnetworks, generating a second weight distribution of the M target processing subnetworks; the second weight distribution includes a second weight value corresponding to each of the target processing subnetworks; The M target processing sub-networks are called to respectively execute the tasks, and based on the second weight distribution and the task execution results respectively obtained by the M target processing sub-networks, the processing results of the tasks are predicted.

2. The method according to claim 1, It is characterized in that The step of selecting M target processing subnetworks for executing the task from the N processing subnetworks based on the amount of the task data includes: Performing dimension transformation processing on the task data to obtain a target threshold gating that matches the data volume of the task data; the target threshold gating is used to indicate: the minimum degree of influence of the task execution result obtained by the processing subnetwork executing the task on the processing result of the task; The target threshold gating is used to perform threshold truncation processing on the first weight distribution, and M target processing subnetworks for executing the task are screened out from the N processing subnetworks.

3. The method according to claim 2, It is characterized in that The performing dimension transformation processing on the task data to obtain a target threshold gating that matches the data volume of the task data includes: Obtaining a distribution characteristic of the first weight distribution; the distribution characteristic is used to indicate the distribution of the N first weight values ​​in the weight range; According to the data volume of the task data and the distribution characteristics of the first weight distribution, the task data is subjected to a first dimensional transformation process to obtain an initial threshold gating; the dimension of the initial threshold gating is one-dimensional; The initial threshold gating is subjected to threshold conversion processing to obtain a target threshold gating that matches the data volume of the task data.

4. The method according to claim 3, It is characterized in that The performing threshold conversion processing on the initial threshold gating to obtain a target threshold gating matching the data volume of the task data includes: Performing threshold compression processing on the initial threshold gating by using a target conversion function to obtain an intermediate threshold gating; the target conversion function is obtained by fine-tuning a normalization function; Screening out a maximum first weight value from the first weight distribution, and comparing the maximum first weight value with the intermediate threshold gating to obtain a comparison result; Determining a target threshold gating that matches the data volume of the task data according to the comparison result; Among them, when the comparison result indicates that the largest first weight value is less than the intermediate threshold gate, the target threshold gate is the largest first weight value; when the comparison result indicates that the largest first weight value is greater than the intermediate threshold gate, the target threshold gate is the intermediate threshold gate.

5. The method according to claim 2, It is characterized in that The step of using the target threshold gating to perform threshold truncation processing on the first weight distribution, and selecting M target processing subnetworks for performing the task from the N processing subnetworks, includes: Subtracting each first weight value in the first weight distribution from the target threshold gate to obtain N subtraction results; Filter out the subtraction results with positive values ​​from the N subtraction results; Selecting a processing subnetwork corresponding to a subtraction result with a positive value from the N processing subnetworks as a target processing subnetwork for executing the task; The number M of networks of the target processing subnetwork is the same as the number of subtraction results that are positive numbers.

6. The method according to claim 1, It is characterized in that The step of generating the second weight distribution of the M target processing subnetworks according to the first weight value corresponding to each target processing subnetwork includes: Obtaining difference information corresponding to each of the target processing subnetworks; the difference information is obtained by subtracting the first weight value corresponding to the corresponding target processing subnetwork from the target threshold gating; Based on the difference information corresponding to each of the target processing sub-networks, a second weight distribution of the M target processing sub-networks is generated.

7. The method according to claim 6, It is characterized in that The step of generating a second weight distribution of the M target processing subnetworks based on the difference information corresponding to each target processing subnetwork includes: The difference information corresponding to the M target processing sub-networks is normalized to generate a second weight distribution of the M target processing sub-networks.

8. The method according to claim 1, It is characterized in that The predicting the processing result of the task based on the second weight distribution and the task execution results respectively obtained by the M target processing sub-networks includes: Using the M target processing sub-networks to perform task processing on the task data of the task respectively, and obtaining a task execution result of each target processing sub-network for the task; Performing weighted operations on the task execution results of each target processing subnetwork for the task and the corresponding second weight values ​​in the second weight distribution, respectively, to obtain weighted results corresponding to each target processing subnetwork; The weighted results corresponding to each target processing subnetwork are summed to obtain the processing result of the task.

9. The method according to claim 1, It is characterized in that Each of the N processing subnetworks has a different processing capability for different tasks; and predicting the first weight distribution of the N processing subnetworks according to the task data of the task to be processed includes: According to the matching relationship between the task and the processing capability of each processing subnetwork in the N processing subnetworks, the task data is subjected to a second dimension transformation process to obtain a gating vector; the dimension of the gating vector is N, and each element in the gating vector corresponds to a processing subnetwork; The gating vector is normalized to obtain a first weight distribution of the N processing sub-networks.

10. The method according to any one of claims 1 to 9, It is characterized in that The task is a task generated when an Internet service has a data processing requirement; the Internet service includes a multimedia data recommendation service and a search service; Wherein, when the Internet service is a multimedia data recommendation service, the task corresponding to the multimedia data recommendation service includes a task of predicting data indicators for multimedia data; the data indicators include at least one of the following: a click rate indicator, a conversion rate indicator, and a resource value estimation indicator; When the task is to predict a click rate indicator for multimedia data, the processing result of the task includes the click rate prediction result corresponding to the multimedia data; or, when the task is to predict a conversion rate indicator for multimedia data, the processing result of the task includes the conversion rate prediction result corresponding to the multimedia data; or, when the task is to predict a resource value indicator for multimedia data, the processing result of the task includes the resource value prediction result corresponding to the multimedia data.

11. A data processing device based on a multi-task model, It is characterized in that The multi-task model includes N processing sub-networks, where N is an integer greater than 1; the device includes: A prediction unit, configured to predict a first weight distribution of N processing subnetworks according to task data of a task to be processed; the first weight distribution includes a first weight value corresponding to each of the processing subnetworks, the first weight value being used to indicate the degree of influence of a task execution result obtained by executing the task by the corresponding processing subnetwork on a processing result of the task; A processing unit, configured to select M target processing subnetworks for executing the task from the N processing subnetworks based on the data volume of the task data; the first weight value corresponding to each of the target processing subnetworks is greater than the target threshold gating; M is a positive integer, and M is less than or equal to N; The processing unit is further used to generate a second weight distribution of the M target processing subnetworks according to the first weight value corresponding to each target processing subnetwork; the second weight distribution includes a second weight value corresponding to each target processing subnetwork; The processing unit is further used to call the M target processing sub-networks to respectively execute the tasks, and predict the processing results of the tasks based on the second weight distribution and the task execution results respectively obtained by the M target processing sub-networks.

12. A computer device, It is characterized in that include: a processor adapted to execute a computer program; A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, and when the computer program is executed by the processor, the data processing method according to any one of claims 1 to 10 is implemented.

13. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the data processing method according to any one of claims 1 to 10.

14. A computer program product, It is characterized in that The computer program product comprises computer instructions, and when the computer instructions are executed by a processor, the data processing method according to any one of claims 1 to 10 is implemented.

Citation Information

Cited By

  • Well logging data interpretation method, device and equipment based on large model fine tuning and storage medium

    CN120430421A