Data processing method, device and computer program product

Through dual-model branch adaptive learning and resampling technology, the processing of long-tail data is optimized, which solves the problem of insufficient recognition ability of deep models in long-tail scenarios, and achieves better long-tail recognition effect and feature expression ability.

CN114841240BActive Publication Date: 2025-05-16ALIBABA (CHINA) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210345642.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-31
Publication Date
2025-05-16
Estimated Expiration
2042-03-31

AI Technical Summary

Technical Problem

In the recommended field, it is difficult for deep models to achieve ideal results in long-tail scenarios. Although existing rebalancing strategies can alleviate the long-tail effect, they can also damage the ability to express deep features.

Method used

Through the dual-model branch adaptive learning ability to control the areas that need to be paid attention to in different stages of models, use resampling technology to obtain long-tail data with high sampling ratio, combine the first and second model branches for weighted summing, and optimize the prediction results.

Benefits of technology

The long-tail recognition ability is improved, and the risks of overfitting long-tail data and underfitting the full amount of data are avoided, ensuring the original feature learning effect while improving the recognition ability of long-tail scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114841240B_ABST
    Figure CN114841240B_ABST
Patent Text Reader

Abstract

The disclosed embodiments disclose a data processing method, device and computer program product. The method includes: obtaining original data; resampling the original data to obtain sampled data, wherein the resampled sampling ratio of the long-tail data in the original data is high; using the original data and the sampled data as input data of the first model branch of the data processing model to obtain a first score for representing the prediction result of the original data; using the training results of a part of the first model branch and the sampled data as input data of the second model branch of the data processing model to obtain a second score for representing the prediction result of the sampled data; determining adaptive parameters based on training batches; based on the adaptive parameters corresponding to the corresponding training batches, weighted summing the first score and the second score to obtain an optimized first score, representing the prediction result after optimizing the long-tail scene data, so that the two branches can be merged to ultimately improve the long-tail recognition capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field, and in particular to a data processing method, device, and computer program product. Background Art

[0002] When discussing issues related to the long tail effect, the protruding part in the middle of the normal curve can be called the "head", and the relatively flat parts on both sides can be called the "tail". From the perspective of analyzing the normal curve of demand, most of the demand will be concentrated in the head, which we can call popular, while the demand distributed in the tail is personalized, scattered and small. This part of differentiated and small demand will form a long "tail" on the demand curve, that is, the long tail.

[0003] The long-tail problem of data distribution has always been a pain point and difficulty in the recommendation field, but it is also an extremely important link. In the recommendation field, the so-called long-tail problem refers to the fact that in the training data, most of the data is relatively concentrated with a large sample size, and a small part of the data is distributed in the tail with a small sample size. Due to the extreme thirst for data by deep models and the imbalanced distribution of long-tail data, it is difficult for deep models to achieve ideal results in long-tail scenarios. Although the existing rebalancing strategy can alleviate the long-tail effect, it will also damage the expressiveness of deep features to a certain extent. Therefore, in the recommendation field, the long-tail problem has always been a challenge faced by many scenarios. Summary of the invention

[0004] In order to solve the problems in the related technology, the embodiments of the present disclosure provide a data processing method, an apparatus and a computer program product, which can control the areas that the model needs to focus on at different stages through the adaptive learning ability of the dual-model branches, so that the two branches are merged to ultimately improve the long-tail recognition ability.

[0005] In a first aspect, an embodiment of the present disclosure provides a data processing method, wherein the method includes:

[0006] Get the original data;

[0007] Resampling the original data to obtain sampled data, wherein the resampled sampling ratio of the long-tail data in the original data is high;

[0008] Using the original data and the sampled data as input data of a first model branch of a data processing model, and having the first model branch predict based on the input data to obtain a first score representing a prediction result of the original data;

[0009] Using a training result of a part of the first model branch and the sampled data as input data of a second model branch of the data processing model, and the second model branch predicts based on the input data to obtain a second score representing a prediction result of the sampled data;

[0010] Determining adaptive parameters based on the training batch;

[0011] Based on the adaptive parameters corresponding to the corresponding training batch, the first score and the second score are weightedly summed to obtain an optimized first score, and the optimized first score represents the prediction result after the original data is optimized for long-tail scenario data.

[0012] In combination with the first aspect, in a first implementation of the first aspect of the present disclosure, the first model branch includes a first neural network and a second neural network, wherein:

[0013] The method of using the original data and the sampled data as input data of a first model branch of a data processing model, and predicting by the first model branch based on the input data to obtain a first score representing a prediction result of the original data, comprises:

[0014] Using the first neural network, training is performed based on the bias characteristics of the data, the original data, and the sampled data to obtain a bias score of the original data and a bias score of the sampled data;

[0015] Using the second neural network, training is performed based on the main features of the data, the original data, and the sampled data to obtain the main score of the original data;

[0016] Performing a weighted summation on the bias score of the original data and the main score of the original data to obtain the score of the original data;

[0017] The score of the original data is calculated using a normalized exponential function to obtain the first score.

[0018] In combination with the first implementation manner of the first aspect, in a second implementation manner of the first aspect of the present disclosure, the training result of a part of the first model branch and the sampled data are used as input data of a second model branch of the data processing model, and the second model branch predicts based on the input data to obtain a second score representing a prediction result of the sampled data, including:

[0019] Obtaining an embedding vector of the sampled data based on the long-tail scene features of the data and the sampled data;

[0020] Using a part of the second neural network to perform training based on the sampled data to obtain a training result of the sampled data;

[0021] Learning the embedding vector of the sampled data based on the training result of the sampled data to obtain the principal score of the sampled data;

[0022] Performing a weighted summation on the bias score of the sampled data and the main score of the sampled data to obtain the score of the sampled data;

[0023] The score of the sampled data is calculated using a normalized exponential function to obtain the second score.

[0024] In combination with the first implementation or the second implementation of the first aspect, in a third implementation of the first aspect of the present disclosure, the first neural network has a plurality of multilayer perceptrons MLPs connected in series, and the second neural network has a plurality of multilayer perceptrons MLPs connected in series.

[0025] In combination with an implementation manner of the first aspect, in a fourth implementation manner of the first aspect of the present disclosure, the bias feature of the data is a feature that represents the impact of the data on user selection when it is presented in different positions, and the main feature of the data is a universal feature that represents the data when it is generated.

[0026] In combination with the second implementation manner of the first aspect, in a fifth implementation manner of the first aspect of the present disclosure, the long-tail scene feature of the data is a feature representing the long-tail scene to which the data belongs.

[0027] In combination with the first aspect, in a sixth implementation manner of the first aspect of the present disclosure, based on the adaptive parameters corresponding to the corresponding training batches, weighted summing the first score and the second score to obtain an optimized first score, wherein the optimized first score represents a prediction result after long-tail scenario data optimization is performed on the original data, including:

[0028] In the training batch, the adaptive parameter corresponding to the corresponding training batch is adjusted from large to small to gradually weight the first score, and at the same time, the adaptive parameter is subtracted from 1 to weight the second score.

[0029] In combination with the first aspect, in a seventh implementation manner of the first aspect of the present disclosure, the original data is traffic data, and the long-tail scene data is data of traffic scenes whose occurrence probability in the traffic data is lower than a preset threshold.

[0030] In a second aspect, an embodiment of the present disclosure provides a data processing device, wherein the device includes:

[0031] An acquisition module, configured to acquire raw data;

[0032] A sampling data acquisition module is configured to resample the original data to obtain sampled data, wherein the resampled sampling ratio of the long-tail data in the original data is high;

[0033] A first score calculation module is configured to use the original data and the sampled data as input data of a first model branch of a data processing model, and the first model branch predicts based on the input data to obtain a first score representing a prediction result of the original data;

[0034] A second score calculation module is configured to use a training result of a part of the first model branch and the sample data as input data of a second model branch of the data processing model, and the second model branch predicts based on the input data to obtain a second score representing a prediction result of the sample data;

[0035] An adaptive parameter determination module, configured to determine an adaptive parameter based on a training batch;

[0036] The first score optimization module is configured to perform weighted summation of the first score and the second score based on the adaptive parameters corresponding to the corresponding training batch to obtain an optimized first score, wherein the optimized first score represents the prediction result after the original data is optimized for long-tail scenario data.

[0037] The technical solution provided by the embodiments of the present disclosure may have the following beneficial effects:

[0038] According to the technical solution provided by the embodiment of the present disclosure, by obtaining the original data; resampling the original data to obtain the sampled data, wherein the resampling sampling ratio of the long-tail data in the original data is high; using the original data and the sampled data as the input data of the first model branch of the data processing model, the first model branch predicts based on the input data to obtain a first score for representing the prediction result of the original data; using the training result of a part of the first model branch and the sampled data as the input data of the second model branch of the data processing model, the second model branch predicts based on the input data to obtain a second score for representing the prediction result of the sampled data; determining the adaptive parameters based on the training batch; based on the adaptive parameters corresponding to the corresponding training batch, the first score and the second score are weighted and summed to obtain an optimized first score, the optimized first score represents the prediction result of the original data after the long-tail scene data is optimized, and the adaptive learning ability of the dual model branch can be used to control the areas that need to be paid attention to in the model at different stages, so that the two branches are merged to finally improve the long-tail recognition ability. Moreover, the embodiment of the present disclosure uses the full amount of data and the sampled data as the original data for simultaneous training, which to a certain extent avoids the risk of overfitting of the long-tail data and underfitting of the full amount of data.

[0039] According to the technical solution provided by the embodiment of the present disclosure, the first model branch includes a first neural network and a second neural network, wherein the original data and the sampled data are used as input data of the first model branch of the data processing model, and the first model branch predicts the first score for representing the prediction result of the original data based on the input data, including: using the first neural network to perform training based on the bias features of the data, the original data and the sampled data to obtain the bias score of the original data and the bias score of the sampled data; using the second neural network to perform training based on the main features of the data, the original data and the sampled data to obtain the main score of the original data; weighted summing the bias score of the original data and the main score of the original data to obtain the score of the original data; and using a normalized exponential function to calculate the score of the original data to obtain the first score. The adaptive learning ability of the dual model branches can be used to control the areas that the models at different stages need to focus on, so that the fusion of the two branches ultimately improves the long-tail recognition ability.

[0040] According to the technical solution provided by the embodiment of the present disclosure, by using the training result of a part of the first model branch and the sampled data as the input data of the second model branch of the data processing model, the second model branch predicts the second score for representing the prediction result of the sampled data based on the input data, including: obtaining the embedding vector of the sampled data based on the long-tail scene features of the data and the sampled data; using a part of the second neural network to perform training based on the sampled data to obtain the training result of the sampled data; learning the embedding vector of the sampled data based on the training result of the sampled data to obtain the main score of the sampled data; performing weighted summation of the bias score of the sampled data and the main score of the sampled data to obtain the score of the sampled data; and using a normalized exponential function to calculate the score of the sampled data to obtain the second score. The adaptive learning ability of the dual model branches can be used to control the areas that the models at different stages need to focus on, so that the fusion of the two branches ultimately improves the long-tail recognition ability.

[0041] According to the technical solution provided by the embodiment of the present disclosure, the first neural network has multiple multi-layer perceptrons MLP connected in series, and the second neural network has multiple multi-layer perceptrons MLP connected in series, so that the training of the full amount of data can be better completed.

[0042] According to the technical solution provided by the embodiments of the present disclosure, the bias features of the data are features that represent the impact of the data on user selection when presented in different positions, and the main features of the data are features that represent the common features of the data when it is generated, so that the training of the full amount of data can be better completed.

[0043] According to the technical solution provided by the embodiments of the present disclosure, the long-tail scene features of the data are used as features representing the long-tail scene to which the data belongs, so that the training of the long-tail scene data can be better completed.

[0044] According to the technical solution provided by the embodiment of the present disclosure, the first score and the second score are weighted and summed based on the adaptive parameters corresponding to the corresponding training batch to obtain an optimized first score, and the optimized first score represents the prediction result after the original data is optimized for long-tail scenario data, including: in the training batch, the adaptive parameters corresponding to the corresponding training batch are adjusted from large to small to gradually weight the first score, and at the same time, the adaptive parameters are subtracted from 1 to weight the second score. The adaptive learning ability of the dual-model branch can be used to control the areas that the model needs to pay attention to at different stages, so that the two branches are merged to ultimately improve the long-tail recognition ability. Moreover, the embodiment of the present disclosure uses the full data and sampled data as the original data for simultaneous training, which to a certain extent avoids the risks of overfitting of long-tail data and underfitting of the full data.

[0045] According to the technical solution provided by the embodiments of the present disclosure, the original data is traffic data, and the long-tail scene data is data of traffic scenes whose probability of appearing in the traffic data is lower than a preset threshold. The adaptive learning ability of the dual-model branches can be used to control the areas that the models need to focus on at different stages, so that the two branches are fused to ultimately improve the long-tail traffic scene recognition capability.

[0046] According to the technical solution provided by the embodiment of the present disclosure, an acquisition module is configured to acquire original data; a sampling data acquisition module is configured to resample the original data to acquire sampling data, wherein the resampling ratio of long-tail data in the original data is high; a first score calculation module is configured to use the original data and the sampling data as input data of a first model branch of a data processing model, and the first model branch predicts a first score representing a prediction result of the original data based on the input data; a second score calculation module is configured to use a training result of a part of the first model branch and the sampling data as input data of the data processing model The input data of the second model branch is used, and the second model branch predicts the second score for representing the prediction result of the sampled data based on the input data; the adaptive parameter determination module is configured to determine the adaptive parameters based on the training batch; the first score optimization module is configured to perform weighted summation of the first score and the second score based on the adaptive parameters corresponding to the corresponding training batch to obtain the optimized first score, and the optimized first score represents the prediction result of the original data after the long-tail scene data is optimized. The adaptive learning ability of the dual model branches can be used to control the areas that the model needs to pay attention to at different stages, so that the two branches are merged to ultimately improve the long-tail recognition ability. Moreover, the disclosed embodiment uses the full data and the sampled data as the original data for simultaneous training, which to a certain extent avoids the risks of overfitting of the long-tail data and underfitting of the full data.

[0047] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Other features, objectives and advantages of the present disclosure will become more apparent through the following detailed description of non-limiting embodiments in conjunction with the accompanying drawings. In the accompanying drawings:

[0049] Figure 1 A flowchart showing a data processing method according to an embodiment of the present disclosure;

[0050] Figure 2 An exemplary schematic diagram showing an implementation process of a data processing method according to an embodiment of the present disclosure;

[0051] Figure 3 A schematic diagram showing a change process of an adaptive parameter used in a data processing method according to an embodiment of the present disclosure;

[0052] Figure 4 A structural block diagram of a data processing device according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0053] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so that those skilled in the art can easily implement them. In addition, for the sake of clarity, parts not related to the description of the exemplary embodiments are omitted in the accompanying drawings.

[0054] In the present disclosure, it should be understood that terms such as "include" or "have" are intended to indicate the presence of labels, numbers, steps, behaviors, components, parts, or combinations thereof disclosed in the present specification, and are not intended to exclude the possibility that one or more other labels, numbers, steps, behaviors, components, parts, or combinations thereof exist or are added.

[0055] It should also be noted that, in the absence of conflict, the embodiments and labels in the embodiments of the present disclosure can be combined with each other. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0056] In the field of recommendation, the long-tail problem has always been a challenge in many scenarios. For example, the fields facing the long-tail problem include product recommendations in online shopping, information recommendations in search engines, and route recommendations in applications with map navigation functions. For the long-tail problem, relevant solutions include re-weighting and re-sampling.

[0057] Reweighting scheme: weight the sample loss.

[0058] Advantages: Network training can be adjusted, which can promote the learning of classifiers to a certain extent.

[0059] Disadvantages: Directly changing or flipping the frequency of occurrence of the original data distorts the original data and damages the learning ability of deep features to a certain extent.

[0060] Resampling scheme: Samples are sampled within small batches.

[0061] Advantages: It can balance data distribution and ensure the learning of various data.

[0062] Disadvantages: There is a risk of overfitting long-tail data and underfitting the full data.

[0063] In the solution of the embodiment of the present disclosure, a new model is proposed to focus on the long-tail problem, and the model is adaptively considered to learn both original features and long-tail scenarios through the dual-branch fusion idea, which greatly improves the long-tail recognition ability while ensuring the original feature learning effect. Moreover, the risk of overfitting of long-tail data and underfitting of full data can be avoided.

[0064] According to the technical solution provided by the embodiment of the present disclosure, by obtaining the original data; resampling the original data to obtain the sampled data, wherein the resampling sampling ratio of the long-tail data in the original data is high; using the original data and the sampled data as the input data of the first model branch of the data processing model, the first model branch predicts based on the input data to obtain a first score for representing the prediction result of the original data; using the training result of a part of the first model branch and the sampled data as the input data of the second model branch of the data processing model, the second model branch predicts based on the input data to obtain a second score for representing the prediction result of the sampled data; determining the adaptive parameters based on the training batch; based on the adaptive parameters corresponding to the corresponding training batch, the first score and the second score are weighted and summed to obtain an optimized first score, the optimized first score represents the prediction result of the original data after the long-tail scene data is optimized, and the adaptive learning ability of the dual model branch can be used to control the areas that need to be paid attention to in the model at different stages, so that the two branches are merged to finally improve the long-tail recognition ability. Moreover, the embodiment of the present disclosure uses the full amount of data and the sampled data as the original data for simultaneous training, which to a certain extent avoids the risk of overfitting of the long-tail data and underfitting of the full amount of data.

[0065] In order to solve the above problems, the present disclosure proposes a data processing method, an apparatus and a computer program product.

[0066] Figure 1 A flow chart of a data processing method according to an embodiment of the present disclosure is shown. Figure 1 As shown, the data processing method includes steps S101, S102, S103, S104, S105, and S106.

[0067] In step S101, original data is acquired.

[0068] In step S102, the original data is resampled to obtain sampled data, wherein the resampled sampling ratio of the long-tail data in the original data is high.

[0069] In step S103, the original data and the sampled data are used as input data of a first model branch of a data processing model, and the first model branch predicts a first score representing a prediction result of the original data based on the input data.

[0070] In step S104, a training result of a part of the first model branch and the sampling data are used as input data of a second model branch of the data processing model, and the second model branch predicts a second score representing a prediction result of the sampling data based on the input data.

[0071] In step S105 , adaptive parameters are determined based on the training batch.

[0072] In step S106, based on the adaptive parameters corresponding to the corresponding training batch, the first score and the second score are weightedly summed to obtain an optimized first score, and the optimized first score represents the prediction result after the long-tail scenario data is optimized for the original data.

[0073] In one embodiment of the present disclosure, the original data may also be considered as full data. For example, the navigation history data of all users in an application with a map navigation function. The original data is resampled to obtain sampled data, wherein the resampled sampling ratio of the long-tail data in the original data is high. In other words, sampled data that is biased towards the long-tail scene data in the original data can be obtained by resampling based on the original data, that is, data obtained by proportionally sampling the original data in a manner biased towards the long-tail scene data using a resampling scheme. Therefore, the proportion of long-tail scene data in the sampled data is higher than the proportion of long-tail scene data in the original data. In one embodiment of the present disclosure, the original data and the sampled data are used to train the data processing model at the same time, that is, a dual-branch data processing model optimized for long-tail scene data.

[0074] In one embodiment of the present disclosure, a dual-branch data processing model includes a first model branch and a second model branch. In one embodiment of the present disclosure, the first model branch may be a representation learning model, and the second model branch may be a classifier learning model. The first model branch is trained based on the original data and the sampled data to obtain a first score, and the training result of a part of the first model branch and the second model branch are used to train based on the sampled data to obtain a second score. In one embodiment of the present disclosure, the adaptive parameters are determined based on the training batch, that is, the adaptive parameters may be different in each training batch. Based on the adaptive parameters corresponding to the corresponding training batch, the first score and the second score are weighted and summed to obtain an optimized first score, and the optimized first score represents the prediction result after the original data is optimized for long-tail scene data. That is, after sufficient training using the first model branch and the second model branch, a dual-branch data processing model optimized for long-tail scene data can be obtained. In one embodiment of the present disclosure, the training of the dual-branch data processing model requires iterative training through multiple training batches. The total loss function of the model is obtained by weighted calculation of the loss functions of the two branches according to the adaptive parameters determined based on the training batches, so that the model focuses on learning the original feature distribution in the early stage and tends to learn long-tail data in the later stage. When the original features are learned to a certain extent, the long-tail recognition ability of the model is improved.

[0075] The following use Figure 2 The example of the implementation process of the data processing method shown is used to further describe the data processing method in the embodiment of the present disclosure. Figure 2 An exemplary schematic diagram showing the implementation process of a data processing method according to an embodiment of the present disclosure.

[0076] Figure 2 A dual-branch data processing model for implementing data processing is shown. Figure 2 As shown, with the vertical dotted line in the figure as the boundary, the left branch part in the dual-branch data processing model of the data processing method of the embodiment of the present disclosure can be called the first model branch, and the branch part on the right side of the dotted line can be called the second model branch.

[0077] like Figure 2 As shown, the first model branch includes a first neural network 220 and a second neural network 240. Figure 2 middle, Figure 1Step S103 in the method includes: using the first neural network 220, training based on the bias feature 221 of the data, the original data 210 and the sampled data 240 to obtain the bias score bias_op_u 225 of the original data 210 and the bias score bias_op_s 244 of the sampled data 240. Using the second neural network 230, training based on the main feature 231 of the data, the original data 210 and the sampled data 240 to obtain the main score main_op_u 235 of the original data 210. Weighted summing the bias score bias_op_u 225 of the original data 210 and the main score main_op_u 235 of the original data 210 is performed to obtain the score score_u 226 of the original data 210. Using the normalized exponential function softmax_u 227 to calculate the score score_u 226 of the original data 210 to obtain the first score loss_u 238.

[0078] In one embodiment of the present disclosure, loss represents a loss function, which reflects the degree of fit of the model to the data. The larger the loss value, the worse the fit, and vice versa.

[0079] In one embodiment of the present disclosure, the bias feature of the data is a feature that indicates the influence of the data on the user's selection when the data is presented at different positions, and the main feature of the data is a universal feature that the data possesses when it is generated.

[0080] According to the technical solution provided by the embodiments of the present disclosure, the bias features of the data are features that represent the impact of the data on user selection when presented in different positions, and the main features of the data are features that represent the common features of the data when it is generated, so that the training of the full amount of data can be better completed.

[0081] In one embodiment of the present disclosure, the setting of the bias feature 221 can correctly classify the original data and the sampled data, can better fit the data, and can represent the sending position bias, similarity bias, etc. For example, in an application with a map navigation function, the original data and the sampled data include various navigation information, such as a navigation recommended route. In this case, the bias feature 221 may include the location where the application with the map navigation function presents the navigation recommended route. For example, the application with the map navigation function allows 3 navigation recommended routes to be presented on the interface for user selection, and in addition, there are hidden navigation recommended routes that are not presented on the interface. In this case, this bias feature may include 4 values, which respectively represent the navigation recommended route presented at the first position of the interface, the navigation recommended route presented at the second position of the interface, the navigation recommended route presented at the third position of the interface, and the navigation recommended route not presented on the interface. The purpose of introducing such a bias feature to the original data and the sampled data using the first model branch is to exclude the influence of the presentation position of the navigation recommended route on the interface on the user's choice when recommending navigation routes to users for long-tail scenarios using the trained dual-branch data processing model. When the data processing method of the embodiment of the present disclosure is applied to recommend news, advertisements, etc. on the Internet, the bias feature also plays a similar role.

[0082] In one embodiment of the present disclosure, the bias score bias_op_u 225 of the original data 210 represents the prediction result of the original data 210 based on the bias feature 221 of the data using the first neural network 220. The bias score bias_op_u 225 of the sampled data 240 represents the prediction result of the original data 210 based on the bias feature 221 of the data using the first neural network 220. The bias score bias_op_s 244 of the sampled data 240 represents the prediction result of the sampled data 240 based on the bias feature 221 of the data using the first neural network 220.

[0083] In one embodiment of the present disclosure, training is performed offline using the first neural network 220 based on the bias feature 221 of the data, the original data 210, and the sampled data 240 to obtain the bias score bias_op_u 225 of the original data 210 and the bias score bias_op_s 244 of the sampled data 240, that is, the bias score bias_op_u 225 of the original data 210 and the bias score bias_op_s 244 of the sampled data 240 are values ​​calculated offline.

[0084] In one embodiment of the present disclosure, the setting of the main feature 231 can correctly classify the original data and the sampled data, and can better fit. For example, in an application with a map navigation function, the main feature can be a data route feature, a statistical feature, a user's historical preference, and other features. The original data and the sampled data include various navigation information, such as a navigation recommended route. In this case, the main feature 231 may include the user settings when generating a navigation recommended route in an application with a map navigation function, for example, the user's historical preference features set by the user in an application with a map navigation function about highways, toll roads, and avoiding congestion, and for another example, the main feature may include data route features such as lane width and number of traffic lights in the navigation route, and for another example, the main feature may include statistical features such as the number of users passing through a certain route in the navigation route. In an embodiment of the present disclosure, the second neural network 230 is used to train based on the main feature 231 of the data, the original data 210, and the sampled data 240 to obtain the main score main_op_u 235 of the original data 210. The main score main_op_u 235 of the original data 210 represents the prediction result of the original data 210 based on the main feature 231 of the data using the second neural network 230.

[0085] In one embodiment of the present disclosure, a weighted sum is performed on the bias score bias_op_u 225 of the original data 210 and the main score main_op_u 235 of the original data 210 to obtain a score score_u 226 of the original data 210. The score score_u 226 of the original data 210 represents the prediction result of the original data 210 based on the bias feature 221 and the main feature 231 of the data using the first model branch.

[0086] In one embodiment of the present disclosure, the value of the score_u 226 of the original data 210 will be in a larger value range, and the difference between different score_u 226 is large, and it is difficult to maintain a value suitable for comparison in the global step (global_step). Therefore, the score_u 226 of the original data 210 is calculated using the normalized exponential function softmax_u 227 to obtain the first score loss_u 238, that is, the first score loss_u 238 is a value in the range of 0-1. In one embodiment of the present disclosure, loss reflects the degree of fit of the model to the data. The larger the loss value, the worse the fit, and vice versa. In one embodiment of the present disclosure, for an application with a map navigation function, the actual coverage rate of the recommended navigation route recommended by the application with a map navigation function to the user each time for the actual route walked by the user and the coverage rate of the recommended navigation route recommended by the application with a map navigation function to the user based on the dual-branch data processing model The actual walking route of the user is the difference between the actual coverage rate of the actual walking route of the user. In one embodiment of the present disclosure, for the original data 210 , the gap is the first score loss_u 238 .

[0087] According to the technical solution provided by the embodiment of the present disclosure, the first model branch includes a first neural network and a second neural network, wherein the original data and the sampled data are used as input data of the first model branch of the data processing model, and the first model branch predicts the first score for representing the prediction result of the original data based on the input data, including: using the first neural network to perform training based on the bias features of the data, the original data and the sampled data to obtain the bias score of the original data and the bias score of the sampled data; using the second neural network to perform training based on the main features of the data, the original data and the sampled data to obtain the main score of the original data; weighted summing the bias score of the original data and the main score of the original data to obtain the score of the original data; and using a normalized exponential function to calculate the score of the original data to obtain the first score. The adaptive learning ability of the dual model branches can be used to control the areas that the models at different stages need to focus on, so that the fusion of the two branches ultimately improves the long-tail recognition ability.

[0088] Reference Figure 2 In one embodiment of the present disclosure, step S104 includes: obtaining an embedding vector 242 of the sampled data 240 based on the long-tail scene feature 241 of the data and the sampled data 240. Using a part of the second neural network 230 ( Figure 2The first two multi-layer perceptrons (MLPs 232 and 233) shown in FIG. 2 are trained based on the sample data 240 to obtain training results of the sample data 240. The training results (MLPs 232 and 233) based on the sample data 240 are trained based on the sample data 240 to obtain training results of the sample data 240. Figure 2 The multi-layer perceptron MLP 233 shown learns the embedding vector 242 of the sampled data 240 (the output of the second model branch) to obtain the main score main_op_s 245 of the sampled data 240. The bias score bias_op_s 244 of the sampled data 240 and the main score main_op_s 245 of the sampled data 240 are weighted summed to obtain the score score_s 246 of the sampled data 240. The score score_s 246 of the sampled data 240 is calculated using the normalized exponential function softmax_s 246 to obtain the second score loss_s 248.

[0089] In one embodiment of the present disclosure, the long-tail scenario feature of the data is a feature representing the long-tail scenario to which the data belongs.

[0090] According to the technical solution provided by the embodiments of the present disclosure, the long-tail scene features of the data are used as features representing the long-tail scene to which the data belongs, so that the training of the long-tail scene data can be better completed.

[0091] In one embodiment of the present disclosure, in an application with a map navigation function, various features in the navigation recommended routes in long-tail scenarios such as commuting, stations, intra-city and inter-city, and familiar roads, such as specific landmark buildings, can be used as long-tail scenario features of the data. For example, specific landmark buildings may include specific bridges, specific stations, specific buildings, and so on in the navigation recommended routes in long-tail scenarios. In one embodiment of the present disclosure, embedding refers to a way of converting discrete variables into continuous vector representations. In neural networks, embedding can not only reduce the spatial dimension of discrete variables, but also represent the variables meaningfully. In one embodiment of the present disclosure, the discrete representations of different scenarios will have an embedding vector after embedding, and this vector will be continuously learned and updated during the training process.

[0092] In one embodiment of the present disclosure, a portion of the second neural network 230 in the first model branch ( Figure 2The first two multi-layer perceptrons MLP 232 and 233 shown are trained based on the sampled data 240 to obtain the training results of the sampled data 240 as an input of the second model branch. In one embodiment of the present disclosure, the embedding vector 242 of the sampled data 240 is learned based on the training results of the sampled data 240 from the multi-layer perceptron MLP 233 to obtain the main score main_op_s 245 of the sampled data 240. The main score main_op_s 245 of the sampled data 240 represents the prediction result of the sampled data 240 based on the long-tail scene feature 241 of the data using the second model branch.

[0093] In one embodiment of the present disclosure, a bias score bias_op_s 244 of the sampled data 240 output from the first neural network 220 and a main score main_op_s 245 of the sampled data 240 are weighted summed to obtain a score score_s 246 of the sampled data 240. The score score_s 246 of the sampled data 240 represents a prediction result of the sampled data 240 based on the bias feature 221 and the main feature 231 of the data using the second model branch.

[0094] In one embodiment of the present disclosure, the value of the score_s 246 of the sampled data 240 is in a relatively large value range, and the difference between different score_s 246 is relatively large, and it is difficult to maintain a value suitable for comparison in the global step (global_step). Therefore, the score_s 246 of the sampled data 240 is calculated using the normalized exponential function softmax_s 246 to obtain the second score loss_s 248, that is, the second score loss_s 248 is a value in the range of 0-1.

[0095] According to the technical solution provided by the embodiment of the present disclosure, by using the training result of a part of the first model branch and the sampled data as the input data of the second model branch of the data processing model, the second model branch predicts the second score for representing the prediction result of the sampled data based on the input data, including: obtaining the embedding vector of the sampled data based on the long-tail scene features of the data and the sampled data; using a part of the second neural network to perform training based on the sampled data to obtain the training result of the sampled data; learning the embedding vector of the sampled data based on the training result of the sampled data to obtain the main score of the sampled data; performing weighted summation of the bias score of the sampled data and the main score of the sampled data to obtain the score of the sampled data; and using a normalized exponential function to calculate the score of the sampled data to obtain the second score. The adaptive learning ability of the dual model branches can be used to control the areas that the models at different stages need to focus on, so that the fusion of the two branches ultimately improves the long-tail recognition ability.

[0096] In one embodiment of the present disclosure, the first neural network 220 has a plurality of multi-layer perceptrons MLPs 222 , 223 , and 224 connected in series, and the second neural network 230 has a plurality of multi-layer perceptrons MLPs 232 , 233 , and 234 connected in series.

[0097] In the related art, a multilayer perceptron (MLP) is also called an artificial neural network (ANN). In addition to the input and output layers, there may be multiple hidden layers in the middle. The simplest multilayer perceptron MLP contains only one hidden layer, that is, a three-layer structure. Its specific form can be obtained from the related art, and this disclosure will not elaborate on this. In one embodiment of the present disclosure, Figure 2 As shown, the first neural network 220 has three multi-layer perceptrons MLP 222, 223, 224 connected in series, and the second neural network 230 has three multi-layer perceptrons MLP 232, 233, 234 connected in series. However, this is only an example, and the first neural network 220 and the second neural network 230 may have more or fewer multi-layer perceptrons MLP connected in series. Figure 2The part of the second neural network 230 shown, i.e., the first two multi-layer perceptrons MLP232 and 233, are trained based on the sample data 240 to obtain the training results of the sample data 240 for example only. In the case where the second neural network 230 includes four or more multi-layer perceptrons MLP connected in series, the part of the second neural network 230 may include three or more multi-layer perceptrons MLP to be trained based on the sample data 240 to obtain the training results of the sample data 240 as an input of the second model branch.

[0098] According to the technical solution provided by the embodiment of the present disclosure, the first neural network has multiple multi-layer perceptrons MLP connected in series, and the second neural network has multiple multi-layer perceptrons MLP connected in series, so that the training of the full amount of data can be better completed.

[0099] In one embodiment of the present disclosure, step S106 includes: in the training batch, adjusting the adaptive parameter corresponding to the corresponding training batch from large to small to gradually weight the first score loss_u 228, and at the same time subtracting the adaptive parameter from 1 to weight the second score loss_s 248.

[0100] In one embodiment of the present disclosure, the training of the dual-branch data processing model needs to be iteratively trained through multiple training batches. In one embodiment of the present disclosure, the training batch can be represented by global_step, that is, the number of iterations of the current training.

[0101] like Figure 2 As shown, through the formula shown in box 260, in each training batch, the adaptive parameter alpha corresponding to the corresponding training batch is adjusted from large to small to gradually weight the first score loss_u 228, and correspondingly 1-alpha is adjusted from small to large to gradually weight the second score loss_s 248.

[0102] The following reference Figure 3 The adaptive parameters used in the data processing method according to one embodiment of the present disclosure are described. Figure 3 A schematic diagram showing a change process of an adaptive parameter used in a data processing method according to an embodiment of the present disclosure.

[0103] In one embodiment of the present disclosure, the adaptive parameter alpha refers to an automatic parameter used to control the learning of the first model branch and the second model branch. The specific calculation method is to specify an upper limit training batch global_step, for example, called the maximum training batch max_step. According to the preset formula, the training batch global_step of the model changes from 1 to max_step. At the same time, the adaptive parameter alpha can be achieved as follows: Figure 3 As shown, it decreases smoothly as the training batch global_step increases until it reaches 0.

[0104] exist Figure 2 The formula alpha=1-(g / g max ) 2 In the example, g refers to the current training batch global_step, g max Refers to the maximum training batch max_step. By using the adaptive parameter alpha to weight the first score loss_u 228 as the training batch increases, and using 1-alpha to weight the second score loss_s 248. As the training batch global_step gradually increases, the adaptive parameter alpha can be gradually increased from a small value to make the dual-branch data processing model focus on learning the feature distribution of the original data in the early stage (learning through the first model branch), and tend to learn the features of the long-tail scene data in the later stage (learning through the second model branch). When the features of the original data are learned to a certain extent, the long-tail scene recognition ability of the dual-branch data processing model is improved.

[0105] In one embodiment of the present disclosure, Figure 2As shown, in each training batch, the adaptive parameter alpha corresponding to the corresponding training batch is adjusted from large to small to gradually weight the first score loss_u 228 to obtain the weighted first score loss_u 229, and 1-alpha is adjusted from small to large to gradually weight the second score loss_s 248 to obtain the weighted second score loss_s 249. The weighted first score loss_u 229 and the weighted second score loss_s 249 are weighted and summed to calculate the optimized first score loss_u 270 to represent the prediction result of the original data 210 that has been optimized based on the long-tail scene data 240. That is, based on the trained dual-branch data processing model, the recognition ability of the long-tail scene in the original data can be improved. In one embodiment of the present disclosure, by back-propagating the optimized first score loss_u 270, the dual-branch data processing model can adjust its internal parameters, so that the dual-branch data processing model optimizes its learning effect. The specific method of back-propagating the optimized first score loss_u 270 can be understood from the relevant technology, and the present disclosure will not elaborate on it.

[0106] According to the technical solution provided by the embodiment of the present disclosure, the first score and the second score are weighted and summed based on the adaptive parameters corresponding to the corresponding training batch to obtain an optimized first score, and the optimized first score represents the prediction result after the original data is optimized for long-tail scenario data, including: in the training batch, the adaptive parameters corresponding to the corresponding training batch are adjusted from large to small to gradually weight the first score, and at the same time, the adaptive parameters are subtracted from 1 to weight the second score. The adaptive learning ability of the dual-model branch can be used to control the areas that the model needs to pay attention to at different stages, so that the two branches are merged to ultimately improve the long-tail recognition ability. Moreover, the embodiment of the present disclosure uses the full data and sampled data as the original data for simultaneous training, which to a certain extent avoids the risks of overfitting of long-tail data and underfitting of the full data.

[0107] In one embodiment of the present disclosure, the original data 210 is traffic data, and the long-tail scene data 240 is data of traffic scenes whose occurrence probability is lower than a preset threshold in the traffic data 210 .

[0108] In one embodiment of the present disclosure, in the personalized recommendation scenario of driving routes in an application with a map navigation function, there are many long-tail scenarios with sparse data but important, such as commuting, stations, inter-city, unfamiliar roads, etc. The diversity of scenarios also puts forward higher requirements and challenges for personalized recommendation of routes. The needs of users may vary due to different scenarios. In this way, the user's preferences for speed, time, tolls, highways, congestion, etc. may be completely different. This requires the model to have a certain ability to capture, learn and distinguish the user group behavior of each long-tail scenario, and meet the needs of users in multiple scenarios while ensuring the quality of route recommendations. Therefore, the data processing scheme of the embodiment of the present disclosure is suitable for processing traffic scene data and improving the ability to recognize long-tail traffic scenes.

[0109] According to the technical solution provided by the embodiments of the present disclosure, the original data is traffic data, and the long-tail scene data is data of traffic scenes whose probability of appearing in the traffic data is lower than a preset threshold. The adaptive learning ability of the dual-model branches can be used to control the areas that the models need to focus on at different stages, so that the two branches are fused to ultimately improve the long-tail traffic scene recognition capability.

[0110] Those skilled in the art can understand that, according to the above teachings of the embodiments of the present disclosure, the data processing scheme of the embodiments of the present disclosure can be applied to long-tail scene data recognition in other application fields.

[0111] The following reference Figure 4 A data processing device according to an embodiment of the present disclosure is described. Figure 4 A structural block diagram of a data processing device 400 according to an embodiment of the present disclosure is shown.

[0112] like Figure 4 As shown, the data processing device 400 includes:

[0113] An acquisition module 401 is configured to acquire original data;

[0114] The sampling data acquisition module 402 is configured to resample the original data to obtain the sampling data, wherein the resampled sampling ratio of the long tail data in the original data is high;

[0115] A first score calculation module 403 is configured to use the original data and the sampled data as input data of a first model branch of a data processing model, and the first model branch predicts based on the input data to obtain a first score representing a prediction result of the original data;

[0116] A second score calculation module 404 is configured to use a training result of a part of the first model branch and the sample data as input data of a second model branch of the data processing model, and the second model branch predicts based on the input data to obtain a second score representing a prediction result of the sample data;

[0117] An adaptive parameter determination module 405, configured to determine an adaptive parameter based on a training batch;

[0118] The first score optimization module 406 is configured to perform a weighted summation of the first score and the second score based on the adaptive parameters corresponding to the corresponding training batch to obtain an optimized first score, wherein the optimized first score represents the prediction result after the original data is optimized for long-tail scenario data.

[0119] According to the technical solution provided by the embodiment of the present disclosure, an acquisition module is configured to acquire original data; a sampling data acquisition module is configured to resample the original data to acquire sampling data, wherein the resampling ratio of long-tail data in the original data is high; a first score calculation module is configured to use the original data and the sampling data as input data of a first model branch of a data processing model, and the first model branch predicts a first score representing a prediction result of the original data based on the input data; a second score calculation module is configured to use a training result of a part of the first model branch and the sampling data as input data of the data processing model The input data of the second model branch is used, and the second model branch predicts the second score for representing the prediction result of the sampled data based on the input data; the adaptive parameter determination module is configured to determine the adaptive parameters based on the training batch; the first score optimization module is configured to perform weighted summation of the first score and the second score based on the adaptive parameters corresponding to the corresponding training batch to obtain the optimized first score, and the optimized first score represents the prediction result of the original data after the long-tail scene data is optimized. The adaptive learning ability of the dual model branches can be used to control the areas that the model needs to pay attention to at different stages, so that the two branches are merged to ultimately improve the long-tail recognition ability. Moreover, the disclosed embodiment uses the full data and the sampled data as the original data for simultaneous training, which to a certain extent avoids the risks of overfitting of the long-tail data and underfitting of the full data.

[0120] Those skilled in the art will understand that referring to Figure 4 The technical solutions described can be compared with Figures 1 to 3 The described embodiments are combined to provide a reference Figures 1 to 3 The technical effects achieved by the described embodiments. Figures 1 to 3The specific content of the description will not be repeated here.

[0121] In particular, according to an embodiment of the present disclosure, the method described above with reference to the accompanying drawings can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, including computer instructions, which implement the method described with reference to the accompanying drawings when the computer instructions are executed by a processor. The computer program product includes a computer program tangibly contained on a readable medium thereof, and the computer program includes a program code for executing the method in the accompanying drawings. In such an embodiment, the computer program can be downloaded and installed from a network through a communication part, and / or installed from a removable medium. For example, an embodiment of the present disclosure includes a readable storage medium on which computer instructions are stored, and when the computer instructions are executed by a processor, the program code for executing the method in the accompanying drawings is implemented.

[0122] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the methods, devices and computer program products according to various embodiments of the present disclosure. In this regard, each box in the road map or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0123] The units or modules involved in the embodiments described in the present disclosure may be implemented by software or hardware. The units or modules described may also be set in a processor, and the names of these units or modules do not constitute limitations on the units or modules themselves in some cases.

[0124] As another aspect, the present disclosure further provides a computer-readable storage medium, which may be a computer-readable storage medium included in the node described in the above embodiment; or a computer-readable storage medium that exists independently and is not installed in a device. The computer-readable storage medium stores one or more programs, and the programs are used by one or more processors to execute the method described in the present disclosure.

[0125] The above description is only a preferred embodiment of the present disclosure and an explanation of the technical principles used. Those skilled in the art should understand that the scope of the invention involved in the present disclosure is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the inventive concept. For example, the above features are replaced with the technical features with similar functions disclosed in the present disclosure (but not limited to) by each other.

Claims

1. A data processing method, wherein: The method comprises: Get the original data; Resampling the original data to obtain sampled data, wherein the resampled sampling ratio of the long-tail data in the original data is high; Using the original data and the sampled data as input data of a first model branch of a data processing model, and having the first model branch predict based on the input data to obtain a first score representing a prediction result of the original data; Using a training result of a part of the first model branch and the sampled data as input data of a second model branch of the data processing model, and the second model branch predicts based on the input data to obtain a second score representing a prediction result of the sampled data; Determining adaptive parameters based on the training batch; Based on the adaptive parameters corresponding to the corresponding training batch, the first score and the second score are weightedly summed to obtain an optimized first score, wherein the optimized first score represents a prediction result after the original data is optimized for long-tail scenario data. The original data is traffic data, and the long-tail scene data is data of traffic scenes whose occurrence probability in the traffic data is lower than a preset threshold.

2. The method according to claim 1, wherein: The first model branch includes a first neural network and a second neural network, wherein: The method of using the original data and the sampled data as input data of a first model branch of a data processing model, and predicting by the first model branch based on the input data to obtain a first score representing a prediction result of the original data, comprises: Using the first neural network, training is performed based on the bias characteristics of the data, the original data, and the sampled data to obtain a bias score of the original data and a bias score of the sampled data; Using the second neural network, training is performed based on the main features of the data, the original data, and the sampled data to obtain the main score of the original data; Performing a weighted summation on the bias score of the original data and the main score of the original data to obtain the score of the original data; The score of the original data is calculated using a normalized exponential function to obtain the first score.

3. The method according to claim 2, wherein: The method of using the training result of a part of the first model branch and the sampled data as input data of a second model branch of the data processing model, and predicting by the second model branch based on the input data to obtain a second score representing a prediction result of the sampled data, comprises: Obtaining an embedding vector of the sampled data based on the long-tail scene features of the data and the sampled data; Using a part of the second neural network to perform training based on the sampled data to obtain a training result of the sampled data; Learning the embedding vector of the sampled data based on the training result of the sampled data to obtain the principal score of the sampled data; Performing a weighted summation on the bias score of the sampled data and the main score of the sampled data to obtain the score of the sampled data; The score of the sampled data is calculated using a normalized exponential function to obtain the second score.

4. The method according to claim 2 or 3, wherein: The first neural network has a plurality of multi-layer perceptrons MLP connected in series, and the second neural network has a plurality of multi-layer perceptrons MLP connected in series.

5. The method according to claim 2, wherein: The bias feature of the data is a feature that represents the influence of the data on the user's selection when the data is presented at different positions, and the main feature of the data is a universal feature that the data possesses when it is generated.

6. The method according to claim 3, wherein: The long-tail scene feature of the data is a feature representing the long-tail scene to which the data belongs.

7. The method according to claim 1, wherein: The first score and the second score are weightedly summed based on the adaptive parameters corresponding to the corresponding training batch to obtain an optimized first score, wherein the optimized first score represents a prediction result after the original data is optimized for long-tail scenario data, including: In the training batch, the adaptive parameter corresponding to the corresponding training batch is adjusted from large to small to gradually weight the first score, and at the same time, the adaptive parameter is subtracted from 1 to weight the second score.

8. A data processing device, wherein: The device comprises: An acquisition module, configured to acquire raw data; A sampling data acquisition module is configured to resample the original data to obtain sampled data, wherein the resampled sampling ratio of the long-tail data in the original data is high; A first score calculation module is configured to use the original data and the sampled data as input data of a first model branch of a data processing model, and the first model branch predicts based on the input data to obtain a first score representing a prediction result of the original data; A second score calculation module is configured to use a training result of a part of the first model branch and the sample data as input data of a second model branch of the data processing model, and the second model branch predicts based on the input data to obtain a second score representing a prediction result of the sample data; An adaptive parameter determination module, configured to determine an adaptive parameter based on a training batch; A first score optimization module is configured to perform a weighted summation of the first score and the second score based on the adaptive parameters corresponding to the corresponding training batch to obtain an optimized first score, wherein the optimized first score represents a prediction result after the original data is optimized for long-tail scenario data. The original data is traffic data, and the long-tail scene data is data of traffic scenes whose occurrence probability in the traffic data is lower than a preset threshold.

9. A computer program product, comprising computer instructions, which, when executed by a processor, implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Student model training method based on pre-training language model and text classification system

    CN115526332A