Data Expansion Method, Apparatus and Storage Medium
By acquiring and analyzing the data sets of target business scenarios and associated business scenarios, determining feature data and selecting target extension data, the problem of low data expansion efficiency in the existing technology is solved, and fast and efficient data expansion is achieved.
Patent Information
- Application Number
- CN202210625648.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-02
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-06-02
AI Technical Summary
Existing data scaling methods are inefficient and time-consuming when processing large amounts of data, and cannot quickly and effectively scale data.
By obtaining the reference data set in the target business scenario and the data set to be expanded for the associated business scenario, the reference feature data and target feature data of each data in the reference data set are determined, and the data to be expanded with the target feature data is selected as the target extension data.
It realizes rapid and efficient data expansion, reducing the time-consuming and improving the efficiency of data expansion.
Smart Images

Figure CN115017145B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a data expansion method, apparatus, and storage medium. Background Art
[0002] In the related art, when expanding data in a target business scenario, it is usually necessary to calculate the data similarity between the data in the target business scenario and the data in other business scenarios. Then, the data to be expanded is selected from other business scenarios according to the data similarity. This method requires full-scale calculation of the data in the scenario. Once the data volume in the scenario is large, the existing data expansion method not only takes a lot of time, but also has the problem of low data expansion efficiency. Summary of the Invention
[0003] The present invention provides a data expansion method, apparatus, and storage medium to achieve more rapid and effective data expansion, thereby reducing the time consumption of data expansion and further improving the efficiency of data expansion.
[0004] According to one aspect of the present invention, a data expansion method is provided, and the method includes:
[0005] Obtain a reference data set in a target business scenario and a data set to be expanded in an associated business scenario associated with the target business scenario;
[0006] Determine the reference feature data of each data in the reference data set, and determine the target feature data in the reference feature data;
[0007] Select first data to be expanded having the target feature data from the data set to be expanded, and use the first data to be expanded as the target expansion data of the target business scenario.
[0008] Optionally, determining the target feature data in the reference feature data includes:
[0009] Determine the positive feature data associated with the target business scenario in the reference feature data according to the importance coefficient of the reference feature data;
[0010] Determine the target feature data in the positive feature data based on the time interval between the current moment and the target update moment of the positive feature data.
[0011] Optionally, the determining the target feature data in the positive feature data based on the time interval between the current moment and the target update moment includes:
[0012] Calculate an update metric value of the positive feature data according to the time interval between the current moment and the target update moment, where the update metric value is used to measure the update interval duration of the positive feature data;
[0013] Select, from the update metric values, the update metric values that meet the preset metric conditions as target metric values, and use the positive feature data corresponding to the target metric values as target feature data.
[0014] Optionally, the calculating an update metric value of the positive feature data according to the time interval between the current moment and the target update moment includes:
[0015] Calculate the update metric value of the positive feature data according to the following formula:
[0016]
[0017] where X i ' represents the update metric value of the positive feature data, and X i represents the time interval between the current moment and the update moment of the positive feature data.
[0018] Optionally, the method further includes:
[0019] Determine the to-be-expanded feature data of each data in the to-be-expanded dataset;
[0020] Select second expansion data of the target business scenario from the to-be-expanded dataset according to each to-be-expanded feature data;
[0021] The using the first expansion data as the target expansion data of the target business scenario includes:
[0022] Using all users in the first expansion data and the second expansion data as the target expansion data of the target business scenario; or,
[0023] Using the common users in the first expansion data and the second expansion data as the target expansion data of the target business scenario.
[0024] Optionally, the selecting second expansion data of the target business scenario from the to-be-expanded dataset according to each to-be-expanded feature data includes:
[0025] Perform clustering processing on the target feature data to obtain a clustering center;
[0026] Calculate the data distance between each to-be-expanded feature data and the clustering center;
[0027] Select the second expansion data of the target business scenario from the to-be-expanded dataset based on the distances of the respective data.
[0028] Optionally, the step of selecting the second expansion data of the target business scenario from the to-be-expanded dataset according to the respective to-be-expanded feature data includes:
[0029] Input the to-be-expanded feature data into data expansion models for data expansion for different expansion dimensions respectively, to obtain data expansion values output by the respective data expansion models;
[0030] Perform weighted average processing on the respective data expansion values to obtain a target expansion value, and select the second expansion data in the target business scenario from the to-be-expanded dataset according to the target expansion value.
[0031] Optionally, the method further includes:
[0032] For the initial network model of each expansion dimension, obtain sample data and expected output data corresponding to the sample data, where the sample data includes positive sample data and negative sample data, the positive sample data includes the reference feature data, and the negative sample data is the feature data of other business scenarios except the target business scenario;
[0033] Input the sample data into the initial network model to obtain the actual output data of the initial network model;
[0034] Adjust the parameters of the initial network model according to the actual output data and the expected output data of the sample data, so as to obtain a data expansion model.
[0035] According to another aspect of the present invention, there is provided a data expansion device. The device includes:
[0036] A dataset acquisition module, configured to acquire a reference dataset in a target business scenario and a to-be-expanded dataset of an associated business scenario associated with the target business scenario;
[0037] A feature data determination module, configured to determine reference feature data of each data in the reference dataset and determine target feature data in the reference feature data;
[0038] A target expansion data acquisition module, configured to select first to-be-expanded data having the target feature data from the to-be-expanded dataset and use the first to-be-expanded data as the target expansion data of the target business scenario.
[0039] According to another aspect of the present invention, there is provided an electronic device, the electronic device includes:
[0040] At least one processor; and
[0041] A memory communicatively connected to the at least one processor; wherein,
[0042] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the data expansion method according to any embodiment of the present invention.
[0043] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the data expansion method according to any embodiment of the present invention when executed.
[0044] The technical solution of the embodiment of the present invention can more quickly determine the scope of data expansion by obtaining a reference data set in a target business scenario and a data set to be expanded in an associated business scenario associated with the target business scenario. Furthermore, the reference feature data of each data in the reference data set can be determined, and the target feature data in the reference feature data can be determined, so that the conditions for expanding the data can be obtained more accurately. After determining the target feature data, the first data to be expanded having the target feature data can be selected from the data set to be expanded, and the first data to be expanded is used as the target expansion data of the target business scenario. The technical solution in the embodiment of the present invention solves the problems that the existing data expansion method not only takes a lot of time, but also has low data expansion efficiency, realizes more rapid and effective data expansion, thereby reducing the time consumed by data expansion and further improving the efficiency of data expansion.
[0045] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is a schematic flowchart of a data expansion method provided in Embodiment 1 of the present invention;
[0048] Figure 2 It is a schematic flowchart of a data expansion method provided in Embodiment 2 of the present invention;
[0049] Figure 3 A flowchart of a data expansion method provided in Embodiment 3 of the present invention;
[0050] Figure 4 A structural diagram of a data expansion device provided in Embodiment 4 of the present invention;
[0051] Figure 5 A structural diagram of an electronic device provided in Embodiment 5 of the present invention. Detailed implementation manners
[0052] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0053] It should be understood that the steps recorded in the method embodiments of the present disclosure can be executed in different orders and / or executed in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0054] It can be understood that the data involved in the technical solution of the present invention (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the corresponding laws, regulations and related provisions.
[0055] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0056] Embodiment 1
[0057] Figure 1FIG. 0 is a schematic flowchart of a data expansion method provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of data expansion. The method can be executed by a data expansion device, which can be implemented in the form of hardware and / or software, and can be configured in an electronic device such as a computer or a server.
[0058] As Figure 1 shown, the method of this embodiment includes:
[0059] S110. Obtain a reference data set in a target business scenario and a data set to be expanded in an associated business scenario associated with the target business scenario.
[0060] Among them, the target business scenario can be understood as the business scenario that currently needs data expansion. The associated business scenario can be understood as the business scenario associated with the target business scenario. The number of associated business scenarios can be one, two or more. The reference data set can be understood as a set of all or part of the data in the target business scenario. Optionally, the reference data set can be a user set in the target business scenario. The data set to be expanded can be understood as a set of all or part of the data in the associated business scenario.
[0061] Optionally, the associated business scenario associated with the target business scenario is determined by the following method:
[0062] Determine each business processing scenario except the target business scenario; for each business processing scenario except the target business scenario, determine whether the business processing scenario except the target business scenario is an associated business scenario according to the scenario characteristics of the business processing scenario except the target business scenario and the scenario characteristics of the target business scenario.
[0063] Specifically, determining whether the business processing scenario except the target business scenario is an associated business scenario according to the scenario characteristics of the business processing scenario except the target business scenario and the scenario characteristics of the target business scenario includes: calculating the scenario similarity between the scenario characteristics of the business processing scenario except the target business scenario and the scenario characteristics of the target business scenario; if the scenario similarity meets the preset scenario similarity threshold, the business processing scenario except the target business scenario can be used as an associated business scenario; if the scenario similarity does not meet the preset scenario similarity threshold, it can be determined that the business processing scenario except the target business scenario is not an associated business scenario. Among them, the scenario similarity threshold can be set according to actual needs, such as 0.85, 0.9 or 0.95, etc., and no specific limitation is made here.
[0064] S120. Determine the reference feature data of each data in the reference data set and determine the target feature data in the reference feature data.
[0065] Among them, the reference feature data can be the feature data obtained by extracting features from each data in the reference data (for example, the feature data such as the age, region, gender, business behavior, etc. of each user in the user set). The number of reference feature data can be one, two, or more than two. The target feature data can be one or more feature data among the reference feature data.
[0066] Specifically, in the obtained reference data set, data feature extraction is performed on each data in the reference data set. Thus, the reference feature data of each data in the reference data set can be obtained. After obtaining each reference feature data, based on the preset feature data selection condition, the feature data that meets the preset feature data selection condition can be selected from each reference feature data, and the selected feature data is used as the target feature data.
[0067] In the embodiment of the present invention, the reference feature data of each data in the reference data set is determined by the following method:
[0068] Obtain the data labels of each data in the reference data set. Furthermore, each data label can be parsed, so that the label content included in each data label can be obtained, that is, the data features of the data corresponding to each data label can be obtained, that is, the reference feature data of each data in the reference data set can be obtained. It can be understood that the data labels of each data in the reference data set can include
[0069] Based on the above embodiment, the data expansion method provided by the embodiment of the present invention further includes: adding data labels to each data in the reference data set, which can avoid the problem of data islands. There are various ways to add data labels to each data in the reference data set.
[0070] As an optional implementation manner in the embodiment of the present invention, adding data labels to each data in the reference data set includes: performing data feature extraction on each data included in the data set, and then the data features of each data in the reference data set can be obtained. Thus, corresponding data labels can be added to each data in the reference data set according to the data features of each data in the reference data set.
[0071] As another optional implementation manner in the embodiment of the present invention, adding data labels to each data in the reference data set includes: inputting each data in the reference data set into a pre-trained data classification model, and the data labels of each data in the reference data set can be obtained. Furthermore, data labels are added to each data in the reference data set.
[0072] Among them, the data classification model can be a hybrid model obtained by combining a word vector module (such as Word2Vec) and a text convolutional neural network module (such as TextCNN). In the embodiments of the present invention, Word2Vec can improve the effect of data augmentation. TextCNN can improve the fitting speed and has a good fitting effect. Compared with the prior art, in the embodiments of the present invention, the data classification model including the Word2Vec-TextCNN hybrid model has better applicability and stronger extensibility, and can improve the effect of data classification.
[0073] In the embodiments of the present invention, the data classification model can be obtained by the following method:
[0074] Obtain sample data and expected data corresponding to the sample data. Among them, the sample data can be data in multiple business scenarios, and the expected data can be the expected data labels of the data in each business scenario;
[0075] Input the sample data into a pre-constructed data classification model to obtain the actual data labels output by the data classification model; based on the expected data labels and the actual data labels, adjust the model parameters of the pre-constructed data classification model to obtain a trained data classification model.
[0076] S130. Select the first data to be extended with the target feature data from the data set to be extended, and use the first data to be extended as the target extended data of the target business scenario.
[0077] Among them, the first data to be extended can be understood as the data in the data set to be extended that has the target feature data. The target extended data can be understood as the data that meets the data extension conditions in the data set to be extended, that is, the first data to be extended. Among them, the data extension condition is the data with the target feature data.
[0078] Specifically, after determining the target feature data, compare the feature data of each data in the data set to be extended with the target feature data. Thus, a comparison result can be obtained. If the comparison result is consistent, the data corresponding to the feature data that is consistent with the target feature data in the data set to be extended can be determined, that is, the first data to be extended is determined. After determining the data to be extended, the first data to be extended can be used as the target extended data of the target business scenario.
[0079] It should be noted that in the embodiment of the present invention, the data table for storing data is obtained by using the spark distributed computing platform and the python development framework for data processing. After obtaining the data table, the erroneous data can be cleaned by rule verification, and the missing data can be filled by data backfilling, so as to obtain the data in the reference data set in the target business scenario and the data in the to-be-expanded data set of the associated business scenario associated with the target business scenario.
[0080] The technical solution of the embodiment of the present invention can more quickly determine the scope of data expansion by obtaining a reference data set in the target business scenario and a data set to be expanded of an associated business scenario associated with the target business scenario. Furthermore, the reference feature data of each data in the reference data set can be determined, and the target feature data in the reference feature data can be determined, so that the conditions for expanding the data can be obtained more accurately. After determining the target feature data, the first data to be expanded having the target feature data can be selected from the data set to be expanded, and the first data to be expanded can be used as the target expansion data for the target business scenario. The technical solution in the embodiment of the present invention solves the problem that the existing data expansion method not only takes a lot of time, but also has the problem of low efficiency of data expansion, and realizes faster and more effective data expansion, thereby reducing the time consumption of data expansion, further improving the efficiency of data expansion, and laying the foundation for subsequent data drainage.
[0081] Embodiment 2
[0082] Figure 2 A flow chart of a data expansion method provided for Embodiment 2 of the present invention, based on the aforementioned embodiment, optionally, determining target feature data in the reference feature data, including: determining positive feature data associated with the target business scenario in the reference feature data based on an importance coefficient of the reference feature data; determining target feature data in the positive feature data based on an interval between a current moment and a target update moment of the positive feature data, wherein technical terms identical or corresponding to those in the aforementioned embodiment are not repeated here.
[0083] like Figure 2 As shown, the method of this embodiment specifically includes:
[0084] S210: Acquire a reference data set in a target business scenario and a data set to be expanded of an associated business scenario associated with the target business scenario.
[0085] S220: Determine reference feature data for each data in the reference data set.
[0086] S230. Determine the positive feature data associated with the target business scenario in the reference feature data according to the importance coefficient of the reference feature data.
[0087] Among them, the importance coefficient can be used to represent the importance of the reference feature data. Optionally, the importance coefficient can be the TGI (Target Group Index) index. The positive feature data can be understood as the feature data with a relatively high degree of association with the target business scenario in the reference feature data. The number of positive feature data can be one, two or more.
[0088] Specifically, obtain the importance coefficient of each reference feature data. Furthermore, based on a preset coefficient threshold, determine the reference feature data corresponding to the importance coefficient exceeding the preset coefficient threshold, and use it as the positive feature data associated with the target business scenario. Among them, the preset coefficient threshold can be set according to the actual data expansion demand and will not be specifically limited here.
[0089] It can be understood that the larger the value of the importance coefficient of the reference feature data, the higher the importance of the reference feature data. On the contrary, the smaller the value of the importance coefficient of the reference feature data, the lower the importance of the reference feature data.
[0090] S240. Determine the target feature data in the positive feature data based on the interval duration between the current moment and the target update moment of the positive feature data.
[0091] Among them, the target update moment can be the moment when the positive feature data is updated for the last time, or it can be the moment when the positive feature data is added for the first time.
[0092] Specifically, the target update moment of the positive feature data can be used as the target update moment of the positive feature data. Furthermore, calculate the interval duration between the target update moment and the current moment. If the interval duration does not exceed the preset interval duration threshold, the positive feature data corresponding to the interval duration not exceeding the preset interval duration threshold can be used as the target feature data. It can be understood that if the interval duration exceeds the preset interval duration threshold, it can be determined that the positive feature data corresponding to the interval duration exceeding the preset interval duration is historical positive feature data. The advantage of doing this is that it can effectively select the target feature data in the positive feature data and further improve the timeliness of data expansion.
[0093] In an embodiment of the present invention, determining the target feature data in the positive feature data based on the time interval between the current moment and the target update moment includes: calculating an update metric value of the positive feature data according to the time interval between the current moment and the target update moment; selecting, from the update metric values, the update metric values that meet a preset metric condition as the target metric values, and using the positive feature data corresponding to the target metric values as the target feature data. Wherein, the update metric value is used to measure the update time interval of the positive feature data. The update metric value can reflect the novelty of the positive feature data. The preset metric condition can be set according to actual requirements. The update metric values corresponding to different preset metric conditions can be the same or different. The target metric value can be an update metric value that meets the preset metric condition.
[0094] Optionally, the update metric value of the positive feature data can be calculated according to the following formula:
[0095]
[0096] Wherein, X i ' represents the update metric value of the positive feature data, and X i represents the time interval between the current moment and the update moment of the positive feature data.
[0097] S250. Select the first data to be expanded with the target feature data from the data set to be expanded, and use the first data to be expanded as the target expansion data of the target business scenario.
[0098] The technical solution of the embodiment of the present invention realizes more accurately determining the target feature data in the positive feature data and ensures the accuracy of data expansion by determining the positive feature data associated with the target business scenario in the reference feature data according to the importance coefficient of the reference feature data, and determining the target feature data in the positive feature data based on the time interval between the current moment and the target update moment of the positive feature data.
[0099] Embodiment III
[0100] Figure 3A flowchart of a data expansion method provided in the third embodiment of the present invention. Optionally, based on the foregoing embodiments, the data expansion method implemented by the present invention further includes: determining the to-be-expanded feature data of each data in the to-be-expanded dataset; selecting the second expansion data of the target business scenario from the to-be-expanded dataset according to each to-be-expanded feature data; The step of using the first expansion data as the target expansion data of the target business scenario includes: using all users in the first expansion data and the second expansion data as the target expansion data of the target business scenario; or, using the public users in the first expansion data and the second expansion data as the target expansion data of the target business scenario. Wherein, the same or corresponding technical terms as those in the above embodiments will not be elaborated herein.
[0101] As Figure 3 shown, the method of this embodiment specifically includes:
[0102] S310. Obtain a reference dataset in the target business scenario and a to-be-expanded dataset of an associated business scenario associated with the target business scenario.
[0103] S320. Determine the reference feature data of each data in the reference dataset and determine the target feature data in the reference feature data.
[0104] S330. Select the first to-be-expanded data having the target feature data from the to-be-expanded dataset.
[0105] S340. Determine the to-be-expanded feature data of each data in the to-be-expanded dataset.
[0106] Wherein, the to-be-expanded feature data can be the feature data obtained by performing data feature extraction on the data in the to-be-expanded dataset.
[0107] S350. Select the second expansion data of the target business scenario from the to-be-expanded dataset according to each to-be-expanded feature data.
[0108] Wherein, the second expansion data can be the data selected from the to-be-expanded dataset based on the to-be-expanded feature data of the data in the to-be-expanded dataset.
[0109] In the embodiment of the present invention, there are various ways to select the second expansion data of the target business scenario from the to-be-expanded dataset according to each to-be-expanded feature data.
[0110] As an alternative implementation manner in the embodiments of the present invention, the step of selecting the second expansion data of the target business scenario from the to-be-expanded data set according to each to-be-expanded feature data includes: performing clustering processing on the target feature data to obtain a clustering center; calculating the data distance between each to-be-expanded feature data and the clustering center; and based on each data distance, selecting the second expansion data of the target business scenario from the to-be-expanded data set. Herein, the data distance can be understood as the data similarity between each to-be-expanded feature data and the clustering center.
[0111] Among them, the step of selecting the second expansion data of the target business scenario from the to-be-expanded data set based on each data distance includes: sorting each data distance in descending order, and determining the to-be-expanded feature data corresponding to the data distance exceeding a preset data distance threshold as the second expansion data. The preset data distance threshold can be set according to actual requirements to obtain data within a certain range from each clustering center as the to-be-expanded data.
[0112] As another alternative implementation manner in the embodiments of the present invention, the step of selecting the second expansion data of the target business scenario from the to-be-expanded data set according to each to-be-expanded feature data includes: respectively inputting the to-be-expanded feature data into data expansion models for data expansion in different expansion dimensions to obtain data expansion values output by each data expansion model; performing weighted average processing on each data expansion value to obtain a target expansion value, and according to the target expansion value, selecting the second expansion data in the target business scenario from the to-be-expanded data set.
[0113] In the embodiments of the present invention, the data expansion model is obtained through the following manner:
[0114] For the initial network model of each expansion dimension, obtaining sample data and expected output data corresponding to the sample data; inputting the sample data into the initial network model to obtain the actual output data of the initial network model; and adjusting the parameters of the initial network model according to the actual output data and the expected output data of the sample data to obtain the data expansion model.
[0115] Among them, the sample data includes positive sample data and negative sample data. The positive sample data includes the reference feature data, and the negative sample data is the feature data of other business scenarios except the target business scenario. To avoid the problem of unbalanced sample data between positive sample data and negative sample data, in the embodiments of the present invention, random undersampling processing can be performed on the positive sample data and the negative sample data to balance the positive sample data and the negative sample data, so as to achieve sample equalization. On this basis, to avoid the loss of spatial information of the sample data, random undersampling can be performed multiple times during the model training process to obtain the sampling result. Further, the mean value of the sampling result can be taken.
[0116] It should be noted that in the embodiments of the present invention, by training the initial network model for each extended dimension, data extension models of different dimensions can be obtained to support the requirements of multi-business scenario data extension. Machine learning algorithms such as logistic regression, random forest, XGBoost, support vector machine, and neural network can be used to integrate the models, and the weighted average method is used for fusion. Weights are generated according to the accuracy of each technology, so as to obtain the final prediction result. In the embodiments of the present invention, the data extension is implemented by fusing multiple machine learning algorithms, which effectively solves the problems of algorithm stability and universality in different application scenarios, helps to solidify the data extension technology, and improves the prediction accuracy.
[0117] S360. Obtain target extended data based on the first extended data and the second extended data.
[0118] In the embodiments of the present invention, there are multiple ways to obtain target extended data based on the first extended data and the second extended data.
[0119] As an optional implementation manner in the embodiments of the present invention, obtaining target extended data based on the first extended data and the second extended data includes: using all users in the first extended data and the second extended data as the target extended data of the target business scenario; as another optional implementation manner in the embodiments of the present invention, obtaining target extended data based on the first extended data and the second extended data includes: using the public users in the first extended data and the second extended data as the target extended data of the target business scenario.
[0120] Based on the above embodiments, the technical solution of the embodiments of the present invention determines the to-be-expanded feature data of each data in the to-be-expanded data set, and selects the second expanded data of the target business scenario from the to-be-expanded data set according to each to-be-expanded feature data. Furthermore, all users in the first expanded data and the second expanded data can be used as the target expanded data of the target business scenario; or, the common users in the first expanded data and the second expanded data can be used as the target expanded data of the target business scenario, so as to effectively expand the data in the target business scenario and further meet the user requirements.
[0121] Embodiment 4
[0122] Figure 4 FIG. 4 is a schematic structural diagram of a data expansion device provided in Embodiment 4 of the present invention. As shown in FIG. 4, the device includes: a data set acquisition module 410, a feature data determination module 420, and a target expansion data acquisition module 430.
[0123] Among them, the data set acquisition module 410 is configured to acquire a reference data set in a target business scenario and a to-be-expanded data set of an associated business scenario associated with the target business scenario;
[0124] The feature data determination module 420 is configured to determine the reference feature data of each data in the reference data set and determine the target feature data in the reference feature data;
[0125] The target expansion data acquisition module 430 is configured to select first to-be-expanded data having the target feature data from the to-be-expanded data set and use the first to-be-expanded data as the target expansion data of the target business scenario.
[0126] Based on the technical solution of the embodiments of the present invention, the data set acquisition module acquires a reference data set in the target business scenario and a to-be-expanded data set of an associated business scenario associated with the target business scenario, so that the scope of data expansion can be determined more quickly. Furthermore, the feature data determination module can determine the reference feature data of each data in the reference data set and determine the target feature data in the reference feature data, so that the conditions for expanding the data can be obtained more accurately. After determining the target feature data, the target expansion data acquisition module can select first to-be-expanded data having the target feature data from the to-be-expanded data set and use the first to-be-expanded data as the target expansion data of the target business scenario. The technical solution in the embodiments of the present invention solves the problems that the existing data expansion method not only takes a lot of time but also has low data expansion efficiency, realizes more rapid and effective data expansion, thereby reducing the time consumed by data expansion and further improving the efficiency of data expansion.
[0127] Optionally, the feature data determination module 420 includes a forward feature data determination unit and a target feature data determination unit; wherein,
[0128] The forward feature data determination unit is configured to determine, according to the importance coefficient of the reference feature data, the forward feature data associated with the target service scenario in the reference feature data;
[0129] The target feature data determination unit is configured to determine the target feature data in the forward feature data based on the time interval between the current moment and the target update moment of the forward feature data.
[0130] Optionally, the target feature data determination unit is configured to:
[0131] Calculate an update metric value of the forward feature data according to the time interval between the current moment and the target update moment, where the update metric value is used to measure the update time interval of the forward feature data;
[0132] Select the update metric values that meet the preset metric conditions from the update metric values as the target metric values, and use the forward feature data corresponding to the target metric values as the target feature data.
[0133] Optionally, the target feature data determination unit is configured to:
[0134] Calculate the update metric value of the forward feature data according to the following formula:
[0135]
[0136] where X i ' represents the update metric value of the forward feature data, and X i represents the time interval between the current moment and the update moment of the forward feature data.
[0137] Optionally, the device further includes: a second extended data selection module, configured to:
[0138] Determine the to-be-extended feature data of each data in the to-be-extended data set;
[0139] Select the second extended data of the target service scenario from the to-be-extended data set according to each to-be-extended feature data;
[0140] The target extended data acquisition module 430 includes:
[0141] Use all users in the first extended data and the second extended data as the target extended data of the target service scenario; or,
[0142] Use the public users in the first extended data and the second extended data as the target extended data for the target business scenario.
[0143] Optionally, the second extended data selection module is specifically configured to:
[0144] Perform clustering processing on the target feature data to obtain a clustering center;
[0145] Calculate the data distance between each feature data to be extended and the clustering center;
[0146] Based on each data distance, select the second extended data for the target business scenario from the data set to be extended.
[0147] Optionally, the second extended data selection module is specifically configured to:
[0148] Input the feature data to be extended into data extension models for data extension in different extended dimensions respectively, and obtain the data extension values output by each data extension model;
[0149] Perform weighted average processing on each data extension value to obtain a target extension value, and based on the target extension value, select the second extended data in the target business scenario from the data set to be extended.
[0150] Optionally, the apparatus further includes: a model training module, configured to:
[0151] For the initial network model of each extended dimension, obtain sample data and expected output data corresponding to the sample data, where the sample data includes positive sample data and negative sample data, the positive sample data includes the reference feature data, and the negative sample data is the feature data of other business scenarios except the target business scenario;
[0152] Input the sample data into the initial network model to obtain the actual output data of the initial network model;
[0153] Adjust the parameters of the initial network model according to the actual output data and the expected output data of the sample data to obtain a data extension model.
[0154] The data extension apparatus provided by the embodiments of the present invention can execute the data extension method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.
[0155] It should be noted that the various units and modules included in the above data expansion device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the embodiments of the present invention.
[0156] Embodiment Five
[0157] Figure 5 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described herein and / or claimed.
[0158] As Figure 5 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0159] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0160] Processor 11 may be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 executes the various methods and processes described above, such as the data expansion method.
[0161] In some embodiments, the data expansion method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the data expansion method described above may be executed. Alternatively, in other embodiments, the processor 11 may be configured to execute the data expansion method by any other suitable means (e.g., by means of firmware).
[0162] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0163] The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer program can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0164] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0165] To provide for interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0166] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0167] A computing system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0168] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0169] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A data expansion method, characterized in that: include: Acquire a reference data set in a target business scenario and a data set to be expanded of an associated business scenario associated with the target business scenario; Determine reference feature data of each data in the reference data set, and determine target feature data in the reference feature data; Selecting first extended data having the target characteristic data from the to-be-extended data set, and using the first extended data as target extended data for the target business scenario; Wherein, the determining the target feature data in the reference feature data comprises: determining the positive feature data associated with the target business scenario in the reference feature data according to the importance coefficient of the reference feature data; determining the target feature data in the positive feature data based on the interval between the current moment and the target update moment of the positive feature data; Among them, based on the interval duration between the current moment and the target update moment, the target feature data in the positive feature data is determined, including: according to the interval duration between the current moment and the target update moment, an update metric value of the positive feature data is calculated, wherein the update metric value is used to measure the update interval duration of the positive feature data; an update metric value that meets the preset metric condition is selected from the update metric values as the target metric value, and the positive feature data corresponding to the target metric value is used as the target feature data.
2. The method according to claim 1, characterized in that The step of calculating the update metric value of the forward feature data according to the interval between the current time and the target update time includes: The updated metric value of the forward feature data is calculated according to the following formula: ; in, represents an updated metric value of the forward feature data, Indicates the interval between the current time and the update time of the forward feature data.
3. The method according to claim 1, characterized in that The method further comprises: Determine the feature data to be expanded of each data in the data set to be expanded; According to each feature data to be extended, selecting second extended data of the target business scenario from the data set to be extended; The using the first extended data as target extended data for the target business scenario includes: taking all users in the first extension data and the second extension data as target extension data for the target business scenario; or, The public users in the first extended data and the second extended data are used as target extended data for the target business scenario.
4. The method according to claim 3, characterized in that The selecting the second extended data of the target business scenario from the data set to be extended according to each feature data to be extended includes: Clustering the target feature data to obtain a cluster center; Calculating the data distance between each feature data to be expanded and the cluster center; Based on each data distance, second extended data of the target business scenario is selected from the to-be-extended data set.
5. The method according to claim 3, characterized in that, The selecting the second extended data of the target business scenario from the data set to be extended according to each feature data to be extended includes: Inputting the feature data to be expanded into data expansion models for data expansion in different expansion dimensions respectively, and obtaining data expansion values output by each data expansion model; Perform weighted average processing on each data expansion value to obtain a target expansion value, and select second expansion data in the target business scenario from the to-be-expanded data set according to the target expansion value.
6. The method according to claim 5, characterized in that The method further includes: For the initial network model of each expansion dimension, obtain sample data and expected output data corresponding to the sample data, where the sample data includes positive sample data and negative sample data, the positive sample data includes the reference feature data, and the negative sample data is feature data of other business scenarios except the target business scenario; Input the sample data into the initial network model to obtain the actual output data of the initial network model; Adjust the parameters of the initial network model according to the actual output data and expected output data of the sample data to obtain a data expansion model.
7. A data expansion device, characterized in that: It includes: A data set acquisition module, configured to acquire a reference data set in a target business scenario and a to-be-expanded data set of an associated business scenario associated with the target business scenario; A feature data determination module, configured to determine reference feature data of each data in the reference data set and determine target feature data in the reference feature data; A target expansion data acquisition module, configured to select first expansion data having the target feature data from the to-be-expanded data set and use the first expansion data as the target expansion data of the target business scenario; Among them, determining the target feature data in the reference feature data includes: determining positive feature data associated with the target business scenario in the reference feature data according to the importance degree coefficient of the reference feature data; determining the target feature data in the positive feature data based on the time interval between the current moment and the target update moment of the positive feature data; Among them, determining the target feature data in the positive feature data based on the time interval between the current moment and the target update moment includes: calculating an update metric value of the positive feature data according to the time interval between the current moment and the target update moment, where the update metric value is used to measure the update interval of the positive feature data; selecting an update metric value that meets a preset metric condition from the update metric values as a target metric value, and using the positive feature data corresponding to the target metric value as the target feature data.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a processor to implement the data expansion method according to any one of claims 1-6 when executed.
Citation Information
Patent Citations
Data set selection method and device based on multi-task learning
CN111062484A
Target service execution method and system, server and storage medium
CN112559095A