Data processing method, system and related device based on federated learning

By introducing short video interactive data into the insurance industry and performing data alignment and LightGBM model training, the issues of data security and data availability were resolved, improving the accuracy and efficiency of ad placement.

CN117010529BActive Publication Date: 2026-04-10INSIGHT TECHNOLOGY (XIONGAN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSIGHT TECHNOLOGY (XIONGAN) CO LTD
Filing Date
2023-08-07
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Traditional federated learning-based data processing methods in insurance industry advertising suffer from limited data availability, low relevance to user purchase intentions, and high costs. Furthermore, directly integrating raw data from different institutions poses data security risks.

Method used

We acquire short video interaction data from data providers and convert it into a word vector library. We then align the data between the initiator and the data provider to identify common samples and generate sentence vectors. We train and adjust the model parameters using the LightGBM model and obtain the target prediction results to determine the processing strategy.

Benefits of technology

While ensuring data security, the data dimensions available for modeling have been enriched, improving modeling efficiency and accuracy, and enhancing the precision of ad targeting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117010529B_ABST
    Figure CN117010529B_ABST
Patent Text Reader

Abstract

The application discloses a data processing method and system based on federal learning and related equipment, the method comprises the following steps: a data provider obtains first interaction data under a preset scene, converts the first interaction data into a word vector library, a modeling initiator determines modeling samples, the modeling initiator and the data provider perform data alignment on the first data and second data, and determine common samples, the data provider determines feature data of the common samples, determines sentence vectors according to the common samples and the word vector library, obtains k sentence vectors, the modeling initiator divides the modeling samples to obtain a training set and a test set, pre-processes the k sentence vectors to obtain target feature data, trains a preset federal LightGBM model according to the target feature data, the training set and the test set, adjusts model parameters of the preset federal LightGBM model, and obtains a target federal LightGBM model. The application embodiment can improve the advertising effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of privacy computing and the technical field of computer technology, in particular to a data processing method and system based on federated learning and related equipment. BACKGROUND

[0002] With the continuous development of economy and the improvement of people's living standards, short videos have become an indispensable part of modern people, and short video platforms have become an important channel for various brands and enterprises to carry out advertising. Through publishing interesting and attractive content on short video platforms, brands can interact with users, increase the number of fans, establish brand loyalty, and promote users to become potential consumers.

[0003] With the explosive growth of the short video industry, the insurance industry, for example, has also begun to actively explore short video marketing. Traditional data processing methods based on federated learning largely depend on media channels and market research data. Due to limited available data and low correlation with user purchase intentions, there are often problems of poor effectiveness and high cost. Therefore, how to improve the effectiveness of advertising needs to be solved. SUMMARY

[0004] The embodiments of the present application provide a data processing method and system based on federated learning and related equipment, which can improve the effectiveness of advertising.

[0005] In a first aspect, the embodiments of the present application provide a data processing method based on federated learning, applied to a two-party computing system, the two-party computing system comprising a modeling initiator and a data provider; the modeling initiator corresponds to first data, and the data provider corresponds to second data; the method comprises:

[0006] obtaining first interaction data in a preset scene through the data provider, the first interaction data comprising at least one of the following: bullet screen, comment, whether to like, whether to collect, and viewing time; and converting the first interaction data into a word vector library;

[0007] determining a modeling sample through the modeling initiator;

[0008] aligning the first data and the second data through the modeling initiator and the data provider, and determining common samples of the modeling initiator and the data provider;

[0009] determining feature data of the common samples through the data provider, determining sentence vectors according to the common samples and the word vector library, obtaining k sentence vectors, and k is a positive integer;

[0010] The modeling initiator divides the modeling sample to obtain a training set and a test set, pre-processes the k sentence vectors to obtain target feature data, trains a preset federated LightGBM model according to the target feature data, the training set and the test set, and adjusts model parameters of the preset federated LightGBM model to obtain a target federated LightGBM model;

[0011] The modeling initiator obtains target user data, inputs the target user data into the target federated LightGBM model to obtain a target prediction result, and determines a target processing strategy corresponding to the target prediction result.

[0012] In a second aspect, an embodiment of the present application provides a two-party computing system, which comprises a modeling initiator and a data provider; the modeling initiator corresponds to first data, and the data provider corresponds to second data; wherein,

[0013] The data provider is configured to obtain first interaction data in a preset scenario, the first interaction data comprising at least one of the following: a barrage, a comment, whether to like, whether to collect, and a viewing time length; and convert the first interaction data into a word vector library.

[0014] The modeling initiator is configured to determine a modeling sample.

[0015] The modeling initiator and the data provider are configured to perform data alignment on the first data and the second data, and determine a common sample of the modeling initiator and the data provider.

[0016] The data provider is configured to determine feature data of the common sample, determine sentence vectors according to the common sample and the word vector library, and obtain k sentence vectors, k being a positive integer.

[0017] The modeling initiator is configured to divide the modeling sample to obtain a training set and a test set, pre-process the k sentence vectors to obtain target feature data, train a preset federated LightGBM model according to the target feature data, the training set and the test set, and adjust model parameters of the preset federated LightGBM model to obtain a target federated LightGBM model; obtain target user data, input the target user data into the target federated LightGBM model to obtain a target prediction result, and determine a target processing strategy corresponding to the target prediction result.

[0018] In a third aspect, an electronic device is provided and includes a processor, a memory, a communication interface, and one or more programs. The one or more programs are stored in the memory and configured to be executed by the processor. The programs include instructions for performing the steps in the first aspect of the embodiments.

[0019] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program for electronic data exchange. The computer program causes a computer to perform some or all of the steps described in the first aspect of the embodiments.

[0020] In a fifth aspect, a computer program product is provided. The computer program product includes a non-transitory computer-readable storage medium storing a computer program. The computer program is operable to cause a computer to perform some or all of the steps described in the first aspect of the embodiments. The computer program product can be a software installation package.

[0021] By implementing the embodiments of the present application, the following beneficial effects are achieved:

[0022] It can be seen that the data processing method, system and related equipment based on federated learning described in the embodiments of the application are applied to a two-party computing system, the two-party computing system includes a modeling initiator and a data provider; the modeling initiator corresponds to first data, and the data provider corresponds to second data; first interaction data in a preset scene is obtained through the data provider, the first interaction data includes at least one of the following: a barrage, a comment, whether to like, whether to collect, and a viewing time length; the first interaction data is converted into a word vector library, a modeling sample is determined through the modeling initiator, the first data and the second data are data-aligned through the modeling initiator and the data provider, and common samples of the modeling initiator and the data provider are determined, feature data of the common samples is determined through the data provider, a sentence vector is determined according to the common samples and the word vector library, k sentence vectors are obtained, k is a positive integer, the modeling sample is divided through the modeling initiator to obtain a training set and a test set, the k sentence vectors are preprocessed to obtain target feature data, a preset federated LightGBM model is trained according to the target feature data, the training set and the test set, and model parameters of the preset federated LightGBM model are adjusted to obtain a target federated LightGBM model, target user data is obtained through the modeling initiator, the target user data is input into the target federated LightGBM model to obtain a target prediction result; a target processing strategy corresponding to the target prediction result is determined, on the one hand, the interaction data is introduced to enrich the data dimension available for modeling, and on the other hand, the data is aligned, and the corresponding sentence vector is determined, and then the corresponding feature data is determined based on the sentence vector to adjust the model parameters of the model, so that the overall modeling efficiency and accuracy are greatly improved, thereby improving the precision of the delivery under the premise of ensuring data security, and the advertisement delivery strategy can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0024] Figure 1 is an architecture schematic diagram of a two-party computing system for implementing a data processing method based on federated learning provided by an embodiment of the present application;

[0025] Figure 2 is a flow schematic diagram of a data processing method based on federated learning provided by an embodiment of the present application;

[0026] Figure 3 is a flow schematic diagram of another data processing method based on federated learning provided by an embodiment of the present application;

[0027] Figure 4 FIG. 1 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0028] In order to enable persons skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative labor fall within the scope of protection of the present application.

[0029] The terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product, or device.

[0030] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of the phrase in various places in the specification does not necessarily all refer to the same embodiment, nor is it necessarily mutually exclusive of other embodiments. It is explicitly and implicitly understood that the embodiments described herein can be combined with other embodiments.

[0031] In the related art, in order to realize accurate marketing, user interaction data such as comments, likes, collections, and bullet screens of short videos or live broadcasts in insurance and medical scenarios can be introduced when building a prediction model. These data have higher relevance to insurance, can more greatly mine users interested in insurance, improve marketing accuracy, and thus realize accurate marketing. However, the following problems often exist in the process of building a model: User interaction data are often non-numeric, and cannot be directly used when modeling, and need to be processed by manual labeling and the like. However, the manual processing method not only consumes time and effort, but is also relatively rough and cannot discover hidden information in the data. Insurance institutions need to introduce short video interaction data for in-depth mining and analysis, but insurance marketing samples and short video interaction data are distributed in different institutions. If the original data of the two parties is directly put together for fusion analysis, there will be a problem of data security.

[0032] To address the shortcomings of related technologies, this application provides a data processing method based on federated learning, applied to a two-party computing system, wherein the two-party computing system includes: a modeling initiator and a data provider; the modeling initiator corresponds to first data, and the data provider corresponds to second data; the method includes:

[0033] The data provider obtains first interactive data under a preset scenario, the first interactive data including at least one of the following: bullet comments, comments, whether it is liked, whether it is favorited, and viewing time; the first interactive data is converted into a word vector library;

[0034] The modeling sample is determined by the modeling initiator.

[0035] The modeling initiator and the data provider perform data alignment on the first data and the second data, and determine the common samples of the modeling initiator and the data provider;

[0036] The feature data of the common samples are determined by the data provider, and sentence vectors are determined based on the common samples and the word vector library to obtain k sentence vectors, where k is a positive integer.

[0037] The modeling initiator divides the modeling samples to obtain a training set and a test set. The k sentence vectors are preprocessed to obtain target feature data. A preset federated LightGBM model is trained based on the target feature data, the training set, and the test set, and the model parameters of the preset federated LightGBM model are adjusted to obtain the target federated LightGBM model.

[0038] The target user data is obtained by the modeling initiator, and the target user data is input into the target federated LightGBM model to obtain the target prediction result; the target processing strategy corresponding to the target prediction result is determined.

[0039] In this embodiment, a federal insurance marketing model is built based on privacy-preserving computation technology, which solves the data security problem of joint modeling with external data agencies. Simultaneously, user interaction data from short videos is introduced, enriching the data dimensions available for modeling. Word2vec and pre-training methods address the difficulties of manual annotation and the scarcity of data information, greatly improving the overall modeling efficiency and accuracy. This allows for increased targeting precision while ensuring data security.

[0040] The embodiments of this application will be described in detail below.

[0041] Please see Figure 1 , Figure 1The embodiment of the application provides a kind of for realizing the architecture schematic diagram of two-party computing system based on federated learning data processing method, as shown in the figure, this two-party computing system includes: modeling initiator and data provider;The modeling initiator corresponds to first data, and the data provider corresponds to second data;Based on the two-party computing system, the following functions can be realized:

[0042] The two-party computing system includes a modeling initiator and a data provider.

[0043] The data provider is configured to obtain first interaction data in a preset scenario, wherein the first interaction data includes at least one of the following: bullet screen, comment, whether to like, whether to collect, and viewing time length.

[0044] The modeling initiator is configured to determine a modeling sample.

[0045] The modeling initiator and the data provider are configured to perform data alignment on the first data and the second data, and determine common samples of the modeling initiator and the data provider.

[0046] The data provider is configured to determine feature data of the common samples, determine sentence vectors based on the common samples and the word vector library, obtain k sentence vectors, and k is a positive integer.

[0047] The modeling initiator is configured to divide the modeling sample to obtain a training set and a test set, pre-process the k sentence vectors to obtain target feature data, train a preset federated LightGBM model based on the target feature data, the training set and the test set, adjust model parameters of the preset federated LightGBM model to obtain a target federated LightGBM model, obtain target user data, input the target user data into the target federated LightGBM model to obtain a target prediction result, and determine a target processing strategy corresponding to the target prediction result.

[0048] Optionally, in the aspect of converting the first interaction data into the word vector library, the data provider is specifically configured to:

[0049] Tokenize and remove stop words from non-numeric data in the first interaction data to obtain second interaction data.

[0050] Convert the second interaction data into the word vector library using word2vec technology.

[0051] Optionally, in the data alignment of the first data and the second data, the modeling initiator and the data provider are specifically used for:

[0052] aligning the first data and the second data based on a privacy intersection technique by the modeling initiator and the data provider.

[0053] Optionally, in the training of the preset federated LightGBM model according to the target feature data, the training set and the test set, and the adjustment of the model parameters of the preset federated LightGBM model, a target federated LightGBM model is obtained, which includes:

[0054] training the preset federated LightGBM model based on the training set and the test set to obtain a first reference federated LightGBM model;

[0055] adjusting the model parameters of the first reference federated LightGBM model according to the target feature data to obtain the target federated LightGBM model;

[0056] or,

[0057] inputting the target feature data into the preset federated LightGBM model to adjust the model parameters of the preset federated LightGBM model to obtain a second reference federated LightGBM model;

[0058] training the second reference federated LightGBM model based on the training set and the test set to obtain the target federated LightGBM model.

[0059] Optionally, in the preprocessing of the k sentence vectors to obtain target feature data, it includes:

[0060] calculating the relevance and / or IV value of each sentence vector in the k sentence vectors to obtain k calculation results;

[0061] screening the k calculation results to obtain a screening result, and taking the screening result as the target feature data.

[0062] Please refer to Figure 2 , Figure 2 is a flowchart of a data processing method based on federated learning provided by the present application, applied to a two-party computing system, the two-party computing system including a modeling initiator and a data provider; the modeling initiator corresponds to first data, and the data provider corresponds to second data; as shown in the figure, the present data processing method based on federated learning includes:

[0063] 201, obtaining first interaction data in a preset scene through the data provider, the first interaction data including at least one of the following: a barrage, a comment, whether to like, whether to collect, a viewing time length; and converting the first interaction data into a word vector library.

[0064] In the embodiments of the present application, the modeling initiator can be understood as an advertisement provider, and the advertisement provider can include at least one of the following: an insurance company, an automobile sales company, a real estate company, a medical company, a supermarket, a tourism company, an electronic commodity company, and the like, without limitation. The data provider is a party that provides data. The preset scene can be pre-set or system default, and the preset scene can include at least one of the following: insurance, medical treatment, tourism, electronic commodities, automobiles, food, and the like, without limitation.

[0065] The first data can include related information of a customer group, and the related information can include at least one of the following: age, gender, region, occupation, family status, and the like, without limitation. The second data can include at least one of the following: age, gender, region, occupation, family status, consumption situation, barrage, comment, whether to like, whether to collect, viewing time length, and the like, without limitation. The second data can include the first interaction data.

[0066] In specific implementation, the data provider can obtain first interaction data in a preset scene, and the first interaction data can include at least one of the following: barrage, comment, whether to like, whether to collect, viewing time length, and the like, without limitation, and then the first interaction data can be converted into a word vector library.

[0067] Optionally, the step 201 of converting the first interaction data into a word vector library can include the following steps:

[0068] The non-numeric data in the first interaction data is segmented and stop words are removed to obtain second interaction data; and the second interaction data is converted into the word vector library by using a word2vec technology.

[0069] In specific implementation, the non-numeric data in the first interaction data can be pre-trained, specifically, the first interaction data is segmented and stop words are removed to obtain second interaction data, and then the second interaction data is converted into a word vector library by using a word2vec technology.

[0070] In the embodiments of the present application, the non-numeric data generally refers to text data such as comments and barrages, and the pre-training operation refers to first segmenting the non-numeric data in the preset scene, that is, breaking a sentence into words, and then removing stop words, and then processing the data to generate a word vector library by using a word2vec technology.

[0071] Among them, word segmentation is to obtain the words contained in the text data, in order to learn and train the hidden meaning and intention of different words in the preset scene. Stop words are some words in the text data, which have little to do with the preset scene (such as "ah", "haha", "of", "we"), and removing stop words is to make the word2vec technology generate better word vector library. In specific implementation, different stop word libraries can be established based on different scenes, so that the corresponding stop word library can be selected for stop word removal based on different scenes.

[0072] For example, in the embodiment of the application, the modeling initiator can include an insurance company, and the data provider can include a short video data agency. Specifically, the short video data agency obtains interactive data of short videos in the insurance and medical scenarios, such as bullet screen, comments, whether to like, whether to collect, viewing time, etc. For non-numeric data such as user comments and bullet screen data, a word vector library is generated by pre-training using a word2vec model. The word2vec technology solves the defects of interactive data that cannot be used by traditional models and artificial annotation difficulty, expands the data dimension, and improves the efficiency and accuracy of the model.

[0073] 202. Determine the modeling sample through the modeling initiator.

[0074] In the embodiment of the application, the modeling initiator can determine the corresponding modeling sample based on the modeling demand.

[0075] 203. Align the first data and the second data through the modeling initiator and the data provider, and determine the common sample of the modeling initiator and the data provider.

[0076] In the embodiment of the application, the modeling initiator and the data provider align the first data and the second data, and determine the common sample (aligned sample) of the modeling initiator and the data provider. Specifically, the modeling initiator and the data provider align the data, and confirm the sample common to both parties for subsequent federated modeling.

[0077] Optionally, the step 203 of aligning the first data and the second data through the modeling initiator and the data provider can include the following steps:

[0078] Align the first data and the second data through the modeling initiator and the data provider based on the privacy intersection technology.

[0079] In specific implementation, the modeling initiator and the data provider can perform secure data alignment on the first data and the second data based on a privacy intersection technique. Specifically, the modeling initiator determines modeling samples and sample label definitions, and then performs secure sample alignment with the short video data institution based on the privacy intersection technique to determine modeling samples common to both parties, thereby ensuring data security, ensuring that original data is not leaked, and solving the security problem of joint modeling of data of two parties.

[0080] 204. Determine feature data of the common samples by the data provider, determine sentence vectors according to the common samples and the word vector library, and obtain k sentence vectors, k being a positive integer.

[0081] In the embodiments of the present application, the data provider can determine the feature data of the common samples, and then determine the sentence vectors according to the common samples and the word vector library to obtain k sentence vectors, k being a positive integer.

[0082] In specific implementation, the short video data institution can prepare feature data of the common samples. The short video data institution then processes non-numeric data in the samples, which can include processing such as word segmentation and removing stop words. The word vector library generated by the insurance interactive data is used to generate sentence vectors of sample comments for subsequent modeling. Specifically, the data provider and the modeling initiator respectively train a light gradient boosting machine (LightGBM) model, and combine the local models to establish a federated model.

[0083] In the embodiments of the present application, the sample features that the data provider can provide in the common sample data include numeric and non-numeric data. The non-numeric data of the common samples is processed using a pre-trained sentence vector library to generate sentence vectors. The sentence vectors and the numeric data are referred to as modeling features of the common samples, that is, feature data.

[0084] 205. Divide the modeling samples by the modeling initiator to obtain a training set and a test set, pre-process the k sentence vectors to obtain target feature data, train a preset federated LightGBM model according to the target feature data, the training set and the test set, and adjust model parameters of the preset federated LightGBM model to obtain a target federated LightGBM model.

[0085] In the embodiments of the present application, the word vector library is obtained according to the non-numerical data of the preset scene, that is, the first interaction data, and then the word vector library is used to process the non-numerical data such as comments of the common sample (aligned sample) to obtain the sentence vector of the common data for modeling. In addition to the sentence vector generated by the non-numerical data, the modeling features can also include some data value type data provided by the data provider. The data value type data is a number by name.

[0086] The target feature data can include the sentence vector obtained by processing the word vector library, and the remaining numerical data of the data provider.

[0087] The preset federated LightGBM model can be pre-set or system default. The LightGBM model has the advantages of high efficiency, scalability, accuracy, flexibility and ease of use.

[0088] In the embodiments of the present application, the modeling initiator can divide the modeling sample to obtain the training set and the test set. Specifically, the modeling sample can be randomly divided, or divided based on user selection to obtain the training set and the test set.

[0089] Then, the k sentence vectors can be preprocessed to obtain target feature data. The preset federated LightGBM model is trained according to the target feature data, the training set and the test set, and the model parameters of the preset federated LightGBM model are adjusted to obtain the target federated LightGBM model, so as to improve the accuracy of the target federated LightGBM model. The model parameters can include at least one of the following: learning rate, tree depth, training times, column sampling, row sampling, etc. without limitation.

[0090] Optionally, the step 205 of preprocessing the k sentence vectors to obtain the target feature data can include the following steps:

[0091] The relevance and / or IV value of each sentence vector in the k sentence vectors is calculated to obtain k calculation results. The k calculation results are screened to obtain a screening result, and the screening result is taken as the target feature data.

[0092] In the embodiments of the present application, the IV value is used to represent the contribution degree of the feature to the target prediction. Generally, the feature with IV < 0.02 is considered as a useless feature. The relevance generally refers to the correlation degree of two variables. Generally, the feature with a correlation coefficient > 0.8 is considered to have strong correlation.

[0093] In the embodiments of the present application, the relevance of each sentence vector in the k sentence vectors and / or the IV value can be calculated to obtain k calculation results, and the k calculation results can be screened to obtain a screening result, and the screening result can be taken as the target feature data.

[0094] Optionally, the step 205 can train a preset federated LightGBM model according to the target feature data, the training set and the test set, and adjust model parameters of the preset federated LightGBM model to obtain a target federated LightGBM model, including:

[0095] The training set and the test set are used to train the preset federated LightGBM model to obtain a first reference federated LightGBM model; model parameters of the first reference federated LightGBM model are adjusted according to the target feature data to obtain the target federated LightGBM model; or the target feature data is input into the preset federated LightGBM model to adjust model parameters of the preset federated LightGBM model to obtain a second reference federated LightGBM model; and the second reference federated LightGBM model is trained according to the training set and the test set to obtain the target federated LightGBM model.

[0096] In the embodiments of the present application, the training set and the test set can be used to train a preset federated LightGBM model to obtain a first reference federated LightGBM model, and then the model parameters of the first reference federated LightGBM model can be adjusted according to the target feature data to obtain the target federated LightGBM model, so that the model capability of the federated LightGBM model can be improved. Alternatively, the target feature data can be input into the preset federated LightGBM model to adjust the model parameters of the preset federated LightGBM model to obtain a second reference federated LightGBM model, and then the second reference federated LightGBM model can be trained according to the training set and the test set to obtain the target federated LightGBM model, so that the model capability of the federated LightGBM model can be improved.

[0097] In specific implementations, when constructing a federated marketing model, first, the modeling data is divided into a training set and a test set, then the relevance, IV value and the like of the modeling features are calculated and pre-screened, and finally the preprocessed feature data is used to construct a federated LightGBM model, and the model is trained by adjusting appropriate parameters.

[0098] The target federated LightGBM model can include a federated insurance marketing model.

[0099] 206. Obtain target user data through the modeling initiator, input the target user data into the target federated LightGBM model to obtain target prediction results; determine the target processing strategy corresponding to the target prediction results.

[0100] In this embodiment of the application, the target user data can be user data related to a preset scenario. The target user data can include at least one of the following: age, gender, region, occupation, family status, user interaction data, etc., which are not limited here. The user interaction data can include at least one of the following: comments, likes, favorites, bullet comments, etc. of short videos or live broadcasts, which are not limited here.

[0101] In practice, different prediction results can correspond to different processing strategies. That is, a pre-defined mapping relationship between prediction results and processing strategies can be set, and then the target processing strategy corresponding to the target prediction result can be determined based on this mapping relationship. The target processing strategy can be an advertising placement strategy, which may include at least one of the following: target audience, target region, target time period, target location, etc., without limitation here.

[0102] In practice, based on the established federal insurance marketing model, the model can predict the target customer groups and output the model scores of all target customer groups. Through analysis, marketing strategies can be formulated, and target customer groups can be selected for marketing campaigns.

[0103] In this embodiment, a federal insurance marketing model is built based on privacy-preserving computation technology, which solves the data security problem of joint modeling with external data agencies. Simultaneously, user interaction data from short videos is introduced, enriching the data dimensions available for modeling. Word2vec and pre-training methods address the difficulties of manual annotation and the scarcity of data information, greatly improving the overall modeling efficiency and accuracy. This allows for increased targeting precision while ensuring data security.

[0104] For example, the main steps to build a precise insurance targeting model using short video data are as follows:

[0105] 1. Short video data agencies acquire interactive data from insurance and medical scenarios for pre-training, generating a word vector library;

[0106] 2. Confirm the modeling sample data used by the modeling initiator and align the data with the data provider;

[0107] 3. The data provider prepares the modeling feature data for the intersection samples;

[0108] 4. Construct a federated LightGBM model;

[0109] 5. Application of the federal insurance marketing model.

[0110] In practice, firstly, short video data agencies acquire interactive data from short videos in insurance and medical scenarios, such as bullet comments, comments, likes, favorites, and viewing duration. For non-numerical data, such as user comments and bullet comment data, a word2vec model is used for pre-training to generate a word vector library. Next, the modeling initiator and data provider align their data to confirm shared samples for subsequent federated modeling. Then, the short video data provider prepares sample features. The short video data agency uses the word vector library to generate sentence vectors of recent interaction data for modeling. Then, a vertical federated model is trained using data from all parties, building the final marketing model while ensuring the data from all parties remains within their respective domains. Finally, based on the constructed insurance marketing model, the model predicts the target customer group, formulates marketing strategies, and then launches marketing campaigns.

[0111] Let me give another example, such as Figure 3 As shown, taking an insurance company as the model initiator and a short video data agency as the data provider as an example, the specific implementation process is as follows: Modeling Sample Preparation: The model initiator determines the modeling samples, backtracking dates, and tags; Short Video Data Agency Trains Insurance Scene Corpus: The short video data agency acquires insurance scene interaction data and uses technologies such as word segmentation and word2vec to generate an insurance scene corpus for subsequent training; Sample Alignment: The model initiator and the short video data agency perform secure sample alignment based on privacy-preserving intersection technology, generating intersection samples for subsequent modeling use, and the aligned samples are returned to the data provider; Short Video Data Agency Modeling Feature Preparation: The short video data agency prepares the intersection sample features... The system utilizes an insurance scenario corpus to process interactive data from samples, generating sentence vectors for modeling. These features are then uploaded to the insurance company. Modeling feature preprocessing involves first dividing the modeling samples into training and testing sets, followed by calculating and pre-screening the relevance and IV values ​​of the modeling features. Federated algorithm model training involves training the preprocessed feature data using a federated LightGBM model, adjusting appropriate parameters. Federated insurance marketing model application: Based on the constructed federated insurance marketing model, model prediction is performed on the target customer groups, outputting model scores for all target customer groups. Through analysis, marketing strategies are formulated, and target customer groups are selected for marketing campaigns.

[0112] In a specific implementation, the scheme in the embodiment of the present application can be roughly divided into five parts: data pre-training, confirming training samples, preparing modeling features, constructing an insurance marketing model based on federated learning technology, and federated insurance marketing model application. Data pre-training: first, the short video data institution obtains the interactive data of users in the insurance and medical scenarios, such as comments, bullet screens, and non-numerical data of short videos in live broadcast rooms, and the like. Based on the above obtained data, the non-numerical data is pre-trained, and after word segmentation and stop word removal, the word vector library of the comment and bullet screen data is generated by using the word2vec technology. Determine the modeling sample and perform data alignment: the modeling initiator determines the modeling sample and sample label definition. Then, based on the privacy intersection technology, the short video data institution and the insurance institution perform safe sample alignment to determine the modeling sample owned by both parties. Modeling feature preparation: the short video data institution prepares feature data of the common sample. At the same time, the short video data institution needs to process the non-numerical data in the sample, generate a sentence vector of the sample comment by using the word vector library generated by the insurance interactive data, and use it for subsequent modeling. Training model based on federated algorithm: in order to complete the joint training of the model safely and privately, the data of each party needs to be jointly modeled without domain exposure and original data exposure. Constructing a federated marketing model: first, divide the modeling data into training set and test set, then calculate and pre-screen the correlation and IV value of the modeling features; finally, construct a federated LightGBM model by using the preprocessed feature data, and adjust the appropriate parameters to train the model. Application of federated insurance marketing model: based on the constructed federated insurance marketing model, the model is predicted for the marketing customer group, and the model score of all marketing customer groups is output. Through analysis, a marketing strategy is formulated, and a target customer group is selected for marketing investment.

[0113] In the embodiment of the present application, the method of word2vec+LightGBM realizes automatic processing and modeling of short video interactive data, effectively solving the defect that traditional modeling methods cannot use interactive data; under the premise of safety and compliance, the precision of the model is improved by increasing the data of two parties, so that the intended population of the insurance industry can be found more accurately, and the marketing cost is reduced. That is, for the short video industry, the precision of advertisement investment can be improved.

[0114] It can be seen that the data processing method based on federated learning described in the embodiments of the application is applied to a two-party computing system, the two-party computing system comprising: a modeling initiator and a data provider; the modeling initiator corresponds to first data, and the data provider corresponds to second data; first interaction data in a preset scene is obtained through the data provider, the first interaction data comprising at least one of: a barrage, a comment, whether to like, whether to collect, and a viewing time length; the first interaction data is converted into a word vector library; modeling samples are determined through the modeling initiator; the first data and the second data are data-aligned through the modeling initiator and the data provider, and common samples of the modeling initiator and the data provider are determined; feature data of the common samples is determined through the data provider; sentence vectors are determined according to the common samples and the word vector library, and k sentence vectors are obtained, k being a positive integer; the modeling samples are divided through the modeling initiator, and a training set and a test set are obtained; the k sentence vectors are preprocessed, and target feature data is obtained; a preset federated LightGBM model is trained according to the target feature data, the training set and the test set, and model parameters of the preset federated LightGBM model are adjusted, and a target federated LightGBM model is obtained; target user data is obtained through the modeling initiator, the target user data is input into the target federated LightGBM model, and a target prediction result is obtained; a target processing strategy corresponding to the target prediction result is determined, on the one hand, the introduction of interaction data enriches the data dimension available for modeling, and on the other hand, the data is aligned, and the corresponding sentence vectors are determined, and then the corresponding feature data is determined based on the sentence vectors, so as to adjust the model parameters of the model, so that the overall modeling efficiency and accuracy are greatly improved, thereby improving the precision of the delivery under the premise of ensuring data security, and the advertisement delivery strategy can be improved.

[0115] Consistent with the above embodiment, please refer to Figure 4 , Figure 4 is a structural schematic diagram of an electronic device provided by the embodiments of the application, as shown in the figure, the electronic device comprises a processor, a memory, a communication interface and one or more programs, and is applied to a two-party computing system, the two-party computing system comprising: a modeling initiator and a data provider; the modeling initiator corresponds to first data, and the data provider corresponds to second data; the above one or more programs are stored in the above memory and are configured to be executed by the above processor, in the embodiments of the application, the program comprises instructions for executing the following steps:

[0116] first interaction data in a preset scene is obtained through the data provider, the first interaction data comprising at least one of: a barrage, a comment, whether to like, whether to collect, and a viewing time length; the first interaction data is converted into a word vector library;

[0117] modeling samples are determined through the modeling initiator;

[0118] aligning the first data and the second data by the modeling initiator and the data provider, and determining common samples of the modeling initiator and the data provider;

[0119] determining feature data of the common samples by the data provider, determining sentence vectors according to the common samples and the word vector library, obtaining k sentence vectors, k being a positive integer;

[0120] dividing the modeling samples by the modeling initiator to obtain a training set and a test set, preprocessing the k sentence vectors to obtain target feature data, training a preset federated LightGBM model according to the target feature data, the training set and the test set, and adjusting model parameters of the preset federated LightGBM model to obtain a target federated LightGBM model;

[0121] obtaining target user data by the modeling initiator, inputting the target user data into the target federated LightGBM model to obtain a target prediction result, and determining a target processing strategy corresponding to the target prediction result.

[0122] Optionally, in the converting the first interaction data into a word vector library, the above program includes instructions for performing the following steps:

[0123] segmenting and removing stop words from non-numeric data in the first interaction data to obtain second interaction data;

[0124] converting the second interaction data into the word vector library by using a word2vec technology.

[0125] Optionally, in the aligning the first data and the second data by the modeling initiator and the data provider, the above program includes instructions for performing the following steps:

[0126] aligning the first data and the second data by the modeling initiator and the data provider based on a privacy intersection technology.

[0127] Optionally, in the training a preset federated LightGBM model according to the target feature data, the training set and the test set, and adjusting model parameters of the preset federated LightGBM model to obtain a target federated LightGBM model, the above program includes instructions for performing the following steps:

[0128] training the preset federated LightGBM model by using the training set and the test set to obtain a first reference federated LightGBM model;

[0129] adjusting model parameters of the first reference federated LightGBM model according to the target feature data to obtain the target federated LightGBM model;

[0130] or,

[0131] inputting the target feature data into the preset federated LightGBM model to adjust model parameters of the preset federated LightGBM model to obtain a second reference federated LightGBM model;

[0132] training the second reference federated LightGBM model according to the training set and the test set to obtain the target federated LightGBM model.

[0133] Optionally, in the aspect of preprocessing the k sentence vectors to obtain target feature data, the above program includes instructions for performing the following steps:

[0134] calculating the relevance and / or IV value of each of the k sentence vectors to obtain k calculation results;

[0135] screening the k calculation results to obtain a screening result, and taking the screening result as the target feature data.

[0136] It can be seen that the electronic device described in the embodiment of the application is applied to a two-party computing system, the two-party computing system comprising: a modeling initiator and a data provider; the modeling initiator corresponds to first data, and the data provider corresponds to second data; first interaction data in a preset scene is obtained through the data provider, the first interaction data comprising at least one of the following: a barrage, a comment, whether to like, whether to collect, and a viewing time length; the first interaction data is converted into a word vector library, modeling samples are determined through the modeling initiator, the first data and the second data are data-aligned through the modeling initiator and the data provider, and common samples of the modeling initiator and the data provider are determined, feature data of the common samples is determined through the data provider, a sentence vector is determined according to the common samples and the word vector library, k sentence vectors are obtained, k is a positive integer, the modeling samples are divided through the modeling initiator, a training set and a test set are obtained, the k sentence vectors are preprocessed, target feature data is obtained, a preset federated LightGBM model is trained according to the target feature data, the training set and the test set, and model parameters of the preset federated LightGBM model are adjusted, a target federated LightGBM model is obtained, target user data is obtained through the modeling initiator, the target user data is input into the target federated LightGBM model, and a target prediction result is obtained; a target processing strategy corresponding to the target prediction result is determined, on the one hand, the interaction data is introduced, and the data dimension available for modeling is enriched, and on the other hand, the data is aligned, and corresponding sentence vectors are determined, and then corresponding feature data is determined based on the sentence vectors, so that the model parameters of the model are adjusted, the overall modeling efficiency and accuracy are greatly improved, and therefore the precision of the delivery is improved under the premise of ensuring data security, and the advertisement delivery strategy can be improved.

[0137] The embodiment of the application further provides a computer storage medium, wherein the computer storage medium stores a computer program for electronic data exchange, and the computer program causes a computer to execute part or all of the steps of any method described in the above method embodiments, and the computer comprises an electronic device.

[0138] The embodiment of the application further provides a computer program product, and the computer program product comprises a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute part or all of the steps of any method described in the above method embodiments. The computer program product can be a software installation package, and the computer comprises an electronic device.

[0139] It should be noted that, for the foregoing method embodiments, the sequences of the described actions are not necessarily required to achieve the objects of the application, and certain steps can be performed in other sequences or even concurrently. Additionally, the described embodiments are merely provided as examples, and not all of the actions described are necessarily required to achieve desired results.

[0140] In the above embodiments, the description of each embodiment is focused on different aspects, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.

[0141] In several embodiments provided in the present application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely schematic. For example, the division of the above units is merely a logical function division. In actual implementation, another division manner can be adopted. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical or other forms.

[0142] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e. can be located in one place, or can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0143] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0144] If the above integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable memory. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a memory and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the above-mentioned method of each embodiment of the present application. The aforementioned memory includes: a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0145] A person of ordinary skill in the art can understand that all or part of the steps in the above-mentioned embodiments can be completed by instructing the relevant hardware through a program, which can be stored in a computer readable memory. The memory can include: a flash disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0146] The embodiments of the present application are described in detail above, and the specific examples are applied to the principles and implementation modes of the present application. The above embodiment description is only used to help understand the method of the present application and its core idea; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed; in summary, the content of the specification should not be understood as a limitation of the present application.

Claims

1. A data processing method based on federated learning, characterized in that, The application is applied to a two-party computing system, the two-party computing system comprises a modeling initiator and a data provider; the modeling initiator corresponds to first data, and the data provider corresponds to second data; the modeling initiator comprises an advertiser, and the advertiser comprises at least one of an insurance company, an automobile sales company, a real estate company, a medical company, a supermarket, a tourism company and an electronic commodity company; the method comprises the following steps: obtaining first interaction data in a preset scene through the data provider, wherein the first interaction data comprises at least one of the following: a barrage, a comment, whether to like, whether to collect and a viewing time length; and converting the first interaction data into a word vector library; determining a modeling sample through the modeling initiator; aligning the first data and the second data through the modeling initiator and the data provider, and determining a common sample of the modeling initiator and the data provider; determining feature data of the common sample through the data provider, determining a sentence vector according to the common sample and the word vector library, obtaining k sentence vectors, and k is a positive integer; dividing the modeling sample through the modeling initiator to obtain a training set and a test set, preprocessing the k sentence vectors to obtain target feature data, training a preset federated LightGBM model according to the target feature data, the training set and the test set, adjusting model parameters of the preset federated LightGBM model, and obtaining a target federated LightGBM model; obtaining target user data through the modeling initiator, inputting the target user data into the target federated LightGBM model to obtain a target prediction result, and determining a target processing strategy corresponding to the target prediction result.

2. The method of claim 1, wherein, The conversion of the first interaction data into the word vector library comprises the following steps: performing word segmentation and eliminating stop words on non-numeric data in the first interaction data to obtain second interaction data; converting the second interaction data into the word vector library by using a word2vec technology.

3. The method according to claim 1 or 2, characterized in that, The data alignment of the first data and the second data through the modeling initiator and the data provider comprises the following steps: aligning the first data and the second data based on a privacy intersection technology through the modeling initiator and the data provider.

4. The method according to claim 1 or 2, characterized in that, The training of the preset federated LightGBM model according to the target feature data, the training set and the test set and the adjustment of the model parameters of the preset federated LightGBM model to obtain the target federated LightGBM model comprise the following steps: training the preset federated LightGBM model by using the training set and the test set to obtain a first reference federated LightGBM model; adjusting the model parameters of the first reference federated LightGBM model according to the target feature data to obtain the target federated LightGBM model; or, input the target feature data into the preset federated LightGBM model to adjust model parameters of the preset federated LightGBM model, to obtain a second reference federated LightGBM model; train the second reference federated LightGBM model according to the training set and the test set, to obtain the target federated LightGBM model.

5. The method of claim 4, wherein, The preprocessing of the k sentence vectors to obtain target feature data comprises: calculating the relevance and / or IV value of each sentence vector in the k sentence vectors to obtain k calculation results; screening the k calculation results to obtain a screening result, and taking the screening result as the target feature data.

6. A two-party computation system, characterized by The two-party computing system comprises a modeling initiator and a data provider; the modeling initiator corresponds to first data, and the data provider corresponds to second data; the modeling initiator comprises an advertisement publisher, and the advertisement publisher comprises at least one of an insurance company, an automobile sales company, a real estate company, a medical company, a supermarket, a tourism company and an electronic commodity company; wherein The data provider is configured to obtain first interaction data in a preset scene, and the first interaction data comprises at least one of the following: a barrage, a comment, whether to like, whether to collect, and a viewing time length; and convert the first interaction data into a word vector library. The modeling initiator is configured to determine a modeling sample. The modeling initiator and the data provider are configured to perform data alignment on the first data and the second data, and determine a common sample of the modeling initiator and the data provider. The data provider is configured to determine feature data of the common sample, determine sentence vectors according to the common sample and the word vector library, and obtain k sentence vectors, k being a positive integer. The modeling initiator is configured to divide the modeling sample to obtain a training set and a test set, preprocess the k sentence vectors to obtain target feature data, train a preset federated LightGBM model according to the target feature data, the training set and the test set, and adjust model parameters of the preset federated LightGBM model to obtain a target federated LightGBM model; obtain target user data, input the target user data into the target federated LightGBM model to obtain a target prediction result, and determine a target processing strategy corresponding to the target prediction result.

7. The system of claim 6, wherein, In the aspect of converting the first interaction data into a word vector library, the data provider is specifically configured to: perform word segmentation and stop word elimination on non-numeric data in the first interaction data to obtain second interaction data; convert the second interaction data into the word vector library by using a word2vec technology.

8. The system of claim 6 or 7, wherein, In the aspect of performing data alignment on the first data and the second data, the modeling initiator and the data provider are specifically configured to: perform data alignment on the first data and the second data based on a privacy intersection technology by the modeling initiator and the data provider.

9. An electronic device, comprising: It includes a processor and a memory, the memory being used to store one or more programs and configured to be executed by the processor, the programs including instructions for performing the steps of the method as described in any one of claims 1-5.

10. A computer-readable storage medium, characterized in that, A computer program for storing electronic data interchange is provided, wherein the computer program causes a computer to perform the method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Ultrasonic examination follow-up patient screening method based on machine learning

    CN111524570A

  • Federal learning content pushing method and device based on multi-party multi-model privacy intersection

    CN115907043A