Method, device, medium and electronic equipment for generating marketing auxiliary information
By extracting and converting data from the big data platform, screening variables to establish a predictive model, generating deep triple information and constructing a knowledge graph, the problem of marketers obtaining shallow information is solved, the generation of personalized marketing auxiliary information is achieved, and marketing efficiency and success rate are improved.
Patent Information
- Application Number
- CN202110987853.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-26
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2041-08-26
AI Technical Summary
In the existing technology, the information obtained by marketers is relatively simple and superficial, and cannot provide in-depth personalized marketing assistance, resulting in unsatisfactory marketing results.
By extracting structured data from the big data platform, converting it into triple information, screening variables to establish a basic classification prediction model, determining feature explanation weights, generating deep triple information, and constructing a knowledge graph, marketing auxiliary information is provided.
It achieves efficient and accurate generation of marketing auxiliary information, can meet individualized marketing needs, and improve marketing efficiency and success rate.
Smart Images

Figure CN115730046B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of auxiliary marketing technology, and in particular to a method, device, medium and electronic device for generating auxiliary marketing information. Background Art
[0002] When marketing, frontline marketers typically query a single customer information system and then conduct customer marketing based on the information retrieved. The drawback of this approach is that the information obtained is relatively simple and represented in data. Therefore, the information obtained is shallow and not very helpful in supporting marketing. Summary of the Invention
[0003] In the field of auxiliary marketing technology, in order to solve the above technical problems, the purpose of this application is to provide a method, device, medium and electronic device for generating marketing auxiliary information.
[0004] According to one aspect of the present application, a method for generating marketing auxiliary information is provided, the method comprising:
[0005] Extracting structured data from a big data platform, the structured data including user-related data and product-related data, the big data platform including multiple information management systems;
[0006] Converting the structured data into triple information to obtain shallow triple information;
[0007] Obtaining a data table from the big data platform, the data table including multiple business data, the business data including variables and variable values corresponding to the variables, each of the business data being associated with a user;
[0008] Performing a variable screening operation on the data table to remove at least one variable in the data table to obtain a final data table;
[0009] Establishing a basic classification prediction model using the final data table;
[0010] Taking the variables in the final data table as features, and based on the business data associated with each user in the final data table, determining the average marginal contribution of each feature in the business data to the prediction result of the basic classification prediction model as the explanatory weight of each feature for the business data;
[0011] For each business data, determine the target feature according to the explanatory weight of each feature on the business data, and generate deep triple information according to the variable value corresponding to the target feature, wherein the explanatory weight of the target feature is greater than that of other features;
[0012] Constructing a knowledge graph using the shallow triple information and the deep triple information;
[0013] When a business question retrieval request is received, the retrieval results are obtained by querying the knowledge graph, and marketing auxiliary information is generated based on the retrieval results.
[0014] According to another aspect of the present application, a device for generating marketing auxiliary information is provided, the device comprising:
[0015] an extraction module configured to extract structured data from a big data platform, wherein the structured data includes user-related data and product-related data, and the big data platform includes multiple information management systems;
[0016] a conversion module, configured to convert the structured data into triple information to obtain shallow triple information;
[0017] an acquisition module configured to acquire a data table from the big data platform, the data table including multiple business data items, the business data including variables and variable values corresponding to the variables, each of the business data items being associated with a user;
[0018] a removal module configured to perform a variable screening operation on the data table to remove at least one variable in the data table to obtain a final data table;
[0019] An establishing module configured to establish a basic classification prediction model using the final data table;
[0020] a determination module configured to use the variables in the final data table as features and, based on the business data associated with each user in the final data table, determine an average marginal contribution of each feature in the business data to the prediction result of the basic classification prediction model as an explanatory weight of each feature for the business data;
[0021] The first generating module is configured to determine, for each piece of business data, a target feature according to the explanatory weights of each feature on the business data, and generate deep triplet information according to the variable value corresponding to the target feature, wherein the explanatory weight of the target feature is greater than that of other features;
[0022] A construction module is configured to construct a knowledge graph using the shallow triple information and the deep triple information;
[0023] The second generation module is configured to obtain retrieval results by querying the knowledge graph when a business question retrieval request is received, and generate marketing auxiliary information based on the retrieval results.
[0024] According to another aspect of the present application, a computer-readable program medium is provided, which stores computer program instructions. When the computer program instructions are executed by a computer, the computer is caused to execute the method described above.
[0025] According to another aspect of the present application, an electronic device is provided, comprising:
[0026] processor;
[0027] A memory having computer-readable instructions stored thereon, wherein the computer-readable instructions, when executed by the processor, implement the method described above.
[0028] The technical solutions provided by the embodiments of the present application may have the following beneficial effects:
[0029] According to the method, device, medium and electronic device for generating marketing auxiliary information provided in this application, shallow triple information and deep triple information are obtained respectively, and a knowledge graph is constructed using the shallow triple information and deep triple information, and the deep triple information is generated based on the variable value of the target feature, and the target feature is generated based on the feature's interpretation weight of the business data. Therefore, by combining shallow and deep knowledge, more comprehensive and profound business knowledge is formed. After the marketing personnel inputs a business question, the corresponding query result can be obtained by querying the knowledge graph, making the knowledge feedback to the marketing personnel more accurate and intuitive, and easy to understand, thereby assisting front-line marketing personnel to improve marketing efficiency and success rate; at the same time, due to the use of a one-stop intelligent question-and-answer method to feedback answers to questions to marketing personnel, marketing personnel can obtain knowledge more conveniently and accurately.
[0030] It should be understood that the foregoing general description and the following detailed description are merely illustrative and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0032] Figure 1 is a schematic diagram of a system architecture of a method for generating marketing auxiliary information according to an exemplary embodiment;
[0033] Figure 2 is a flow chart showing a method for generating marketing auxiliary information according to an exemplary embodiment;
[0034] Figure 3 is a schematic diagram of a data table based on which structured data is extracted according to an exemplary embodiment;
[0035] Figure 4A-4B is a schematic diagram showing shallow triplet information according to an exemplary embodiment;
[0036] Figure 5 is a schematic diagram showing a portion of a data table obtained from a big data platform according to an exemplary embodiment;
[0037] Figure 6 is a schematic diagram of a final data table according to an exemplary embodiment;
[0038] Figure 7 is a schematic diagram of an interpretable model based on machine learning according to an exemplary embodiment;
[0039] Figure 8 1 is a schematic diagram of a calculation process of an average contribution margin according to an exemplary embodiment;
[0040] Figure 9 is a schematic diagram showing the influence of various features visually outputted based on the DKMM model on the prediction result according to an exemplary embodiment;
[0041] Figure 10 is a schematic diagram illustrating construction of a knowledge graph corresponding to deep triple information and use of the knowledge graph according to an exemplary embodiment;
[0042] Figure 11 is a schematic diagram of a knowledge graph constructed based on deep triple information according to an exemplary embodiment;
[0043] Figure 12 is a schematic diagram of a knowledge graph constructed based on shallow triple information and deep triple information according to an exemplary embodiment;
[0044] Figure 13 is a schematic diagram illustrating generating an answer to a question based on an input question according to an exemplary embodiment;
[0045] Figure 14 is a schematic diagram showing answer information corresponding to a query question according to a method for generating marketing auxiliary information according to an exemplary embodiment;
[0046] Figure 15 is a block diagram of a device for generating marketing auxiliary information according to an exemplary embodiment;
[0047] Figure 16 is a block diagram illustrating an example of an electronic device for implementing the above-mentioned method for generating auxiliary marketing information according to an exemplary embodiment;
[0048] Figure 17 A program product is shown according to an exemplary embodiment for implementing the above-mentioned method for generating marketing auxiliary information. DETAILED DESCRIPTION
[0049] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0050] The accompanying drawings are merely schematic illustrations of the present application and are not necessarily drawn to scale. Identical reference numerals in the drawings represent identical or similar parts, and thus their repeated descriptions will be omitted. Some of the blocks shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically separate entities.
[0051] In related technologies, frontline marketing personnel traditionally use machine learning classification to identify high-probability potential users as target customers when conducting precision marketing for 5G services, and then assign these targets to the frontline for marketing. Because machine learning classification models only provide predictions and cannot explain the reasons for the high probability, frontline personnel can only obtain superficial knowledge by querying multiple systems (knowledge bases, CRMs, product libraries, tag libraries, etc.). This is inefficient, and since this knowledge is targeted at specific groups, it is impossible to obtain deeper knowledge, such as the key factors that most influence each user. This makes it impossible to carry out differentiated marketing campaigns tailored to each individual user, resulting in suboptimal marketing results.
[0052] While this technology can retrieve a significant amount of information, it's fragmented, disorganized, and highly discrete, making it difficult to quickly retrieve accurate answers. Furthermore, the explanations obtained are group-oriented and cannot meet the needs of individuals seeking personalized support for their conversational skills. The model can only output results, without any interpretable information. The process is a closed black box, and humans cannot understand the causal relationships reflected in its output.
[0053] To this end, this application first provides a method for generating marketing auxiliary information, which can overcome the above-mentioned defects, realize efficient and accurate generation of marketing auxiliary information, and meet the individual's needs for speech assistance in different ways for different people.
[0054] The implementation terminal of this application can be any device with computing, processing and communication functions, which can be connected to an external device for receiving or sending data. Specifically, it can be a portable mobile device, such as a smart phone, tablet computer, laptop computer, PDA (Personal Digital Assistant), etc., or a fixed device, such as a computer device, field terminal, desktop computer, server, workstation, etc. It can also be a collection of multiple devices, such as the physical infrastructure or server cluster of cloud computing. Optionally, the implementation terminal of this application can be a server or the physical infrastructure of cloud computing.
[0055] The solutions of the embodiments of this application can be applied to marketing scenarios of various products or services. These products can be physical products or virtual products. Physical products can be products such as smartphones and household appliances; virtual products can be products such as operator packages and insurance policies.
[0056] Figure 1 FIG. 1 is a schematic diagram of a system architecture of a method for generating marketing auxiliary information according to an exemplary embodiment. Figure 1 As shown, the system architecture of the system includes three components, namely the system application model, the DKMM algorithm model and the data acquisition module, wherein DKMM (deep knowledge minning model) is a model proposed in the embodiment of the present application that can execute the method for generating marketing auxiliary information. The system has four system roles, namely electronic outbound call personnel, account manager direct sales personnel, customer service marketing personnel and store marketing personnel, that is, personnel in these four system roles can use the system to obtain marketing auxiliary information. Figure 1In the DKMM algorithm model, the data collection module collects various types of data, including CRM data, knowledge base data, tag data, and product data. This data can be obtained from multiple systems, such as the CRM, knowledge base, tag library, and product library. The data extracted by the data collection module is used to construct the DKMM algorithm model. The DKMM algorithm model consists of a business-wide table, a graph database built based on the business-wide table, multiple functional modules, and a capability layer. The graph database can include a knowledge graph generated using the business-wide table. The functional modules of the DKMM algorithm model include a business model, a classification model, an interpretation module, knowledge reasoning, and a knowledge retrieval engine. The capability layer of the DKMM algorithm model provides the capabilities of the DKMM algorithm model to the system application module. The system application module includes a unified knowledge retrieval query portal. Various system roles in the system can use this unified knowledge retrieval query portal to retrieve the required marketing support information from the DKMM algorithm model in a one-stop manner. This marketing support information comprehensively, accurately, and easily understood reflects marketing knowledge. For example, it can provide reasons why a marketing target needs a product, thereby helping marketers improve marketing efficiency and success rates.
[0057] Figure 2 This is a flow chart of a method for generating marketing auxiliary information according to an exemplary embodiment. The method for generating marketing auxiliary information provided in this embodiment can be executed by a server and can be applied to the marketing activities of operators upgrading 4G packages to 5G packages. Figure 2 As shown, the following steps are included:
[0058] Step 210: extract structured data from the big data platform.
[0059] The structured data includes user-related data and product-related data, and the big data platform includes multiple information management systems.
[0060] Specifically, in the communications field, an information management system can be any of the following: CRM, knowledge base, tag library, or product library. CRM stands for Customer Relationship Management. Of course, in other fields, information management systems can take on other types. Structured data is data stored in a database. A database contains tables that record data corresponding to fields. This data is structured data.
[0061] Figure 3 FIG2 is a schematic diagram of a data table based on which structured data is extracted according to an exemplary embodiment. Figure 3, which shows multiple data tables, so data can be extracted from multiple data tables in the big data platform; each data table includes fields and the serial numbers and field meanings corresponding to the fields, such as Figure 3 The data table in the upper left corner contains the "NET_NUM" field, which indicates whether the device is online. Each data table contains different types of fields, so data can be extracted from multiple tables on the big data platform.
[0062] Specifically, triple information includes two entities and a relationship or attribute used to establish a connection between the two entities. In order to convert structured data into triple information, it is necessary to extract structured data from the perspective of triples.
[0063] For example, in the field of 5G marketing, data extraction can be performed based on the knowledge graph Schema framework, encompassing nine entity types: user, package, terminal, package area, developer, package instance, customer, account, and tag. This includes eight relationships: user-package, user-terminal, package-package instance, user-developer, user-customer, user-account, and user-package instance. Data related to 5G business information is obtained through a big data platform, primarily from basic user attributes, subscription relationship data, and personalized interest data. Table data associations are used to extract data for eight entity types: user (163 fields), package (10 fields), terminal (156 fields), package area (8 fields), developer (3 fields), package instance (34 fields), customer (5 fields), and account (7 fields), along with eight relationships between them. This generates 16 entity and relationship tables. Tags can indicate whether a user has subscribed to a 5G package. In this context, a customer refers to the owner of a telecommunications product under the same registered name. Customer communication information includes information about all products under an organization. An account is a customer's account, and a customer can have one or multiple accounts. A user is an entity that represents a telecommunications product and the value-added services it provides, such as mobile or broadband services. The relationship between customers, accounts, and users can be summarized as follows: a customer can have multiple accounts, and an account can contain multiple users.
[0064] Step 220: Convert the structured data into triple information to obtain shallow triple information.
[0065] Specifically, all fields in the entity table are used as attributes of the corresponding entity, and the relational table is converted into a triple in the form of (user ID, subscription package, package ID), (package ID, attribute, free traffic) (label, attribute, label explanation).
[0066] Figure 4A-4B FIG. 1 is a schematic diagram showing shallow triple information according to an exemplary embodiment. Figure 4A As shown, the shallow triplet information shown is (user ID, terminal, terminal model); Figure 4B In the figure, not only the shallow triple information (user ID, terminal, terminal model) is shown, but also other forms of shallow triple information are shown. For example, it shows the shallow triple information (user ID, package, package name). It is easy to understand that since the triple information corresponds to the knowledge graph, Figure 4A and Figure 4B The shallow triple information shown can also be understood as a knowledge graph built based on the shallow triple information.
[0067] Step 230: Obtain a data table from the big data platform.
[0068] The data table includes multiple pieces of business data, each of which includes variables and variable values corresponding to the variables. Each piece of business data is associated with a user.
[0069] The data table here can be the same as or different from the data table from which the structured data was extracted. The variables here are equivalent to the fields mentioned above. A business data item is a collection of data corresponding to each field.
[0070] Obtain a data table from the big data platform. The data table contains 101 fields, including package, package value, whether it is integrated, whether it is roaming, user preferences, etc. Figure 5 is a schematic diagram showing a portion of a data table obtained from a big data platform according to an exemplary embodiment. Figure 5 Multiple fields and the data corresponding to the fields are shown, and each row of data can be a piece of business data.
[0071] Step 240: Perform a variable screening operation on the data table to remove at least one variable in the data table to obtain a final data table.
[0072] In one embodiment, before performing a variable screening operation on the data table, the method further includes: for each variable, determining the proportion of business data in all business data whose variable value corresponding to the variable is null; and eliminating the variable based on the data proportion being greater than a predetermined ratio threshold.
[0073] The predetermined ratio threshold can be set arbitrarily based on experience, for example, it can be set to 80%. In the embodiment of the present application, the validity of the data is ensured by eliminating variables with a high proportion of null values.
[0074] You can also use the data in the data table to further process it to further improve the validity of the data. For example, for discrete variables, use the WOE (Weight of Evidence) method to transform the features to form variables that the program can recognize; for continuous variables, use the z-score method to standardize the features and remove the units of the data.
[0075] In one embodiment, performing a variable screening operation on the data table to eliminate at least one variable in the data table to obtain a final data table includes: sequentially performing a variable screening operation on the data table using a chi-square test method, a correlation coefficient calculation method, and an information value evaluation method to eliminate at least one variable in the data table to obtain a final data table.
[0076] Specifically, the chi-square test method can be used to test the difference of the independent variable on the qualitative dependent variable, retain the variables with obvious differences, such as retaining the variables with differences greater than the predetermined threshold, and eliminate other variables; use the correlation coefficient method to calculate the correlation coefficient between variables, verify whether there is multicollinearity between the variables, perform variable screening, select variables with a correlation coefficient above 0.5, and eliminate other variables; use the IV (Information Value) method to evaluate the predictive ability of the variables, perform variable screening, and thus obtain the final data table. Figure 6 FIG. 1 is a schematic diagram of a final data table according to an exemplary embodiment. The final data table can be as follows: Figure 6 As shown, it includes the corresponding field name, field type, Chinese annotation and sample data, where the field name is a variable retained after variable elimination, and the sample data is the variable value corresponding to the variable.
[0077] For example, fk_cnt is a variable, the Chinese annotation of which is the number of secondary cards, and its corresponding variable value is 2.
[0078] In the embodiment of the present application, variables that are helpful in generating marketing auxiliary information are obtained by performing variable elimination.
[0079] Step 250: Establish a basic classification prediction model using the final data table.
[0080] Each item of business data in the final data table can be used as a sample to build a basic classification prediction model. The business data can include a label indicating whether the user has upgraded to a 5G package. Specifically, the final data table can be divided into a training set, a test set, and a validation set according to a predetermined ratio. The training set is then used to train the basic classification prediction model, the test set is used to test the model, and the validation set is used to verify the model. The predetermined ratio can be 8:1:1.
[0081] In one embodiment, the use of the final data table to establish a basic classification prediction model includes: using the final data table to train a logistic regression model, a CART model, and an Xgboost model respectively; establishing a basic classification prediction model based on the logistic regression model, the CART model, and the Xgboost model, wherein the prediction result of the basic classification prediction model is a weighted calculation result obtained by weighted calculation of the output results of the logistic regression model, the CART model, and the Xgboost model.
[0082] The logistic regression (LR) model is a generalized linear regression analysis model that applies a logistic function to linear regression. The CART (Classification And Regression Trees) model is a decision tree-based algorithm that can be used for both classification and regression. The Xgboost model is an ensemble learning model built by integrating multiple base learners.
[0083] In other words, the basic classification prediction model may include a weighted calculation module for outputting a weighted calculation result, and the weighted calculation module is connected to the output ends of the logistic regression model, the CART model and the Xgboost model respectively.
[0084] In an embodiment of the present application, by integrating multiple models into the established basic classification prediction model, the final model output result is obtained by weighted calculation of the output results of multiple models, thereby further improving the accuracy of model prediction.
[0085] The basic classification prediction model can output the success rate of marketing, for example, the success rate of users upgrading to 5G packages.
[0086] Step 260, using the variables in the final data table as features, and based on the business data associated with each user in the final data table, determining the average marginal contribution of each feature in the business data to the prediction results of the basic classification prediction model as the explanatory weight of each feature for the business data.
[0087] The variable value corresponding to the feature is equivalent to the feature value.
[0088] In one embodiment, the method uses the variables in the final data table as features and, based on the business data associated with each user in the final data table, determines the average marginal contribution of each feature in the business data to the prediction result of the basic classification prediction model as the explanatory weight of each feature for the business data, including: dividing the business data in the final data table into multiple layers according to a predetermined rule, each layer including multiple business data; selecting a business data from the final data table as the selected business data, and iteratively executing the step of determining the marginal contribution value for each feature until a predetermined number of times is executed, wherein the step of determining the marginal contribution includes: randomly generating a feature order, and sorting the selected business data and the business data in the final data table according to the feature order; selecting a layer from an unselected layer, and randomly selecting a business data from the business data of the layer as the constructed business data corresponding to the selected business data; and selecting the selected business data and the constructed business data respectively according to the selected business data and the constructed business data. Construct first instance business data and second instance business data, wherein the first instance business data includes the variable values corresponding to the feature and the feature before the feature in the selected business data, and the variable values corresponding to the feature after the feature in the constructed business data, and the second instance business data includes the variable values corresponding to the feature before the feature in the selected business data, and the variable values corresponding to the feature and the feature after the feature in the constructed business data; input the selected business data and the constructed business data into the basic classification prediction model respectively, and obtain a first prediction result corresponding to the selected business data and a second prediction result corresponding to the constructed business data respectively; determine the marginal contribution value corresponding to the feature based on the first prediction result and the second prediction result; for each feature, determine the average marginal contribution value based on the marginal contribution values obtained by executing the step of determining the marginal contribution value for the feature, as the explanatory weight of the feature for the business data.
[0089] In an embodiment of the present application, by stratifying the business data, selecting and constructing the business data based on the stratification, and constructing two instance business data and second instance business data respectively based on the constructed business data, the marginal contribution value is determined based on the first prediction result and the second prediction result, thereby ensuring the accuracy of the determined marginal contribution value. The marginal contribution value can be equivalent to the weighted score of each feature to the model prediction result, thereby improving the interpretability of the system.
[0090] Figure 7 FIG is a schematic diagram of an interpretable model based on machine learning according to an exemplary embodiment. The interpretable model based on machine learning is the core part of the DKMM algorithm model. Figure 7 As shown in Figure 1, the machine learning-based interpretable model is established through the following process:
[0091] 1. Use data to build machine learning classification algorithms
[0092] The data is divided into training set, test set, and validation set; and according to the business scenario requirements, three machine learning classification algorithms, namely LR model, CART model, and XGBOOST model, are established respectively, and the parameters of each model are adjusted to the optimal level.
[0093] 2. Prediction of basic classification prediction model
[0094] The results of the LR, CART, and XGBOOST models are retrained and weighted to form the final prediction result, which is used as the prediction result of the basic classification prediction model that integrates the LR, CART, and XGBOOST models.
[0095] 3. Building an interpretable module
[0096] Construct an interpretable algorithm by randomly sampling instances x, z.
[0097] 4. Explain the output.
[0098] Output the contribution of each feature to each user. By analyzing each user, we can calculate the weighted score of each feature on the model prediction result.
[0099] Figure 8 4 is a schematic diagram of a calculation process of an average marginal contribution according to an exemplary embodiment. Figure 8 The embodiment shown can be applied to the field of telecommunication business marketing, such as the marketing of 5G packages. The process of determining the average marginal contribution in the above embodiment can be specifically achieved by Figure 8 Implementation. Figure 8 As shown, the calculation process includes the following steps:
[0100] 1. Stratified data sampling
[0101] Because telecom user package amounts have distinct tiers, we use statistical stratified sampling techniques to categorize users into several tiers based on their spending levels. For example, monthly spending amounts can be categorized into tiers of 0-50, 51-100, 101-200, 201-300, and above 300.
[0102] 2. Generate a random sequence of features o.
[0103] Randomly sort the features to be calculated. This step can reduce the impact of the feature sorting in the original table on the results.
[0104] 3. Generate random order instances.
[0105] First generate random order instances x o=(x1, .., x j ,..,x p ), and then generate random order instances z o =(z1, .., z j ,…,z p ), where x o Corresponding to the selected service data in the previous embodiment, z o This corresponds to the construction of business data in the previous embodiment.
[0106] For example, take any user x o (such as 41 years old, male, 59 package, etc.), select any random user z in a layer in step 1 o (20 years old, female, 99 package, etc.).
[0107] 4. Construct feature instances
[0108] Construct an instance x with feature j +j =(x1,...,x j-1 , x j , z j+1 ,...,z p ), construct an instance x without feature j -j =(x1,...,x j-1 , z j , z j+1 ,...,z p ), they are all based on x o and z o Among them, x +j Corresponding to the first instance of business data in the previous embodiment, x -j Corresponding to the second instance business data in the previous embodiment. For example, the extracted instance is reconstructed to form a new feature user x +j (41 years old, male, 99 yuan set meal, etc.), x -j (41 years old, female, 99 yuan set meal, etc.)
[0109] 5. Calculate marginal contribution
[0110] The new feature user x +j 、x -j Substitute into the basic classification prediction model and calculate the marginal contribution of a single feature (such as gender) to the prediction result. Specifically, This formula calculates the contribution margin where As the basic classification prediction model.
[0111] 6. Calculate the average marginal contribution
[0112] For instance x +j, randomly extract different users z from different layers, calculate the marginal contribution value with different users z each time, iterate M times, and finally, use This formula calculates the average value as the marginal contribution of feature j, that is, the average marginal contribution, which is the average marginal contribution of feature j to x. o The explanatory weight of the sample prediction results.
[0113] In one embodiment, the method further includes: outputting the interpretation weight of each feature on the business data in a visual manner according to a user request. Specifically, the interpretation weight can be output in a graphical manner.
[0114] Figure 9 The figure is a schematic diagram showing the influence of various features visually outputted based on the DKMM model on the prediction result according to an exemplary embodiment. Figure 9 The embodiment is applied to the marketing process of upgrading 4G to 5G. Figure 9 As shown in Figure 2, the DKMM model can be used to observe the impact of various features in the telecommunications data set on the prediction results, intuitively present the decision-making process of the final prediction results, and assist in finding the reasons for a customer's promotion, such as Figure 8 As shown in the figure, the nodes to the right of f(x) are negatively impacted features, while the nodes to the left of f(x) are positively impacted features. The user's initial score was 23.4. The ofr_cdma_nums (number of phones in a plan) feature increased the user's score by 8.17 points, while the avg_flux_3m (average traffic) feature decreased the user's score by 5.39 points. After traversing all features, the user's final score was 64.03, indicating a 64.03% probability of promotion. In formal applications, positive factors are mapped to user labels, matched with sales pitches, and ultimately outputting explainable reasons for the user's actions, assisting frontline marketing efforts and improving success rates.
[0115] Step 270 : For each piece of business data, determine the target feature according to the interpretation weight of each feature on the business data, and generate deep triplet information according to the variable value corresponding to the target feature.
[0116] Among them, the explanation weight of the target feature is greater than that of other features.
[0117] In one embodiment, for each business data, the target feature is determined based on the interpretation weight of each feature on the business data, including: for each business data, sorting the interpretation weight of each feature on the business data from large to small; and taking the features that are ranked first by a predetermined number as the target features.
[0118] In one embodiment, for each business data, the target feature is determined based on the interpretation weight of each feature for the business data, including: for each business data, obtaining the feature whose absolute value of the interpretation weight of the business data is greater than a predetermined interpretation weight threshold as the target feature.
[0119] The explanatory weight of a feature on business data can be positive or negative.
[0120] In one embodiment, the method further includes generating deep triplet information based on the prediction results of the basic classification prediction model on the business data. Specifically, for example, the deep triplet information may include whether a user is willing to upgrade to a 5G package. This allows for more accurate and accessible marketing assistance information.
[0121] Step 280: construct a knowledge graph using the shallow triple information and the deep triple information.
[0122] In practical applications, shallow triple information and deep triple information can be imported into the Neo4j graph database to construct a 5G business knowledge graph.
[0123] Figure 11 FIG is a schematic diagram of a knowledge graph constructed based on deep triple information according to an exemplary embodiment. Figure 11 As shown in Figure 2, the knowledge graph constructed based on deep triple information includes various knowledge such as the number of process overflows in the past three months corresponding to the user's unique ID.
[0124] Figure 12 FIG is a schematic diagram of a knowledge graph constructed based on shallow triple information and deep triple information according to an exemplary embodiment. Figure 12 As shown in the figure, the constructed knowledge graph includes three types of information, namely shallow knowledge, explainable factors and basic attribute factors. For example, the age corresponding to the user ID is 36, which is the basic attribute factor, the traffic saturation corresponding to the user ID is high, which is the explainable factor, and the free traffic corresponding to Enjoy 5G189 is 40G, which is shallow knowledge. The shallow knowledge here is the shallow triple information, and the explainable factor is the deep triple information.
[0125] Step 290: When a business question retrieval request is received, the retrieval results are obtained by querying the knowledge graph, and marketing auxiliary information is generated based on the retrieval results.
[0126] In one embodiment, after generating the marketing auxiliary information according to the search result, the method further includes: returning the marketing auxiliary information to the sender of the business question search request.
[0127] Specifically, a unified knowledge retrieval query portal may be provided, and the knowledge retrieval query portal may be set in the form of a text entry box; a user inputs a business question in the knowledge retrieval query portal to submit a business question retrieval request.
[0128] Figure 10 FIG is a schematic diagram showing how to construct a knowledge graph corresponding to deep triple information and use the knowledge graph according to an exemplary embodiment. Figure 10 As shown, first, in the model prediction stage, the possibility of users upgrading from 4G to 5G is predicted, and the result is a high probability; then, in the feature weighting stage, the interpretable module in the aforementioned embodiment is used to interpret the model and obtain the interpretation weight corresponding to each feature; then, in the graph construction stage, the knowledge graph is constructed based on the interpretation results; next, in the knowledge translation stage, knowledge translation and retrieval are performed; finally, in the knowledge output stage, knowledge output is realized, which can not only output the probability of users upgrading from 4G to 5G, but also output the reasons behind the result.
[0129] In one embodiment, the use of the shallow triple information and the deep triple information to construct a knowledge graph includes: importing the shallow triple information and the deep triple information into a graph database to form a knowledge graph in the graph database; when a business problem retrieval request is received, obtaining retrieval results by querying the knowledge graph, and generating marketing auxiliary information based on the retrieval results, including: when a business problem retrieval request is received, obtaining the business problem in the business problem retrieval request; determining a result template that matches the business problem; querying the knowledge graph in the graph database to obtain retrieval results corresponding to the business problem; filling the retrieval results into the result template to obtain marketing auxiliary information.
[0130] In one embodiment, determining the result template that matches the business problem includes: determining the category to which the business problem belongs through a preset classification model; determining a question template that matches the business problem from the question templates corresponding to the category through a first syntactic analysis model, as a target question template; obtaining a result template corresponding to the target question template, as a result template that matches the business problem; querying the knowledge graph in the graph database to obtain a retrieval result corresponding to the business problem includes: extracting the query object in the business problem through a second syntactic analysis model; and querying the knowledge graph in the graph database based on the query object to obtain a retrieval result corresponding to the business problem.
[0131] The query object can be the entity, attribute, relationship and other information to be queried in the business problem. Multiple question templates and result templates corresponding to each question template can be pre-set.
[0132] Specifically, the preset classification model may be a naive Bayes model, the second syntactic analysis model and the first syntactic analysis model may be the same model, and the second syntactic analysis model and the first syntactic analysis model may be implemented through a Bert model.
[0133] Figure 13 The figure is a schematic diagram showing generating an answer to a question based on an input question according to an exemplary embodiment. Figure 13 The following process is shown: first, data is extracted from the 5G information database and preprocessed; then, the preprocessed data is imported into the DKMM model; after the question is input, in order to query the answer, the question needs to be input into the DKMM model first, and the Chinese word segmentation operation is first performed on the question to obtain the word segmentation result. The input question will also be classified by the Bayesian classifier, and then the question matching is performed through the standard question library to determine the matching Cypher template file; then, the word segmentation result is syntactically analyzed to realize entity recognition and obtain the entity; then, the triple search library is used to search and obtain the corresponding retrieval result; next, Cypher construction is performed based on the retrieval result, and combined with the previously determined Cypher template file, and the combined result is provided to the Neoj service as the answer; the Neoj service returns the answer to generate the answer.
[0134] Figure 14 FIG. 1 is a schematic diagram showing answer information corresponding to a query question according to a method for generating marketing auxiliary information according to an exemplary embodiment. Figure 14 As shown, no matter what questions the marketer enters, the method for generating marketing auxiliary information provided by the embodiment of the present application can generate matching and accurate answers. For example, when the question "Why is user 1811966XXXX a high-probability 5G user?" is entered, the answer "User 1811966XXXX has insufficient data, loves playing games, and is a young group" can be automatically generated. Therefore, the key factors of users with a high probability of upgrading to 5G can be understood.
[0135] The present application also provides a device for generating marketing auxiliary information. The following is an embodiment of the device of the present application.
[0136] Figure 15 FIG. 1 is a block diagram of a device for generating marketing auxiliary information according to an exemplary embodiment. Figure 15As shown, the device 1500 includes: an extraction module 1510, configured to extract structured data from a big data platform, wherein the structured data includes user-related data and product-related data, and the big data platform includes multiple information management systems; a conversion module 1520, configured to convert the structured data into triple information to obtain shallow triple information; an acquisition module 1530, configured to obtain a data table from the big data platform, wherein the data table includes multiple business data, wherein the business data includes variables and variable values corresponding to the variables, and each of the business data is associated with a user; a removal module 1540, configured to perform a variable screening operation on the data table to remove at least one variable in the data table to obtain a final data table; an establishment module 1550, configured to establish a basic classification prediction model using the final data table; a determination module 156 0, is configured to use the variables in the final data table as features, and according to the business data associated with each user in the final data table, determine the average marginal contribution of each feature in the business data to the prediction result of the basic classification prediction model as the explanatory weight of each feature for the business data; the first generation module 1570, is configured to determine the target feature for each business data according to the explanatory weight of each feature for the business data, and generate deep triple information according to the variable value corresponding to the target feature, wherein the explanatory weight of the target feature is greater than that of other features; the construction module 1580, is configured to construct a knowledge graph using the shallow triple information and the deep triple information; the second generation module 1590, is configured to obtain the retrieval result by querying the knowledge graph when receiving a business question retrieval request, and generate marketing auxiliary information according to the retrieval result.
[0137] According to the third aspect of the present application, an electronic device capable of implementing the above method is also provided. Those skilled in the art will understand that various aspects of the present application can be implemented as systems, methods or program products. Therefore, various aspects of the present application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation that combines hardware and software aspects, which can be collectively referred to as "circuits", "modules" or "systems". Refer to the following Figure 16 16 is a diagram to describe an electronic device 1600 according to this embodiment of the present application. Figure 16 The electronic device 1600 shown is only an example and should not limit the functions and scope of use of the embodiments of the present application. Figure 16As shown, electronic device 1600 is implemented as a general-purpose computing device. Components of electronic device 1600 may include, but are not limited to, the at least one processing unit 1610 described above, the at least one storage unit 1620 described above, and a bus 1630 connecting various system components (including storage unit 1620 and processing unit 1610). The storage unit stores program code that can be executed by the processing unit 1610, causing the processing unit 1610 to perform the steps described in the "Example Method" section above of this specification according to various exemplary embodiments of the present application. The storage unit 1620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 1621 and / or a cache memory unit 1622, and may further include a read-only memory unit (ROM) 1623. The storage unit 1620 may also include a program / utility 1624 having a set (at least one) of program modules 1625. Such program modules 1625 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Bus 1630 can represent one or more of several types of bus structures, including a storage unit bus or storage unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures. Electronic device 1600 can also communicate with one or more external devices 1800 (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), one or more devices that enable a user to interact with electronic device 1600, and / or any device that enables electronic device 1600 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). Such communication can be performed via input / output (I / O) interface 1650, such as with display unit 1640. Furthermore, electronic device 1600 can also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via network adapter 1660. As shown, network adapter 1660 communicates with other modules of electronic device 1600 via bus 1630. It should be understood that although not shown in the figures, other hardware and / or software modules may be used in conjunction with the electronic device 1600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0138] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0139] According to a fourth aspect of the present application, a computer-readable storage medium is further provided, on which is stored a program product capable of implementing the above-described method of this specification. In some possible implementations, various aspects of the present application may also be implemented in the form of a program product, which includes program code. When the program product is executed on a terminal device, the program code is used to cause the terminal device to perform the steps of the various exemplary implementations of the present application described in the "Exemplary Methods" section of this specification.
[0140] refer to Figure 17As shown, a program product 1700 for implementing the above method according to an embodiment of the present application is described. It can use a portable compact disk read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present application is not limited to this. In this document, a readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device or device. The program product can use any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A readable signal medium may also be any readable medium other than a readable storage medium, which may transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing. The program code for performing the operations of this application may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a standalone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., via the Internet using an Internet service provider). Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present application and are not intended to be limiting. It is readily understood that the processes illustrated in the above figures do not indicate or limit the temporal order of these processes.In addition, it is also easy to understand that these processes can be executed synchronously or asynchronously, for example, in multiple modules.
[0141] It should be understood that the present application is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be performed without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A method for generating marketing auxiliary information, characterized in that: The method comprises: Extracting structured data from a big data platform, the structured data including user-related data and product-related data, the big data platform including multiple information management systems; Converting the structured data into triple information to obtain shallow triple information; Obtaining a data table from the big data platform, the data table including multiple business data, the business data including variables and variable values corresponding to the variables, each of the business data being associated with a user; Performing a variable screening operation on the data table to remove at least one variable in the data table to obtain a final data table; Establishing a basic classification prediction model using the final data table; Dividing the business data in the final data table into multiple layers according to a predetermined rule, each layer including multiple business data; Selecting one business data from the final data table as selected business data; Taking the variables in the final data table as features, for each feature, iteratively executing the step of determining the marginal contribution value until a predetermined number of times are executed, wherein the step of determining the marginal contribution value includes: randomly generating a feature order, and sorting the selected business data and the business data in the final data table according to the feature order; selecting a layer from an unselected layer, and randomly selecting a business data from the business data of the layer as the constructed business data corresponding to the selected business data; constructing first instance business data and second instance business data respectively according to the selected business data and the constructed business data, wherein the first instance business data includes the selected business data The variable values corresponding to the feature and the feature preceding the feature in the selected business data and the variable values corresponding to the feature following the feature in the constructed business data are respectively input into the basic classification prediction model to obtain a first prediction result corresponding to the selected business data and a second prediction result corresponding to the constructed business data; and the marginal contribution value corresponding to the feature is determined based on the first prediction result and the second prediction result; For each feature, determining an average contribution margin based on the respective contribution margin values obtained by performing the step of determining the contribution margin value for the feature, as an explanatory weight of the feature for the business data; For each business data, determine the target feature according to the explanatory weight of each feature on the business data, and generate deep triple information according to the variable value corresponding to the target feature, wherein the explanatory weight of the target feature is greater than that of other features; Constructing a knowledge graph using the shallow triple information and the deep triple information; When a business question retrieval request is received, the retrieval results are obtained by querying the knowledge graph, and marketing auxiliary information is generated based on the retrieval results.
2. The method according to claim 1, characterized in that The constructing of a knowledge graph using the shallow triple information and the deep triple information includes: Importing the shallow triple information and the deep triple information into a graph database to form a knowledge graph in the graph database; When a business question retrieval request is received, retrieval results are obtained by querying the knowledge graph, and marketing auxiliary information is generated based on the retrieval results, including: When a business problem search request is received, obtaining the business problem in the business problem search request; Determine a result template that matches the business problem; Querying the knowledge graph in the graph database to obtain search results corresponding to the business question; The search results are filled into the result template to obtain marketing auxiliary information.
3. The method according to claim 2, characterized in that The determining of a result template matching the business problem includes: Determine the category to which the business problem belongs through a preset classification model; Determine, by using a first syntax analysis model, a question template matching the business question from the question templates corresponding to the category, as a target question template; Obtaining a result template corresponding to the target problem template as a result template matching the business problem; Querying the knowledge graph in the graph database to obtain a search result corresponding to the business question includes: Extracting the query object in the business problem through the second syntax analysis model; According to the query object, the knowledge graph is queried in the graph database to obtain a retrieval result corresponding to the business question.
4. The method according to claim 1, wherein The method further includes: outputting the interpretation weight of each feature on the business data in a visual manner according to a user request.
5. The method according to claim 1, characterized in that The method of establishing a basic classification prediction model using the final data table includes: The final data table is used to train a logistic regression model, a CART model, and an Xgboost model respectively; A basic classification prediction model is established based on the logistic regression model, the CART model and the Xgboost model, wherein the prediction result of the basic classification prediction model is a weighted calculation result obtained by weighted calculation of the output results of the logistic regression model, the CART model and the Xgboost model.
6. The method according to claim 1, characterized in that The performing a variable screening operation on the data table to remove at least one variable in the data table to obtain a final data table includes: The chi-square test method, the correlation coefficient calculation method and the information value evaluation method are used in sequence to perform variable screening operations on the data table to eliminate at least one variable in the data table to obtain a final data table.
7. A device for generating marketing auxiliary information, characterized in that: The device comprises: an extraction module configured to extract structured data from a big data platform, wherein the structured data includes user-related data and product-related data, and the big data platform includes multiple information management systems; a conversion module, configured to convert the structured data into triple information to obtain shallow triple information; an acquisition module configured to acquire a data table from the big data platform, the data table including multiple business data items, the business data including variables and variable values corresponding to the variables, each of the business data items being associated with a user; a removal module configured to perform a variable screening operation on the data table to remove at least one variable in the data table to obtain a final data table; An establishing module configured to establish a basic classification prediction model using the final data table; a determination module configured to use the variables in the final data table as features and, based on the business data associated with each user in the final data table, determine an average marginal contribution of each feature in the business data to the prediction result of the basic classification prediction model as an explanatory weight of each feature for the business data; The first generating module is configured to determine, for each piece of business data, a target feature according to the explanatory weights of each feature on the business data, and generate deep triplet information according to the variable value corresponding to the target feature, wherein the explanatory weight of the target feature is greater than that of other features; A construction module, configured to construct a knowledge graph using the shallow triple information and the deep triple information; The second generating module is configured to, upon receiving a business question retrieval request, obtain retrieval results by querying the knowledge graph, and generate marketing auxiliary information based on the retrieval results; The determination module is further configured to: divide the business data in the final data table into multiple layers according to a predetermined rule, each layer including multiple business data; select a business data from the final data table as the selected business data; for each feature, iteratively execute the step of determining the marginal contribution value until it is executed a predetermined number of times, wherein the step of determining the marginal contribution includes: randomly generating a feature sequence, and sorting the selected business data and the business data in the final data table according to the feature sequence; selecting a layer from the layers that have not been selected, and randomly selecting a business data from the business data of the layer as the constructed business data corresponding to the selected business data; constructing the first instance business data and the second instance business data according to the selected business data and the constructed business data, wherein the first instance business data includes the business data in the selected business data. The variable values corresponding to the feature and the feature before the feature, and the variable values corresponding to the feature after the feature in the constructed business data, the second instance business data includes the variable values corresponding to the feature before the feature in the selected business data, and the variable values corresponding to the feature and the feature after the feature in the constructed business data; the selected business data and the constructed business data are respectively input into the basic classification prediction model to obtain a first prediction result corresponding to the selected business data and a second prediction result corresponding to the constructed business data; the marginal contribution value corresponding to the feature is determined based on the first prediction result and the second prediction result; for each feature, the marginal contribution average is determined based on the marginal contribution values obtained by executing the step of determining the marginal contribution value for the feature, as the explanatory weight of the feature on the business data.
8. A computer-readable program medium, characterized in that The computer program instructions are stored therein, and when the computer program instructions are executed by a computer, the computer is caused to execute the method according to any one of claims 1 to 6.
9. An electronic device, characterized in that: The electronic device comprises: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Domain knowledge graph construction method and system based on big data driving
CN109597855A
Knowledge graph Schema design method based on enterprise data
CN111191043A