A method and apparatus for identifying key behavior data
By configuring behavior judgment rules in the rules engine and filtering and optimizing rule sets with machine learning or causal inference models, we automatically identify and store key behavior data, and solve the problems of low manual screening efficiency and limited database storage for business personnel, and achieve efficient and accurate identification and storage of key clues.
Patent Information
- Application Number
- CN202111423413.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-26
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-11-26
AI Technical Summary
In the prior art, business personnel need to manually screen out valuable key behavioral data from a large amount of customer information, which is inefficient and has limited database storage capabilities, making it difficult to effectively identify key customer acquisition clues.
By receiving behavior judgment rule configuration in the rule engine, the optimized rule set is filtered out from the rule set using machine learning, causal inference or interpretation neural network models, automatically identify key behavior data, and collect and store these data through unified data acquisition standards.
It improves the identification efficiency and accuracy of key behavior data, enhances the storage capabilities of the database, realizes automated identification and storage of key clues, and improves the customer acquisition efficiency of business personnel.
Smart Images

Figure CN114297234B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method and device for identifying key behavior data. Background Art
[0002] With the rapid development of Internet technology, Internet insurance has gradually transferred the sale and service process of insurance products from the traditional offline to the online. The online behavior of customers is increasing day by day. Facing the daily increasing user behavior information, how to identify key customer acquisition clues poses new challenges to business personnel for online customer acquisition. Using traditional database collection and message push, through the method of collecting from the business database and manual screening, all customer behavior data is pushed to business personnel in real time and in full volume for independent selection. Business personnel need to identify valuable clues for customer acquisition from numerous messages. In addition, when reporting behavior data to the business database in a hard-coded manner, the limitation problem of the data volume of the traditional database needs to be considered, and the storage capacity is limited, resulting in the inability to collect key clues.
[0003] In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art:
[0004] Business personnel need to manually screen valuable key behavior data from a large amount of customer information, which has poor usability and low efficiency, and the storage capacity of the database for behavior data is limited. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a method and device for identifying key behavior data, which can automatically load a rule engine to identify key behavior data, with high usability and improved clue identification efficiency. By combining an algorithm model to optimize the rule engine, the accuracy of clue identification is improved. Behavior data can be collected through a unified behavior data collection standard for storage, improving the database storage capacity.
[0006] To achieve the above object, according to one aspect of the embodiments of the present invention, a method for identifying key behavior data is provided.
[0007] A method for identifying key behavior data includes: receiving the configuration of behavior judgment rules in a rule engine, obtaining a first rule set according to the configured behavior judgment rules; screening out a second rule set from the first rule set through a preset algorithm model; using the second rule set as the rule set actually used by the rule engine, and inputting the collected dataset to be identified into the rule engine to identify the key behavior dataset in the dataset to be identified, where the dataset to be identified includes the behavior data of business personnel and customers.
[0008] Optionally, the preset algorithm model is a machine learning algorithm model, and the machine learning algorithm model is a logistic regression model or a gradient boosting model; screening the second rule set from the first rule set through the preset algorithm model includes: constructing input features of the machine learning algorithm model, where the input features of the machine learning algorithm model include features obtained based on the first rule set; outputting a feature score corresponding to each input feature through the machine learning algorithm model, and classifying the behavior judgment rules in the first rule set corresponding to the feature scores greater than a preset threshold into the second rule set.
[0009] Optionally, the preset algorithm model is a causal inference model; screening the second rule set from the first rule set through the preset algorithm model includes: using the causal inference model to determine a screening metric in the following manner: obtaining a first behavior data sequence of a customer in a first customer set before the current moment, where the first customer set is a set composed of customers who have not yet generated a behavior result; selecting one or more target customers from a second customer set whose behavior data sequence before the current moment is similar to the first behavior data sequence, and obtaining a second behavior data sequence of each target customer from after the current moment to before generating the behavior result, where the second customer set is a set composed of customers known to have the behavior result; determining the screening metric according to the behavior data with the number of occurrences greater than a first number threshold in each second behavior data sequence, and / or the behavior data with the total number of occurrences greater than a second number threshold in all second behavior data sequences; using the screening metric to screen behavior judgment rules from the first rule set, and classifying the screened behavior judgment rules into the second rule set.
[0010] Optionally, the preset algorithm model is an interpretive neural network model; screening the second rule set from the first rule set through the preset algorithm model includes: constructing input features of the interpretive neural network model, where the input features of the interpretive neural network model include features obtained based on the first rule set; obtaining an interpretation result of the neural network through an interpretation module of the interpretive neural network model, where the interpretation result includes text about behavior data; matching the interpretation result with the input features of the interpretive neural network model, and obtaining behavior judgment rules from the matched input features and classifying them into the second rule set.
[0011] Optionally, using an event model as the data collection standard for collecting the dataset to be recognized, collecting the behavior data of the business personnel and the behavior data of the customers respectively, where the event model includes a user entity and an event entity, at least some fields of the user entity and the event entity are different, and the user is a business personnel or a customer.
[0012] Optionally, inputting the to-be-identified data set to be collected into the rule engine to identify the key behavior data set in the to-be-identified data set includes: performing metric processing on the to-be-identified data set to extract metrics for behavior judgment and corresponding metric values; determining, by the rule engine, whether the extracted metrics and corresponding metric values conform to the behavior judgment rules in the actually used rule set. If they conform, the behavior data corresponding to the metrics in the to-be-identified data set is used as key behavior data and added to the key behavior data set.
[0013] Optionally, after identifying the key behavior data set in the to-be-identified data set, it includes: storing the key behavior data set in a big data cluster or a structured database, and providing a real-time data query service through a data service interface for the client of the business personnel to query the behavior data in the key behavior data set; or providing a message channel push service through the data service interface to push the behavior data in the key behavior data set to the client of the business personnel through the message channel.
[0014] According to another aspect of the embodiments of the present invention, a device for identifying key behavior data is provided.
[0015] A device for identifying key behavior data includes: a first rule set configuration module, configured to receive the configuration of the behavior judgment rules in the rule engine and obtain a first rule set according to the configured behavior judgment rules; a second rule set screening module, configured to screen out a second rule set from the first rule set through a preset algorithm model; a key behavior data identification module, configured to use the second rule set as the actually used rule set of the rule engine, input the to-be-identified data set to be collected into the rule engine to identify the key behavior data set in the to-be-identified data set, and the to-be-identified data set includes the behavior data of business personnel and customers.
[0016] Optionally, the preset algorithm model is a machine learning algorithm model, and the machine learning algorithm model is a logistic regression model or a gradient boosting model; the second rule set screening module is further configured to: construct the input features of the machine learning algorithm model, and the input features of the machine learning algorithm model include the features obtained based on the first rule set; output the feature scores corresponding to each input feature through the machine learning algorithm model, and classify the behavior judgment rules in the first rule set corresponding to the feature scores greater than a preset threshold into the second rule set.
[0017] Optionally, the preset algorithm model is a causal inference model; the second rule set screening module is further configured to: use the causal inference model to determine screening metrics in the following manner: obtain a first behavior data sequence of customers in the first customer set before the current moment, where the first customer set is a set composed of customers who have not yet generated behavior results; select one or more target customers from the second customer set whose behavior data sequences before the current moment are similar to the first behavior data sequence, and obtain a second behavior data sequence of each target customer from after the current moment until the generation of the behavior result, where the second customer set is a set composed of customers known to have the behavior result; determine the screening metrics according to the behavior data with the occurrence times greater than the first frequency threshold in each second behavior data sequence, and / or the behavior data with the total occurrence times greater than the second frequency threshold in all second behavior data sequences; use the screening metrics to screen behavior judgment rules from the first rule set, and classify the screened behavior judgment rules into the second rule set.
[0018] Optionally, the preset algorithm model is an interpretive neural network model; the second rule set screening module is further configured to: construct input features of the interpretive neural network model, where the input features of the interpretive neural network model include features obtained based on the first rule set; obtain an interpretation result of the neural network through an interpretation module of the interpretive neural network model, where the interpretation result includes text about behavior data; match the interpretation result with the input features of the interpretive neural network model, and obtain behavior judgment rules from the matched input features and classify them into the second rule set.
[0019] Optionally, it further includes a data collection module, configured to use an event model as the data collection standard for collecting the data set to be recognized, and respectively collect the behavior data of the business personnel and the behavior data of the customers. The event model includes a user entity and an event entity, and at least some fields of the user entity and the event entity are different, and the user is a business personnel or a customer.
[0020] Optionally, the key behavior data recognition module is further configured to: perform metric processing on the data set to be recognized to extract metrics for behavior judgment and corresponding metric values; determine whether the extracted metrics and corresponding metric values conform to the behavior judgment rules in the actually used rule set through the rule engine. If they conform, use the behavior data corresponding to the metric in the data set to be recognized as key behavior data and add it to the key behavior data set.
[0021] Optionally, it further includes a data storage and provision module for: storing the critical behavior data set into a big data cluster or a structured database, and providing real-time data query services through a data service interface for the client of the business personnel to query the behavior data in the critical behavior data set; or providing a message channel push service through the data service interface to push the behavior data in the critical behavior data set to the client of the business personnel through a message channel.
[0022] According to another aspect of the embodiments of the present invention, an electronic device is provided.
[0023] An electronic device includes: one or more processors; a memory for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the method for identifying critical behavior data provided by the embodiments of the present invention.
[0024] According to another aspect of the embodiments of the present invention, a computer-readable medium is provided.
[0025] A computer-readable medium has a computer program stored thereon, and when the program is executed by a processor, it implements the method for identifying critical behavior data provided by the embodiments of the present invention.
[0026] One embodiment of the above invention has the following advantages or beneficial effects: receiving the configuration of the behavior judgment rules in the rule engine, obtaining a first rule set according to the configured behavior judgment rules; screening out a second rule set from the first rule set through a preset algorithm model; using the second rule set as the rule set actually used by the rule engine, and inputting the collected data set to be identified into the rule engine to identify the critical behavior data set in the data set to be identified, where the data set to be identified includes the behavior data of business personnel and customers. It can automatically load the rule engine to identify critical behavior data, has high usability and improves the lead identification efficiency, optimizes the rule engine in combination with the algorithm model, improves the accuracy of lead identification, and can collect behavior data through a unified behavior data collection standard for storage, improving the database storage capacity.
[0027] The further effects of the above non-conventional optional methods will be described in conjunction with specific embodiments below. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:
[0029] Figure 1 is a schematic diagram of the main steps of the method for identifying critical behavior data according to an embodiment of the present invention;
[0030] Figure 2Schematic diagram of the real-time acquisition process of behavior data according to an embodiment of the present invention;
[0031] Figure 3 Schematic diagram of the main modules of the recognition device for key behavior data according to an embodiment of the present invention;
[0032] Figure 4 Exemplary system architecture diagram to which the embodiments of the present invention can be applied;
[0033] Figure 5 Schematic diagram of the structure of a computer system of a server suitable for implementing the embodiments of the present invention. Detailed implementation manners
[0034] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0035] Figure 1 Schematic diagram of the main steps of the recognition method for key behavior data according to an embodiment of the present invention. As Figure 1 shown, the recognition method for key behavior data in an embodiment of the present invention mainly includes the following steps S101 to step S103.
[0036] Step S101: Receive the configuration of the behavior judgment rules in the rule engine, and obtain the first rule set according to the configured behavior judgment rules;
[0037] Step S102: Screen out the second rule set from the first rule set through a preset algorithm model;
[0038] Step S103: Use the second rule set as the rule set actually used by the rule engine, and input the collected data set to be recognized into the rule engine to identify the key behavior data set in the data set to be recognized. The data set to be recognized includes the behavior data of business personnel and customers. It should be understood that these behavior data and the customer information and behavior information mentioned in the text are data and information that can be collected, stored, and used for subsequent recommendation and other applications with user authorization.
[0039] The configuration of the behavior judgment rules in the rule engine can be manually input into the rule engine.
[0040] In one embodiment, the preset algorithm model is a machine learning algorithm model, and specifically, the machine learning algorithm model can be a logistic regression model or a gradient boosting model. Screening the second rule set from the first rule set through the preset algorithm model, the specific steps may include: constructing the input features of the machine learning algorithm model. The input features of the machine learning algorithm model may include features obtained based on the first rule set, and may also include features obtained through manual screening or algorithm screening, etc., which will be introduced in detail below; outputting the feature scores corresponding to each input feature through the machine learning algorithm model, and classifying the behavior judgment rules in the first rule set with the corresponding feature scores greater than the preset threshold into the second rule set.
[0041] In one embodiment, the preset algorithm model is a causal inference model. Screening the second rule set from the first rule set through the preset algorithm model, the specific steps may include: using the causal inference model to determine the screening index in the following way: obtaining the first behavior data sequence of the customers in the first customer set before the current moment, for example, the sequence (t1, t2,..., t n ), which includes the 1st to nth behavior data. The first customer set is a set composed of customers who have not yet generated behavior results. The behavior result refers to the behavior that the key behavior data will cause. For example, browsing a certain insurance product A for t seconds is the key behavior data, and the behavior it will cause is underwriting, so underwriting is the behavior result; selecting one or more target customers from the second customer set whose behavior data sequence before the current moment is similar to the first behavior data sequence, and obtaining the second behavior data sequence of each target customer from the current moment until the behavior result is generated, for example, the sequence (t n+1 , t n+2 ,..., t n+m ), where the second customer set is a set composed of customers who are known to have behavior results, such as a set composed of customers who are known to be underwritten; determining the screening index according to the behavior data with the occurrence times greater than the first occurrence threshold in each second behavior data sequence, and / or the behavior data with the total occurrence times greater than the second occurrence threshold in all second behavior data sequences; using the screening index to screen the behavior judgment rules from the first rule set, and classifying the screened behavior judgment rules into the second rule set.
[0042] In one embodiment, the preset algorithm model is an interpretive neural network model. Screening the second rule set from the first rule set through the preset algorithm model, the specific steps may include: constructing the input features of the interpretive neural network model. The input features of the interpretive neural network model include features obtained based on the first rule set; obtaining the interpretation result of the neural network through the interpretation module of the interpretive neural network model. The interpretation result includes the text about the behavior data; matching the interpretation result with the input features of the interpretive neural network model, and obtaining the behavior judgment rules from the matched input features and classifying them into the second rule set.
[0043] In one embodiment, an event model can be used as the data acquisition standard for collecting the dataset to be recognized, and the behavior data of business personnel and the behavior data of customers are collected respectively. The event model includes a user entity and an event entity, and at least some fields of the user entity and the event entity are different. The user is a business person or a customer.
[0044] The collected dataset to be recognized is input into a rule engine to identify the key behavior dataset in the dataset to be recognized. Specifically, it can include: performing metric processing on the dataset to be recognized to extract metrics for behavior judgment and corresponding metric values; judging whether the extracted metrics and corresponding metric values conform to the behavior judgment rules in the actually used rule set through the rule engine. If they conform, the behavior data corresponding to the metric in the dataset to be recognized is used as key behavior data and added to the key behavior dataset. For example, perform metric processing on the dataset to be recognized, extract the metric of product browsing duration from a certain behavior data in the dataset to be recognized, and the corresponding metric value is t seconds. If the behavior judgment rule is whether the product browsing duration reaches t seconds, then this behavior data conforms to the behavior judgment rule and is added to the key behavior dataset.
[0045] After identifying the key behavior dataset in the dataset to be recognized, the key behavior dataset can be stored in a big data cluster or a structured database, and real-time data query services can be provided through a data service interface for the client of business personnel to query the behavior data in the key behavior dataset; or, a message channel push service can be provided through the data service interface to push the behavior data in the key behavior dataset to the client of business personnel through the message channel. Business personnel can use the key behavior data queried or pushed over as the key clues for business expansion.
[0046] Taking the insurance field as an example, the method for identifying key behavior data in the embodiments of the present invention is further introduced below. When collecting the dataset to be recognized in the embodiments of the present invention, the behavior data of business personnel and customers are collected through a unified data acquisition standard, that is, an event model. The event model includes two entities: a user and an event. The user entity includes fields such as attribute English variable name, attribute display name, attribute value type, attribute value description or example, etc. The event entity includes fields such as user ID (identifier), event name, affiliated channel, page name, behavior time, etc. The content of the user entity table is shown in Table 1 as an example, and the content of the event entity table is shown in Table 2 as an example.
[0047] Table 1
[0048]
[0049] Table 2
[0050]
[0051] The embodiments of the present invention also set public attributes such as team ID, product line, business line, and business line version, all of which are dynamic public attributes. Each event in the event entity table can include public attributes, and events from different team IDs, product lines, business lines, business line versions, etc. can be identified through the public attributes. The public attributes can be as shown in Table 3.
[0052] Table 3
[0053]
[0054]
[0055] The behavior data of insurance business personnel and customers at each end are reported in the above format, which is convenient for data processing, analysis, and storage.
[0056] For the collected dataset to be identified, the rule engine is used to identify the key behavior dataset therein. The key behavior dataset is a collection of key behavior data, and the key behavior data can be used as the key clues for insurance business personnel to expand customers.
[0057] Different rule sets for key clues can be set according to scenarios. For example, behavior judgment rules are generated based on product list browsing, product clicking, product browsing duration, immediate insurance purchase clicking, etc.
[0058] For example, if the number of product list browsing and product clicking exceeds N times (which can be dynamically adjusted according to the effect) and the product browsing duration exceeds t seconds, then the behavior data is considered as key behavior data, and the information such as the last time, number of times, time, etc. when the event of the number of product clicks exceeding N times and the product browsing duration exceeding t seconds occurs is stored in the structured database, and the identified key behavior data can be sent to insurance business personnel in real time.
[0059] The embodiment of the present invention provides an automatic setting system for a rule engine, which can dynamically set the behavior judgment rules (hereinafter referred to as rules) in the rule engine according to business needs. The set of configured behavior judgment rules constitutes the configured rule set of the rule engine. After the setting is completed, select to publish the set rules. The ETL module (ETL is the abbreviation of Extract-Transform-Load, that is, the process of data extraction, transformation, and loading. The specific role of the ETL module will be further introduced in the following embodiments) can automatically load the rule engine and identify the dataset to be recognized. When the behavior data in the dataset to be recognized meets the rule conditions in the rule engine, the behavior data is identified as key behavior data, and the key behavior data is stored in the database or triggers a message channel to be sent to business personnel in real time. When the behavior data in the dataset to be recognized does not meet the rule conditions in the rule engine, the behavior data is stored in the Redis cache database, and a TTL (expiration time) is set for the behavior data in the Redis cache. The behavior data that has not met the rule conditions for a long time is deleted, and it is considered that the behavior data cannot become a key clue for insurance business personnel to expand customers.
[0060] In one embodiment, an algorithm model can be used to optimize the rule engine. Specifically, the above-mentioned configured rule set for the rule engine can be called the first rule set, and a second rule set can be selected from the first rule set through a preset algorithm model. The second rule set is a subset of the first rule set, and the second rule set is used as the rule set actually used by the rule engine, that is, the rule engine uses the second rule set to identify key behavior data from the dataset to be recognized.
[0061] The preset algorithm model can be one or more of a machine learning algorithm model, a causal inference model, and an interpretive neural network model. When the preset algorithm model adopts multiple models, the union of the rule sets obtained through each model can be used as the second rule set.
[0062] First, the machine learning algorithm model of the embodiments of the present invention will be introduced. Specifically, it can be LR (Logistic Regression Model) or XGBoost (Gradient Boosting Model). The process of screening out the second rule set from the first rule set through the machine learning algorithm model includes: First, the selection and processing of behavioral data features (i.e., the input features of the preset algorithm model) can be carried out. The selection of behavioral data features can adopt one or more methods such as manual screening, rule engine screening, and algorithm screening. Manual screening means artificially selecting behavioral features with strong association with the training objective (i.e., whether the customer is insured) from the buried point events of business personnel and customers. For example, feature feature1 is: whether product A is clicked, feature2 is: the daily click count of product A, feature3 is: whether product A is clicked to apply for insurance immediately, feature4 is: whether the product browsing duration is greater than t seconds, and so on. Rule engine screening means taking the rules in the rule engine set by business personnel as the behavioral data features of the preset algorithm model, inputting them into the preset algorithm model, and through the output of the preset algorithm model (i.e., the feature score), it can be verified whether the rules in the rule engine set by business personnel are high-value key behavioral clues. Algorithm screening can select algorithms for feature correlation analysis such as covariance, Pearson coefficient, etc., and only retain the variables with a larger Pearson coefficient with the label. For example, if feature feature1 and feature feature4 are strongly correlated, but feature feature4 has a larger Pearson coefficient with the training objective, then feature feature4 is selected as the input feature of the preset algorithm model to participate in the training. The purpose is to make only one of the two strongly correlated features participate in the training. The embodiments of the present invention can be implemented by using the commonly used methods for calculating covariance and Pearson coefficient in feature correlation analysis, and the embodiments of the present invention will not be introduced in detail here. The processing of behavioral data features can perform corresponding numerical processing according to the type of feature setting. For example, if the feature is whether it is clicked, click is set to 1 and not clicked is set to 0. If the feature is the click count, then the daily click count is statistically calculated, and so on.
[0063] Then, the cleaning of feature data can be carried out. The null values and missing values in the features processed above are processed, and default values 0 or other values are filled. For continuous features, the features with a large numerical span are selected for normalization processing. Since the embodiments of the present invention adopt the LR or XGBoost model, preferably, the continuous features are discretized. At the same time, for the features with a large proportion of null values, they can be deleted.
[0064] In the model training stage, the behavioral data features X of each customer can be associated with whether the customer is insured Y to form sample data, and the ratio of positive and negative samples in the sample can be counted. If the sample is imbalanced, a sample sampling method can be used to screen the sample data. The finally processed sample data is divided into a training set, a test set, and a validation set in a certain ratio, such as 7:2:1, for model training. In the model optimization stage, the parameters L1, L2, and learning rate α in the model can be selected to optimize the model effect, or the model effect can be improved by optimizing the features. The specific model optimization can refer to the general machine learning model optimization method, which will not be introduced in detail in the embodiments of the present invention. Among them, optimizing the features can remove the features with weak feature importance, or process and combine features. For example, the click count feature2 and the click to immediately insure feature4 are combined to form a combined feature and added to the model for training.
[0065] Use the LR or XGBoost algorithm model to optimize the rule set in the rule engine. Specifically, the input features of the algorithm model can be constructed based on the first rule set configured in the rule engine. The first rule set can be used as a subset of the input feature set of the algorithm model, and other features in the input feature set can be obtained through the manual screening, algorithm screening, etc. introduced above. Input the input features into the LR or XGBoost algorithm model, and the score corresponding to each input feature, that is, the feature score, can be output. Sort the features according to the score. For example:
[0066] [feature3: 0.2026578, feature4: 0.17109634, feature1: 0.089701, feature2: 0.04651163,......]. The input features that have a score greater than 0.15 (a preset threshold, which can be determined according to empirical values) and belong to the first rule set are attributed to the second rule set, thereby optimizing the rule set in the rule engine to the second rule set.
[0067] The causal inference model of the embodiments of the present invention is introduced below. The specific implementation steps for screening out the second rule set from the first rule set through the causal inference model are as follows: First, the behavioral data sequence t (t1, t2,..., t n )(n is a positive integer greater than 0) of a single customer in the first customer set can be determined. This behavioral data sequence is the first behavioral data sequence. The current moment is after the behavior corresponding to t n is completed, but before t n+1The moment corresponding to the behavior. The first customer set is the set composed of customers who have not yet produced a behavior result. Taking the insurance field as an example, the behavior result can be underwriting, and the customers in the first customer set are the set composed of customers who have not yet determined whether they will be underwritten. For example, the behavior data sequence t of a certain customer A (t1: browse the home page, t2: browse information A, t3: browse the product list, t4: click on product A, t5: browse product A for t seconds) (n = 5), the behavior data sequence t of a single customer (t1, t2,..., t n ) is namely t1, t2,..., t5. Select one or more target customers from the second customer set whose behavior data sequences before the current moment are similar to the first behavior data sequence. The second customer set is the set composed of customers who are known to have a behavior result. Taking the insurance field as an example, the second customer set is the set composed of customers who are known to be underwritten. The process of selecting target customers from the second customer set is to select the customer set S that is similar (including the same) to the sequence t (t1, t2,..., t n ) and is underwritten. The customers in the customer set S are the target customers. Obtain the second behavior data sequence of each target customer from the current moment until before the behavior result is generated. For this example, it is to obtain (or construct) the behavior data sequence between t n and underwriting in the customer set S (t n+1 , t n+2 ,..., t n+m ) (m is a positive integer greater than 0), (t n+1 , t n+2 ,..., t n+m ) is the second behavior data sequence. When there are multiple behavior data, the behavior data with the most occurrences can be taken. If there are multiple behavior data with the most occurrences, one of them can be randomly taken.
[0068] The screening index can be determined according to the behavior data in a single second behavior data sequence whose occurrence times are greater than the first occurrence threshold, and / or the behavior data in all second behavior data sequences whose total occurrence times are greater than the second occurrence threshold. That is: taking the behavior data in the behavior data sequence (t n+1 , t n+2 ,..., t n+m ) whose occurrence times are greater than the first occurrence threshold as the screening index, and / or taking the behavior data in all behavior data sequences (t n+1 , t n+2 ,..., t n+m ) in the customer set S whose total occurrence times are greater than the second occurrence threshold as the screening index. Through the above screening index, the behavior judgment rules can be screened from the first rule set. For example, a certain behavior data sequence (t n+1 , t n+2 ,..., t n+m) If the number of times of browsing Product A for more than t seconds is greater than 5 times, then whether the browsing time of Product A is greater than t seconds can be used as a behavior judgment rule. Or, all behavior data sequences (t n+1 , t n+2 ,..., t n+m ) If the total number of times of browsing Product A for more than t seconds is greater than 5k times (k can be the total number of target customers in the customer set S), then whether the browsing time of Product A is greater than t seconds can be used as a behavior judgment rule. The behavior judgment rules screened from the first rule set through the above screening indicators are attributed to the second rule set.
[0069] The following introduces the explanatory neural network model of the embodiments of the present invention. The explanatory neural network model (Explaining Deep Neural Networks) can implant a module for generating prediction explanations in the model to explain the results predicted by the model. The model can judge the explanations generated by itself to obtain the results of neural network explanations. Since the explanatory neural network model itself is an existing model, the embodiments of the present invention do not introduce the model itself in detail. The specific process of screening the second rule set from the first rule set by the explanatory neural network model in the embodiments of the present invention is as follows: First, the input features of the explanatory neural network model can be constructed, specifically, the selection and processing of behavior data features (i.e., input features) and the cleaning of feature data are performed. The specific methods are the same as those of feature selection, processing, and cleaning in LR (logistic regression model) or XGBoost (gradient boosting model) introduced above, which will not be elaborated here. In the model training stage, in addition to generating sample data, the explanatory neural network model needs to implant a module for generating prediction explanations in the model to obtain the results of neural network explanations, and the explanation results include texts about behavior data. In the model optimization stage, the effects of the model are adjusted by adjusting hyperparameters such as the number of layers of the neural network, dropout (abandonment), or learning rate. The explanation results are matched with the input features of the explanatory neural network model, and the behavior judgment rules are obtained from the matched input features, and the obtained behavior judgment rules are attributed to the second rule set. For example, the explanation result includes a text about whether the product browsing duration is greater than t seconds: product browsing duration. An input feature of the explanatory neural network model is whether the product browsing duration is greater than t seconds. Then, by matching the two, this input feature can be used as the matched input feature, that is, the obtained behavior judgment rule is whether the product browsing duration is greater than t seconds, and this rule is attributed to the second rule set.
[0070] In one embodiment, two or more preset algorithm models can be selected from all the above-mentioned preset algorithm models. A rule set can be obtained respectively through the selected two or more preset algorithm models, and then the rule sets obtained by the selected preset algorithm models are attributed to the second rule set. Specifically, the union of the rule sets output by the selected preset algorithm models can be taken to obtain the second rule set, that is, the rule set actually used by the rule engine.
[0071] In one embodiment of the present invention, preset algorithm models such as machine learning algorithm models, causal inference models, and interpretive neural network models can be directly used to identify key behavior data sets in the data set to be identified in addition to being used to optimize the rule set of the rule engine. The recognition results of each model are used as a supplement to the recognition results of the rule engine. For example, the first key behavior data set identified from the data set to be identified through the preset algorithm model and the second key behavior data set identified from the data set to be identified through the rule engine are combined, and the union of the two key behavior data sets is taken to obtain the key behavior data set in the data set to be identified finally recognized. When using a machine learning algorithm model, such as the LR or XGBoost algorithm model, to identify the key behavior data set, the input features are the features corresponding to the behavior data in the data set to be identified. For the specific form of the input features, refer to the above introduction to the machine learning algorithm model. The feature scores are output, and sorted according to the feature scores. The behavior data with a feature score greater than 0.15 (preset threshold) corresponding to the features in the data set to be identified can be used as key behavior data. For example, if a behavior data is whether the product browsing duration is greater than t seconds, and the feature score obtained through the machine learning algorithm model of the embodiment of the present invention is greater than 0.15, then this behavior data can be used as key behavior data. When using the causal inference model to identify key behavior data, referring to the above introduction, the data set to be identified can be used as the first customer set, and the screening indicators obtained through the causal inference model can be used to identify the key behavior data in the data set to be identified. For example, the screening indicator is whether the browsing time of product A is greater than t seconds. If there is behavior data with a browsing time of product A greater than t seconds in the data set to be identified, it is used as key behavior data. When using the interpretive neural network model to identify the key behavior data set, the text included in the interpretation result of the neural network about the behavior data can be used to identify the key behavior data in the data set to be identified. For example, the interpretation result includes text about whether the product browsing duration is greater than t seconds: product browsing duration. If the data set to be identified includes behavior data with a product browsing duration greater than t seconds, then this behavior data can be used as key behavior data.
[0072] Figure 2 is a schematic diagram of the real-time acquisition process of behavior data according to an embodiment of the present invention. As Figure 2As shown, the server and the client respectively report the behavior data of business personnel (reported by the server) and the behavior data of customers (by the client) in real time according to the standard template of the embodiments of the present invention. The standard template is the unified data collection standard, that is, the event model. The receiving end (specifically, it can be the data collection module of the embodiments of the present invention, see the following embodiments for details) receives the reported behavior data in real time, adds the receiving time to the received behavior data, and sends the behavior data to kafka (a high-throughput distributed publish-subscribe messaging system) in real time. After kafka receives the behavior data, it will retain the data of the most recent 7 days.
[0073] The ETL module consumes the behavior data in kafka in real time and stores the processed behavior data in hive (a data warehouse tool based on Hadoop) or a structured database in the big data cluster. The ETL module mainly uses SparkStreaming or Flink (an open-source stream processing framework) to consume the behavior data in kafka in real time and perform parsing. The parsing mainly includes the structured storage of detailed data into the database, the processing of metrics, and the identification of key behavior data. The detailed data refers to the behavior data in the form of events collected through the unified data collection standard, that is, the dataset to be identified. It needs to be converted into structured detailed data for storage in the database, which can be used for subsequent visualization reports in the reporting system and for offline processing of subsequent requirement metrics. The offline processing of metrics is the non-real-time processing of metrics, that is, first stored in the big data cluster, etc., and then the metrics are processed later. The offline processing of metrics can be the offline processing of metrics for the dataset to be identified to extract the metrics and corresponding metric values for behavior judgment. The processing of metrics can include but is not limited to the extraction and calculation of core metric values such as pv (page view, also called page browsing volume, click volume), uv (unique visitor), conversion rate, click-through rate, etc.
[0074] The big data cluster can store detailed data, metric data, and key behavior data. The metric data is the metrics and corresponding metric values extracted by processing the dataset to be identified. The key behavior data identified based on the offline processed metrics and metric values, that is, the offline key clues, and the offline metrics and offline key clues can be processed and identified based on the big data cluster. The structured database can be mainly used to store real-time and offline metric data and key behavior data, and provide high-concurrency data query services.
[0075] The data service interface provides real-time data query services or message channel push services for the application layer (such as application layer modules such as policy view). Business personnel can query the clue information of customers in time through the message notification function or function module of the application, which helps business personnel understand and master customer dynamics in time and effectively improve the success rate of customer acquisition.
[0076] In one embodiment, in the policy review scenario, the businessperson forwards the policy review report to the customer. After the customer views the policy review report, the customer's viewing information is pushed to the businessperson in real time. The businessperson is, for example, a policy review agent (hereinafter referred to as the agent). The behavior data of the policy review agent and the customer is collected. For the event that the customer views the policy review report, its event ID is: bdjs_bdjsbg_view, and the event attributes are shown in Table 4:
[0077] Table 4
[0078]
[0079]
[0080] When the customer generates the behavior of viewing the policy review report, the behavior data is reported according to the event ID and attributes. The client reports the event that the user views the policy review report, and the receiving end receives the behavior event and adds the receiving date. Due to the situation of delayed reporting, the receiving date and the reporting date may differ greatly. The behavior data of the policy review agent is also reported in the same way. After receiving the behavior data, the receiving end sends the behavior data to kafka. The ETL module loads the rule engine to process the behavior data in real time to identify the key behavior data and store it in the database (big data / structured database). The application layer queries the key behavior data from the data service interface and sends it to the message channel in real time, and then sends it to the businessperson through the message channel. By the rule engine, rules can be added. For example, the behavior that the customer views the policy review report is considered a key clue of the behavior data. Then the ETL module judges whether the event ID is bdjs_bdjsbg_view. If so, it is stored in the database. According to the message notification pushed by the message channel, the businessperson communicates with the customer specifically and forwards the content of interest, thereby improving the pertinence of service and content push.
[0081] Based on the characteristics of the customer's behavior data on different terminals, the embodiment of the present invention adopts a unified data collection standard according to different product lines and business lines, comprehensively collects the behavior data of each touchpoint (such as mini programs, official WeChat accounts, shopping malls, APPs, etc.), and uses a big data processing platform to realize the real-time collection of massive data, with the ability to collect and process massive data in real time. Based on business rules and algorithm models, the key clues of the customer's behavior data are identified, and the business logic of insurance is adopted. For example, based on the behavior data such as the insurance application process and product browsing, a set of key behavior clues is obtained, and the customer's key clues are effectively pushed to the businessperson in real time. Adopting causal inference or interpretive neural network algorithm models helps to optimize the rule engine. Using the message channel mechanism, the customer's key clues are sent to the businessperson in real time, realizing the rapid conversion of customers and improving the agent's customer acquisition ability.
[0082] Figure 3 It is a schematic diagram of the main modules of an identification device for key behavior data according to an embodiment of the present invention. As Figure 3 shown, the identification device 300 for key behavior data according to an embodiment of the present invention mainly includes: a first rule set configuration module 301, a second rule set screening module 302, and a key behavior data identification module 303.
[0083] The first rule set configuration module 301 is configured to receive the configuration of the behavior judgment rules in the rule engine, and obtain a first rule set according to the configured behavior judgment rules;
[0084] The second rule set screening module 302 is configured to screen out a second rule set from the first rule set through a preset algorithm model;
[0085] The key behavior data identification module 303 is configured to use the second rule set as the rule set actually used by the rule engine, input the collected data set to be identified into the rule engine, so as to identify the key behavior data set in the data set to be identified, and the data set to be identified includes the behavior data of business personnel and customers.
[0086] In one embodiment, the preset algorithm model is a machine learning algorithm model, and the machine learning algorithm model is specifically a logistic regression model or a gradient boosting model.
[0087] Specifically, the second rule set screening module 302 may be configured to: construct the input features of the machine learning algorithm model, and the input features of the machine learning algorithm model include the features obtained based on the first rule set; output the feature scores corresponding to each input feature through the machine learning algorithm model, and classify the behavior judgment rules in the first rule set with the feature scores greater than the preset threshold into the second rule set.
[0088] In one embodiment, the preset algorithm model is a causal inference model. Specifically, the second rule set screening module 302 may be configured to determine the screening index by the following method using the causal inference model: obtain the first behavior data sequence of the customers in the first customer set before the current moment, where the first customer set is the set of customers who have not yet generated behavior results; select one or more target customers from the second customer set whose behavior data sequence before the current moment is similar to the first behavior data sequence, and obtain the second behavior data sequence of each target customer from after the current moment to before generating the behavior result, where the second customer set is the set of customers known to have behavior results; determine the screening index according to the behavior data with the number of occurrences greater than the first number threshold in each second behavior data sequence, and / or the behavior data with the total number of occurrences greater than the second number threshold in all second behavior data sequences; use the screening index to screen the behavior judgment rules from the first rule set, and classify the screened behavior judgment rules into the second rule set.
[0089] In one embodiment, the preset algorithm model is an interpretive neural network model. The second rule set screening module 302 can specifically be used to: construct the input features of the interpretive neural network model, where the input features of the interpretive neural network model include the features obtained based on the first rule set; obtain the interpretation result of the neural network through the interpretation module of the interpretive neural network model, and the interpretation result includes the text regarding the behavior data; match the interpretation result with the input features of the interpretive neural network model, and obtain the behavior judgment rules from the matched input features and classify them into the second rule set.
[0090] The key behavior data recognition device 300 may further include a data collection module, which is used to use the event model as the data collection standard for collecting the dataset to be recognized, and respectively collect the behavior data of business personnel and the behavior data of customers. The event model includes a user entity and an event entity, and at least some fields of the user entity and the event entity are different, and the user is a business personnel or a customer.
[0091] The key behavior data recognition module 303 can specifically be used to: perform metric processing on the dataset to be recognized to extract the metrics and corresponding metric values for behavior judgment; determine whether the extracted metrics and corresponding metric values conform to the behavior judgment rules in the actually used rule set through the rule engine. If they conform, the behavior data corresponding to the metric in the dataset to be recognized is used as the key behavior data and added to the key behavior dataset.
[0092] The key behavior data recognition device 300 may further include a data storage and provision module, which is used to: store the key behavior dataset in a big data cluster or a structured database, and provide real-time data query services through a data service interface for the business personnel's client to query the behavior data in the key behavior dataset; or provide a message channel push service through the data service interface to push the behavior data in the key behavior dataset to the business personnel's client through the message channel.
[0093] In addition, the specific implementation content of the key behavior data recognition device in the embodiments of the present invention has been described in detail in the above key behavior data recognition method, so the repeated content will not be described here.
[0094] Figure 4 An exemplary system architecture 400 is shown to which the key behavior data recognition method or the key behavior data recognition device of the embodiments of the present invention can be applied.
[0095] Such as Figure 4As shown, the system architecture 400 may include terminal devices 401, 402, 403, a network 404, and a server 405. The network 404 is used to provide a medium for communication links between the terminal devices 401, 402, 403 and the server 405. The network 404 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0096] Users can use the terminal devices 401, 402, 403 to interact with the server 405 via the network 404 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 401, 402, 403, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (for example only).
[0097] The terminal devices 401, 402, 403 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0098] The server 405 may be a server that provides various services, such as a background management server that supports shopping websites browsed by users using the terminal devices 401, 402, 403 (for example only). The background management server may analyze and process data such as product information query requests received, and feedback the processing results (such as target push information, product information - for example only) to the terminal devices.
[0099] It should be noted that the method for identifying key behavior data provided in the embodiments of the present invention is generally executed by the server 405. Correspondingly, the device for identifying key behavior data is generally set in the server 405.
[0100] It should be understood that Figure 4 the numbers of terminal devices, networks, and servers in
[0101] are merely illustrative. According to the implementation requirements, there may be any number of terminal devices, networks, and servers. Figure 5 is a schematic structural diagram of a computer system 500 of a server suitable for use in implementing the embodiments of the present application. Figure 5 The server shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0102] As Figure 5As shown, computer system 500 includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 502 or a program loaded from a storage section 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the system 500 are also stored. The CPU 501, ROM 502, and RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0103] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. A drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed so that a computer program read from it can be installed into the storage section 508 as needed.
[0104] Specifically, according to the embodiments disclosed in the present invention, the process described above with reference to the main step schematic diagram can be implemented as a computer software program. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the method shown in the main step schematic diagram. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, the above functions defined in the system of the present application are executed.
[0105] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0106] The schematic diagrams of the main steps and block diagrams in the drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the schematic diagram of the main steps or block diagram can represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the drawings. For example, two consecutively shown blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or schematic diagram of the main steps, and the combination of blocks in the block diagram or schematic diagram of the main steps, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0107] The modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes a first rule set configuration module, a second rule set screening module, and a key behavior data recognition module. Among them, the names of these modules do not constitute a limitation to the module itself in some cases. For example, the first rule set configuration module can also be described as "a module for receiving the configuration of the behavior judgment rules in the rule engine and obtaining the first rule set according to the configured behavior judgment rules".
[0108] As another aspect, the present invention also provides a computer-readable medium, which can be included in the device described in the above embodiments; or can exist alone without being assembled into the device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by the device, the device includes: receiving the configuration of the behavior judgment rules in the rule engine and obtaining the first rule set according to the configured behavior judgment rules; screening out the second rule set from the first rule set through a preset algorithm model; using the second rule set as the rule set actually used by the rule engine, and inputting the collected data set to be recognized into the rule engine to recognize the key behavior data set in the data set to be recognized, where the data set to be recognized includes the behavior data of business personnel and customers.
[0109] According to the technical solution of the embodiments of the present invention, the first rule set is obtained according to the configured behavior judgment rules in the rule engine, the second rule set is screened out from the first rule set through a preset algorithm model, the second rule set is used as the rule set actually used by the rule engine, and the collected data set to be recognized is input into the rule engine to recognize the key behavior data set in the data set to be recognized. It can automatically load the rule engine to recognize key behavior data, with high availability and improved lead recognition efficiency. By combining the algorithm model to optimize the rule engine, the accuracy of lead recognition is improved. Behavior data can be collected according to a unified behavior data collection standard for different product lines and business lines for storage, improving the database storage capacity.
[0110] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for identifying key behavior data, characterized in that Including: Taking the event model as the data collection standard for collecting the dataset to be recognized, collecting the behavior data of business personnel and the behavior data of customers respectively. The event model includes a user entity and an event entity. At least some fields of the user entity and the event entity are different. The user is a business personnel or a customer. Each event in the event entity table includes dynamic common attributes; Receiving the configuration of the behavior judgment rules in the rule engine, and obtaining the first rule set according to the configured behavior judgment rules; Filtering out the second rule set from the first rule set through a preset algorithm model; The preset algorithm model is one or more of a machine learning algorithm model, a causal inference model, and an interpretive neural network model; When the preset algorithm model adopts multiple models, taking the union of the rule sets obtained by each model as the second rule set; Taking the second rule set as the rule set actually used by the rule engine, and inputting the collected dataset to be recognized into the rule engine to identify the key behavior dataset in the dataset to be recognized. The dataset to be recognized includes the behavior data of business personnel and customers.
2. The method according to claim 1, characterized in that The preset algorithm model is a machine learning algorithm model, and the machine learning algorithm model is a logistic regression model or a gradient boosting model; The filtering out the second rule set from the first rule set through the preset algorithm model includes: Constructing the input features of the machine learning algorithm model. The input features of the machine learning algorithm model include the features obtained based on the first rule set; Outputting the feature scores corresponding to each input feature through the machine learning algorithm model, and attributing the behavior judgment rules in the first rule set corresponding to the feature scores greater than the preset threshold to the second rule set.
3. The method according to claim 1, wherein The preset algorithm model is a causal inference model; The filtering out the second rule set from the first rule set through the preset algorithm model includes: Using the causal inference model to determine the screening index in the following way: obtaining the first behavior data sequence of the customers in the first customer set before the current moment. The first customer set is a set composed of customers who have not yet produced behavior results; selecting one or more target customers from the second customer set whose behavior data sequence before the current moment is similar to the first behavior data sequence, and obtaining the second behavior data sequence of each target customer from after the current moment to before producing the behavior result. The second customer set is a set composed of customers who are known to have the behavior result; determining the screening index according to the behavior data with the occurrence times greater than the first number threshold in each second behavior data sequence, and / or the behavior data with the total occurrence times greater than the second number threshold in all the second behavior data sequences; Using the screening index to screen the behavior judgment rules from the first rule set, and attributing the screened behavior judgment rules to the second rule set.
4. The method according to claim 1, wherein The preset algorithm model is an interpretive neural network model; The filtering out the second rule set from the first rule set through the preset algorithm model includes: Construct the input features of the interpretive neural network model, where the input features of the interpretive neural network model include features obtained based on the first rule set; Obtain the interpretation result of the interpretive neural network model through the interpretation module of the interpretive neural network model, where the interpretation result includes text regarding the behavior data; Match the interpretation result with the input features of the interpretive neural network model, and obtain the behavior judgment rule from the matched input features and classify it into the second rule set.
5. The method according to claim 1, characterized in that, The step of inputting the collected dataset to be recognized into the rule engine to recognize the key behavior dataset in the dataset to be recognized includes: Perform metric processing on the dataset to be recognized to extract metrics for behavior judgment and corresponding metric values; Determine whether the extracted metrics and corresponding metric values conform to the behavior judgment rules in the actually used rule set through the rule engine. If they conform, use the behavior data corresponding to the metric in the dataset to be recognized as the key behavior data and add it to the key behavior dataset.
6. The method according to claim 1, characterized in that, After recognizing the key behavior dataset in the dataset to be recognized, it includes: Store the key behavior dataset in a big data cluster or a structured database, and provide real-time data query services through a data service interface for the business personnel's client to query the behavior data in the key behavior dataset; or provide a message channel push service through the data service interface to push the behavior data in the key behavior dataset to the business personnel's client through the message channel.
7. An identification device for key behavior data, characterized in that, It includes: A data collection module, which uses an event model as the data collection standard for collecting the dataset to be recognized, and collects the behavior data of business personnel and the behavior data of customers respectively. The event model includes a user entity and an event entity, at least some fields of the user entity and the event entity are different, the user is a business personnel or a customer, and each event in the event entity table includes dynamic common attributes; A first rule set configuration module, which is used to receive the configuration of the behavior judgment rules in the rule engine and obtain the first rule set according to the configured behavior judgment rules; A second rule set screening module, which is used to screen out the second rule set from the first rule set through a preset algorithm model; The preset algorithm model is one or more of a machine learning algorithm model, a causal inference model, and an interpretive neural network model; When the preset algorithm model adopts multiple models, the union of the rule sets obtained by each model is used as the second rule set; A key behavior data recognition module, which uses the second rule set as the actually used rule set of the rule engine, inputs the collected dataset to be recognized into the rule engine, and recognizes the key behavior dataset in the dataset to be recognized. The dataset to be recognized includes the behavior data of business personnel and customers.
8. An electronic device, characterized in that, It includes: One or more processors; A memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.
9. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Data analysis method and device, computer equipment and storage medium
CN110738559A
Business-oriented risk control method and device
CN112651619A