Association Rule-Based Clue Analysis Method and System
The clue feature map network is constructed through the Apriori algorithm and graph network algorithm, and the clue data analysis is performed in combination with the LSTM algorithm, which solves the problem of inaccurate clue feature recognition and prediction in the existing technology, and realizes efficient and intelligent clue analysis and prediction, and supports business decision-making.
Patent Information
- Application Number
- CN202510439349.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-09
AI Technical Summary
It is difficult for the existing technology to effectively identify high-correlation clue characteristics on the business management platform, and the ability to dynamically analyze clue characteristics over time, resulting in insufficient prediction of clue development trends, which affects the scientificity and timeliness of business decisions.
The Apriori algorithm is used to determine the high correlation feature group of clue data, build a clue feature map network, and use the LSTM algorithm to perform development prediction, and combine the graph network algorithm and long-term short-term memory network (LSTM) for clue data analysis.
It improves the efficiency and intelligence of business clue analysis, improves the reliability and accuracy of clue prediction, and provides effective data support for business decisions.
Smart Images

Figure CN119940982B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a clue analysis method and system based on association rules. Background Art
[0002] Some business management platforms provide services for uploading clue information. For example, on police management platforms, users can upload reporting clues, or on business management platforms, users can upload business cooperation clues. The analysis of business clues usually relies on manual experience or matching methods based on fixed rules, and mainly classifies and mines clue data through preset keywords, screening conditions, etc. However, with the complexity of the business environment and the growth of data scale, this method has great limitations in identifying clue features with high correlation, and it is difficult to effectively discover the deep-level correlation between clues. In addition, the existing methods lack the ability to dynamically analyze the changes of clue features over time, resulting in inaccurate prediction results of clue development trends, affecting the scientificity and timeliness of business decisions. It can be seen that the existing technology has defects that need to be solved urgently. Summary of the invention
[0003] The technical problem to be solved by the present invention is to provide a clue analysis method and system based on association rules, which can improve the efficiency and intelligence of business clue analysis, improve the reliability and accuracy of clue prediction, and provide effective data support for business decision-making.
[0004] In order to solve the above technical problems, the first aspect of the present invention discloses a clue analysis method based on association rules, the method comprising:
[0005] Obtain lead data related to the current business uploaded by multiple users;
[0006] Based on the Apriori algorithm, determining multiple high-correlation clue feature groups corresponding to all the clue data;
[0007] Based on the graph network algorithm, construct a clue feature graph network corresponding to all the clue feature groups;
[0008] According to the clue feature graph network and the LSTM algorithm, the development prediction result corresponding to any clue data is determined.
[0009] As an optional implementation, in the first aspect of the present invention, the clue data includes at least one of clue type, clue level, clue content, clue object and clue recommendation processing method.
[0010] As an optional implementation, in the first aspect of the present invention, the determining of a plurality of high-correlation clue feature groups corresponding to all the clue data based on the Apriori algorithm includes:
[0011] For each of the said clue data, based on a preset hash value rule, calculate the feature hash corresponding to each data feature in the clue data;
[0012] Map the feature hash corresponding to each data feature in the clue data to a preset data feature hash table to obtain a mapped hash table;
[0013] Based on the Apriori algorithm and the mapped hash table, determine multiple highly correlated clue feature groups corresponding to all the said clue data.
[0014] As an optional implementation manner, in the first aspect of the present invention, the determining, based on the Apriori algorithm and the mapped hash table, multiple highly correlated clue feature groups corresponding to all the said clue data includes:
[0015] Based on the mapped hash table and a hash comparison algorithm, perform matching counting on the occurrence times of each of the said data features in all the said clue data to obtain a counting result;
[0016] Based on the Apriori algorithm and a preset minimum support threshold and / or minimum confidence threshold, perform feature mining on all the said clue data according to the counting result to obtain multiple highly correlated clue feature groups.
[0017] As an optional implementation manner, in the first aspect of the present invention, the constructing, based on a graph network algorithm, a clue feature graph network corresponding to all the said clue feature groups includes:
[0018] Determine the data features belonging to any of the said clue feature groups as a corresponding graph node;
[0019] For any two of the said graph nodes, determine the clue feature groups in which the data features corresponding to the two graph nodes exist simultaneously to obtain at least one associated feature group;
[0020] Calculate the weighted summation average of the supports corresponding to all the said associated feature groups to obtain the association degree parameter corresponding to the two graph nodes;
[0021] Based on all the said graph nodes and the corresponding association degree parameters, establish a corresponding clue feature graph network.
[0022] As an optional implementation manner, in the first aspect of the present invention, when calculating the weighted summation average of the supports corresponding to all the said associated feature groups, the weighted calculation weight corresponding to the support of each of the said associated feature groups is proportional to the feature importance degree corresponding to the associated feature group; the feature importance degree is calculated through the following steps:
[0023] For each of the other data features in the associated feature group except for the data features corresponding to the two graph nodes respectively, calculate the number of clue feature groups in which the other data feature coexists with the data features corresponding to the two graph nodes respectively, to obtain the feature similarity corresponding to the other data feature;
[0024] Screen out the features with the feature similarity greater than the similarity threshold from all the other data features, to obtain at least one high-similarity feature;
[0025] Calculate the ratio of the number of all the other data features to the total number of features, to obtain the first quantity ratio;
[0026] Calculate the ratio of the number of all the high-similarity features to the number of all the other data features, to obtain the second quantity ratio;
[0027] Calculate the quantity weight inversely proportional to the second quantity ratio;
[0028] Calculate the product of the quantity weight and the first quantity ratio;
[0029] Calculate the reciprocal of the product, to obtain the feature importance degree corresponding to the associated feature group.
[0030] As an optional implementation manner, in the first aspect of the present invention, the determining the development prediction result corresponding to any one of the clue data according to the clue feature graph network and the LSTM algorithm includes:
[0031] Determine the clue data that needs to be predicted for clue prediction as the to-be-predicted clue data;
[0032] Determine the graph node information related to the to-be-predicted clue data according to the clue feature graph network;
[0033] Input the graph node information into the trained LSTM algorithm to obtain the development prediction result corresponding to the to-be-predicted clue data; the LSTM algorithm is trained by a training data set including a plurality of training clue development sequences and corresponding clue graph node information annotations.
[0034] As an optional implementation manner, in the first aspect of the present invention, the determining the graph node information related to the to-be-predicted clue data according to the clue feature graph network includes:
[0035] In the clue feature graph network, extract the graph node information related to the to-be-predicted clue data; the graph node information includes graph nodes related to a plurality of data features, at least one adjacent node directly connected, and the association degree parameter between the graph node and the adjacent node.
[0036] In the second aspect of the embodiments of the present invention, a clue analysis system based on association rules is disclosed, and the system includes:
[0037] An acquisition module, configured to acquire clue data related to the current business uploaded by multiple users;
[0038] A determination module, configured to determine multiple highly correlated clue feature groups corresponding to all the clue data based on the Apriori algorithm;
[0039] A construction module, configured to construct a clue feature graph network corresponding to all the clue feature groups based on the graph network algorithm;
[0040] A prediction module, configured to determine a development prediction result corresponding to any one of the clue data according to the clue feature graph network and the LSTM algorithm.
[0041] As an optional implementation manner, in the second aspect of the present invention, the clue data includes at least one of a clue type, a clue level, a clue content, a clue object, and a clue recommendation processing method.
[0042] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the determination module determines multiple highly correlated clue feature groups corresponding to all the clue data based on the Apriori algorithm includes:
[0043] For each piece of the clue data, based on a preset hash value rule, calculate a feature hash corresponding to each data feature in the clue data;
[0044] Map the feature hash corresponding to each data feature in the clue data to a preset data feature hash table to obtain a mapped hash table;
[0045] Based on the Apriori algorithm and the mapped hash table, determine multiple highly correlated clue feature groups corresponding to all the clue data.
[0046] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the determination module determines multiple highly correlated clue feature groups corresponding to all the clue data based on the Apriori algorithm and the mapped hash table includes:
[0047] Based on the mapped hash table and a hash comparison algorithm, perform matching counting on the number of occurrences of each data feature in all the clue data to obtain a counting result;
[0048] Based on the Apriori algorithm and a preset minimum support threshold and / or minimum confidence threshold, perform feature mining on all the clue data according to the counting result to obtain multiple highly correlated clue feature groups.
[0049] As an optional implementation manner, in the second aspect of the present invention, the specific manner in which the construction module constructs a clue feature graph network corresponding to all the clue feature groups based on the graph network algorithm includes:
[0050] Determine the data features belonging to any one of the clue feature groups as a corresponding graph node;
[0051] For any two of the graph nodes, determine the clue feature groups in which the data features corresponding to the two graph nodes coexist, to obtain at least one associated feature group;
[0052] Calculate the weighted summation average of the supports corresponding to all the associated feature groups to obtain the association degree parameter corresponding to the two graph nodes;
[0053] Based on all the graph nodes and the corresponding association degree parameters, establish a corresponding clue feature graph network.
[0054] As an optional implementation manner, in the second aspect of the present invention, when calculating the weighted summation average of the supports corresponding to all the associated feature groups, the weighted calculation weight corresponding to the support of each associated feature group is proportional to the feature importance degree corresponding to the associated feature group; the feature importance degree is calculated through the following steps:
[0055] For each other data feature in the associated feature group except for the data features corresponding to the two graph nodes respectively, calculate the number of clue feature groups in which the other data feature coexists with the data features corresponding to the two graph nodes respectively, to obtain the feature similarity corresponding to the other data feature;
[0056] Screen out the features with the feature similarity greater than the similarity threshold from all the other data features to obtain at least one high-similarity feature;
[0057] Calculate the ratio of the number of all the other data features to the total number of features to obtain a first quantity ratio;
[0058] Calculate the ratio of the number of all the high-similarity features to the number of all the other data features to obtain a second quantity ratio;
[0059] Calculate a quantity weight inversely proportional to the second quantity ratio;
[0060] Calculate the product of the quantity weight and the first quantity ratio;
[0061] Calculate the reciprocal of the product to obtain the feature importance degree corresponding to the associated feature group.
[0062] As an alternative embodiment, in the second aspect of the present invention, the specific manner in which the prediction module determines the development prediction result corresponding to any one of the clue data according to the clue feature map network and the LSTM algorithm includes:
[0063] Determine the clue data for which clue prediction is required as the clue data to be predicted;
[0064] Determine the graph node information related to the clue data to be predicted according to the clue feature map network;
[0065] Input the graph node information into the trained LSTM algorithm to obtain the development prediction result corresponding to the clue data to be predicted; the LSTM algorithm is trained by a training data set including a plurality of training clue development sequences and corresponding clue graph node information annotations.
[0066] As an alternative embodiment, in the second aspect of the present invention, the specific manner in which the prediction module determines the graph node information related to the clue data to be predicted according to the clue feature map network includes:
[0067] In the clue feature map network, extract the graph node information related to the clue data to be predicted; the graph node information includes a plurality of graph nodes related to data features, at least one directly connected adjacent node, and the association degree parameter between the graph node and the adjacent node.
[0068] The third aspect of the present invention discloses another clue analysis system based on association rules, and the system includes:
[0069] A memory storing executable program code;
[0070] A processor coupled to the memory;
[0071] The processor calls the executable program code stored in the memory and executes some or all of the steps in the clue analysis method based on association rules disclosed in the first aspect of the present invention.
[0072] The fourth aspect of the present invention discloses a computer storage medium, and the computer storage medium stores computer instructions, which are used to execute some or all of the steps in the clue analysis method based on association rules disclosed in the first aspect of the present invention when called.
[0073] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0074] Based on the lead data related to the current business uploaded by multiple users, the present invention uses the Apriori algorithm to determine multiple highly correlated lead feature groups corresponding to all the lead data, constructs a lead feature graph network corresponding to the lead feature groups based on the graph network algorithm, and determines the development prediction result corresponding to any lead data according to the lead feature graph network and the LSTM algorithm, so as to improve the efficiency and intelligence of business lead analysis, enhance the reliability and accuracy of lead prediction, and provide effective data support for business decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0076] Figure 1 FIG. is a schematic flowchart of a lead analysis method based on association rules disclosed in an embodiment of the present invention.
[0077] Figure 2 FIG. is a schematic structural diagram of a lead analysis system based on association rules disclosed in an embodiment of the present invention.
[0078] Figure 3 FIG. is a schematic structural diagram of another lead analysis system based on association rules disclosed in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0079] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0080] The terms "first", "second", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or equipment.
[0081] References to "embodiments" in this specification mean that the particular features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive of other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0082] The present invention discloses a clue analysis method and system based on association rules. Based on clue data related to the current business uploaded by multiple users, the Apriori algorithm is used to determine multiple highly correlated clue feature groups corresponding to all the clue data, and a clue feature graph network corresponding to the clue feature groups is constructed based on the graph network algorithm. According to the clue feature graph network and the LSTM algorithm, the development prediction result corresponding to any clue data is determined, thereby being able to improve the efficiency and intelligence of business clue analysis, enhance the reliability and accuracy of clue prediction, and provide effective data support for business decision-making. The following will be described in detail respectively.
[0083] Embodiment 1
[0084] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a clue analysis method based on association rules disclosed in an embodiment of the present invention. Among them, Figure 1 the described clue analysis method based on association rules can be applied to a data processing system / data processing device / data processing server (wherein, the server includes a local processing server or a cloud processing server). As Figure 1 shown, the clue analysis method based on association rules may include the following operations:
[0085] 101. Obtain clue data related to the current business uploaded by multiple users.
[0086] 102. Based on the Apriori algorithm, determine multiple highly correlated clue feature groups corresponding to all the clue data.
[0087] 103. Based on the graph network algorithm, construct a clue feature graph network corresponding to all the clue feature groups.
[0088] 104. According to the clue feature graph network and the LSTM algorithm, determine the development prediction result corresponding to any clue data.
[0089] It can be seen that based on the clue data related to the current business uploaded by multiple users, the above-mentioned invention embodiment uses the Apriori algorithm to determine multiple highly correlated clue feature groups corresponding to all the clue data, constructs a clue feature graph network corresponding to the clue feature groups based on the graph network algorithm, and determines the development prediction result corresponding to any clue data according to the clue feature graph network and the LSTM algorithm, so as to improve the efficiency and intelligence of business clue analysis, improve the reliability and accuracy of clue prediction, and provide effective data support for business decision-making.
[0090] As an optional embodiment, in the above steps, the clue data includes at least one of clue type, clue level, clue content, clue object, and clue recommendation processing method.
[0091] It can be seen that through the above optional embodiment, the content of the clue data is defined to comprehensively represent the relevant features of the clue, so as to facilitate subsequent accurate feature mining and clue prediction, and assist in improving the efficiency and intelligence of business clue analysis, improving the reliability and accuracy of clue prediction, and providing effective data support for business decision-making.
[0092] As an optional embodiment, in the above steps, based on the Apriori algorithm, determining multiple highly correlated clue feature groups corresponding to all the clue data includes:
[0093] For each clue data, calculate the feature hash corresponding to each data feature in the clue data based on a preset hash value rule;
[0094] Map the feature hash corresponding to each data feature in the clue data to a preset data feature hash table to obtain a mapped hash table;
[0095] Based on the Apriori algorithm and the mapped hash table, determine multiple highly correlated clue feature groups corresponding to all the clue data.
[0096] It can be seen that through the above optional embodiment, for each clue data, calculate the feature hash corresponding to each data feature therein based on a preset hash value rule, map all the feature hashes to a preset data feature hash table to obtain a mapped hash table, and then based on the Apriori algorithm and the mapped hash table, determine multiple highly correlated clue feature groups corresponding to all the clue data, so as to improve the efficiency and accuracy of clue feature extraction, assist in improving the efficiency and intelligence of business clue analysis, improve the reliability and accuracy of clue prediction, and provide effective data support for business decision-making.
[0097] As an alternative embodiment, in the above steps, based on the Apriori algorithm and the mapping hash table, determining multiple highly correlated clue feature groups corresponding to all clue data, including:
[0098] Based on the mapping hash table and the hash comparison algorithm, perform matching counting on the occurrence times of each data feature in all clue data to obtain a counting result;
[0099] Based on the Apriori algorithm and a preset minimum support threshold and / or minimum confidence threshold, perform feature mining on all clue data according to the counting result to obtain multiple highly correlated clue feature groups.
[0100] It can be seen that through the above alternative embodiment, based on the mapping hash table and the hash comparison algorithm, perform matching counting on the occurrence times of each data feature in all clue data to obtain a counting result, and use the Apriori algorithm in combination with a preset minimum support threshold and / or minimum confidence threshold to perform feature mining on all clue data according to the counting result to obtain multiple highly correlated clue feature groups, thereby improving the efficiency and accuracy of clue feature mining, assisting in improving the efficiency and intelligence of business clue analysis, enhancing the reliability and accuracy of clue prediction, and providing effective data support for business decision-making.
[0101] As an alternative embodiment, in the above steps, based on the graph network algorithm, construct a clue feature graph network corresponding to all clue feature groups, including:
[0102] Determine the data features belonging to any clue feature group as a corresponding graph node;
[0103] For any two graph nodes, determine the clue feature groups in which the data features corresponding to the two graph nodes exist simultaneously to obtain at least one associated feature group;
[0104] Calculate the weighted summation average of the support degrees corresponding to all associated feature groups to obtain the association degree parameter corresponding to the two graph nodes;
[0105] Based on all graph nodes and the corresponding association degree parameters, establish a corresponding clue feature graph network.
[0106] It can be seen that through the above optional embodiments, the data features belonging to any clue feature group are determined as corresponding graph nodes. For any two graph nodes, the clue feature groups in which the data features corresponding to the two graph nodes coexist are determined, and at least one associated feature group is obtained. The weighted sum average of the support degrees corresponding to all the associated feature groups is calculated to obtain the association degree parameter corresponding to the two graph nodes. Based on all the graph nodes and the corresponding association degree parameters, a clue feature graph network is established, thereby realizing the visual expression of the association relationship between the clue features, improving the structured analysis ability of the clue data, and providing more accurate data support for subsequent feature mining and development trend prediction.
[0107] As an optional embodiment, in the above steps, when calculating the weighted sum average of the support degrees corresponding to all the associated feature groups, the weighted calculation weight corresponding to the support degree of each associated feature group is proportional to the feature importance degree corresponding to the associated feature group; the feature importance degree is calculated through the following steps:
[0108] For each other data feature in the associated feature group except the data features corresponding to the two graph nodes respectively, calculate the number of clue feature groups in which the other data feature coexists with the data features corresponding to the two graph nodes respectively, and obtain the feature similarity degree corresponding to the other data feature;
[0109] Screen out the features with feature similarity degrees greater than the similarity threshold from all the other data features to obtain at least one high-similarity feature;
[0110] Calculate the ratio of the number of all the other data features to the total number of features to obtain the first quantity ratio;
[0111] Calculate the ratio of the number of all the high-similarity features to the number of all the other data features to obtain the second quantity ratio;
[0112] Calculate the quantity weight inversely proportional to the second quantity ratio;
[0113] Calculate the product of the quantity weight and the first quantity ratio;
[0114] Calculate the reciprocal of the product to obtain the feature importance degree corresponding to the associated feature group.
[0115] It can be seen that through the above optional embodiments, the proportion degree of the data features corresponding to the two graph nodes in each associated feature group is considered when calculating the association degree parameter, so that the relative importance of the features can be measured according to the distribution of the features in different clue feature groups, assisting in realizing the visual expression of the association relationship between the clue features and improving the structured analysis ability of the clue data.
[0116] As an alternative embodiment, in the above steps, determining the development prediction result corresponding to any clue data according to the clue feature map network and the LSTM algorithm includes:
[0117] Determine the clue data for which clue prediction is required as the clue data to be predicted;
[0118] Determine the graph node information related to the clue data to be predicted according to the clue feature map network;
[0119] Input the graph node information into the trained LSTM algorithm to obtain the development prediction result corresponding to the clue data to be predicted; the LSTM algorithm is trained by a training data set including multiple training clue development sequences and corresponding clue graph node information annotations.
[0120] It can be seen that through the above alternative embodiment, based on the clue feature map network, the relevant graph node information of the clue data to be predicted is extracted, and the trained LSTM algorithm is used to analyze it to generate the clue development prediction result, so as to combine the temporal change trend of the clue features, improve the accuracy and reliability of the prediction, and enhance the intelligent judgment ability of business clues.
[0121] As an alternative embodiment, in the above steps, determining the graph node information related to the clue data to be predicted according to the clue feature map network includes:
[0122] In the clue feature map network, extract the graph node information related to the clue data to be predicted; the graph node information includes multiple graph nodes related to data features, at least one directly connected adjacent node, and the correlation parameter between the graph node and the adjacent node.
[0123] It can be seen that through the above alternative embodiment, based on the clue feature map network, the graph node information related to the clue data to be predicted is extracted, including the graph nodes corresponding to data features and their directly connected adjacent nodes, and combined with the correlation parameters between nodes, comprehensively reflecting the correlation relationship between clue features, so as to provide more accurate input data for subsequent clue development prediction and improve the accuracy and rationality of the prediction.
[0124] Embodiment Two
[0125] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a clue analysis system based on association rules disclosed in an embodiment of the present invention. Among them, Figure 2 The described clue analysis system based on association rules can be applied to a data processing system / data processing device / data processing server (wherein, the server includes a local processing server or a cloud processing server). As Figure 2 shown, the clue analysis system based on association rules can include:
[0126] An acquisition module 201, configured to acquire lead data related to the current business uploaded by multiple users.
[0127] A determination module 202, configured to determine multiple highly correlated lead feature groups corresponding to all lead data based on the Apriori algorithm.
[0128] A construction module 203, configured to construct a lead feature graph network corresponding to all lead feature groups based on the graph network algorithm.
[0129] A prediction module 204, configured to determine a development prediction result corresponding to any lead data according to the lead feature graph network and the LSTM algorithm.
[0130] It can be seen that based on the lead data related to the current business uploaded by multiple users, the above invention embodiments use the Apriori algorithm to determine multiple highly correlated lead feature groups corresponding to all lead data, construct a lead feature graph network corresponding to the lead feature groups based on the graph network algorithm, and determine a development prediction result corresponding to any lead data according to the lead feature graph network and the LSTM algorithm, so as to improve the efficiency and intelligence of business lead analysis, improve the reliability and accuracy of lead prediction, and provide effective data support for business decision-making.
[0131] As an optional embodiment, the lead data includes at least one of a lead type, a lead level, a lead content, a lead object, and a lead recommendation processing method.
[0132] It can be seen that through the above optional embodiment, the content of the lead data is defined to comprehensively represent the relevant features of the lead, so as to facilitate subsequent accurate feature mining and lead prediction, and assist in improving the efficiency and intelligence of business lead analysis, improving the reliability and accuracy of lead prediction, and providing effective data support for business decision-making.
[0133] As an optional embodiment, the specific manner in which the determination module determines multiple highly correlated lead feature groups corresponding to all lead data based on the Apriori algorithm includes:
[0134] For each lead data, calculate a feature hash corresponding to each data feature in the lead data based on a preset hash value rule;
[0135] Map the feature hash corresponding to each data feature in the lead data to a preset data feature hash table to obtain a mapped hash table;
[0136] Based on the Apriori algorithm and the mapped hash table, determine multiple highly correlated lead feature groups corresponding to all lead data.
[0137] It can be seen that through the above optional embodiments, based on each piece of clue data, the characteristic hash corresponding to each data characteristic is calculated using a preset hash value rule, and all characteristic hashes are mapped to a preset data characteristic hash table to obtain a mapped hash table. Subsequently, based on the Apriori algorithm and the mapped hash table, multiple highly correlated clue feature groups corresponding to all clue data are determined, thereby improving the efficiency and accuracy of clue feature extraction, assisting in enhancing the efficiency and intelligence of business clue analysis, improving the reliability and accuracy of clue prediction, and providing effective data support for business decision-making.
[0138] As an optional embodiment, the specific manner in which the determination module determines multiple highly correlated clue feature groups corresponding to all clue data based on the Apriori algorithm and the mapped hash table includes:
[0139] Based on the mapped hash table and the hash comparison algorithm, the occurrence times of each data characteristic in all clue data are matched and counted to obtain a counting result;
[0140] Based on the Apriori algorithm and a preset minimum support threshold and / or minimum confidence threshold, feature mining is performed on all clue data according to the counting result to obtain multiple highly correlated clue feature groups.
[0141] It can be seen that through the above optional embodiments, based on the mapped hash table and the hash comparison algorithm, the occurrence times of each data characteristic in all clue data are matched and counted to obtain a counting result. Using the Apriori algorithm in combination with a preset minimum support threshold and / or minimum confidence threshold, feature mining is performed on all clue data according to the counting result to obtain multiple highly correlated clue feature groups, thereby improving the efficiency and accuracy of clue feature mining, assisting in enhancing the efficiency and intelligence of business clue analysis, improving the reliability and accuracy of clue prediction, and providing effective data support for business decision-making.
[0142] As an optional embodiment, the specific manner in which the construction module constructs a clue feature graph network corresponding to all clue feature groups based on the graph network algorithm includes:
[0143] Determine the data characteristics belonging to any clue feature group as a corresponding graph node;
[0144] For any two graph nodes, determine the clue feature groups in which the data characteristics corresponding to the two graph nodes exist simultaneously to obtain at least one associated feature group;
[0145] Calculate the weighted summation average of the supports corresponding to all associated feature groups to obtain the association degree parameter corresponding to the two graph nodes;
[0146] Based on all graph nodes and corresponding correlation degree parameters, a corresponding clue feature graph network is established.
[0147] It can be seen that through the above optional embodiments, the data features belonging to any clue feature group are determined as corresponding graph nodes. For any two graph nodes, the clue feature groups in which the data features corresponding to the two graph nodes exist simultaneously are determined to obtain at least one associated feature group. The weighted sum average of the support degrees corresponding to all associated feature groups is calculated to obtain the correlation degree parameter corresponding to the two graph nodes. Based on all graph nodes and corresponding correlation degree parameters, a clue feature graph network is established, thereby realizing the visual expression of the association relationship between clue features, improving the structured analysis ability of clue data, and providing more accurate data support for subsequent feature mining and development trend prediction.
[0148] As an optional embodiment, when calculating the weighted sum average of the support degrees corresponding to all associated feature groups, the weighted calculation weight corresponding to the support degree of each associated feature group is proportional to the feature importance degree corresponding to the associated feature group; the feature importance degree is calculated through the following steps:
[0149] For each other data feature in the associated feature group except the data features corresponding to the two graph nodes respectively, calculate the number of clue feature groups in which the other data feature coexists with the data features corresponding to the two graph nodes respectively to obtain the feature similarity corresponding to the other data feature;
[0150] Screen out the features with feature similarity greater than the similarity threshold from all other data features to obtain at least one high-similarity feature;
[0151] Calculate the ratio of the number of all other data features to the total number of features to obtain the first quantity ratio;
[0152] Calculate the ratio of the number of all high-similarity features to the number of all other data features to obtain the second quantity ratio;
[0153] Calculate the quantity weight inversely proportional to the second quantity ratio;
[0154] Calculate the product of the quantity weight and the first quantity ratio;
[0155] Calculate the reciprocal of the product to obtain the feature importance degree corresponding to the associated feature group.
[0156] It can be seen that through the above optional embodiments, the proportion degree of the data features corresponding to the two graph nodes in each associated feature group is considered when calculating the correlation degree parameter, so that the relative importance of the features can be measured according to the distribution of the features in different clue feature groups, assisting in realizing the visual expression of the association relationship between clue features and improving the structured analysis ability of clue data.
[0157] As an alternative embodiment, the specific manner in which the prediction module determines the development prediction result corresponding to any clue data according to the clue feature map network and the LSTM algorithm includes:
[0158] Determine the clue data for which clue prediction needs to be performed as the clue data to be predicted;
[0159] Determine the graph node information related to the clue data to be predicted according to the clue feature map network;
[0160] Input the graph node information into the trained LSTM algorithm to obtain the development prediction result corresponding to the clue data to be predicted; the LSTM algorithm is trained through a training data set including multiple training clue development sequences and corresponding clue graph node information annotations.
[0161] It can be seen that through the above alternative embodiment, based on the clue feature map network, the relevant graph node information of the clue data to be predicted is extracted, and the trained LSTM algorithm is used to analyze it to generate the clue development prediction result, so as to combine the temporal change trend of the clue features, improve the accuracy and reliability of the prediction, and enhance the intelligent judgment ability of business clues.
[0162] As an alternative embodiment, the specific manner in which the prediction module determines the graph node information related to the clue data to be predicted according to the clue feature map network includes:
[0163] In the clue feature map network, extract the graph node information related to the clue data to be predicted; the graph node information includes multiple graph nodes related to data features, at least one directly connected adjacent node, and the correlation degree parameter between the graph node and the adjacent node.
[0164] It can be seen that through the above alternative embodiment, based on the clue feature map network, the graph node information related to the clue data to be predicted is extracted, including the graph nodes corresponding to the data features and their directly connected adjacent nodes, and combined with the correlation degree parameter between the nodes, comprehensively reflecting the correlation relationship between the clue features, so as to provide more accurate input data for the subsequent clue development prediction and improve the accuracy and rationality of the prediction.
[0165] Embodiment III
[0166] Please refer to Figure 3 , Figure 3 which is another clue analysis system based on association rules disclosed in the embodiments of the present invention. Figure 3 The described clue analysis system based on association rules is applied to a data processing system / data processing device / data processing server (wherein, the server includes a local processing server or a cloud processing server). As Figure 3As shown, the clue analysis system based on association rules may include:
[0167] A memory 301 storing executable program codes;
[0168] A processor 302 coupled to the memory 301;
[0169] Wherein, the processor 302 invokes the executable program codes stored in the memory 301 to execute the steps of the clue analysis method based on association rules described in Embodiment 1.
[0170] Embodiment 4
[0171] An embodiment of the present invention discloses a computer-readable storage medium storing a computer program for electronic data exchange, wherein the computer program causes a computer to execute the steps of the clue analysis method based on association rules described in Embodiment 1.
[0172] Embodiment 5
[0173] An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the steps of the clue analysis method based on association rules described in Embodiment 1.
[0174] The above describes specific embodiments of this specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily have to be performed in the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0175] The systems, devices, modules, or units illustrated in the above embodiments may be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.
[0176] For convenience of description, when describing the above devices, they are described as various units according to functions. Of course, when implementing this specification, the functions of each unit may be implemented in one or more software and / or hardware.
[0177] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, the embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0178] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of this specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.
[0179] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.
[0180] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.
[0181] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0182] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0183] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0184] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0185] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0186] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0187] Finally, it should be noted that: The method and system for clue analysis based on association rules disclosed in the embodiments of the present invention only disclose the preferred embodiments of the present invention. It is only used to illustrate the technical solutions of the present invention, rather than limiting them; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A clue analysis method based on association rules, characterized in that, The method includes: Obtaining clue data related to the current business uploaded by multiple users; Based on the Apriori algorithm, determining multiple highly correlated clue feature groups corresponding to all the clue data; Based on the graph network algorithm, constructing a clue feature graph network corresponding to all the clue feature groups, including: Determining the data feature belonging to any one of the clue feature groups as a corresponding graph node; For any two of the graph nodes, determining the clue feature groups in which the data features corresponding to the two graph nodes exist simultaneously, to obtain at least one associated feature group; Calculating the weighted summation average of the support degrees corresponding to all the associated feature groups to obtain the association degree parameter corresponding to the two graph nodes; when calculating the weighted summation average of the support degrees corresponding to all the associated feature groups, the weighted calculation weight corresponding to the support degree of each associated feature group is proportional to the feature importance degree corresponding to the associated feature group; the feature importance degree is calculated through the following steps: For each other data feature in the associated feature group except the data features corresponding to the two graph nodes respectively, calculating the number of clue feature groups in which the other data feature coexists with the data features corresponding to the two graph nodes respectively, to obtain the feature similarity corresponding to the other data feature; Screening out the features with the feature similarity greater than the similarity threshold from all the other data features to obtain at least one highly similar feature; Calculating the ratio of the number of all the other data features to the total number of features to obtain a first quantity ratio; Calculating the ratio of the number of all the highly similar features to the number of all the other data features to obtain a second quantity ratio; Calculating a quantity weight inversely proportional to the second quantity ratio; Calculating the product of the quantity weight and the first quantity ratio; Calculating the reciprocal of the product to obtain the feature importance degree corresponding to the associated feature group; Based on all the graph nodes and the corresponding association degree parameters, establishing a corresponding clue feature graph network; According to the clue feature graph network and the LSTM algorithm, determining the development prediction result corresponding to any one of the clue data.
2. The clue analysis method based on association rules according to claim 1, wherein The clue data includes at least one of clue type, clue level, clue content, clue object, and clue recommendation processing method.
3. The clue analysis method based on association rules according to claim 1, characterized in that The determining, based on the Apriori algorithm, of multiple highly correlated clue feature groups corresponding to all the clue data includes: For each of the clue data, calculating the feature hash corresponding to each data feature in the clue data based on a preset hash value rule; Mapping the feature hash corresponding to each data feature in the clue data to a preset data feature hash table to obtain a mapped hash table; Based on the Apriori algorithm and the mapped hash table, determining multiple highly correlated clue feature groups corresponding to all the clue data.
4. The clue analysis method based on association rules according to claim 3, wherein The determining, based on the Apriori algorithm and the mapped hash table, of multiple highly correlated clue feature groups corresponding to all the clue data includes: Based on the mapping hash table and the hash comparison algorithm, match and count the occurrence times of each data feature in all the clue data to obtain a counting result; Based on the Apriori algorithm and a preset minimum support threshold and / or minimum confidence threshold, perform feature mining on all the clue data according to the counting result to obtain multiple clue feature groups with high correlation.
5. The method for clue analysis based on association rules according to claim 1, characterized in that, The determining of the development prediction result corresponding to any one of the clue data according to the clue feature graph network and the LSTM algorithm includes: Determine the clue data for which clue prediction is to be performed as the clue data to be predicted; Determine the graph node information related to the clue data to be predicted according to the clue feature graph network; Input the graph node information into the trained LSTM algorithm to obtain the development prediction result corresponding to the clue data to be predicted; the LSTM algorithm is trained by a training data set including multiple training clue development sequences and corresponding clue graph node information annotations.
6. The clue analysis method based on association rules according to claim 5, wherein The determining of the graph node information related to the clue data to be predicted according to the clue feature graph network includes: In the clue feature graph network, extract the graph node information related to the clue data to be predicted; the graph node information includes multiple graph nodes related to data features, at least one adjacent node directly connected thereto, and the correlation degree parameter between the graph node and the adjacent node.
7. A clue analysis system based on association rules, characterized in that, The system includes: An acquisition module, configured to acquire multiple pieces of clue data related to the current service uploaded by users; A determination module, configured to determine multiple clue feature groups with high correlation corresponding to all the clue data based on the Apriori algorithm; A construction module, configured to construct a clue feature graph network corresponding to all the clue feature groups based on the graph network algorithm, including: Determine the data features belonging to any one of the clue feature groups as a corresponding graph node; For any two of the graph nodes, determine the clue feature groups in which the data features corresponding to the two graph nodes exist simultaneously to obtain at least one associated feature group; Calculate the weighted summation average of the supports corresponding to all the associated feature groups to obtain the correlation degree parameter corresponding to the two graph nodes; when calculating the weighted summation average of the supports corresponding to all the associated feature groups, the weighted calculation weight corresponding to the support of each associated feature group is proportional to the feature importance degree corresponding to the associated feature group; the feature importance degree is calculated through the following steps: For each other data feature in the associated feature group except the data features corresponding to the two graph nodes respectively, calculate the number of clue feature groups in which the other data feature coexists with the data features corresponding to the two graph nodes respectively to obtain the feature similarity corresponding to the other data feature; Screen out the features with the feature similarity greater than the similarity threshold from all the other data features to obtain at least one high-similarity feature; Calculate the ratio of the number of all the other data features to the total number of features to obtain a first quantity ratio; Calculate the ratio of the number of all the highly similar features to the number of all the other data features to obtain a second quantity ratio; Calculate a quantity weight that is inversely proportional to the second quantity ratio; Calculate the product of the quantity weight and the first quantity ratio; Calculate the reciprocal of the product to obtain the feature importance degree corresponding to the associated feature group; Based on all the graph nodes and the corresponding association degree parameters, establish a corresponding clue feature graph network; A prediction module, configured to determine a development prediction result corresponding to any one of the clue data according to the clue feature graph network and the LSTM algorithm.
8. A clue analysis system based on association rules, characterized in that, The system includes: A memory storing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory and executes the clue analysis method based on association rules according to any one of claims 1-6.
Citation Information
Patent Citations
Pediatric patient electronic health recording system
CN117854665A
PVB production process energy efficiency optimization system based on Apriori algorithm and edge cloud collaboration
CN119338330A