Association rule-based clue analysis method and system
By applying the Apriori algorithm and graph network algorithm to determine the clue feature group with high correlation and combining the LSTM algorithm to predict clues, the problems of inaccurate clue analysis and inaccurate prediction in the existing technology are solved, and more efficient and accurate clue analysis and prediction are achieved.
Patent Information
- Application Number
- CN202510439349.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-09
AI Technical Summary
The existing technology is difficult to identify high-correlation clue characteristics in business clue analysis, and lacks the ability to dynamically analyze clue characteristics over time, which leads to inaccurate clue prediction and affects the scientificity and timeliness of business decisions.
A clue analysis method based on association rules is adopted, and a clue feature group with high correlation is determined through the Apriori algorithm, and a clue feature map network is constructed using the graph network algorithm, and a clue development prediction is carried out in combination with the LSTM algorithm.
It improves the efficiency and intelligence of business clue analysis, enhances the reliability and accuracy of clue prediction, and provides effective data support for business decisions.
Smart Images

Figure CN119940982A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular to a clue analysis method and system based on association rules. Background Art
[0002] Some business management platforms provide services for uploading clue information. For example, on police management platforms, users can upload reporting clues, or on business management platforms, users can upload business cooperation clues. The analysis of business clues usually relies on manual experience or matching methods based on fixed rules, and mainly classifies and mines clue data through preset keywords, screening conditions, etc. However, with the complexity of the business environment and the growth of data scale, this method has great limitations in identifying clue features with high correlation, and it is difficult to effectively discover the deep-level correlation between clues. In addition, the existing methods lack the ability to dynamically analyze the changes of clue features over time, resulting in inaccurate prediction results of clue development trends, affecting the scientificity and timeliness of business decisions. It can be seen that the existing technology has defects that need to be solved urgently. Summary of the invention
[0003] The technical problem to be solved by the present invention is to provide a clue analysis method and system based on association rules, which can improve the efficiency and intelligence of business clue analysis, improve the reliability and accuracy of clue prediction, and provide effective data support for business decision-making.
[0004] In order to solve the above technical problems, the first aspect of the present invention discloses a clue analysis method based on association rules, the method comprising: Obtain lead data related to the current business uploaded by multiple users; Based on the Apriori algorithm, determining multiple high-correlation clue feature groups corresponding to all the clue data; Based on the graph network algorithm, construct a clue feature graph network corresponding to all the clue feature groups; According to the clue feature graph network and the LSTM algorithm, the development prediction result corresponding to any clue data is determined.
[0005] As an optional implementation, in the first aspect of the present invention, the clue data includes at least one of clue type, clue level, clue content, clue object and clue recommendation processing method.
[0006] As an optional implementation, in the first aspect of the present invention, the determining of a plurality of high-correlation clue feature groups corresponding to all the clue data based on the Apriori algorithm includes: For each clue data, based on a preset hash value rule, a feature hash corresponding to each data feature in the clue data is calculated; Mapping the feature hash corresponding to each data feature in the clue data to a preset data feature hash table to obtain a mapping hash table; Based on the Apriori algorithm and the mapping hash table, multiple high-correlation clue feature groups corresponding to all the clue data are determined.
[0007] As an optional implementation, in the first aspect of the present invention, the determining of a plurality of high-correlation clue feature groups corresponding to all the clue data based on the Apriori algorithm and the mapping hash table includes: Based on the mapping hash table and the hash comparison algorithm, the number of occurrences of each data feature in all the clue data is matched and counted to obtain a counting result; Based on the Apriori algorithm and a preset minimum support threshold and / or a minimum confidence threshold, feature mining is performed on all the clue data according to the calculation results to obtain a plurality of clue feature groups with high correlation.
[0008] As an optional implementation, in the first aspect of the present invention, the graph network algorithm is based on which a clue feature graph network corresponding to all the clue feature groups is constructed, including: Determine the data feature belonging to any of the clue special investigation groups as a corresponding graph node; For any two of the graph nodes, determining the clue feature groups in which the data features respectively corresponding to the two graph nodes exist simultaneously, and obtaining at least one associated feature group; Calculate the weighted average of the support corresponding to all the associated feature groups to obtain the association parameter corresponding to the two graph nodes; Based on all the graph nodes and the corresponding association parameters, a corresponding clue feature graph network is established.
[0009] As an optional implementation, in the first aspect of the present invention, when calculating the weighted sum average of the support corresponding to all the associated feature groups, the weighted calculation weight corresponding to the support corresponding to each of the associated feature groups is proportional to the feature importance corresponding to the associated feature group; the feature importance is calculated by the following steps: For each other data feature in the associated feature group except the data features corresponding to the two graph nodes respectively, calculate the number of clue feature groups in which the other data feature and the data features corresponding to the two graph nodes respectively exist together, and obtain the feature similarity corresponding to the other data feature; Filter out the features whose feature similarity is greater than a similarity threshold from all the other data features to obtain at least one high-similarity feature; Calculate the ratio of the number of all other data features to the total number of features to obtain a first quantity ratio; Calculating the ratio of the number of all the high-similarity features to the number of all the other data features to obtain a second quantity ratio; calculating a quantity weight inversely proportional to the second quantity ratio; Calculating the product of the quantity weight and the first quantity ratio; The reciprocal of the product is calculated to obtain the feature importance corresponding to the associated feature group.
[0010] As an optional implementation, in the first aspect of the present invention, determining the development prediction result corresponding to any clue data according to the clue feature graph network and the LSTM algorithm includes: Determining the clue data that needs to be predicted as clue data to be predicted; Determining graph node information related to the clue data to be predicted according to the clue feature graph network; The graph node information is input into the trained LSTM algorithm to obtain the development prediction result corresponding to the clue data to be predicted; the LSTM algorithm is trained by a training data set including multiple training clue development sequences and corresponding clue graph node information annotations.
[0011] As an optional implementation, in the first aspect of the present invention, determining the graph node information related to the clue data to be predicted according to the clue feature graph network includes: In the clue feature graph network, graph node information related to the clue data to be predicted is extracted; the graph node information includes graph nodes related to multiple data features and at least one directly connected adjacent node and a correlation parameter between the graph node and the adjacent node.
[0012] A second aspect of an embodiment of the present invention discloses a clue analysis system based on association rules, the system comprising: The acquisition module is used to obtain the lead data related to the current business uploaded by multiple users; A determination module, used for determining a plurality of highly correlated clue feature groups corresponding to all the clue data based on an Apriori algorithm; A construction module, used for constructing a clue feature graph network corresponding to all the clue feature groups based on a graph network algorithm; The prediction module is used to determine the development prediction result corresponding to any clue data based on the clue feature graph network and the LSTM algorithm.
[0013] As an optional implementation, in the second aspect of the present invention, the clue data includes at least one of clue type, clue level, clue content, clue object and clue recommendation processing method.
[0014] As an optional implementation, in the second aspect of the present invention, the specific manner in which the determination module determines the multiple high-correlation clue feature groups corresponding to all the clue data based on the Apriori algorithm includes: For each clue data, based on a preset hash value rule, a feature hash corresponding to each data feature in the clue data is calculated; Mapping the feature hash corresponding to each data feature in the clue data to a preset data feature hash table to obtain a mapping hash table; Based on the Apriori algorithm and the mapping hash table, multiple high-correlation clue feature groups corresponding to all the clue data are determined.
[0015] As an optional implementation, in the second aspect of the present invention, the specific manner in which the determination module determines the multiple high-correlation clue feature groups corresponding to all the clue data based on the Apriori algorithm and the mapping hash table includes: Based on the mapping hash table and the hash comparison algorithm, the number of occurrences of each data feature in all the clue data is matched and counted to obtain a counting result; Based on the Apriori algorithm and a preset minimum support threshold and / or a minimum confidence threshold, feature mining is performed on all the clue data according to the calculation results to obtain a plurality of clue feature groups with high correlation.
[0016] As an optional implementation, in the second aspect of the present invention, the specific manner in which the construction module constructs the clue feature graph network corresponding to all the clue feature groups based on the graph network algorithm includes: Determine the data feature belonging to any of the clue special investigation groups as a corresponding graph node; For any two of the graph nodes, determining the clue feature groups in which the data features respectively corresponding to the two graph nodes exist simultaneously, and obtaining at least one associated feature group; Calculate the weighted average of the support corresponding to all the associated feature groups to obtain the association parameter corresponding to the two graph nodes; Based on all the graph nodes and the corresponding association parameters, a corresponding clue feature graph network is established.
[0017] As an optional implementation, in the second aspect of the present invention, when calculating the weighted sum average of the support corresponding to all the associated feature groups, the weighted calculation weight corresponding to the support corresponding to each of the associated feature groups is proportional to the feature importance corresponding to the associated feature group; the feature importance is calculated by the following steps: For each other data feature in the associated feature group except the data features corresponding to the two graph nodes respectively, calculate the number of clue feature groups in which the other data feature and the data features corresponding to the two graph nodes respectively exist together, and obtain the feature similarity corresponding to the other data feature; Filter out the features whose feature similarity is greater than a similarity threshold from all the other data features to obtain at least one high-similarity feature; Calculate the ratio of the number of all other data features to the total number of features to obtain a first quantity ratio; Calculating the ratio of the number of all the high-similarity features to the number of all the other data features to obtain a second quantity ratio; calculating a quantity weight inversely proportional to the second quantity ratio; Calculating the product of the quantity weight and the first quantity ratio; The reciprocal of the product is calculated to obtain the feature importance corresponding to the associated feature group.
[0018] As an optional implementation, in the second aspect of the present invention, the prediction module determines the specific manner of the development prediction result corresponding to any clue data according to the clue feature graph network and the LSTM algorithm, including: Determining the clue data that needs to be predicted as clue data to be predicted; Determining graph node information related to the clue data to be predicted according to the clue feature graph network; The graph node information is input into the trained LSTM algorithm to obtain the development prediction result corresponding to the clue data to be predicted; the LSTM algorithm is trained by a training data set including multiple training clue development sequences and corresponding clue graph node information annotations.
[0019] As an optional implementation, in the second aspect of the present invention, the prediction module determines the specific manner of the graph node information related to the clue data to be predicted according to the clue feature graph network, including: In the clue feature graph network, graph node information related to the clue data to be predicted is extracted; the graph node information includes graph nodes related to multiple data features and at least one directly connected adjacent node and a correlation parameter between the graph node and the adjacent node.
[0020] The third aspect of the present invention discloses another clue analysis system based on association rules, the system comprising: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute part or all of the steps in the clue analysis method based on association rules disclosed in the first aspect of the present invention.
[0021] The fourth aspect of the present invention discloses a computer storage medium, which stores computer instructions. When the computer instructions are called, they are used to execute some or all of the steps in the clue analysis method based on association rules disclosed in the first aspect of the present invention.
[0022] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: The present invention is based on clue data related to the current business uploaded by multiple users, uses the Apriori algorithm to determine multiple high-correlation clue feature groups corresponding to all clue data, and constructs a clue feature graph network corresponding to the clue feature groups based on the graph network algorithm. According to the clue feature graph network and the LSTM algorithm, the development prediction result corresponding to any clue data is determined, thereby improving the efficiency and intelligence of business clue analysis, improving the reliability and accuracy of clue prediction, and providing effective data support for business decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0024] Figure 1 It is a flowchart of a clue analysis method based on association rules disclosed in an embodiment of the present invention.
[0025] Figure 2 It is a structural diagram of a clue analysis system based on association rules disclosed in an embodiment of the present invention.
[0026] Figure 3 It is a structural diagram of another clue analysis system based on association rules disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0027] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0028] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, device, product or equipment that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units that are not listed, or may optionally include other steps or units that are inherent to these processes, methods, products or equipment.
[0029] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present invention. The appearance of the phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0030] The present invention discloses a clue analysis method and system based on association rules. Based on clue data related to the current business uploaded by multiple users, the Apriori algorithm is used to determine multiple clue feature groups with high correlation corresponding to all clue data, and a clue feature graph network corresponding to the clue feature group is constructed based on a graph network algorithm. According to the clue feature graph network and the LSTM algorithm, the development prediction result corresponding to any clue data is determined, thereby improving the efficiency and intelligence of business clue analysis, improving the reliability and accuracy of clue prediction, and providing effective data support for business decision-making. The following are detailed descriptions.
[0031] Embodiment 1 See also Figure 1 , Figure 1 : is a flowchart of a clue analysis method based on association rules disclosed in an embodiment of the present invention. Figure 1 The described clue analysis method based on association rules can be applied to a data processing system / data processing device / data processing server (wherein the server includes a local processing server or a cloud processing server). Figure 1 As shown, the clue analysis method based on association rules may include the following operations: 101. Obtain clue data related to the current business uploaded by multiple users.
[0032] 102. Based on the Apriori algorithm, multiple high-correlation clue feature groups corresponding to all clue data are determined. 103. Based on the graph network algorithm, a clue feature graph network corresponding to all clue feature groups is constructed. 104. According to the clue feature graph network and LSTM algorithm, the development prediction result corresponding to any clue data is determined.
[0033] It can be seen that the above-mentioned embodiment of the invention is based on the clue data related to the current business uploaded by multiple users, and uses the Apriori algorithm to determine multiple high-correlation clue feature groups corresponding to all clue data, and constructs a clue feature graph network corresponding to the clue feature group based on the graph network algorithm. According to the clue feature graph network and the LSTM algorithm, the development prediction result corresponding to any clue data is determined, thereby improving the efficiency and intelligence of business clue analysis, improving the reliability and accuracy of clue prediction, and providing effective data support for business decision-making.
[0034] As an optional embodiment, in the above steps, the clue data includes at least one of clue type, clue level, clue content, clue object and clue recommendation processing method.
[0035] It can be seen that through the above optional embodiments, the content of the clue data is limited to comprehensively characterize the relevant features of the clues, so as to facilitate subsequent accurate feature mining and clue prediction, assist in improving the efficiency and intelligence of business clue analysis, improve the reliability and accuracy of clue prediction, and provide effective data support for business decisions.
[0036] As an optional embodiment, in the above step, based on the Apriori algorithm, determining multiple high-correlation clue feature groups corresponding to all clue data includes: For each clue data, based on the preset hash value rule, calculate the feature hash corresponding to each data feature in the clue data; Mapping the feature hash corresponding to each data feature in the clue data to a preset data feature hash table to obtain a mapping hash table; Based on the Apriori algorithm and the mapping hash table, multiple high-correlation clue feature groups corresponding to all clue data are determined.
[0037] It can be seen that through the above optional embodiments, based on each clue data, the preset hash value rule is used to calculate the feature hash corresponding to each data feature, and all feature hashes are mapped to the preset data feature hash table to obtain a mapping hash table, and then based on the Apriori algorithm and the mapping hash table, multiple high-correlation clue feature groups corresponding to all clue data are determined, thereby improving the efficiency and accuracy of clue feature extraction, assisting in improving the efficiency and intelligence of business clue analysis, improving the reliability and accuracy of clue prediction, and providing effective data support for business decisions.
[0038] As an optional embodiment, in the above steps, based on the Apriori algorithm and the mapping hash table, determining multiple high-correlation clue feature groups corresponding to all clue data includes: Based on the mapping hash table and hash comparison algorithm, the number of occurrences of each data feature in all clue data is matched and counted to obtain the counting result; Based on the Apriori algorithm and a preset minimum support threshold and / or minimum confidence threshold, feature mining is performed on all clue data according to the calculation results to obtain multiple clue feature groups with high correlation.
[0039] It can be seen that through the above optional embodiments, based on the mapping hash table and the hash comparison algorithm, the number of occurrences of each data feature in all clue data is matched and counted to obtain a counting result, and the Apriori algorithm is used in combination with the preset minimum support threshold and / or minimum confidence threshold to perform feature mining on all clue data according to the counting results to obtain multiple highly correlated clue feature groups, thereby improving the efficiency and accuracy of clue feature mining, assisting in improving the efficiency and intelligence of business clue analysis, improving the reliability and accuracy of clue prediction, and providing effective data support for business decisions.
[0040] As an optional embodiment, in the above steps, constructing a clue feature graph network corresponding to all clue feature groups based on a graph network algorithm includes: Determine the data feature belonging to any clue special investigation group as a corresponding graph node; For any two graph nodes, determine a clue feature group where the data features corresponding to the two graph nodes respectively exist simultaneously, and obtain at least one associated feature group; Calculate the weighted sum average of the support corresponding to all associated feature groups to obtain the association parameter corresponding to the two graph nodes; Based on all graph nodes and corresponding relevance parameters, a corresponding clue feature graph network is established.
[0041] It can be seen that through the above optional embodiments, the data features belonging to any clue feature group are determined as corresponding graph nodes. For any two graph nodes, the clue feature groups in which the data features corresponding to the two graph nodes respectively exist simultaneously are determined to obtain at least one associated feature group. The weighted sum average of the support corresponding to all associated feature groups is calculated to obtain the association parameters corresponding to the two graph nodes. Based on all graph nodes and the corresponding association parameters, a clue feature graph network is established to achieve a visual expression of the association relationship between clue features, thereby improving the structured analysis capabilities of clue data and providing more accurate data support for subsequent feature mining and development trend prediction.
[0042] As an optional embodiment, in the above steps, when calculating the weighted sum average of the support corresponding to all associated feature groups, the weighted calculation weight corresponding to the support corresponding to each associated feature group is proportional to the feature importance corresponding to the associated feature group; the feature importance is calculated by the following steps: For each other data feature in the associated feature group except the data features corresponding to the two graph nodes respectively, calculate the number of clue feature groups in which the other data feature and the data features corresponding to the two graph nodes respectively exist together, and obtain the feature similarity corresponding to the other data feature; Filter out features whose feature similarity is greater than a similarity threshold from all other data features to obtain at least one high-similarity feature; Calculate the ratio of the number of all other data features to the total number of features to obtain a first quantity ratio; Calculate the ratio of the number of all high-similarity features to the number of all other data features to obtain a second quantity ratio; calculating a quantity weight inversely proportional to the second quantity ratio; Calculate the product of the quantity weight and the first quantity ratio; Calculate the reciprocal of the product to obtain the feature importance corresponding to the associated feature group.
[0043] It can be seen that through the above optional embodiments, the proportion of the data features corresponding to the two graph nodes in each associated feature group is taken into account when calculating the correlation parameter, so that the relative importance of the features can be measured according to their distribution in different clue feature groups, thereby assisting in the visual expression of the correlation relationship between clue features and improving the structured analysis capabilities of clue data.
[0044] As an optional embodiment, in the above steps, determining the development prediction result corresponding to any clue data according to the clue feature graph network and the LSTM algorithm includes: Determine the clue data that needs to be predicted as the clue data to be predicted; Determine graph node information related to the clue data to be predicted according to the clue feature graph network; The graph node information is input into the trained LSTM algorithm to obtain the development prediction result corresponding to the clue data to be predicted; the LSTM algorithm is trained by a training data set including multiple training clue development sequences and corresponding clue graph node information annotations.
[0045] It can be seen that through the above optional embodiments, based on the clue feature graph network, the relevant graph node information of the clue data to be predicted is extracted, and it is analyzed using the trained LSTM algorithm to generate the clue development prediction results, which can combine the time series change trend of the clue characteristics, improve the accuracy and reliability of the prediction, and enhance the intelligent judgment ability of business clues.
[0046] As an optional embodiment, in the above step, determining the graph node information related to the clue data to be predicted according to the clue feature graph network includes: In the clue feature graph network, graph node information related to the clue data to be predicted is extracted; the graph node information includes graph nodes related to multiple data features and at least one directly connected adjacent node and correlation parameters between the graph nodes and the adjacent nodes.
[0047] It can be seen that through the above optional embodiments, based on the clue feature graph network, the graph node information related to the clue data to be predicted is extracted, including the graph nodes corresponding to the data features and their directly connected adjacent nodes, and combined with the correlation parameters between the nodes, the correlation relationship between the clue features is comprehensively reflected, thereby providing more accurate input data for subsequent clue development predictions and improving the accuracy and rationality of the predictions.
[0048] Embodiment 2 See also Figure 2 , Figure 2 is a schematic diagram of a clue analysis system based on association rules disclosed in an embodiment of the present invention. Figure 2 The described clue analysis system based on association rules can be applied to a data processing system / data processing device / data processing server (wherein the server includes a local processing server or a cloud processing server). Figure 2 As shown, the clue analysis system based on association rules may include: The acquisition module 201 is used to acquire clue data related to the current business uploaded by multiple users.
[0049] The determination module 202 is used to determine a plurality of high-correlation clue feature groups corresponding to all clue data based on the Apriori algorithm. The construction module 203 is used to construct a clue feature graph network corresponding to all clue feature groups based on a graph network algorithm. The prediction module 204 is used to determine the development prediction result corresponding to any clue data based on the clue feature graph network and the LSTM algorithm.
[0050] It can be seen that the above-mentioned embodiment of the invention is based on the clue data related to the current business uploaded by multiple users, and uses the Apriori algorithm to determine multiple high-correlation clue feature groups corresponding to all clue data, and constructs a clue feature graph network corresponding to the clue feature group based on the graph network algorithm. According to the clue feature graph network and the LSTM algorithm, the development prediction result corresponding to any clue data is determined, thereby improving the efficiency and intelligence of business clue analysis, improving the reliability and accuracy of clue prediction, and providing effective data support for business decision-making.
[0051] As an optional embodiment, the clue data includes at least one of clue type, clue level, clue content, clue object and clue recommendation processing method.
[0052] It can be seen that through the above optional embodiments, the content of the clue data is limited to comprehensively characterize the relevant features of the clues, so as to facilitate subsequent accurate feature mining and clue prediction, assist in improving the efficiency and intelligence of business clue analysis, improve the reliability and accuracy of clue prediction, and provide effective data support for business decisions.
[0053] As an optional embodiment, the specific manner in which the determination module determines multiple high-correlation clue feature groups corresponding to all clue data based on the Apriori algorithm includes: For each clue data, based on the preset hash value rule, calculate the feature hash corresponding to each data feature in the clue data; Mapping the feature hash corresponding to each data feature in the clue data to a preset data feature hash table to obtain a mapping hash table; Based on the Apriori algorithm and the mapping hash table, multiple high-correlation clue feature groups corresponding to all clue data are determined.
[0054] It can be seen that through the above optional embodiments, based on each clue data, the preset hash value rule is used to calculate the feature hash corresponding to each data feature, and all feature hashes are mapped to the preset data feature hash table to obtain a mapping hash table, and then based on the Apriori algorithm and the mapping hash table, multiple high-correlation clue feature groups corresponding to all clue data are determined, thereby improving the efficiency and accuracy of clue feature extraction, assisting in improving the efficiency and intelligence of business clue analysis, improving the reliability and accuracy of clue prediction, and providing effective data support for business decisions.
[0055] As an optional embodiment, the specific manner in which the determination module determines multiple high-correlation clue feature groups corresponding to all clue data based on the Apriori algorithm and the mapping hash table includes: Based on the mapping hash table and hash comparison algorithm, the number of occurrences of each data feature in all clue data is matched and counted to obtain the counting result; Based on the Apriori algorithm and a preset minimum support threshold and / or minimum confidence threshold, feature mining is performed on all clue data according to the calculation results to obtain multiple clue feature groups with high correlation.
[0056] It can be seen that through the above optional embodiments, based on the mapping hash table and the hash comparison algorithm, the number of occurrences of each data feature in all clue data is matched and counted to obtain a counting result, and the Apriori algorithm is used in combination with the preset minimum support threshold and / or minimum confidence threshold to perform feature mining on all clue data according to the counting results to obtain multiple highly correlated clue feature groups, thereby improving the efficiency and accuracy of clue feature mining, assisting in improving the efficiency and intelligence of business clue analysis, improving the reliability and accuracy of clue prediction, and providing effective data support for business decisions.
[0057] As an optional embodiment, the construction module constructs the clue feature graph network corresponding to all clue feature groups based on the graph network algorithm, including: Determine the data feature belonging to any clue special investigation group as a corresponding graph node; For any two graph nodes, determine a clue feature group where the data features corresponding to the two graph nodes respectively exist simultaneously, and obtain at least one associated feature group; Calculate the weighted sum average of the support corresponding to all associated feature groups to obtain the association parameter corresponding to the two graph nodes; Based on all graph nodes and corresponding relevance parameters, a corresponding clue feature graph network is established.
[0058] It can be seen that through the above optional embodiments, the data features belonging to any clue feature group are determined as corresponding graph nodes. For any two graph nodes, the clue feature groups in which the data features corresponding to the two graph nodes respectively exist simultaneously are determined to obtain at least one associated feature group. The weighted sum average of the support corresponding to all associated feature groups is calculated to obtain the association parameters corresponding to the two graph nodes. Based on all graph nodes and the corresponding association parameters, a clue feature graph network is established to achieve a visual expression of the association relationship between clue features, thereby improving the structured analysis capabilities of clue data and providing more accurate data support for subsequent feature mining and development trend prediction.
[0059] As an optional embodiment, when calculating the weighted sum average of the support corresponding to all associated feature groups, the weighted calculation weight corresponding to the support corresponding to each associated feature group is proportional to the feature importance corresponding to the associated feature group; the feature importance is calculated by the following steps: For each other data feature in the associated feature group except the data features corresponding to the two graph nodes respectively, calculate the number of clue feature groups in which the other data feature and the data features corresponding to the two graph nodes respectively exist together, and obtain the feature similarity corresponding to the other data feature; Filter out features whose feature similarity is greater than a similarity threshold from all other data features to obtain at least one high-similarity feature; Calculate the ratio of the number of all other data features to the total number of features to obtain a first quantity ratio; Calculate the ratio of the number of all high-similarity features to the number of all other data features to obtain a second quantity ratio; calculating a quantity weight inversely proportional to the second quantity ratio; Calculate the product of the quantity weight and the first quantity ratio; Calculate the reciprocal of the product to obtain the feature importance corresponding to the associated feature group.
[0060] It can be seen that through the above optional embodiments, the proportion of the data features corresponding to the two graph nodes in each associated feature group is taken into account when calculating the correlation parameter, so that the relative importance of the features can be measured according to their distribution in different clue feature groups, thereby assisting in the visual expression of the correlation relationship between clue features and improving the structured analysis capabilities of clue data.
[0061] As an optional embodiment, the prediction module determines the specific method of the development prediction result corresponding to any clue data according to the clue feature graph network and the LSTM algorithm, including: Determine the clue data that needs to be predicted as the clue data to be predicted; Determine graph node information related to the clue data to be predicted according to the clue feature graph network; The graph node information is input into the trained LSTM algorithm to obtain the development prediction result corresponding to the clue data to be predicted; the LSTM algorithm is trained by a training data set including multiple training clue development sequences and corresponding clue graph node information annotations.
[0062] It can be seen that through the above optional embodiments, based on the clue feature graph network, the relevant graph node information of the clue data to be predicted is extracted, and it is analyzed using the trained LSTM algorithm to generate the clue development prediction results, which can combine the time series change trend of the clue characteristics, improve the accuracy and reliability of the prediction, and enhance the intelligent judgment ability of business clues.
[0063] As an optional embodiment, the prediction module determines the specific manner of the graph node information related to the clue data to be predicted according to the clue feature graph network, including: In the clue feature graph network, graph node information related to the clue data to be predicted is extracted; the graph node information includes graph nodes related to multiple data features and at least one directly connected adjacent node and correlation parameters between the graph nodes and the adjacent nodes.
[0064] It can be seen that through the above optional embodiments, based on the clue feature graph network, the graph node information related to the clue data to be predicted is extracted, including the graph nodes corresponding to the data features and their directly connected adjacent nodes, and combined with the correlation parameters between the nodes, the correlation relationship between the clue features is comprehensively reflected, thereby providing more accurate input data for subsequent clue development predictions and improving the accuracy and rationality of the predictions.
[0065] Embodiment 3 See also Figure 3 , Figure 3 It is another clue analysis system based on association rules disclosed in an embodiment of the present invention. Figure 3 The described clue analysis system based on association rules is applied to a data processing system / data processing device / data processing server (wherein the server includes a local processing server or a cloud processing server). Figure 3 As shown, the clue analysis system based on association rules may include: A memory 301 storing executable program codes; a processor 302 coupled to the memory 301; The processor 302 calls the executable program code stored in the memory 301 to execute the steps of the clue analysis method based on association rules described in the first embodiment.
[0066] Embodiment 4 An embodiment of the present invention discloses a computer-readable storage medium storing a computer program for electronic data exchange, wherein the computer program enables a computer to execute the steps of the clue analysis method based on association rules described in the first embodiment.
[0067] Embodiment 5 An embodiment of the present invention discloses a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to enable a computer to execute the steps of the clue analysis method based on association rules described in the first embodiment.
[0068] The above describes specific embodiments of the present specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily have to be performed in the specific order or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0069] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0070] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0071] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, the embodiments of this specification may be in the form of complete hardware embodiments, complete software embodiments, or embodiments in combination with software and hardware. Moreover, the embodiments of this specification may be in the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0072] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0073] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0074] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0075] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0076] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0077] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0078] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0079] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0080] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0081] Finally, it should be noted that the clue analysis method and system based on association rules disclosed in the embodiments of the present invention only disclose the preferred embodiments of the present invention, which are only used to illustrate the technical scheme of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, it should be understood by those skilled in the art that the technical schemes described in the aforementioned embodiments can still be modified, or some of the technical features therein can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical schemes from the spirit and scope of the technical schemes of the embodiments of the present invention.
Claims
1. A clue analysis method based on association rules, characterized in that: The method comprises: Obtain lead data related to the current business uploaded by multiple users; Based on the Apriori algorithm, determining multiple high-correlation clue feature groups corresponding to all the clue data; Based on the graph network algorithm, construct a clue feature graph network corresponding to all the clue feature groups; According to the clue feature graph network and the LSTM algorithm, the development prediction result corresponding to any clue data is determined.
2. The clue analysis method based on association rules according to claim 1 is characterized in that: The clue data includes at least one of clue type, clue level, clue content, clue object and clue recommendation processing method.
3. The clue analysis method based on association rules according to claim 1 is characterized in that: The method of determining multiple high-correlation clue feature groups corresponding to all clue data based on the Apriori algorithm includes: For each clue data, based on a preset hash value rule, a feature hash corresponding to each data feature in the clue data is calculated; Mapping the feature hash corresponding to each data feature in the clue data to a preset data feature hash table to obtain a mapping hash table; Based on the Apriori algorithm and the mapping hash table, multiple high-correlation clue feature groups corresponding to all the clue data are determined.
4. The clue analysis method based on association rules according to claim 3 is characterized in that: The determining of a plurality of high-correlation clue feature groups corresponding to all the clue data based on the Apriori algorithm and the mapping hash table includes: Based on the mapping hash table and the hash comparison algorithm, the number of occurrences of each data feature in all the clue data is matched and counted to obtain a counting result; Based on the Apriori algorithm and a preset minimum support threshold and / or a minimum confidence threshold, feature mining is performed on all the clue data according to the calculation results to obtain a plurality of clue feature groups with high correlation.
5. The clue analysis method based on association rules according to claim 1 is characterized in that: The graph network algorithm is based on which a clue feature graph network corresponding to all the clue feature groups is constructed, including: Determine the data feature belonging to any of the clue special investigation groups as a corresponding graph node; For any two of the graph nodes, determining the clue feature groups in which the data features respectively corresponding to the two graph nodes exist simultaneously, and obtaining at least one associated feature group; Calculate the weighted average of the support corresponding to all the associated feature groups to obtain the association parameter corresponding to the two graph nodes; Based on all the graph nodes and the corresponding association parameters, a corresponding clue feature graph network is established.
6. The clue analysis method based on association rules according to claim 5 is characterized in that: When calculating the weighted sum average of the support corresponding to all the associated feature groups, the weighted calculation weight corresponding to the support corresponding to each associated feature group is proportional to the feature importance corresponding to the associated feature group; the feature importance is calculated by the following steps: For each other data feature in the associated feature group except the data features corresponding to the two graph nodes respectively, calculate the number of clue feature groups in which the other data feature and the data features corresponding to the two graph nodes respectively exist together, and obtain the feature similarity corresponding to the other data feature; Filter out the features whose feature similarity is greater than a similarity threshold from all the other data features to obtain at least one high-similarity feature; Calculate the ratio of the number of all other data features to the total number of features to obtain a first quantity ratio; Calculating the ratio of the number of all the high-similarity features to the number of all the other data features to obtain a second quantity ratio; calculating a quantity weight inversely proportional to the second quantity ratio; Calculating the product of the quantity weight and the first quantity ratio; The reciprocal of the product is calculated to obtain the feature importance corresponding to the associated feature group.
7. The clue analysis method based on association rules according to claim 1 is characterized in that: Determining the development prediction result corresponding to any clue data according to the clue feature graph network and the LSTM algorithm includes: Determining the clue data that needs to be predicted as clue data to be predicted; Determining graph node information related to the clue data to be predicted according to the clue feature graph network; The graph node information is input into the trained LSTM algorithm to obtain the development prediction result corresponding to the clue data to be predicted; the LSTM algorithm is trained by a training data set including multiple training clue development sequences and corresponding clue graph node information annotations.
8. The clue analysis method based on association rules according to claim 7 is characterized in that: The determining the graph node information related to the clue data to be predicted according to the clue feature graph network includes: In the clue feature graph network, graph node information related to the clue data to be predicted is extracted; the graph node information includes graph nodes related to multiple data features and at least one directly connected adjacent node and a correlation parameter between the graph node and the adjacent node.
9. A clue analysis system based on association rules, characterized in that: The system comprises: The acquisition module is used to obtain the lead data related to the current business uploaded by multiple users; A determination module, used for determining a plurality of highly correlated clue feature groups corresponding to all the clue data based on an Apriori algorithm; A construction module, used for constructing a clue feature graph network corresponding to all the clue feature groups based on a graph network algorithm; The prediction module is used to determine the development prediction result corresponding to any clue data based on the clue feature graph network and the LSTM algorithm.
10. A clue analysis system based on association rules, characterized in that: The system comprises: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the clue analysis method based on association rules as described in any one of claims 1-8.
Citation Information
Patent Citations
Large-scale image high-speed retrieval method based on multi-view enhanced depth hash
CN110674333A
Pediatric patient electronic health recording system
CN117854665A
Scientific and technological information recommendation method and device based on graph neural network
CN117951377A
PVB production process energy efficiency optimization system based on Apriori algorithm and edge cloud collaboration
CN119338330A
Efficient incremental method for data mining of a database
US20030217055A1