Banking outlet business mode mining method based on position information
By constructing the spatial index structure of the regional subgraph constructor and refining the geographical location information, the problem of high complexity of the existing frequent subgraph mining algorithm is solved, and efficient frequent subgraph mining is realized to help analyze the business model between bank outlets.
Patent Information
- Application Number
- CN202411846651.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-16
AI Technical Summary
The existing frequent subgraph mining algorithms fail to effectively consider geographical location factors, resulting in high algorithm complexity, long running time and excessive disk space consumption, affecting data processing efficiency and cost-effectiveness.
By constructing the spatial index structure of the regional subgraph constructor, the search problem of the tree structure is initialized using geographical location information, so as to refine the geographical location during frequent subgraph mining and improve mining efficiency.
It realizes efficient positioning of the target search area and obtaining sub-graph information in the area, reducing the scale of candidate sub-graphs, improving the efficiency of frequent sub-graph mining, and helping to analyze the strong business types and business development focus between different outlets.
Smart Images

Figure CN120011417A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a bank branch business model mining method, and more specifically, to a bank branch business model mining method based on location information, and belongs to the field of computer application technology. Background Art
[0002] In recent years, with the rapid development of Internet technology and the widespread popularity of smart mobile devices, banking business has been greatly expanded, showing unprecedented vitality and potential. The data resources generated in this process not only have extremely high commercial value, but also are an important basis for banking business innovation and optimization. Among them, geographic location information, as a key component of user data, has received extensive attention from researchers.
[0003] In the layout of banking business, geographical location plays a decisive role in the business focus of branches. Different geographical regions, due to their unique socio-economic characteristics, have led to significant differences in the key businesses of branches. This difference is not only reflected in the distribution of business types, but also in the business scale and development trend. Therefore, in-depth exploration and understanding of the relationship between geography and business is extremely important for the precise marketing and resource allocation of bank branches.
[0004] Among many data mining algorithms, frequent subgraph mining algorithms are considered to be the essence of graph mining algorithms because they can effectively mine customer behavior subgraphs and habits. However, existing frequent subgraph mining algorithms do not consider geographic location factors, and often generate a large number of candidate subgraphs during the mining process, which not only greatly increases the complexity and running time of the algorithm, but also leads to excessive use of disk space, seriously affecting the efficiency and cost-effectiveness of data processing. Summary of the invention
[0005] The purpose of the present invention is to provide a bank branch business model mining method based on location information, which has the ability to efficiently locate the target search area and obtain the sub-graph information within the area, and through the spatial index structure of the regional sub-graph constructor, the frequent sub-graph mining work can be refined according to the geographical location, thereby helping to analyze the strong business types and business development priorities among different branches and other technical characteristics.
[0006] In order to achieve the above object, the present invention is implemented by the following technical solutions:
[0007] The method for mining the business model of bank outlets based on location information of the present invention is characterized in that the method comprises the following steps:
[0008] Step 1: Initialization of the regional subgraph constructor; Initialize the regional subgraph constructor SubConst according to the geographic location information between nodes, and use this constructor to transform the geographic information retrieval problem into a tree structure search problem, that is, the process from the root node to a leaf node;
[0009] The initialization process of the regional subgraph constructor SubConst is divided into four steps: data preprocessing, leaf node construction, node merging, and upper-level node generation; the subgraph constructor of the banking network graph G(V,E,L) contains a total of n rectangles, where n=|V|, each vertex contains a minimum rectangle, and it is set that the node of each constructor can accommodate at most m rectangles; the banking network graph G(V,E,L) consists of a vertex set V, an edge set E, and a labeling function L, which assigns labels to vertices and edges;
[0010] Step 2: Frequent subgraph mining: input a support threshold τ, and mine all subgraphs with support greater than τ in the graph G; given the initialized regional subgraph constructor SubConst, the latitude and longitude information of the target area, and the frequency threshold τ, return all frequent subgraphs with support greater than τ. The returned results represent all the strong businesses of the outlets in the target area;
[0011] Step 3: Return the result set result.
[0012] Preferably, the initialization of the region subgraph constructor SubConst specifically includes:
[0013] Step 1-1: Data preprocessing: construct an initial rectangle with a side length of 3km for each vertex (the user's main activity range), and place n rectangles in a series of rectangular groups. The number of rectangular groups is Arrange these rectangular groups according to geographical locations, and the rectangular groups explicitly record all the vertices, edge types and their quantities contained in the groups;
[0014] Step 1-2: Leaf node construction; convert the rectangular group generated in step 1-1 into leaf nodes, and each leaf node contains at most m rectangles, representing a business cluster;
[0015] Step 1-3: Node merging: After the leaf nodes are generated, they are merged according to the distance between them, that is, the leaf nodes with closer distance are merged, and the minimum circumscribed rectangle of the merged nodes is used as the rectangle of the merged node. The node merging of the non-leaf node layer is similar, and it should be noted that all the information of the tree nodes needs to be merged when merging;
[0016] Step 1-4: Generate upper-level nodes and generate root nodes; repeat steps 1-1 to 1-3 until the root node root is generated, generate a maximum rectangle covering all rectangles, and realize that the root node contains all additional information of the banking business network graph G(V,E,L).
[0017] Preferably, the specific steps of frequent subgraph mining are as follows:
[0018] Step 2-1: Determine the range of the target area in the regional subgraph constructor SubConst; the process of determining the node range of the target area in the regional subgraph constructor SubConst is a tree search process starting from the root node;
[0019] Step 2-2: Subgraph retrieval: After obtaining all the regional subgraph constructor SubConst nodes corresponding to the target search area, generate the regional subgraph corresponding to the target search area in graph G according to the subgraph information stored in the root node;
[0020] Step 2-3: Initialize the edge queue; initialize an ascending priority queue fEdges to store the frequent edges of the banking area subgraph S;
[0021] Step 2-4: Select the first element e1 of fEdges as the initial subgraph S11;
[0022] Step 2-5: Determine whether the subsequent element can be used to expand the initial subgraph S11; if the element e2 can be used to expand the initial subgraph S11, perform the following steps:
[0023] Step 2-5-1: Generate subgraph S12 by combining the initial subgraph S11 and the element e2;
[0024] Step 2-5-2: Start timer T;
[0025] Step 2-5-3: Calculate the support SG(S12) of subgraph S12. If the support calculation time exceeds the time threshold Tt, discard the subgraph and re-execute step 2-5;
[0026] Step 2-5-4: If the support SG(S12) of subgraph S12 is greater than the threshold τ, then subgraph S12 is added to result and step 2-5 is performed on subgraph S12; steps 2-4 to 2-5 are repeated until the queue fEdges is empty, at which point result is the set of all frequent subgraphs whose support is greater than the threshold τ.
[0027] Preferably, step 2-1 specifically includes:
[0028] Step 2-1-1: Determine whether the target area covers the currently searched tree node;
[0029] Step 2-1-2: If the current node is covered, search for the child nodes of the current node and execute step 2-1-1 again; if it is not covered, return to the upper node; start from the root node and perform steps 2-1-1 to 2-1-2 until all target nodes that meet the conditions are obtained. After completion, update the subgraph information saved by the root node according to the execution status.
[0030] Preferably, step 2-2 specifically includes:
[0031] Step 2-2-1: Get all vertices (v1, v2, ..., vn) stored in the region subgraph constructor SubConst;
[0032] Step 2-2-2: Retrieve the edges connecting the above vertices from graph G;
[0033] Step 2-2-3: Simplify the subgraph and delete the nodes and edges whose frequency of occurrence is lower than the threshold τ; construct the banking business area subgraph S corresponding to the target search area based on the graph node and edge information obtained in steps 2-2-1, 2-2-2, and 2-2-3.
[0034] Preferably, the specific steps of steps 2-3 are as follows:
[0035] Step 2-3-1: Calculate the MNI values of all edges in the edge set E respectively, and select the edges whose MNI values are greater than the threshold τ;
[0036] Step 2-3-2: Put the filtered edges into the priority queue fEdges and the result set result.
[0037] Preferably, the support MNI is calculated based on all instance vertices of edge e in the subgraph S; the specific steps are as follows:
[0038] Step ① Get all instances of e {I1, I2, …, Im} in subgraph S;
[0039] Step ② obtains all vertex information in the instance set and generates sets for vertices with different labels respectively;
[0040] Step ③ count and return the minimum value in the vertex set;
[0041] The value returned in step ③ is the support of edge e. The support calculation of the graph is similar. After the priority queue fEdges is constructed, the subgraph mining work can begin.
[0042] Preferably, support calculation is a constraint satisfaction problem, and the specific steps are as follows:
[0043] Step a: Set a node pool for each vertex of subgraph S12, corresponding to all corresponding vertices in graph S;
[0044] Step b: Search for instances of subgraph S12 in S and mark the corresponding vertices in the node pool;
[0045] Step c: After the instance search is completed, the minimum number of marked vertices in all node pools is the required support;
[0046] Preferably, the lazy search operation steps are as follows:
[0047] Step A: A timer is started before each subgraph starts to calculate the support;
[0048] Step B: If the set time threshold Tt is exceeded during the expansion and calculation process, the process stops and starts the next expansion subgraph.
[0049] Preferably, the steps of the expansion operation of the initial subgraph S11 are as follows:
[0050] Step I: According to the order of the edges stored in the fEdges queue, try to expand the initial subgraph S11 one by one;
[0051] Step II: If the element e2 can be used to expand the initial subgraph S11, that is, there are the same vertices that can be used for expansion, then the initial subgraph S11 is expanded into the subgraph S12.
[0052] Beneficial effects: By constructing a spatial index of a banking business network graph, the geographical location information of the network and the structural information within the region are stored in a tree index structure. After the index structure is constructed, the target search area can be efficiently located and the subgraph information within the region can be obtained. The present invention also introduces a series of strategies for optimizing searches, and proposes an efficient frequent pattern mining method based on this. The creativity of the present invention lies in that through the spatial index structure of the regional subgraph constructor, the frequent subgraph mining work can be refined according to the geographical location, which can help analyze the strong business types and business development priorities between different outlets. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a bank customer network graph G graph including 18 vertices in the embodiment of the present invention.
[0054] Figure 2 It is a regional division diagram of the graph G according to the geographical location in an embodiment of the present invention.
[0055] Figure 3 It is a sub-image corresponding to the target search area in the embodiment of the present invention.
[0056] Figure 4It is a SubConst spatial index structure diagram corresponding to Figure G in an embodiment of the present invention.
[0057] Figure 5 This is a sample diagram of the spatial index structure SubConst search in an embodiment of the present invention.
[0058] Figure 6 It is a flow chart of the present invention. DETAILED DESCRIPTION
[0059] The present invention will be further described below in conjunction with the accompanying drawings, but the present invention is not limited to the following embodiments.
[0060] Principle / technical solution: The present invention is given a banking network graph G (V, E, L), which consists of a vertex set V, an edge set E and a labeling function L, and the labeling function L assigns labels to vertices and edges. Different labels represent different types of business entities, such as outlets, individual customers, corporate customers, government agency customers, etc. Each vertex in V contains its latitude and longitude information. Subgraph S refers to a weakly connected subgraph obtained from graph G. The isomorphic subgraph corresponding to the subgraph in graph G. Frequent subgraph mining refers to inputting a support threshold τ and mining all subgraphs in graph G whose support is greater than τ. The evaluation of the support indicator in the present invention adopts the minimum mapping support (MNI), that is, the minimum value of the vertex type set in all instances of the subgraph.
[0061] In the mining process, it is a long process to construct a subgraph based on the geographic location information of the vertices. This is because the process needs to traverse all the vertices in the graph, which is too time-consuming for large graphs. For this reason, the present invention uses an efficient spatial index SubConst to improve the efficiency of constructing subgraphs, which is subsequently referred to as the regional subgraph constructor. In addition, the present application also uses strategies such as lazy search and constraint satisfaction reuse to accelerate the subgraph mining process.
[0062] The unique creativity of this application: The innovation of this invention includes using spatial index to construct a sub-network corresponding to a specific branch business coverage area, and using the information stored in the index and a series of optimization strategies to significantly reduce the size of the candidate sub-graph, thereby completing efficient mining. The process of the bank branch business model mining method based on location information of the present invention can be based on the attached Figure 6 The process recorded in forms a specific technical solution.
[0063] like Figure 1-6 The figure shows a specific embodiment of the method for mining business patterns of bank branches based on location information.
[0064] Figure 1It is a graph G with 18 vertices, each vertex in the graph represents a business entity. The node label represents the entity type, B represents outlets, P represents corporate customers, C represents individual customers, and G represents government unit customers.
[0065] The specific implementation method is described below using this figure as an example.
[0066] The SubConst index structure of graph G is as follows Figure 3 As shown, the index construction process is shown in step 1 of the present invention, which is specifically the following operations:
[0067] Construct an initial rectangle with a side length of 3km for each of the 18 vertices of graph G, and use these rectangles as the minimum area where the index can operate;
[0068] Figure 2 The regional division of Figure G is obtained by executing step 1-1 of the present invention;
[0069] According to the distance between rectangles, the 18 rectangles are divided into four groups: a1, a2, b1, and b2. Specifically, the four rectangle groups are (v1, v2, v3, v4, v5), (v6, v7, v8, v9, v10, v11), (v12, v13, v14, v15), and (v16, v17, v18).
[0070] The generated four groups are used as four leaf nodes of the SubConst index, that is, four geographically closely related business clusters, and circumscribed rectangles are divided for them, which are obtained by executing steps 1-2 of the present invention;
[0071] Merge the areas represented by the four leaf nodes and divide the minimum circumscribed rectangles for the two generated nodes. Specifically, merge a1 and a2 into A, and merge b1 and b2 into B. This step is obtained by executing steps 1-3 of the present invention;
[0072] The generated two nodes A and B are used as nodes of the second layer of the tree-like spatial index, as a larger business aggregation area; the areas represented by the two nodes A and B are merged and a minimum circumscribed rectangle that can cover the entire graph is generated; the generated final node is used as the root node of the tree-like spatial index, which is obtained by executing steps 1-4 of the present invention;
[0073] If a1 and a2 are used as the target search area, the SubConst index is operated as follows: starting from the root node, it is found that it covers the target search area, and its two child nodes are searched; the area covered by the A node is part of the target search area, and then its leaf nodes are searched; the information of the A node is updated according to the search results of the leaf nodes a1 and a2; the B node does not belong to the target search area, so it is not searched; then the information stored in the root node is updated, and at this time, the SubConst has completed the search and update, and the subgraph retrieval is completed according to the SubConst index, such as step 2-1 to step 2-2 of the present invention;
[0074] Figure 3 The graph S in is the retrieved subgraph corresponding to the target search area, and τ is the support threshold, which is set to 2. Frequent subgraph mining requires executing steps 2-3 to 2-5 of the present invention, and performing the following operations:
[0075] Initialize an ascending priority queue fEdges, sorted by support;
[0076] Calculate the support of all edges in graph S and add them to the queue fEdges. The order of elements in the queue is {BC, BP, BB, BG, PC, PG};
[0077] Delete the edge BC with support lower than τ, which means that the personal business services of the outlets in the target area are relatively few. At this time, the value of fEdges is {BP, BB, BG, PC, PG};
[0078] Add the elements in fEdges to the result set result;
[0079] The first element BP of fEdges is used as the initial subgraph S11;
[0080] The potential expandable edges of S11 are {BP, BB, BG, PC, PG}, generating subgraph S12 = (PBP);
[0081] The support of subgraph S12 is calculated to be 2, which meets the conditions and is added to the result;
[0082] Continue to expand subgraph S12 and generate (PBPB). The calculated support is 0, which is less than τ and does not meet the conditions. Therefore, it is deleted and this business model does not need to be expanded.
[0083] Generate and calculate (PBB, GBP, BPC, BPG) one by one. If they all meet the requirements, add them to result.
[0084] And so on, continue to expand the above subgraph and calculate the support;
[0085] The first element BP of the fEdges team is removed from the team, and the new first element BB is used as the new initial subgraph S21;
[0086] Calculate step by step until the fEdges queue is empty. At this time, result is the result set {BP, BB, BG, PC, PG, PBB, GBP, BPC, BPG, BBG}, which indicates the 10 most frequent business modes in the area.
[0087] Finally, the result is returned.
[0088] Finally, it should be noted that the present invention is not limited to the above embodiments, and there are many variations. All variations that can be directly derived or associated with the content disclosed by ordinary technicians in this field should be considered as the protection scope of the present invention.
Claims
1. A method for mining bank branch business models based on location information, characterized in that The method comprises the following steps: Step 1: Initialization of the regional subgraph constructor; Initialize the regional subgraph constructor SubConst according to the geographic location information between nodes, and use this constructor to transform the geographic information retrieval problem into a tree structure search problem; The initialization process of the regional subgraph constructor SubConst is divided into four steps: data preprocessing, leaf node construction, node merging, and upper-level node generation; the subgraph constructor of the banking network graph G(V,E,L) contains a total of n rectangles, where n=|V|, each vertex contains a minimum rectangle, and it is set that the node of each constructor can accommodate at most m rectangles; the banking network graph G(V,E,L) consists of a vertex set V, an edge set E, and a labeling function L, which assigns labels to vertices and edges; Step 2: Frequent subgraph mining: input a support threshold τ, and mine all subgraphs with support greater than τ in the graph G; given the initialized regional subgraph constructor SubConst, the latitude and longitude information of the target area, and the frequency threshold τ, return all frequent subgraphs with support greater than τ. The returned results represent all the strong businesses of the outlets in the target area; Step 3: Return the result set result.
2. The method for mining bank branch business models based on location information according to claim 1, characterized in that: The initialization of the region subgraph constructor SubConst specifically includes: Step 1-1: Data preprocessing; construct an initial rectangle with a side length of 3km for each vertex, and place n rectangles in a series of rectangular groups. The number of rectangular groups is Arrange these rectangular groups according to geographical locations, and the rectangular groups explicitly record all the vertices, edge types and their quantities contained in the groups; Step 1-2: Leaf node construction; convert the rectangular group generated in step 1-1 into leaf nodes, and each leaf node contains at most m rectangles, representing a business cluster; Step 1-3: Node merging: After the leaf nodes are generated, they are merged according to the distance between them, and the minimum bounding rectangle of the merged nodes is used as the rectangle of the merged node. The node merging of the non-leaf node layer is similar, and it should be noted that all the information of the tree nodes needs to be merged during the merging. Step 1-4: Generate upper-level nodes and generate root nodes; repeat steps 1-1 to 1-3 until the root node root is generated, generate a maximum rectangle covering all rectangles, and realize that the root node contains all additional information of the banking business network graph G(V,E,L).
3. The method for mining bank branch business models based on location information according to claim 1 or 2, characterized in that: The specific steps of frequent subgraph mining are as follows: Step 2-1: Determine the range of the target area in the regional subgraph constructor SubConst; the process of determining the node range of the target area in the regional subgraph constructor SubConst is a tree search process starting from the root node; Step 2-2: Subgraph retrieval: After obtaining all the regional subgraph constructor SubConst nodes corresponding to the target search area, generate the regional subgraph corresponding to the target search area in graph G according to the subgraph information stored in the root node; Step 2-3: Initialize the edge queue; initialize an ascending priority queue fEdges to store the frequent edges of the banking area subgraph S; Step 2-4: Select the first element e1 of fEdges as the initial subgraph S11; Step 2-5: Determine whether the subsequent element can be used to expand the initial subgraph S11; if the element e2 can be used to expand the initial subgraph S11, perform the following steps: Step 2-5-1: Generate subgraph S12 by combining the initial subgraph S11 and the element e2; Step 2-5-2: Start timer T; Step 2-5-3: Calculate the support SG of subgraph S12. If the support calculation time exceeds the time threshold Tt, discard the subgraph and re-execute step 2-5; Step 2-5-4: If the support SG of subgraph S12 is greater than the threshold τ, then subgraph S12 is added to result and step 2-5 is performed on subgraph S12; steps 2-4 to 2-5 are repeated until the queue fEdges is empty, at which point result is the set of all frequent subgraphs whose support is greater than the threshold τ.
4. The method for mining bank branch business models based on location information according to claim 3 is characterized by: Step 2-1 specifically includes: Step 2-1-1: Determine whether the target area covers the currently searched tree node; Step 2-1-2: If the current node is covered, search for the child nodes of the current node and execute step 2-1-1 again; if it is not covered, return to the upper node; start from the root node and perform steps 2-1-1 to 2-1-2 until all target nodes that meet the conditions are obtained. After completion, update the subgraph information saved by the root node according to the execution status.
5. The method for mining bank branch business models based on location information according to claim 3 is characterized by: Step 2-2 specifically includes: Step 2-2-1: Get all vertices (v1, v2, ..., vn) stored in the region subgraph constructor SubConst; Step 2-2-2: Retrieve the edges connecting the above vertices from graph G; Step 2-2-3: Simplify the subgraph and delete the nodes and edges whose frequency of occurrence is lower than the threshold τ; construct the banking business area subgraph S corresponding to the target search area based on the graph node and edge information obtained in steps 2-2-1, 2-2-2, and 2-2-3.
6. The method for mining bank branch business models based on location information according to claim 3, characterized in that: The specific steps for steps 2-3 are as follows: Step 2-3-1: Calculate the MNI values of all edges in the edge set E respectively, and select the edges whose MNI values are greater than the threshold τ; Step 2-3-2: Put the filtered edges into the priority queue fEdges and the result set result.
7. The method for mining bank branch business models based on location information according to claim 6, characterized in that: The support MNI is calculated based on all instance vertices of edge e in subgraph S; the specific steps are as follows: Step ① Get all instances of e {I1, I2, …, Im} in subgraph S; Step ② obtains all vertex information in the instance set and generates sets for vertices with different labels respectively; Step ③ count and return the minimum value in the vertex set; The value returned in step ③ is the support of edge e. The support calculation of the graph is similar. After the priority queue fEdges is constructed, the subgraph mining work can begin.
8. The method for mining bank branch business models based on location information according to claim 3, characterized in that: Support calculation is a constraint satisfaction problem. The specific steps are as follows: Step a: Set a node pool for each vertex of subgraph S12, corresponding to all corresponding vertices in graph S; Step b: Search for instances of subgraph S12 in S and mark the corresponding vertices in the node pool; Step c: After the instance search is completed, the minimum number of marked vertices in all node pools is the required support.
9. The method for mining bank branch business models based on location information according to claim 1, characterized in that: The lazy search operation steps are as follows: Step A: A timer is started before each subgraph starts to calculate the support; Step B: If the set time threshold Tt is exceeded during the expansion and calculation process, the process stops and starts the next expansion subgraph.
10. The method for mining bank branch business models based on location information according to claim 3, 4, 5 or 6, characterized in that: The steps of the expansion operation of the initial subgraph S11 are as follows: Step I: According to the order of the edges stored in the fEdges queue, try to expand the initial subgraph S11 one by one; Step II: If the element e2 can be used to expand the initial subgraph S11, that is, there are the same vertices that can be used for expansion, then the initial subgraph S11 is expanded into the subgraph S12.