Method and device for determining recommendation strategy based on data mining and knowledge graph

By constructing the knowledge graph and cluster analysis of financial institutions, the limitations of formulating recommendation strategies based on general static templates in the existing technology are solved, and the differentiated analysis and dynamic adjustment strategies are achieved for different regions.

CN120013649AActive Publication Date: 2025-05-16BANK OF BEIJING

Patent Information

Application Number
CN202510499414.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-16
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

When formulating recommendation strategies for financial institutions, the existing technology relies on common static templates, making it difficult to conduct differentiated analysis of different regions where different financial institutions are located, resulting in the limitation of the targetedness and effectiveness of recommendation activities.

Method used

By obtaining the original data of the operating area of ​​the financial institution, building a knowledge graph, determining the portrait of the target object, and mapping the portrait to the sub-regions of the operating area for clustering analysis, we determine the recommendation strategy corresponding to the clustering results.

Benefits of technology

It is realized that according to the economic characteristics of different geographical regions, cluster analysis of target objects in specific sub-regions is carried out to determine the most suitable recommendation strategy, thereby improving the targetedness of recommendation activities and dynamically adjusting the strategy according to environmental changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013649A_ABST
    Figure CN120013649A_ABST
Patent Text Reader

Abstract

The invention discloses a recommendation strategy determination method and device based on data mining and a knowledge graph. The method comprises the steps that original data corresponding to an operation area of a financial institution is acquired, and the original data comprises first data corresponding to a target object in the operation area and second data reflecting environment characteristics of the operation area; constructing a knowledge graph corresponding to the financial institution according to the original data, and determining a portrait corresponding to the target object according to the knowledge graph; the portrait is mapped to a sub-region corresponding to the operation region, the operation region corresponds to at least one sub-region, and each sub-region corresponds to different economic characteristics; and performing clustering analysis on the portraits in the sub-regions, and determining a recommendation strategy corresponding to a clustering result. According to the method and the device, the technical problem that recommendation activities are limited due to the fact that different areas where different financial institutions are located are difficult to carry out differentiation analysis when recommendation strategies are formulated based on universal static templates in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and more specifically, to a method and device for determining a recommendation strategy based on data mining and knowledge graphs. Background Art

[0002] In formulating recommendation strategies of financial institutions, traditional practices often rely on pre-defined static templates or rules. These templates are designed based on the industry average or market data of a specific region, while ignoring the specific economic characteristics and market environment differences of the operating regions of financial institutions. For example, the template may preset some common customer stratification standards, product promotion strategies or recommendation channel selection, but these standards do not take into account the unique factors such as the economic development level, industry distribution, customer behavior preferences, etc. of different regions, resulting in the limitation of the pertinence and effectiveness of recommendation activities.

[0003] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0004] The embodiments of the present application provide a method and device for determining a recommendation strategy based on data mining and knowledge graphs, so as to at least solve the technical problem that the related technology formulates recommendation strategies based on general static templates, which makes it difficult to conduct differentiated analysis on different regions where different financial institutions are located, resulting in limited recommendation activities.

[0005] According to one aspect of an embodiment of the present application, a method for determining a recommendation strategy based on data mining and knowledge graph is provided, including: obtaining original data corresponding to an operating area of ​​a financial institution, wherein the original data includes first data corresponding to a target object in the operating area and second data reflecting environmental characteristics of the operating area; constructing a knowledge graph corresponding to the financial institution based on the original data, and determining a portrait corresponding to the target object based on the knowledge graph; mapping the portrait to a sub-region corresponding to the operating area, wherein the operating area corresponds to at least one sub-region and each sub-region corresponds to different economic characteristics; performing cluster analysis on the portrait in the sub-region, and determining a recommendation strategy corresponding to the clustering result.

[0006] In some embodiments of the present application, at least one sub-region corresponding to the operating area is determined in the following manner: dividing the operating area into at least one first sub-region in the first administrative region, wherein the first sub-region is a geographical region whose economic activity intensity and / or industry distribution characteristics meet the first preset condition; dividing the first sub-region into at least one second sub-region in the second administrative region, wherein the second sub-region is a geographical region whose economic activity intensity and / or industry distribution characteristics meet the second preset condition, the indicator threshold in the second preset condition is greater than the corresponding indicator threshold in the first preset condition, and the administrative level of the second administrative region is lower than the administrative level of the first administrative region; dividing the second sub-region into at least one third sub-region in the third administrative region, wherein the third sub-region is a geographical region corresponding to the transportation logistics and / or production supply chain in the second sub-region, and the administrative level of the third administrative region is lower than the administrative level of the second administrative region.

[0007] In some embodiments of the present application, a business area is divided into at least one first sub-area within a first administrative area, including: determining multiple target indicators corresponding to original data, wherein the target indicators are used to quantitatively represent industry attributes and economic activity within the business area; converting the original data into data points on a map of the business area based on the target indicators, wherein the data points include target indicators and geographic location information; and performing cluster analysis on the data points to obtain the first sub-area.

[0008] In some embodiments of the present application, the portrait is mapped to the sub-region corresponding to the business area, including: using a mapping method from high to low administrative levels to sequentially determine the first sub-region, the second sub-region and the third sub-region to which the portrait corresponds respectively; or using a mapping method from low to high administrative levels to sequentially determine the third sub-region, the second sub-region and the first sub-region to which the portrait corresponds respectively.

[0009] In some embodiments of the present application, cluster analysis is performed on portraits in a sub-region, including: determining third data corresponding to the sub-region from second data, wherein the third data is used to reflect the environmental characteristics of the sub-region; determining a clustering algorithm corresponding to the third data and a target parameter corresponding to the clustering algorithm, wherein the target parameter is determined based on the distribution characteristics of all portraits in the sub-region; using a clustering algorithm to perform cluster analysis on all portraits in the sub-region to obtain a clustering result, wherein the clustering result includes multiple groups, each group representing a set of target objects whose similarity is greater than or equal to a preset threshold.

[0010] In some embodiments of the present application, determining a recommendation strategy corresponding to the clustering result includes: determining a target feature corresponding to the target group, wherein the target feature is determined based on features of a portrait of the target group, and the target group is any one of multiple groups; determining the recommendation strategy based on the target feature and third data.

[0011] In some embodiments of the present application, determining a portrait corresponding to a target object based on a knowledge graph includes: obtaining static attributes of the target object from the knowledge graph, wherein the static attributes include information corresponding to first data of the target object; determining dynamic attributes of the target from the knowledge graph, wherein the dynamic attributes include information associated with the target object in second data; fusing the static attributes with the dynamic attributes to obtain a joint representation of the target object; and determining a portrait corresponding to the target object based on the joint representation.

[0012] According to another aspect of an embodiment of the present application, a device for determining a recommendation strategy based on data mining and knowledge graph is also provided, including: an acquisition module for acquiring original data corresponding to the operating area of ​​a financial institution; a determination module for constructing a knowledge graph corresponding to the financial institution based on the original data, and determining a portrait corresponding to the target object based on the knowledge graph; a mapping module for mapping the portrait to a sub-area corresponding to the operating area; and a clustering module for performing cluster analysis on the portrait in the sub-area and determining a recommendation strategy corresponding to the clustering result.

[0013] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including: a memory and a processor, the memory being used to store program instructions; the processor being connected to the memory and being used to execute a method for determining the recommendation strategy based on data mining and knowledge graph.

[0014] According to another aspect of an embodiment of the present application, a non-volatile storage medium is also provided, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the above-mentioned method for determining the recommendation strategy based on data mining and knowledge graph by running the computer program.

[0015] According to another aspect of the embodiments of the present application, a computer program product is also provided, including computer instructions, which, when executed by a processor, implement the above-mentioned method for determining a recommendation strategy based on data mining and knowledge graph.

[0016] In an embodiment of the present application, by obtaining original data corresponding to the operating area of ​​the financial institution, and constructing a knowledge graph corresponding to the financial institution based on the original data, and then determining the portrait corresponding to the target object based on the knowledge graph, after mapping the portrait to the sub-region corresponding to the operating area, the portrait is clustered in the sub-region, and finally the recommendation strategy corresponding to the clustering result is determined. The purpose of clustering the target objects in specific sub-regions based on the economic characteristics of different geographical regions and determining the most suitable recommendation strategy is achieved, thereby achieving the technical effect of improving the targeted nature of recommendation activities and dynamically adjusting strategies according to environmental changes, and further solving the technical problem that the related technology formulates recommendation strategies based on general static templates, and it is difficult to conduct differentiated analysis on different regions where different financial institutions are located, resulting in limited recommendation activities. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0018] Figure 1 It is a hardware structure block diagram of a computer terminal according to a method for determining a recommendation strategy based on data mining and knowledge graph according to an embodiment of the present application;

[0019] Figure 2 is a flow chart of a method for determining a recommendation strategy based on data mining and knowledge graph according to an embodiment of the present application;

[0020] Figure 3 It is a structural diagram of a device for determining a recommendation strategy based on data mining and knowledge graph according to an embodiment of the present application. DETAILED DESCRIPTION

[0021] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0023] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:

[0024] ST-DBSCAN (Spatio-Temporal DBSCAN): A space-time clustering algorithm that extends the traditional DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm. In this application, ST-DBSCAN not only considers the spatial density of data points, but also adds the time dimension, so that the clustering results can reflect the distribution law of data points in time and space, which is used for customer group segmentation.

[0025] Geographic Information System (GIS): A technical system used to collect, store, process, analyze and display data related to geographic location. In this application, GIS is used to analyze the customer's location information, industry distribution and economic activity to support the formulation and implementation of recommendation strategies based on geographic location.

[0026] The recommendation strategy formulation method used by related technologies is difficult to overcome the limitations of static templates, resulting in poor regional adaptability and insufficient customer insights. For example, the operating area of ​​a rural bank may be dominated by the agricultural economy, while the operating area of ​​an urban commercial bank may be dominated by the service industry or high-tech industry. The general recommendation template cannot adapt to this difference, resulting in poor results of recommendation strategies in some areas. At the same time, the behavior patterns and needs of target customer groups in different regions vary significantly, and static templates are often grouped based on the basic attributes of customers (such as age, gender, and occupation), and fail to capture the dynamic attributes of customers, such as changes in transaction behavior, market trend response, and sensitivity to policy impacts, making it difficult for financial institutions to grasp the real-time needs of customers and market dynamics, and the pertinence and effectiveness of recommendation activities are limited.

[0027] In order to solve the above technical problems, the embodiments of the present application provide corresponding solutions, which are described in detail below.

[0028] The method for determining a recommendation strategy based on data mining and knowledge graph provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal for implementing a method for determining a recommendation strategy based on data mining and knowledge graph is shown. Figure 1 As shown, the computer terminal 10 may include one or more (102a, 102b, ..., 102n are used to illustrate) processors (the processor may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions connected via a wired and / or wireless network. In addition, it may also include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, and a BUS bus. Those skilled in the art will understand that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations shown.

[0029] It should be noted that the one or more processors and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10. As described in the embodiments of the present application, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0030] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for determining the recommendation strategy based on data mining and knowledge graph in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, the method for determining the recommendation strategy based on data mining and knowledge graph is realized. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0031] The transmission module 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission module 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission module 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0032] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .

[0033] It should be noted that, in some optional embodiments, the above Figure 1 The computer terminal shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. It should be noted that Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the computer terminal described above.

[0034] In the above-mentioned operating environment, an embodiment of the present application provides an embodiment of a method for determining a recommendation strategy based on data mining and knowledge graph. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0035] Figure 2is a flow chart of a method for determining a recommendation strategy based on data mining and knowledge graph according to an embodiment of the present application, such as Figure 2 As shown, the method comprises the following steps:

[0036] Step S202, obtaining original data corresponding to the operating area of ​​the financial institution, wherein the original data includes first data corresponding to the target object in the operating area and second data reflecting the environmental characteristics of the operating area.

[0037] In the above step S202, the operating area of ​​a financial institution refers to the specific geographical scope of the financial institution's business coverage and services, which can be areas of different levels such as cities, counties, townships, and towns.

[0038] Raw data refers to data that is initially collected without being processed or analyzed, including data from inside and outside the financial institution. For example, internal data may include customer information, transaction records, credit history, etc., and external data may include macroeconomic data, industry reports, geographic information, social media data, etc.

[0039] The first data corresponding to the target object refers to the information directly related to the target customers of the financial institution, such as the basic information of individual customers (age, gender, occupation), the corporate qualifications of corporate customers (registered capital, industry type), the transaction history of customers, credit scores, etc. The second data reflecting the environmental characteristics of the operating area may include but are not limited to the economic indicators (such as GDP, industrial distribution), population structure, geographic information (such as transportation network, infrastructure), policy environment (such as local tax policy, financial policy), etc. of the operating area.

[0040] It should be noted that environmental characteristics refer to a series of external factors and conditions related to the operating area of ​​a financial institution, which have a significant impact on the business operations, customer behavior and market opportunities of the financial institution, and may include but are not limited to:

[0041] (1) Macroeconomic indicators: including GDP growth rate, unemployment rate, inflation rate, interest rate level, etc. These indicators reflect the economic health and business environment of the operating area.

[0042] (2) Industry distribution and characteristics: Different operating regions may have unique industrial layouts. For example, the region where a rural bank is located may be dominated by agriculture and small-scale manufacturing, while an urban financial institution may serve high-tech, real estate or service industries. Industry distribution helps financial institutions identify target customer groups and design products and services that meet industry characteristics.

[0043] (3) Population structure and consumption behavior: including data such as age distribution, gender ratio, education level, income level, and consumption preferences, which are used to understand the needs and preferences of customer groups and develop personalized recommendation strategies.

[0044] (4) Geographic and infrastructure information: geographical location, transportation network, communication facilities, infrastructure construction, etc. of the operating area, which have a direct impact on the physical branch layout and service channel selection of financial institutions.

[0045] (5) Policy and regulatory environment: local tax policies, financial regulatory policies, credit guidelines, industry incentives or restrictions, etc. related to the operating area. These policy environment factors will affect the operating strategies and risk management of financial institutions.

[0046] In order to overcome the problem of insufficient information utilization caused by data silos and different formats, for example, the original data of financial institutions may contain a large amount of unstructured information, such as the handwritten notes of account managers or the verbal descriptions of farmers. When this information is used directly for analysis, it is inefficient and difficult to accurately interpret. Therefore, ETL (Extract, Transform, Load) tools can be used in combination with entity matching algorithms to integrate structured data (such as credit records) with unstructured data (such as account managers' research notes). At the same time, semantic alignment algorithms are applied to ensure data consistency and comparability. For example, for unstructured data, the BERT-CRFT hybrid model is used for entity-relationship extraction to convert key information in the text into structured data for subsequent analysis.

[0047] The operating areas of financial institutions often cover multiple sub-regions with different economic characteristics, such as agricultural areas, industrial areas or tourist areas. There are significant differences in customer needs and behavior patterns in each sub-region. In order to solve the problem that financial institutions find it difficult to intuitively understand the geographical distribution and industry characteristics of customers during the recommendation process, a map recommendation tool can be built to formulate more targeted recommendation strategies based on the characteristics of the sub-region where the customer is located, thereby improving the success rate and ROI (return on investment) of recommendation activities. For example, using cluster analysis algorithms and data visualization technology, potential customers can be classified and displayed according to their sub-regions, industry attributes and economic activity.

[0048] Step S204: construct a knowledge graph corresponding to the financial institution based on the original data, and determine a portrait corresponding to the target object based on the knowledge graph.

[0049] In the above step S204, the knowledge graph is used to represent the association relationship between multi-source heterogeneous data such as business knowledge, customer information, and operating environment of a financial institution.

[0050] Financial institutions have accumulated a large amount of unstructured texts such as research reports and customer communication records. Although these information contain rich business insights, they are difficult to directly apply to decision support due to the diverse formats and difficulty in information extraction. By converting key information in unstructured texts into structured data through entity recognition and relationship extraction, the efficiency of information utilization can be improved and the dimensions of customer portraits can be enriched. For example, natural language processing technology (such as the BERT-CRFT hybrid model) can be used to analyze unstructured text data in the original data within the operating area to identify important entity information (such as company name, transaction amount, policy keywords, etc.), and determine the relationship between entity information through relationship extraction algorithms (such as company A obtained the support of policy B and conducted transaction C).

[0051] In order to solve the problem of data islands and ensure that all relevant data can be effectively integrated into the knowledge graph, data fusion technology (such as federated learning or data lake technology) can be used to fuse the internal system data of financial institutions (first data) with the external environment data (second data) to obtain a comprehensive feature set that reflects internal customer information and external environment characteristics; the entity information (such as customers, enterprises, policies) in the comprehensive feature set is converted into nodes in the knowledge graph, and the relationships between entities (such as credit relationships, policy benefits, industry affiliations) are converted into edges in the knowledge graph to obtain a knowledge graph. The knowledge graph can reflect the actual position and status of the target object (customer) in its operating environment, as well as the impact of environmental factors on the target object.

[0052] In some embodiments of the present application, the portrait corresponding to the target object can be determined in the following manner: obtaining static attributes of the target object from the knowledge graph, wherein the static attributes include information corresponding to the first data of the target object; determining the dynamic attributes of the target from the knowledge graph, wherein the dynamic attributes include information associated with the target object in the second data; fusing the static attributes with the dynamic attributes to obtain a joint representation of the target object; and determining the portrait corresponding to the target object based on the joint representation.

[0053] Static attributes refer to the characteristics of the target object that are relatively stable and not easy to change within a preset time period, such as the basic information of the customer (age, gender, occupation), the establishment time of the enterprise, the registered capital, etc., which are used to describe the identity and basic status of the target object. In some embodiments of the present application, the entity recognition and relationship extraction algorithms in natural language processing technology can be used to accurately extract the static attributes of the target object from the knowledge graph. For example, the customer information document is structured using models such as BERT-CRF, key entities are identified and their attribute categories (such as "age", "income") are labeled.

[0054] Dynamic attributes refer to characteristics that change over time, such as recent trading activities of customers, real-time changes in the market environment, policy adjustments, etc., which are used to reflect the behavior patterns and market reactions of the target object in the current environment. In some embodiments of the present application, graph traversal algorithms (such as breadth-first search, depth-first search) and graph clustering algorithms (such as community detection algorithms) can be used to extract dynamic attributes related to the target object from the knowledge graph. These algorithms can identify entities closely connected to the target object and their dynamic attributes. For example, all entities related to the target industry are determined through a community detection algorithm, and then the dynamic attributes of these entities are extracted. For dynamic attributes with time series characteristics (such as transaction frequency, market interest rates), time series analysis techniques (such as ARIMA, Prophet) can also be used for trend prediction, and the prediction results are added to dynamic data to develop a more accurate customer portrait.

[0055] It should be noted that dynamic attributes have been integrated into the knowledge graph construction stage. Since dynamic attributes change in real time, the dynamic attributes in the knowledge graph need to be updated in real time. For example, when the knowledge graph is constructed, the dynamic attributes are marked; the first data is obtained from an external data source (such as a financial information API, market data, weather forecast, etc.), and the dynamic attributes are updated using the first data. In some embodiments of the present application, preset rules can be used for updating, and the preset rules can be determined based on the update requirements of the dynamic attributes. For example, when the update requirements of the dynamic attributes meet the first preset condition, the first update frequency is used for updating; when the update requirements of the dynamic attributes meet the second preset condition, the second update frequency is used for updating, the second preset condition is stricter on the update time than the first preset condition, and the second update frequency is greater than the first update frequency.

[0056] All extracted dynamic attributes are fused with the static attributes of the target object to form a joint representation. During the fusion process, the dynamic attention mechanism can be used to adjust the weights of different attributes to reflect their importance to the target object at the current point in time. For example, if the current agricultural product price is at its peak, the weight of the "agricultural product price" attribute is increased because it has a more obvious impact on farmers' loan demand.

[0057] Step S206, mapping the portrait to a sub-region corresponding to the business area, wherein the business area corresponds to at least one sub-region, and each sub-region corresponds to different economic characteristics.

[0058] In the above step S206, the sub-region is a subdivided region within the operating region, and each sub-region has unique economic characteristics, such as agricultural output value, industrial growth rate, resident consumption level, etc. It should be noted that the sub-regions can be divided according to the administrative level, for example, divided into a first preset number of first sub-regions at the first administrative level, and divided into a second preset number of second sub-regions at the second administrative level. The sub-region division methods at different administrative levels can be related to each other, such as further dividing the first sub-region into a second preset number of second sub-regions under the first sub-region. The sub-region division methods at different administrative levels can also be independent of each other, that is, the sub-regions are divided separately under different dimensions, so that customers can conduct separate analysis of different dimensions.

[0059] In order to overcome the problem of unclear sub-region division or insufficient description of economic characteristics, the sub-region to which the customer belongs can be determined by calculating the similarity between the characteristic vector of the customer portrait and the economic characteristic vector of each sub-region, such as using algorithms such as cosine similarity and Jaccard similarity coefficient.

[0060] It should be noted that GIS technology can also be combined with machine learning models (such as random forests and K-nearest neighbor algorithms) to predict the most likely active sub-regions based on the geographic location information and behavioral characteristics in the customer portrait. Specifically: Integrate customer portrait data with the geographic information of the business area; Convert the geographic location information and behavioral characteristics in the customer portrait into feature vectors. For example, the geographic location information can be converted into longitude and latitude coordinates and the distance from a specific economic center point, and the behavioral characteristics can include transaction amount, transaction frequency, shopping preferences, visit frequency, etc.; Input the feature vector into the machine learning model for prediction to obtain the sub-region.

[0061] In order to gradually and deeply understand the economic environment within its operating area from macro to micro, at least one sub-region corresponding to the operating area can be determined in the following manner: dividing the operating area into at least one first sub-region in the first administrative region, wherein the first sub-region is a geographical region whose economic activity intensity and / or industry distribution characteristics meet the first preset condition; dividing the first sub-region into at least one second sub-region in the second administrative region, wherein the second sub-region is a geographical region whose economic activity intensity and / or industry distribution characteristics meet the second preset condition, the indicator threshold in the second preset condition is greater than the corresponding indicator threshold in the first preset condition, and the administrative level of the second administrative region is lower than the administrative level of the first administrative region; dividing the second sub-region into at least one third sub-region in the third administrative region, wherein the third sub-region is a geographical region corresponding to the transportation logistics and / or production supply chain in the second sub-region, and the administrative level of the third administrative region is lower than the administrative level of the second administrative region.

[0062] The first sub-region is a geographical area divided within the highest-level administrative region (such as a province) according to the intensity of economic activities (such as total GDP, per capita income) and industry distribution characteristics (such as the proportion of the service industry and the proportion of heavy industry). These areas meet the first preset condition, that is, they have reached a certain economic scale and industry representativeness.

[0063] In some embodiments of the present application, the following steps can be used to divide the operating area into at least one first sub-area within the first administrative area. Specifically, multiple target indicators corresponding to the original data are determined, wherein the target indicators are used to quantitatively represent the industry attributes and economic activity within the operating area; the original data are converted into data points on a map of the operating area based on the target indicators, wherein the data points include the target indicators and geographic location information; and the data points are clustered to obtain the first sub-area.

[0064] The target indicators are used to quantify the values ​​of industry attributes and economic activity in the operating area, such as GDP, industry output value, number of enterprises, frequency of logistics activities, etc. In some embodiments of the present application, preliminary indicators reflecting industry attributes and economic activity can be determined in combination with industry standards and economic theories. For example, for village banks, agricultural output value, agricultural product trading volume, number of small enterprises, per capita income, etc. can be considered as preliminary indicators; based on the preliminary determined target indicators, through data analysis and machine learning techniques (such as principal component analysis PCA or feature selection algorithm), the indicator combination is optimized to screen out the target indicators most relevant to industry attributes and economic activity.

[0065] It should be noted that the random forest algorithm can be used to dynamically calculate the weight of each target indicator to ensure the effective integration and dynamic adjustment of different industry attributes and economic activity indicators.

[0066] After determining the target indicator, the geographic information system (GIS) technology can be used to convert the unstructured geographic location information (such as address, longitude and latitude) in the original data into standardized geographic coordinates, which together with the target indicator form data points. It should be noted that not all data in the original data have geographic location information. In this case, the non-geographic location information in the original data can be converted into attributes or indicators related to the geographic location information. For example, for economic indicators such as corporate transaction volume and agricultural product output value in the original data, virtual "economic points" or "industry points" can be created based on the correlation between these indicators and specific geographic locations. For example, the values ​​of these indicators can be bound to geographic location information (such as the longitude and latitude of the location of the enterprise and the origin of agricultural products) to form data points.

[0067] In some embodiments of the present application, a default geographical location can be set for each economic indicator or industry attribute, which may be the most common location of the indicator or attribute, or the geographical location most directly related to it. For example, for the indicator of agricultural output value, it can be bound to the geographical coordinates of each village and town in the operating area of ​​the village bank; for the enterprise transaction volume, it can be bound according to the coordinates of the enterprise's registered address or main business location.

[0068] In addition, data analysis or machine learning techniques can also be used to explore the potential correlation between non-geographic information and geographic information, thereby inferring the "virtual location" of non-geographic information. For example, through correlation analysis, it is found that the transaction volume of an enterprise is related to the density of logistics facilities in the area. Then, enterprises with high transaction volumes can be associated with geographical areas with high-density logistics facilities, thereby representing these enterprises on the map.

[0069] Specifically, various types of data are collected and organized, including economic indicators (such as corporate transaction volume, agricultural product output value), industry attributes (such as energy industry, agriculture), and corresponding geographic location information (such as corporate registration location, agricultural product origin); correlation analysis methods (such as Pearson correlation coefficient, Spearman rank correlation coefficient) or supervised learning models (such as regression analysis, deep learning models) are used to establish the connection between non-geographic location information and geographic location information; based on this connection, a reasonable "virtual location" is assigned to the original data without direct geographic location information, and represented as data points on the map together with the actual geographic location information.

[0070] After obtaining the data points, density-based clustering algorithms (such as DBSCAN or HDBSCAN) can be used to automatically identify cluster boundaries based on the distance and density between data points to form the first sub-region. In order to combine geographic boundaries and hierarchical clustering to ensure that the clustering results reflect the similarity of economic activities and respect the division of administrative regions and avoid clustering results across administrative boundaries, hierarchical clustering algorithms (such as Agglomerative Clustering) can also be used to combine administrative and geographic boundary information of the operating area for cluster analysis. This method first treats each data point as an independent cluster, then gradually merges the most similar clusters to form a dendrogram, and finally divides the first sub-region based on the consistency of the clustering results with the geographic boundaries.

[0071] The second sub-region is a further subdivision of the first sub-region within the scope of a lower-level administrative region (such as a city or county), screening out regions with higher economic activity intensity and more significant industry distribution characteristics. The third sub-region is a geographical region divided based on the convenience of transportation and logistics and the continuity of the production supply chain in the second sub-region within the scope of a more basic administrative division (such as a town or village).

[0072] In the process of gradually subdividing the business area into the first sub-area, the second sub-area and the third sub-area, the selected clustering algorithm will vary according to the division target and data characteristics. Specifically:

[0073] (1) The division of the first sub-region aims to identify geographical areas whose economic activity intensity and industry distribution characteristics meet the first preset conditions from a macro perspective. Usually, the geographical scope involved is wide, the amount of data is large, and the economic activities and industry distribution have obvious spatial clustering. At this level, density-based clustering algorithms can be used, such as DBSCAN (Density-Based Spatial Clustering of Applications with Noise) or HDBSCAN (Hierarchical DBSCAN).

[0074] (2) The second sub-region is based on the first sub-region, further screening out geographical areas with higher economic activity intensity and more significant industry distribution characteristics. At this time, the indicator threshold is more stringent. Considering that the division of the second sub-region requires more precise clustering results to ensure the accuracy and effectiveness of high-value market segmentation, more complex clustering algorithms can be used, such as GMM (Gaussian Mixture Model) or Spectral Clustering.

[0075] (3) The division of the third sub-region is based on the characteristics of transportation logistics and production supply chain in the second sub-region. The division of sub-regions at this level focuses more on industry relevance and the continuity of logistics activities. Therefore, algorithms with specific industry knowledge and network structure recognition capabilities can be used, such as Community Detection algorithm or Network Analysis.

[0076] In order to ensure that the recommendation strategy of financial institutions has both a global perspective and takes into account local characteristics, the following methods can be used to map the portraits to the sub-regions corresponding to the operating areas: using a mapping method from high to low administrative levels, determine in turn the first sub-region, the second sub-region and the third sub-region to which the portraits correspond; or using a mapping method from low to high administrative levels, determine in turn the third sub-region, the second sub-region and the first sub-region to which the portraits correspond.

[0077] The process of mapping the portrait to the sub-region is essentially the process of matching the characteristics of the target object with the economic characteristics of a specific geographical area. The use of a high-to-low administrative level mapping method can gradually refine from a macro perspective to specific communities or streets, which is conducive to financial institutions to grasp the overall situation and local details and formulate a comprehensive recommendation plan. Conversely, the low-to-high mapping method starts from the grassroots level and gradually expands the perspective. Such a strategy may be more suitable for financial institutions that have already identified specific market segments. They can first focus on a small range of high-potential markets and then gradually expand to larger areas.

[0078] When mapping from high to low administrative levels, the customer's profile data can be matched with the highest-level business area (such as a province) to identify the first sub-area that best matches the customer's profile; within the first sub-area, the profile data is further matched to determine the second sub-area; within the second sub-area, the profile data is again used to determine the third sub-area with the finest granularity. In the case of a large amount of data, the high-to-low mapping method is used to first match on a large scale, gradually narrowing the search range, reducing computational complexity, and improving matching efficiency.

[0079] When mapping from low to high administrative levels, we can start from the third sub-region with the finest granularity, and gradually expand to the second and first sub-regions based on the similarity matching between the customer profile and other customers in the region, and finally determine the most relevant business sub-region for the customer. This approach can make full use of the detailed information in the customer profile, starting from the geographical environment closest to the customer, and gradually comparing it with the characteristics of the wider area, to ensure that the matched sub-regions more accurately reflect the specific needs of the customer and the market environment.

[0080] In order to solve the problem that the customer portrait information in low-level areas is too specific and it is difficult to find matching areas, a layer-by-layer abstraction and fusion strategy can be used in the mapping process from low to high. That is, more specific portrait features are considered when matching the third sub-area, such as the customer's specific needs and behavior patterns, and when entering the second and first sub-areas, more generalized features are gradually integrated, such as industry categories, economic indicators, etc., to ensure that areas matching the customer portrait can be found in sub-areas of different administrative levels.

[0081] Using GIS technology, the mapped customer portrait data can be converted into a visual display, and the distribution of customer portraits in sub-regions at different administrative levels can be displayed according to demand. For example, financial institutions can start with high-level sub-regions at the first administrative level (such as provincial or prefectural-level cities) and use heat maps to understand the customer distribution profile of the entire operating area; further, they can also refine to sub-regions at the second administrative level (such as county or township level) to conduct a deeper analysis of customer portraits. For example, the distribution of customer portraits in a specific county or township can be displayed through a heat map, and the differences in customer portraits and demand hotspots in different villages and towns can be identified; further, the perspective can be focused on sub-regions at the third administrative level (such as the most basic villages or communities), and the distribution of customer portraits in each sub-region can be understood in detail through heat maps.

[0082] Based on the distribution of portraits in different sub-regions (such as the first sub-region, the second sub-region or the third sub-region), cluster analysis can be performed in a specific sub-region according to the customer's choice, or the corresponding sub-region can be automatically clustered when the portrait distribution meets the preset conditions. It should be noted that the preset conditions for cluster analysis of sub-regions corresponding to different administrative levels are different, and the determination of the preset conditions can be based on factors such as the number of portraits in the sub-region, transaction activity, and geographical location.

[0083] Specifically, for sub-regions of any administrative level, a minimum threshold for the number of portraits (i.e., the number of customers) can be set before cluster analysis to ensure that there is enough customer data in the analyzed sub-region to obtain meaningful clustering results. For example, the first sub-region (such as a provincial or prefectural-level city) may require a higher customer number threshold (such as more than 500) because such sub-regions have a wide coverage area and a large customer base, and a smaller customer base may not be sufficient to reflect general trends. The third sub-region (such as villages and towns) may require a lower customer number threshold (such as more than 20) because the total number of customers in villages and towns is relatively small, and even a small number of customers may form specific clustering characteristics. The preset conditions for the second sub-region are between the first and third sub-regions.

[0084] In addition to the number of portraits, customer transaction activity within a sub-region can also be used as a reference indicator. A threshold for transaction volume or transaction frequency can be set, and cluster analysis can only be performed on sub-regions that meet the preset activity standards. For example, the first sub-region requires a higher average transaction volume (such as more than 500,000 transactions per month), while the third sub-region (towns and villages) only needs to reach a lower average transaction volume (such as more than 500 transactions per month) for cluster analysis. This ensures that the analysis focuses on the real business hotspots, thereby optimizing services in a more targeted manner. The preset conditions for the second sub-region are between the first and third sub-regions.

[0085] Different cluster analysis conditions can also be set according to the geographical location and scope of the sub-region. For example, for the first sub-region (province or city), financial institutions need to pay attention to cross-regional liquidity, so the focus of cluster analysis can be placed on sub-regions that are geographically close to each other but have significantly different customer profile characteristics to explore cross-regional recommendation and service opportunities. For the third sub-region (village and town), more attention can be paid to the similarity of customers within the sub-region. The preset conditions can be areas with relatively concentrated geographical locations and similar customer profile characteristics, so as to facilitate centralized recommendation activities or service optimization. The preset conditions for the second sub-region are between the first and third sub-regions, for example, areas with relatively concentrated geographical locations but with certain cross-liquidity.

[0086] Step S208, performing cluster analysis on the portraits in the sub-regions, and determining a recommendation strategy corresponding to the clustering results.

[0087] In the above step S208, a deep learning model, such as an autoencoder or a generative adversarial network (GAN), can be used to reduce the dimension and extract features of the customer portrait data, and then perform clustering analysis based on the extracted feature vectors, such as using a K-means or DBSCAN algorithm.

[0088] In sub-regions, customer characteristics are often closely related to geographic location. Therefore, geographic information can be added to customer profile data as an additional feature, and clustering analysis can be performed using spatial clustering algorithms (such as ST-DBSCAN) or Geographically Weighted Clustering (GWC). ST-DBSCAN takes into account the proximity of spatial and temporal dimensions, while GWC dynamically adjusts clustering parameters based on geographic weights to make it more suitable for customer characteristics in a specific geographic location.

[0089] In a sub-region, if the customer portrait clustering results are too detailed, the cost of customizing the recommendation strategy for each cluster may be too high; on the contrary, if the clustering results are too extensive, they may not effectively meet the needs of different customer groups, affecting the recommendation effect. To solve this problem, adaptive clustering algorithms can be used, such as hierarchical clustering combined with a strategy optimization method based on cost-benefit analysis. Hierarchical clustering can generate a clustering structure tree, gradually merging from the finest clusters to broader groups. Business personnel can select the most appropriate clustering level based on the cost-benefit analysis results to achieve adaptive optimization of the recommendation strategy.

[0090] In order to reduce the cost and time of customizing recommendation strategies, a set of recommendation strategy templates can be predefined for clustering results at different levels, including product recommendations, preferential strategies, recommendation channel selection, etc.; based on the features extracted from the clustering analysis, such as customer age, occupation, income level, etc., the clustering results are mapped to the corresponding strategy templates to achieve rapid customization and execution of recommendation strategies.

[0091] In order to ensure that the clustering results reflect both the similarity of customer portraits and the differences in environmental characteristics of sub-regions, the portraits can be clustered in the sub-regions through the following steps: determine the third data corresponding to the sub-region from the second data, wherein the third data is used to reflect the environmental characteristics of the sub-region; determine the clustering algorithm corresponding to the third data and the target parameters corresponding to the clustering algorithm, wherein the target parameters are determined based on the distribution characteristics of all portraits in the sub-region; use the clustering algorithm to perform cluster analysis on all portraits in the sub-region to obtain clustering results, wherein the clustering results include multiple groups, each group representing a set of target objects whose similarity is greater than or equal to a preset threshold.

[0092] Specifically, we can first use a clustering method based on multi-dimensional feature fusion to combine customer portrait data with third data reflecting the environmental characteristics of the sub-region to form a comprehensive feature set and conduct clustering analysis; dynamically adjust the target parameters of the clustering algorithm and continuously optimize the clustering results until the preset conditions are met, such as satisfying the similarity requirements of customer portraits while taking into account the differences in environmental characteristics of the sub-regions; based on the final clustering results, combined with the specific market environment and socio-economic conditions of the sub-region, formulate and implement customized recommendation strategies. For example, for sub-regions with higher demand for agricultural credit, farmer customers with similar portraits can be clustered, and environmental characteristics such as soil conditions and irrigation facilities in the sub-region can be considered to form a special agricultural credit recommendation strategy.

[0093] In some embodiments of the present application, a recommendation strategy corresponding to the clustering result can be determined by the following steps: determining a target feature corresponding to a target group, wherein the target feature is determined based on a feature of a portrait of the target group, and the target group is any one of multiple groups; and determining a recommendation strategy based on the target feature and third data.

[0094] The target group is a specific customer group in the cluster analysis results, which has similar image characteristics. The identification of the target group is the core of accurate recommendation. By clustering the customers according to similarity, the bank can adopt differentiated recommendation strategies for customer groups with different needs. The target characteristics are key indicators that reflect the image characteristics and needs of the target group, and are used to guide the formulation of recommendation strategies.

[0095] In order to ensure that the design of the recommendation strategy is highly relevant to the portrait characteristics of the target group, for each target group, the key features of all portraits in the group (i.e., target features) can be extracted, such as age, income level, occupation, credit rating, etc.; the potential needs and preferences of the target group can be analyzed based on the target characteristics, and the target needs of the target group in the context of a specific sub-region can be analyzed in combination with the environmental characteristics of the sub-region (such as the intensity of economic activity, industry distribution, etc.); recommendation strategies corresponding to the target needs can be formulated, such as launching specific loan products, designing personalized financial management plans, optimizing the layout of recommendation channels, etc.

[0096] In a specific embodiment, a financial institution identified a specific target group consisting of young farmers after cluster analysis. The customer profile characteristics of this group show that they have a high rate of digital device usage, prefer online financial services, have a large demand for microcredit, and are highly sensitive to agricultural technology information.

[0097] (1) Determine target characteristics: Target characteristics may include high use of digital devices, preference for digital services, demand for microcredit, and attention to agricultural technology information;

[0098] (2) Environmental characteristics: Based on the third-party data, analyze the current agricultural economic status of the sub-region, such as fluctuations in agricultural product prices, the degree of agricultural technology promotion, and the penetration of rural Internet;

[0099] (3) Recommended strategy design, for example, including:

[0100] 1) Product design: Launch digital small agricultural loan products and provide a fast application and approval process through a mobile APP.

[0101] 2) Pricing strategy: Based on local agricultural product price fluctuations and farmers’ income, design flexible repayment methods and loan interest rates linked to market interest rates.

[0102] 3) Channel selection: Prioritize pushing loan product information through mobile Internet channels, and use social media and agricultural information platforms for in-depth recommendations.

[0103] 4) Promotional activities: Organize online agricultural technology training and loan application guidance, provide online consulting services for the initial review of loan applications, and provide interest rate discounts to farmers who apply for loans for the first time using digital services.

[0104] 5) Communication methods: Maintain high-frequency online communication with the target group through group text messages, in-APP message push, social media interaction, etc., and provide instant loan information and technical support.

[0105] Through the above steps S202 to S208, by obtaining the original data corresponding to the operating area of ​​the financial institution, and constructing the knowledge graph corresponding to the financial institution based on the original data, and then determining the portrait corresponding to the target object based on the knowledge graph, after mapping the portrait to the sub-region corresponding to the operating area, the portrait is clustered in the sub-region, and finally the recommendation strategy corresponding to the clustering result is determined, thereby achieving the purpose of clustering the target objects in specific sub-regions according to the economic characteristics of different geographical regions and determining the most suitable recommendation strategy, thereby achieving the technical effect of improving the pertinence of recommendation activities and dynamically adjusting strategies according to environmental changes, and further solving the technical problem that the related technology formulates recommendation strategies based on general static templates, and it is difficult to conduct differentiated analysis on different regions where different financial institutions are located, resulting in limited recommendation activities.

[0106] Figure 3 is a structural diagram of a device for determining a recommendation strategy based on data mining and knowledge graph according to an embodiment of the present application, such as Figure 3 As shown, the device comprises:

[0107] An acquisition module 302, used to acquire original data corresponding to the operating area of ​​the financial institution;

[0108] Determine 304, for constructing a knowledge graph corresponding to the financial institution based on the original data, and determining a portrait corresponding to the target object based on the knowledge graph;

[0109] A mapping module 306, for mapping the portrait to a sub-area corresponding to the business area;

[0110] The clustering module 308 is used to perform cluster analysis on the portraits in the sub-regions and determine a recommendation strategy corresponding to the clustering results.

[0111] It should be noted that Figure 3 The device for determining the recommendation strategy based on data mining and knowledge graph is used to perform Figure 2 The method for determining the recommendation strategy based on data mining and knowledge graph is shown in Figure 2 The explanations in the method for determining the recommendation strategy based on data mining and knowledge graph in the article also apply to Figure 3 The device for determining the recommendation strategy based on data mining and knowledge graph shown will not be described in detail here.

[0112] An embodiment of the present application also provides an electronic device, which includes a memory and a processor, wherein the memory is used to store program instructions; the processor is connected to the memory and is used to execute the steps of a method for determining a recommendation strategy based on data mining and knowledge graphs in each embodiment of the present application.

[0113] For example, the processor performs the following functions by executing program instructions stored in the memory: obtaining original data corresponding to the operating area of ​​the financial institution, wherein the original data includes first data corresponding to a target object within the operating area and second data reflecting environmental characteristics of the operating area; constructing a knowledge graph corresponding to the financial institution based on the original data, and determining a portrait corresponding to the target object based on the knowledge graph; mapping the portrait to a sub-region corresponding to the operating area, wherein the operating area corresponds to at least one sub-region and each sub-region corresponds to different economic characteristics; performing cluster analysis on the portrait in the sub-region, and determining a recommendation strategy corresponding to the clustering result.

[0114] An embodiment of the present application also provides a non-volatile storage medium, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the steps of the method for determining a recommendation strategy based on data mining and knowledge graph in each embodiment of the present application by running the computer program.

[0115] An embodiment of the present application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the method for determining a recommendation strategy based on data mining and knowledge graph in each embodiment of the present application.

[0116] An embodiment of the present application also provides a computer program, which, when executed by a processor, implements the steps of a method for determining a recommendation strategy based on data mining and knowledge graphs in each embodiment of the present application.

[0117] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0118] In the above embodiments of the present application, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0119] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0120] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0121] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0122] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, disk or optical disk, etc., which can store program code.

[0123] The above is only a preferred implementation of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A method for determining a recommendation strategy based on data mining and knowledge graph, characterized in that: include: Acquire original data corresponding to the operating area of ​​the financial institution, wherein the original data includes first data corresponding to a target object in the operating area and second data reflecting environmental characteristics of the operating area; Constructing a knowledge graph corresponding to the financial institution based on the original data, and determining a portrait corresponding to the target object based on the knowledge graph; Mapping the portrait to a sub-region corresponding to the business region, wherein the business region corresponds to at least one sub-region, and each sub-region corresponds to different economic characteristics; Cluster analysis is performed on the portraits in the sub-areas, and a recommendation strategy corresponding to the clustering results is determined.

2. The method according to claim 1, characterized in that: At least one sub-region corresponding to the business region is determined in the following manner: Dividing the operating area into at least one first sub-area within the first administrative area, wherein the first sub-area is a geographical area whose economic activity intensity and / or industry distribution characteristics meet a first preset condition; Divide the first sub-region into at least one second sub-region within the second administrative region, wherein the second sub-region is a geographical region in which the economic activity intensity and / or industry distribution characteristics within the first sub-region meet a second preset condition, the index threshold in the second preset condition is greater than the corresponding index threshold in the first preset condition, and the administrative level of the second administrative region is lower than the administrative level of the first administrative region; The second sub-region is divided into at least one third sub-region within the third administrative region, wherein the third sub-region is a geographical region corresponding to the transportation logistics and / or production supply chain in the second sub-region, and the administrative level of the third administrative region is lower than that of the second administrative region.

3. The method according to claim 2, characterized in that Dividing the business area into at least one first sub-area within the first administrative area includes: Determine a plurality of target indicators corresponding to the original data, wherein the target indicators are used to quantitatively represent the industry attributes and economic activity in the business area; Converting the raw data into data points on a map of the business area according to the target indicator, wherein the data points include the target indicator and geographic location information; Perform cluster analysis on the data points to obtain the first sub-region.

4. The method according to claim 2, characterized in that: Mapping the portrait to a sub-area corresponding to the business area includes: Adopting a mapping method from high to low administrative levels, determining the first sub-region, the second sub-region and the third sub-region corresponding to the portraits in sequence; or, A mapping method from low to high administrative levels is adopted to determine in sequence the third sub-region, the second sub-region and the first sub-region corresponding to the portraits respectively.

5. The method according to claim 1, characterized in that: Performing cluster analysis on the portrait in the sub-area includes: Determining third data corresponding to the sub-area from the second data, wherein the third data is used to reflect environmental characteristics of the sub-area; Determining a clustering algorithm corresponding to the third data and a target parameter corresponding to the clustering algorithm, wherein the target parameter is determined based on distribution characteristics of all portraits in the sub-region; The clustering algorithm is used to perform cluster analysis on all the portraits in the sub-region to obtain a clustering result, wherein the clustering result includes multiple groups, each group representing a set of target objects whose similarity is greater than or equal to a preset threshold.

6. The method according to claim 5, characterized in that Determine the recommended strategy corresponding to the clustering results, including: Determine a target feature corresponding to a target group, wherein the target feature is determined based on a feature of a portrait of the target group, and the target group is any one of the multiple groups; The recommendation strategy is determined according to the target feature and the third data.

7. The method according to claim 1, characterized in that Determining a portrait corresponding to the target object according to the knowledge graph includes: Acquire static attributes of the target object from the knowledge graph, wherein the static attributes include information corresponding to the first data of the target object; Determining a dynamic attribute of the target from the knowledge graph, wherein the dynamic attribute includes information associated with the target object in the second data; Fusing the static attribute with the dynamic attribute to obtain a joint representation of the target object; A portrait corresponding to the target object is determined based on the joint representation.

8. A device for determining a recommendation strategy based on data mining and knowledge graph, characterized in that: include: An acquisition module, used to acquire original data corresponding to the operating area of ​​the financial institution, wherein the original data includes first data corresponding to the target object in the operating area and second data reflecting the environmental characteristics of the operating area; A determination module, configured to construct a knowledge graph corresponding to the financial institution based on the original data, and determine a portrait corresponding to the target object based on the knowledge graph; A mapping module, used for mapping the portrait to a sub-region corresponding to the business region, wherein the business region corresponds to at least one sub-region, and each sub-region corresponds to different economic characteristics; The clustering module is used to perform cluster analysis on the portraits in the sub-areas and determine a recommendation strategy corresponding to the clustering results.

9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory is used to store program instructions; the processor is connected to the memory and is used to execute the method for determining a recommendation strategy based on data mining and knowledge graphs as described in any one of claims 1 to 7.

10. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the method for determining a recommendation strategy based on data mining and knowledge graph as described in any one of claims 1 to 7 by running the computer program.

11. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by the processor, the method for determining the recommendation strategy based on data mining and knowledge graph as described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Bank financial marketing recommendation method and device

    CN112767144A

  • High-level talent recommendation method and system based on multi-dimensional portraits and medium

    CN116756409A

  • User portrait generation query method based on knowledge graph

    CN119149755A

  • Client portrait construction method and device, storage medium and electronic equipment

    CN119336977A

  • Methods and systems for harnessing location based data for making market recommendations

    US20210192553A1

Cited By

  • Enterprise carbon emission intensity calculation method

    CN120471308A