Method and device for determining a recommendation strategy based on data mining and knowledge graph
By constructing a knowledge graph of financial institutions and conducting sub-region cluster analysis, the problem of insufficient regional differences in the recommendation strategies of traditional financial institutions was solved, and more targeted recommendation strategy adjustments were achieved.
Patent Information
- Application Number
- CN202510499414.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-04-21
AI Technical Summary
The recommendation strategies of traditional financial institutions rely on static templates and fail to take into account the differences in economic characteristics and market environments in different regions, resulting in limited targeting and effectiveness of recommendation activities.
By obtaining the original data of the operating areas of financial institutions, building a knowledge graph, mapping it to sub-areas and performing cluster analysis, we can determine differentiated recommendation strategies.
It realizes cluster analysis of target objects in specific sub-regions based on the economic characteristics of different geographical regions, dynamically adjusts strategies, and improves the pertinence and effectiveness of recommendation activities.
Smart Images

Figure CN120013649B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a method and device for determining a recommendation strategy based on data mining and a knowledge graph. BACKGROUND
[0002] In the formulation of recommendation strategies for financial institutions, the traditional approach often relies on pre-defined static templates or rules. These templates are designed based on industry average levels or market data from a specific region, ignoring the specific economic characteristics and market environment differences of the operating region of the financial institution. For example, the templates may preset some general customer stratification standards, product promotion strategies or recommendation channel selection, but these standards do not take into account unique factors such as the economic development level, industry distribution, customer behavior preferences of different regions, resulting in limited targeting and effectiveness of the recommendation activities.
[0003] To address the above problems, no effective solutions have been proposed so far. SUMMARY
[0004] The embodiments of the present application provide a method and device for determining a recommendation strategy based on data mining and a knowledge graph, to at least solve the technical problem that related technologies formulate recommendation strategies based on general static templates, making it difficult to conduct differentiated analysis of different regions where different financial institutions are located, resulting in limited recommendation activities.
[0005] According to an aspect of an embodiment of the present application, a method for determining a recommendation strategy based on data mining and a knowledge graph is provided, comprising: obtaining original data corresponding to an operating region of a financial institution, wherein the original data includes first data corresponding to a target object in the operating region and second data reflecting environmental characteristics of the operating region; constructing a knowledge graph corresponding to the financial institution according to the original data, and determining a portrait corresponding to the target object according to the knowledge graph; mapping the portrait to a sub-region corresponding to the operating region, wherein the operating region corresponds to at least one sub-region, and each sub-region corresponds to different economic characteristics; performing clustering analysis on the portrait in the sub-region, and determining a recommendation strategy corresponding to the clustering result.
[0006] In some embodiments of the present application, the at least one sub-region corresponding to the operating area is determined by: dividing the operating area into at least one first sub-region in a first administrative region, wherein the first sub-region is a geographic region with economic activity intensity and / or industry distribution characteristics satisfying a first preset condition; dividing the first sub-region into at least one second sub-region in a second administrative region, wherein the second sub-region is a geographic region in the first sub-region with economic activity intensity and / or industry distribution characteristics satisfying a second preset condition, an index threshold in the second preset condition is greater than a corresponding index threshold in the first preset condition, and an administrative level of the second administrative region is less than an administrative level of the first administrative region; and dividing the second sub-region into at least one third sub-region in a third administrative region, wherein the third sub-region is a geographic region corresponding to traffic logistics and / or production supply chain in the second sub-region, and an administrative level of the third administrative region is less than an administrative level of the second administrative region.
[0007] In some embodiments of the present application, the operating area is divided into at least one first sub-region in a first administrative region, comprising: determining a plurality of target indicators corresponding to the raw data, wherein the target indicators are used to quantify the industry attributes and economic activity in the operating area; converting the raw data into data points on a map of the operating area according to the target indicators, wherein the data points include target indicators and geographic location information; and performing cluster analysis on the data points to obtain the first sub-region.
[0008] In some embodiments of the present application, the portrait is mapped to the sub-region corresponding to the operating area, comprising: determining the first sub-region, the second sub-region and the third sub-region corresponding to the portrait in turn by using a mapping method of administrative levels from high to low; or determining the third sub-region, the second sub-region and the first sub-region corresponding to the portrait in turn by using a mapping method of administrative levels from low to high.
[0009] In some embodiments of the present application, the cluster analysis of the portrait in the sub-region comprises: determining third data corresponding to the sub-region from the second data, wherein the third data is used to reflect the environmental characteristics of the sub-region; determining a clustering algorithm corresponding to the third data and a target parameter corresponding to the clustering algorithm, wherein the target parameter is determined based on the distribution characteristics of all portraits in the sub-region; and performing cluster analysis on all portraits in the sub-region by using the clustering algorithm to obtain a clustering result, wherein the clustering result includes a plurality of groups, and each group represents a target object set with a similarity greater than or equal to a preset threshold.
[0010] In some embodiments of the present application, the recommended strategy corresponding to the clustering result is determined by: determining a target feature corresponding to a target group, wherein the target feature is determined based on the features of the portraits of the target group, and the target group is any one of the plurality of groups; and determining the recommended strategy according to the target feature and the third data.
[0011] In some embodiments of the present application, the portrait corresponding to the target object is determined according to the knowledge graph, including: obtaining the static attribute of the target object from the knowledge graph, wherein the static attribute includes information corresponding to the first data of the target object; determining the dynamic attribute of the target from the knowledge graph, wherein the dynamic attribute includes information associated with the target object in the second data; fusing the static attribute and the dynamic attribute to obtain the joint representation of the target object; and determining the portrait corresponding to the target object according to the joint representation.
[0012] According to another aspect of the embodiments of the present application, a device for determining a recommendation strategy based on data mining and a knowledge graph is also provided, including: an acquisition module configured to acquire original data corresponding to an operating area of a financial institution; a determination module configured to construct a knowledge graph corresponding to the financial institution according to the original data, and determine a portrait corresponding to a target object according to the knowledge graph; a mapping module configured to map the portrait to a sub-area corresponding to the operating area; and a clustering module configured to perform clustering analysis on the portrait in the sub-area, and determine a recommendation strategy corresponding to the clustering result.
[0013] According to still another aspect of the embodiments of the present application, an electronic device is also provided, including: a memory and a processor, the memory is configured to store program instructions; the processor is connected with the memory, and is configured to execute the above-mentioned method for determining a recommendation strategy based on data mining and a knowledge graph.
[0014] According to still another aspect of the embodiments of the present application, a non-volatile storage medium is also provided, including a stored computer program, wherein a device in which the non-volatile storage medium is located executes the above-mentioned method for determining a recommendation strategy based on data mining and a knowledge graph by running the computer program.
[0015] According to still another aspect of the embodiments of the present application, a computer program product is also provided, including computer instructions, which, when executed by a processor, implement the above-mentioned method for determining a recommendation strategy based on data mining and a knowledge graph.
[0016] In an embodiment of the present application, by obtaining original data corresponding to the operating area of the financial institution, and constructing a knowledge graph corresponding to the financial institution based on the original data, and then determining the portrait corresponding to the target object based on the knowledge graph, and mapping the portrait to the sub-area corresponding to the operating area, the portrait is clustered in the sub-area, and finally the recommendation strategy corresponding to the clustering result is determined. The purpose of clustering the target objects in specific sub-areas based on the economic characteristics of different geographical areas and determining the most suitable recommendation strategy is achieved, thereby achieving the technical effect of improving the pertinence of recommendation activities and dynamically adjusting strategies according to environmental changes, and further solving the technical problem that the related technology formulates recommendation strategies based on general static templates, and it is difficult to conduct differentiated analysis of different areas where different financial institutions are located, resulting in limited recommendation activities. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0018] Figure 1 This is a hardware structure block diagram of a computer terminal according to a method for determining a recommendation strategy based on data mining and knowledge graph according to an embodiment of the present application;
[0019] Figure 2 This is a flowchart of a method for determining a recommendation strategy based on data mining and knowledge graph according to an embodiment of the present application;
[0020] Figure 3 It is a structural diagram of a device for determining a recommendation strategy based on data mining and knowledge graph according to an embodiment of the present application. DETAILED DESCRIPTION
[0021] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0023] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:
[0024] ST-DBSCAN (Spatio-Temporal DBSCAN): A space-time clustering algorithm that extends the traditional DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm. In this application, ST-DBSCAN not only considers the spatial density of data points but also incorporates the temporal dimension, enabling clustering results to reflect the distribution patterns of data points in space and time, useful for customer segmentation.
[0025] Geographic Information System (GIS): A technical system used to collect, store, process, analyze, and display data related to geographic location. In this application, GIS is used to analyze customer location information, industry distribution, and economic activity to support the formulation and implementation of location-based recommendation strategies.
[0026] The recommendation strategy formulation methods used by related technologies struggle to overcome the limitations of static templates, leading to problems such as poor regional adaptability and insufficient customer insights. For example, a rural bank's operating area may be primarily agricultural, while an urban commercial bank's operating area may be primarily service or high-tech. Universal recommendation templates cannot adapt to these differences, resulting in poor recommendation effectiveness in certain regions. Furthermore, the behavioral patterns and needs of target customer groups vary significantly across regions. Static templates often segment customers based on their basic attributes (such as age, gender, and occupation), failing to capture their dynamic attributes, such as changes in transaction behavior, responsiveness to market trends, and sensitivity to policy impacts. This makes it difficult for financial institutions to grasp their customers' real-time needs and market dynamics, limiting the relevance and effectiveness of recommendation activities.
[0027] In order to solve the above technical problems, the embodiments of the present application provide corresponding solutions, which are described in detail below.
[0028] The method for determining a recommendation strategy based on data mining and knowledge graph provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The hardware structure block diagram of a computer terminal for implementing a method for determining a recommendation strategy based on data mining and knowledge graph is shown. Figure 1 As shown, the computer terminal 10 may include one or more processors (illustrated as 102a, 102b, ..., 102n in the figure) (the processor may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions via a wired and / or wireless network connection. In addition, it may also include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, and a BUS bus. Those skilled in the art will understand that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0029] It should be noted that the one or more processors and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be fully or partially integrated into any of the other components of the computer terminal 10. As discussed in the embodiments of the present application, the data processing circuitry functions as a processor control (e.g., the selection of a variable resistor terminal path connected to an interface).
[0030] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the method for determining the recommendation strategy based on data mining and knowledge graph in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizing the above-mentioned method for determining the recommendation strategy based on data mining and knowledge graph. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0031] The transmission module 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission module 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission module 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0032] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .
[0033] It should be noted that, in some optional embodiments, the above Figure 1 The computer terminal shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the computer terminal described above.
[0034] In the above-mentioned operating environment, an embodiment of the present application provides an embodiment of a method for determining a recommendation strategy based on data mining and knowledge graph. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0035] Figure 2is a flow chart of a method for determining a recommendation strategy based on data mining and knowledge graph according to an embodiment of the present application, such as Figure 2 As shown, the method includes the following steps:
[0036] Step S202 : acquiring original data corresponding to the operating area of the financial institution, wherein the original data includes first data corresponding to the target object in the operating area and second data reflecting the environmental characteristics of the operating area.
[0037] In the above step S202, the operating area of a financial institution refers to the specific geographical scope of the financial institution's business coverage and services, which can be areas at different levels such as cities, counties, townships, and towns.
[0038] Raw data refers to data that is initially collected without being processed or analyzed, including data from inside and outside financial institutions. For example, internal data may include customer information, transaction records, credit history, etc., while external data may include macroeconomic data, industry reports, geographic information, social media data, etc.
[0039] The primary data corresponding to the target object refers to information directly related to the financial institution's target customers, such as basic information of individual customers (age, gender, occupation), corporate customers' business qualifications (registered capital, industry type), customer transaction history, credit score, etc. Secondary data reflecting the characteristics of the operating area environment may include but is not limited to economic indicators (such as GDP, industrial distribution), population structure, geographic information (such as transportation network, infrastructure), and policy environment (such as local tax policies and financial policies).
[0040] It should be noted that environmental characteristics refer to a series of external factors and conditions related to the operating area of a financial institution. These factors and conditions have a significant impact on the financial institution's business operations, customer behavior and market opportunities, and may include but are not limited to:
[0041] (1) Macroeconomic indicators: including GDP growth rate, unemployment rate, inflation rate, interest rate level, etc. These indicators reflect the economic health and business environment of the operating area.
[0042] (2) Industry distribution and characteristics: Different operating regions may have unique industrial layouts. For example, the region where a rural bank is located may be dominated by agriculture and small-scale manufacturing, while urban financial institutions may serve high-tech, real estate or service industries. Industry distribution helps financial institutions identify target customer groups and design products and services that meet industry characteristics.
[0043] (3) Population structure and consumption behavior: including data on age distribution, gender ratio, education level, income level, consumption preferences, etc., used to understand the needs and preferences of customer groups and formulate personalized recommendation strategies.
[0044] (4) Geographic and infrastructure information: the geographic location of the operating area, transportation network, communication facilities, infrastructure construction, and other information that directly affect the physical site layout and service channel selection of financial institutions.
[0045] (5) Policy and regulatory environment: local tax policies, financial regulatory policies, credit guidelines, industry incentives or restrictions related to the operating area, which will affect the operation strategy and risk management of financial institutions.
[0046] In order to overcome the problem of insufficient information utilization caused by data silos and different formats, for example, the original data of financial institutions may contain a large amount of unstructured information, such as customer managers' handwritten notes or farmers' oral descriptions. These information is inefficient and difficult to accurately interpret when used directly for analysis. Therefore, ETL (Extract, Transform, Load) tools can be used in combination with entity matching algorithms to integrate structured data (such as credit records) with unstructured data (such as customer manager's research notes). At the same time, semantic alignment algorithms are applied to ensure data consistency and comparability. For example, for unstructured data, use the BERT-CRFT hybrid model for entity-relation extraction to convert key information in the text into structured data for subsequent analysis.
[0047] The operating area of a financial institution often covers multiple sub-regions with different economic characteristics, such as agricultural, industrial or tourist areas, and the customer demand and behavior patterns of each sub-region are significantly different. In order to solve the problem that financial institutions are difficult to intuitively understand the geographical distribution of customers and industry characteristics in the recommendation process, a map recommendation tool can be constructed to develop more targeted recommendation strategies based on the characteristics of the sub-region where the customer is located, improving the success rate and ROI (Return on Investment) of the recommendation activities. For example, use clustering analysis algorithms and data visualization techniques to classify and display potential customers according to their sub-regional attributes, industry attributes and economic activity levels.
[0048] Step S204, constructing a knowledge graph corresponding to the financial institution according to the original data, and determining the portrait corresponding to the target object according to the knowledge graph.
[0049] In the above step S204, the knowledge graph is used to represent the association between the business knowledge, customer information, operating environment and other multi-source heterogeneous data of the financial institution.
[0050] Financial institutions accumulate a vast amount of unstructured text, such as research reports and customer communication records. While this information contains rich business insights, it is difficult to directly apply to decision support due to its diverse formats and difficulty extracting information. By converting key information in unstructured text into structured data through entity recognition and relationship extraction, we can improve information utilization efficiency and enrich the dimensions of customer profiles. For example, natural language processing techniques (such as the BERT-CRFT hybrid model) can be used to analyze unstructured text data in raw data within the operating area to identify important entity information (such as company names, transaction amounts, and policy keywords). Relationship extraction algorithms can then be used to determine the connections between entity information (e.g., company A received support from policy B and conducted transaction C).
[0051] In order to solve the problem of data silos and ensure that all relevant data can be effectively integrated into the knowledge graph, data fusion technology (such as federated learning or data lake technology) can be used to fuse the internal system data of financial institutions (first data) with the external environment data (second data) to obtain a comprehensive feature set that reflects internal customer information and external environment characteristics; the entity information (such as customers, enterprises, policies) in the comprehensive feature set is converted into nodes in the knowledge graph, and the relationship between entities (such as credit relationships, policy benefits, industry affiliation) is converted into edges in the knowledge graph to obtain a knowledge graph. The knowledge graph can reflect the real position and status of the target object (customer) in its operating environment, as well as the impact of environmental factors on the target object.
[0052] In some embodiments of the present application, the portrait corresponding to the target object can be determined in the following manner: obtaining static attributes of the target object from the knowledge graph, wherein the static attributes include information corresponding to the first data of the target object; determining the dynamic attributes of the target from the knowledge graph, wherein the dynamic attributes include information associated with the target object in the second data; fusing the static attributes with the dynamic attributes to obtain a joint representation of the target object; and determining the portrait corresponding to the target object based on the joint representation.
[0053] Static attributes refer to characteristics of a target object that are relatively stable and resistant to change within a preset time period, such as basic customer information (age, gender, occupation), company establishment date, registered capital, etc., used to describe the target object's identity and basic status. In some embodiments of this application, entity recognition and relationship extraction algorithms in natural language processing technology can be used to accurately extract static attributes of the target object from the knowledge graph. For example, a model such as BERT-CRF can be used to structure customer information documents, identify key entities, and annotate their attribute categories (such as "age" and "income").
[0054] Dynamic attributes refer to characteristics that change over time, such as recent customer transaction activities, real-time changes in the market environment, policy adjustments, etc., and are used to reflect the target object's behavior patterns and market reactions in the current environment. In some embodiments of the present application, graph traversal algorithms (such as breadth-first search, depth-first search) and graph clustering algorithms (such as community detection algorithms) can be used to extract dynamic attributes related to the target object from the knowledge graph. These algorithms can identify entities closely connected to the target object and their dynamic attributes. For example, all entities related to the target industry can be determined through a community detection algorithm, and then the dynamic attributes of these entities can be extracted. For dynamic attributes with time series characteristics (such as transaction frequency and market interest rates), time series analysis techniques (such as ARIMA and Prophet) can also be used for trend prediction, and the prediction results can be added to the dynamic data to develop a more accurate customer profile.
[0055] It should be noted that dynamic attributes have been incorporated into the knowledge graph construction phase. Since dynamic attributes change in real time, the dynamic attributes in the knowledge graph need to be updated in real time. For example, when the knowledge graph is constructed, the dynamic attributes are marked; first data is obtained from an external data source (such as a financial information API, market data, weather forecast, etc.), and the dynamic attributes are updated using the first data. In some embodiments of the present application, preset rules can be used for updating, and the preset rules can be determined based on the update requirements of the dynamic attributes. For example, when the update requirements of the dynamic attributes meet the first preset condition, the first update frequency is used for updating; when the update requirements of the dynamic attributes meet the second preset condition, the second update frequency is used for updating, the second preset condition is stricter on the update time than the first preset condition, and the second update frequency is greater than the first update frequency.
[0056] All extracted dynamic attributes are fused with the static attributes of the target object to form a joint representation. During this fusion process, a dynamic attention mechanism can be used to adjust the weights of different attributes to reflect their importance to the target object at the current point in time. For example, if agricultural product prices are currently at their peak, the weight of the "agricultural product price" attribute will be increased because it has a more significant impact on farmers' loan demand.
[0057] Step S206: Map the portrait to a sub-region corresponding to the business area, wherein the business area corresponds to at least one sub-region, and each sub-region corresponds to different economic characteristics.
[0058] In step S206 above, sub-regions are subdivided areas within the operating area, each with unique economic characteristics, such as agricultural output value, industrial growth rate, and resident consumption level. It should be noted that sub-regions can be divided based on administrative levels. For example, at a first administrative level, the area can be divided into a first preset number of first sub-regions, and at a second administrative level, the area can be divided into a second preset number of second sub-regions. The sub-region division methods at different administrative levels can be interrelated, such as further dividing the first sub-region into a second preset number of second sub-regions within the first sub-region. Sub-region division methods at different administrative levels can also be independent, that is, dividing the sub-regions according to different dimensions to facilitate separate analysis of different dimensions by clients.
[0059] In order to overcome the problem of unclear sub-region division or insufficient description of economic characteristics, the sub-region to which the customer belongs can be determined by calculating the similarity between the characteristic vector of the customer portrait and the economic characteristic vector of each sub-region, such as using algorithms such as cosine similarity and Jaccard similarity coefficient.
[0060] It's important to note that GIS technology can also be combined with machine learning models (such as random forests and K-nearest neighbor algorithms) to predict the most likely active sub-regions based on the geographic location and behavioral characteristics of customer profiles. Specifically, customer profile data is integrated with the geographic information of the operating area; the geographic location and behavioral characteristics in the customer profile are converted into feature vectors. For example, geographic location information can be converted into latitude and longitude coordinates and distance from a specific economic center, while behavioral characteristics can include transaction amount, transaction frequency, shopping preferences, and visit frequency. The feature vectors are then fed into the machine learning model for prediction, resulting in the sub-regions.
[0061] In order to gradually and deeply understand the economic environment within its operating area from macro to micro, at least one sub-region corresponding to the operating area can be determined in the following manner: dividing the operating area into at least one first sub-region in the first administrative region, wherein the first sub-region is a geographical region whose economic activity intensity and / or industry distribution characteristics meet the first preset conditions; dividing the first sub-region into at least one second sub-region in the second administrative region, wherein the second sub-region is a geographical region whose economic activity intensity and / or industry distribution characteristics meet the second preset conditions, the indicator threshold in the second preset conditions is greater than the corresponding indicator threshold in the first preset conditions, and the administrative level of the second administrative region is lower than the administrative level of the first administrative region; dividing the second sub-region into at least one third sub-region in the third administrative region, wherein the third sub-region is a geographical region corresponding to the transportation logistics and / or production supply chain in the second sub-region, and the administrative level of the third administrative region is lower than the administrative level of the second administrative region.
[0062] The first sub-region is a geographical region in the scope of the highest administrative region (e.g., a province), which is divided according to the intensity of economic activities (e.g., total GDP, per capita income) and the distribution characteristics of industries (e.g., the proportion of service industry, the proportion of heavy industry). These regions meet the first preset condition, i.e., they reach a certain economic scale and industry representation.
[0063] In some embodiments of the present application, the operating region in the first administrative region can be divided into at least one first sub-region by the following steps: determining a plurality of target indicators corresponding to the original data, wherein the target indicators are used to quantify the industry attributes and economic activity in the operating region; converting the original data into data points on the map of the operating region according to the target indicators, wherein the data points include target indicators and geographic location information; and performing cluster analysis on the data points to obtain the first sub-region.
[0064] The target indicators are numerical values that quantify the industry attributes and economic activity in the operating region, such as GDP, industry output value, number of enterprises, frequency of logistics activities, etc. In some embodiments of the present application, preliminary indicators reflecting industry attributes and economic activity can be determined in combination with industry standards and economic theories, for example, for village banks, agricultural output value, agricultural product transaction volume, number of small enterprises, per capita income, etc. can be considered as preliminary indicators; according to the preliminary target indicators, the index combination is optimized through data analysis and machine learning techniques (such as principal component analysis PCA or feature selection algorithm), and the target indicators most related to industry attributes and economic activity are selected.
[0065] It should be noted that the random forest algorithm can be used to dynamically calculate the weights of each target indicator to ensure effective integration and dynamic adjustment of different industry attributes and economic activity indicators.
[0066] After determining the target indicators, the non-structured geographic location information (such as address, latitude and longitude) in the original data can be converted into standardized geographic coordinates using geographic information system (GIS) technology, together with the target indicators to form data points. It should be noted that not all data in the original data have geographic location information, in which case the non-geographic location information in the original data can be converted into attributes or indicators related to geographic location information, for example, for economic indicators such as enterprise transaction volume and agricultural product output value in the original data, virtual "economic points" or "industry points" can be created based on the relevance of these indicators to specific geographic locations, such as binding the values of these indicators with geographic location information (such as the latitude and longitude of the location of the enterprise or the origin of agricultural products) to form data points.
[0067] In some embodiments of the present application, a default geographic location can be assigned to each economic indicator or industry attribute. This location might be the most common location for that indicator or attribute, or the location most directly associated with it. For example, for the agricultural output value indicator, it could be tied to the geographic coordinates of each village or town within the village bank's operating area; for corporate transaction volume, it could be tied to the coordinates of the company's registered address or principal place of business.
[0068] In addition, data analysis or machine learning techniques can be used to explore potential correlations between non-geographic information and geographic information, thereby inferring the "virtual location" of non-geographic information. For example, if correlation analysis reveals that a company's transaction volume is related to the density of logistics facilities in the area, then companies with high transaction volumes can be linked to geographic areas with high-density logistics facilities, thereby representing these companies on a map.
[0069] Specifically, various types of data are collected and organized, including economic indicators (such as corporate transaction volume and agricultural product output value), industry attributes (such as energy industry and agriculture), and corresponding geographic location information (such as corporate registration place and agricultural product origin); correlation analysis methods (such as Pearson correlation coefficient and Spearman rank correlation coefficient) or supervised learning models (such as regression analysis and deep learning models) are used to establish the connection between non-geographic location information and geographic location information; based on this connection, a reasonable "virtual location" is assigned to the original data without direct geographic location information, and it is represented as a data point on the map together with the actual geographic location information.
[0070] Once the data points are obtained, density-based clustering algorithms (such as DBSCAN or HDBSCAN) can be used to automatically identify cluster boundaries based on the distance and density between data points, forming the first sub-region. To combine geographic boundaries with hierarchical clustering, ensuring that clustering results reflect similarities in economic activities while respecting administrative divisions and avoiding clustering across administrative boundaries, hierarchical clustering algorithms (such as Agglomerative Clustering) can also be used to perform cluster analysis, combining administrative and geographic boundary information of the operating area. This approach initially treats each data point as an independent cluster, then gradually merges the most similar clusters to form a dendrogram. Finally, the first sub-region is divided based on the consistency of the clustering results with the geographic boundaries.
[0071] The second sub-region is a further subdivision of the first sub-region within a lower-level administrative division (such as a city or county), identifying areas with higher economic activity intensity and more distinct industrial distribution characteristics. The third sub-region is a geographical area demarcated within a lower-level administrative division (such as a town or village) based on the convenience of transportation and logistics and the coherence of the production supply chain within the second sub-region.
[0072] In the process of gradually subdividing the business area into the first sub-area, the second sub-area, and the third sub-area, the selected clustering algorithm will vary depending on the division target and data characteristics. Specifically:
[0073] (1) The first sub-regional division aims to identify geographical areas whose economic activity intensity and industry distribution characteristics meet the first precondition from a macro perspective. This usually involves a wide geographical area, a large amount of data, and obvious spatial clustering of economic activities and industry distribution. At this level, density-based clustering algorithms such as DBSCAN (Density-Based Spatial Clustering of Applications with Noise) or HDBSCAN (Hierarchical DBSCAN) can be used.
[0074] (2) The second sub-region is based on the first sub-region, further screening out geographical areas with higher economic activity intensity and more significant industry distribution characteristics. The indicator threshold at this time is more stringent. Considering that the division of the second sub-region requires more precise clustering results to ensure the accuracy and effectiveness of high-value market segmentation, more complex clustering algorithms such as GMM (Gaussian Mixture Model) or Spectral Clustering can be used.
[0075] (3) The division of the third sub-region is based on the characteristics of transportation logistics and production supply chain in the second sub-region. The division of sub-regions at this level focuses more on industry relevance and the continuity of logistics activities. Therefore, algorithms with specific industry knowledge and network structure recognition capabilities can be used, such as Community Detection algorithm or Network Analysis.
[0076] In order to ensure that the recommendation strategies of financial institutions have both a global perspective and take into account local characteristics, the following methods can be used to map the portraits to the sub-regions corresponding to the operating areas: using a mapping method from high to low administrative levels, determine the first sub-region, second sub-region and third sub-region corresponding to the portraits in turn; or using a mapping method from low to high administrative levels, determine the third sub-region, second sub-region and first sub-region corresponding to the portraits in turn.
[0077] The process of mapping images to sub-regions is essentially a process of matching the characteristics of target objects with the economic characteristics of specific geographic regions. Using a mapping method from high to low administrative level, it can gradually refine from a macro perspective to a specific community or street, which is conducive to financial institutions to grasp from the whole to the local details and develop comprehensive recommendation plans. Conversely, the mapping method from low to high is to start from the grassroots and gradually expand the perspective. Such a strategy may be more suitable for financial institutions that have already identified specific market segments. They can first focus on a small range of high-potential markets and then gradually expand to larger areas.
[0078] Under the mapping method of administrative levels from high to low, the customer's portrait data can be matched with the highest level of operating area (such as province) to identify the first sub-region that best matches the customer's portrait. Within the first sub-region, further match the portrait data to determine the second sub-region. Within the second sub-region, use the portrait data again to determine the third sub-region with the finest granularity. In the case of large amounts of data, using the mapping method from high to low first matches on a large scale, gradually narrows the search range, reduces the computational complexity, and improves the matching efficiency.
[0079] Under the mapping method of administrative levels from low to high, it can start from the third sub-region with the finest granularity and match based on the similarity of the customer's portrait and other customers in the region, gradually expand to the second sub-region and the first sub-region, and finally determine the most relevant operating sub-region for the customer. This approach can make full use of the detailed information in the customer's portrait, starting from the geographic environment closest to the customer and gradually comparing it with the characteristics of broader areas to ensure that the matched sub-region more accurately reflects the customer's specific needs and market environment.
[0080] To solve the problem of too specific customer portrait information in low-level regions and difficulty in finding matching regions, in the mapping process from low to high, a layer-by-layer abstraction and fusion strategy can be used. That is, when matching in the third sub-region, more specific portrait features such as customer's specific needs and behavior patterns are considered, while when entering the second sub-region and the first sub-region, more generalized features such as industry categories and economic indicators are gradually integrated to ensure that a region matching the customer's portrait can be found in sub-regions of different administrative levels.
[0081] Utilizing GIS technology, mapped customer profile data can be transformed into visual displays, allowing users to choose to display the distribution of customer profiles across sub-regions at different administrative levels based on their needs. For example, financial institutions can begin with high-level sub-regions at the first administrative level (e.g., provincial or prefectural-level cities) and use heat maps to understand the customer distribution overview across their entire operating area. Furthermore, they can drill down to sub-regions at the second administrative level (e.g., county or township level) for a deeper analysis of customer profiles. For example, heat maps can be used to display the distribution of customer profiles within a specific county or township, identifying differences in customer profiles and demand hotspots across villages and towns. Furthermore, they can focus on sub-regions at the third administrative level (e.g., the most basic villages, towns, or communities), using heat maps to provide a detailed understanding of the distribution of customer profiles within each sub-region.
[0082] Based on the distribution of profiles across different sub-regions (e.g., the first, second, or third sub-regions), cluster analysis can be performed in a specific sub-region at the customer's discretion. Alternatively, cluster analysis can be automatically performed across all sub-regions if the profile distribution meets pre-set conditions. It should be noted that the pre-set conditions for cluster analysis vary for sub-regions at different administrative levels, and can be determined based on factors such as the number of profiles in the sub-region, transaction activity, and geographic location.
[0083] Specifically, for sub-regions at any administrative level, a minimum threshold for the number of profiles (i.e., the number of customers) can be set before cluster analysis to ensure that there is sufficient customer data within the analyzed sub-region to obtain meaningful clustering results. For example, the first sub-region (such as a provincial or prefecture-level city) may require a higher customer threshold (such as 500 or more) because these sub-regions cover a wide area and have a large customer base, and a smaller customer base may not be sufficient to reflect general trends. The third sub-region (such as a village or town) may require a lower customer threshold (such as 20 or more) because the total number of customers in villages and towns is relatively small, and even a small number of customers can form specific clustering characteristics. The preset conditions for the second sub-region are between the first and third sub-regions.
[0084] In addition to the number of profiles, customer transaction activity within a sub-region can also serve as a reference indicator. A threshold for transaction volume or frequency can be set, and cluster analysis can be performed only on sub-regions that meet the pre-set activity criteria. For example, the first sub-region may require a high average transaction volume (e.g., over 500,000 transactions per month), while the third sub-region (towns and villages) only needs a lower average transaction volume (e.g., over 500 transactions per month) for cluster analysis. This ensures that analysis focuses on true business hotspots, leading to more targeted service optimization. The pre-set criteria for the second sub-region lie between the first and third sub-regions.
[0085] Different cluster analysis conditions can also be set based on the geographic location and scope of the sub-region. For example, for the first sub-region (province or city), financial institutions need to focus on cross-regional liquidity. Therefore, cluster analysis can focus on sub-regions with geographical proximity but significantly different customer profiles to explore cross-regional recommendation and service opportunities. For the third sub-region (townships and villages), more attention can be paid to the similarity of customers within the sub-region. The preset condition can be areas with relatively concentrated geographical locations and similar customer profiles to facilitate centralized recommendation activities or service optimization. The preset condition for the second sub-region lies between the first and third sub-regions, for example, areas with relatively concentrated geographical locations but a certain degree of cross-regional liquidity.
[0086] Step S208: performing cluster analysis on the portraits in the sub-regions and determining a recommendation strategy corresponding to the clustering results.
[0087] In the above step S208, a deep learning model, such as an autoencoder or a generative adversarial network (GAN), can be used to reduce the dimension and extract features of the customer portrait data, and then perform clustering analysis based on the extracted feature vectors, such as using the K-means or DBSCAN algorithm.
[0088] Within a subregion, customer characteristics are often closely linked to geographic location. Therefore, geographic information can be added to customer profile data as an additional feature, allowing for cluster analysis using spatial clustering algorithms such as ST-DBSCAN or Geographically Weighted Clustering (GWC). ST-DBSCAN considers proximity in both spatial and temporal dimensions, while GWC dynamically adjusts clustering parameters based on geographic weights, making it more tailored to the characteristics of customers in a specific location.
[0089] Within a sub-region, if the customer profile clustering results are too granular, the cost of customizing the recommendation strategy for each cluster may be too high. Conversely, if the clustering results are too broad, they may not effectively meet the needs of different customer groups, affecting the effectiveness of recommendations. To address this issue, adaptive clustering algorithms can be used, such as hierarchical clustering combined with a strategy optimization method based on cost-benefit analysis. Hierarchical clustering generates a cluster structure tree, gradually merging from the finest clusters to broader groups. Business personnel can select the most appropriate clustering level based on the cost-benefit analysis results, achieving adaptive optimization of the recommendation strategy.
[0090] To reduce the cost and time of customizing the recommendation strategy, a set of recommendation strategy templates can also be predefined for different levels of clustering results, including product recommendation, discount strategy, recommendation channel selection, etc.; according to the features extracted in the clustering analysis, such as customer age, occupation, income level, etc., the clustering results are mapped to the corresponding strategy templates to realize the rapid customization and execution of the recommendation strategy.
[0091] To ensure that the clustering results reflect both the similarity of the customer portraits and the differences in the environmental characteristics of the sub-regions, the following steps can be taken to perform clustering analysis on the portraits in the sub-regions: determining third data corresponding to the sub-regions from the second data, wherein the third data is used to reflect the environmental characteristics of the sub-regions; determining a clustering algorithm corresponding to the third data and a target parameter corresponding to the clustering algorithm, wherein the target parameter is determined based on the distribution characteristics of all the portraits in the sub-regions; and performing clustering analysis on all the portraits of the sub-regions using the clustering algorithm to obtain a clustering result, wherein the clustering result includes a plurality of groups, and each group represents a target object set with a similarity greater than or equal to a preset threshold.
[0092] Specifically, a multi-dimensional feature fusion clustering method can be used first to combine the customer portrait data with the third data reflecting the environmental characteristics of the sub-regions to form a comprehensive feature set and perform clustering analysis; the target parameter of the clustering algorithm is dynamically adjusted to continuously optimize the clustering result until the preset conditions are met, such as meeting the similarity requirements of the customer portraits and taking into account the differences in the environmental characteristics of the sub-regions; and based on the final clustering result, the specific market environment and social and economic conditions of the sub-regions are combined to develop and execute customized recommendation strategies, for example, for a sub-region with high demand for agricultural credit, farmers with similar portraits can be clustered and the environmental characteristics of the sub-region, such as soil conditions and irrigation facilities, are considered to form a special agricultural credit recommendation strategy.
[0093] In some embodiments of the present application, the recommendation strategy corresponding to the clustering result can be determined by the following steps: determining a target feature corresponding to a target group, wherein the target feature is determined based on the features of the portraits of the target group, and the target group is any one of the plurality of groups; and determining the recommendation strategy according to the target feature and the third data.
[0094] The target group is a specific customer group in the clustering analysis result, with similar portrait features, and the identification of the target group is the core of precision recommendation. Through clustering analysis, customers are grouped by similarity, enabling the bank to adopt differentiated recommendation strategies for different customer groups with different needs. The target feature is a key indicator reflecting the portrait features and needs of the target group, which is used to guide the development of the recommendation strategy.
[0095] In order to ensure that the design of the recommendation strategy is highly relevant to the profile characteristics of the target group, for each target group, the key features of all profiles in the group (i.e., target features) can be extracted, such as age, income level, occupation, credit rating, etc.; based on the target characteristics, the potential needs and preferences of the target group are analyzed, and combined with the environmental characteristics of the sub-region (such as the intensity of economic activity, industry distribution, etc.), the target needs of the target group in the context of a specific sub-region are analyzed; and recommendation strategies corresponding to the target needs are formulated, such as launching specific loan products, designing personalized financial plans, and optimizing the layout of recommendation channels.
[0096] In a specific embodiment, a financial institution, after cluster analysis, identified a specific target group consisting of young farmers. The customer profile characteristics of this group showed that they had a high rate of digital device usage, a preference for online financial services, a large demand for microcredit, and a high sensitivity to agricultural technology information.
[0097] (1) Determine target characteristics: Target characteristics may include high usage of digital devices, preference for digital services, demand for microcredit, and attention to agricultural technology information;
[0098] (2) Environmental characteristics: Based on the third data, analyze the current agricultural economic status of the sub-region, such as fluctuations in agricultural product prices, the degree of agricultural technology promotion, and the penetration of rural Internet;
[0099] (3)推荐策略设计,例如,包括:
[0100] 1) Product Design: Launch a digital micro-agricultural loan product, providing a fast application and approval process through a mobile app.
[0101] 2) Pricing strategy: Based on local agricultural product price fluctuations and farmers' income, design flexible repayment methods and loan interest rates linked to market interest rates.
[0102] 3) Channel selection: Prioritize pushing loan product information through mobile Internet channels, and use social media and agricultural information platforms for in-depth recommendations.
[0103] 4) Promotional Activities: Organize online agricultural technology training and loan application guidance, provide online consulting services for the initial review of loan applications, and provide preferential interest rates for farmers who apply for loans using digital services for the first time.
[0104] 5) Communication methods: Maintain high-frequency online communication with the target group through group text messaging, in-app message push, social media interaction, etc., and provide timely loan information and technical support.
[0105] Through the above steps S202 to S208, by obtaining the original data corresponding to the operating area of the financial institution, and constructing the knowledge graph corresponding to the financial institution based on the original data, and then determining the portrait corresponding to the target object based on the knowledge graph, and mapping the portrait to the sub-area corresponding to the operating area, the portrait is clustered in the sub-area, and finally the recommendation strategy corresponding to the clustering result is determined, thereby achieving the purpose of clustering the target objects in specific sub-areas based on the economic characteristics of different geographical areas and determining the most suitable recommendation strategy, thereby achieving the technical effect of improving the pertinence of recommendation activities and dynamically adjusting strategies according to environmental changes, and thus solving the technical problem that the related technology formulates recommendation strategies based on general static templates, and it is difficult to conduct differentiated analysis of different regions where different financial institutions are located, resulting in limited recommendation activities.
[0106] Figure 3 is a structural diagram of a device for determining a recommendation strategy based on data mining and knowledge graph according to an embodiment of the present application, such as Figure 3 As shown, the device includes:
[0107] An acquisition module 302 is used to acquire original data corresponding to the operating area of the financial institution;
[0108] Determine 304 for constructing a knowledge graph corresponding to the financial institution based on the original data, and determining a profile corresponding to the target object based on the knowledge graph;
[0109] A mapping module 306 is used to map the portrait to the sub-area corresponding to the business area;
[0110] The clustering module 308 is used to perform cluster analysis on the portraits in the sub-regions and determine a recommendation strategy corresponding to the clustering results.
[0111] It should be noted that Figure 3 The device for determining the recommendation strategy based on data mining and knowledge graph is used to perform Figure 2 The method for determining the recommendation strategy based on data mining and knowledge graph is shown in Figure 2 The explanations in the method for determining the recommendation strategy based on data mining and knowledge graphs also apply to Figure 3 The device for determining the recommendation strategy based on data mining and knowledge graph shown will not be described in detail here.
[0112] An embodiment of the present application also provides an electronic device, which includes a memory and a processor, wherein the memory is used to store program instructions; the processor is connected to the memory and is used to execute the steps of a method for determining a recommendation strategy based on data mining and knowledge graphs in each embodiment of the present application.
[0113] For example, the processor performs the following functions by executing program instructions stored in the memory: obtaining original data corresponding to an operating area of a financial institution, wherein the original data includes first data corresponding to a target object in the operating area and second data reflecting environmental characteristics of the operating area; constructing a knowledge graph corresponding to the financial institution according to the original data, and determining a portrait corresponding to the target object according to the knowledge graph; mapping the portrait to a sub-area corresponding to the operating area, wherein the operating area corresponds to at least one sub-area, and each sub-area corresponds to different economic characteristics; performing clustering analysis on the portrait in the sub-area, and determining a recommended strategy corresponding to the clustering result.
[0114] The embodiments of the present application further provide a non-volatile storage medium, which comprises a stored computer program, wherein a device where the non-volatile storage medium is located executes steps of the method for determining a recommended strategy based on data mining and a knowledge graph in various embodiments of the present application by running the computer program.
[0115] The embodiments of the present application further provide a computer program product, which comprises computer instructions, and the computer instructions are executed by a processor to implement steps of the method for determining a recommended strategy based on data mining and a knowledge graph in various embodiments of the present application.
[0116] The embodiments of the present application further provide a computer program, which is executed by a processor to implement steps of the method for determining a recommended strategy based on data mining and a knowledge graph in various embodiments of the present application.
[0117] The above-mentioned serial numbers of the embodiments of the present application are only for description, and do not represent advantages or disadvantages of the embodiments.
[0118] In the above-mentioned embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0119] In the several embodiments provided by the present application, it should be understood that the disclosed technology can be implemented in other ways. Of course, the unit described as the division is only a logic function division, and there can be another division way during actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection between each other can be indirect coupling or communication connection through some interface, unit or module, and can be electrical or other forms.
[0120] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0121] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0122] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program code.
[0123] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for determining a recommendation strategy based on data mining and knowledge graph, characterized in that: include: Acquire original data corresponding to the operating area of the financial institution, wherein the original data includes first data corresponding to a target object within the operating area and second data reflecting environmental characteristics of the operating area; Constructing a knowledge graph corresponding to the financial institution based on the original data, and determining a profile corresponding to the target object based on the knowledge graph; Mapping the portrait to a sub-region corresponding to the business area, wherein the business area corresponds to at least one sub-region, and each sub-region corresponds to different economic characteristics; performing cluster analysis on the portraits in the sub-regions and determining a recommendation strategy corresponding to the clustering results; The method further comprises dividing the business area into at least one first sub-area within the first administrative region, including: determining a plurality of target indicators corresponding to the original data, wherein the target indicators are used to quantitatively represent industry attributes and economic activity within the business area; converting the original data into data points on a map of the business area based on the target indicators, wherein the data points include the target indicators and geographic location information; and performing cluster analysis on the data points to obtain the first sub-area. The raw data is converted into data points on a map of the business area according to the target indicator, including: obtaining actual geographic location information corresponding to the raw data; when the raw data does not have geographic location information, determining the connection between the non-geographic location information and the geographic location information of the raw data, and determining the virtual location information corresponding to the raw data without geographic location information based on the connection; and determining the data points based on the actual geographic location information and the virtual location information.
2. The method according to claim 1, characterized in that At least one sub-region corresponding to the business region is determined by: Dividing the operating area into at least one first sub-area within the first administrative area, wherein the first sub-area is a geographical area whose economic activity intensity and / or industry distribution characteristics meet a first preset condition; Dividing the first sub-region into at least one second sub-region within the second administrative region, wherein the second sub-region is a geographical region in which the economic activity intensity and / or industry distribution characteristics within the first sub-region meet a second preset condition, the indicator threshold in the second preset condition is greater than the corresponding indicator threshold in the first preset condition, and the administrative level of the second administrative region is lower than the administrative level of the first administrative region; The second sub-region is divided into at least one third sub-region within the third administrative region, wherein the third sub-region is a geographical region corresponding to the transportation logistics and / or production supply chain in the second sub-region, and the administrative level of the third administrative region is lower than that of the second administrative region.
3. The method according to claim 2, characterized in that Mapping the portrait to a sub-area corresponding to the business area includes: Adopting a mapping method from high to low administrative levels, determining the first sub-region, the second sub-region and the third sub-region corresponding to the portrait respectively; or, A mapping method of administrative levels from low to high is adopted to determine the third sub-region, the second sub-region and the first sub-region corresponding to the portrait respectively.
4. The method according to claim 1, wherein Performing cluster analysis on the portrait in the sub-region includes: Determining third data corresponding to the sub-area from the second data, wherein the third data is used to reflect environmental characteristics of the sub-area; determining a clustering algorithm corresponding to the third data and a target parameter corresponding to the clustering algorithm, wherein the target parameter is determined based on distribution characteristics of all portraits in the sub-region; The clustering algorithm is used to perform cluster analysis on all portraits in the sub-region to obtain a clustering result, wherein the clustering result includes multiple groups, each group representing a set of target objects with a similarity greater than or equal to a preset threshold.
5. The method according to claim 4, characterized in that Determine the recommended strategy corresponding to the clustering results, including: determining a target feature corresponding to a target group, wherein the target feature is determined based on a feature of a portrait of the target group, and the target group is any one of the multiple groups; The recommendation strategy is determined according to the target feature and the third data.
6. The method according to claim 1, characterized in that Determining a portrait corresponding to the target object based on the knowledge graph includes: Acquire static attributes of the target object from the knowledge graph, wherein the static attributes include information corresponding to the first data of the target object; Determining dynamic attributes of the target from the knowledge graph, wherein the dynamic attributes include information associated with the target object in the second data; Fusing the static attributes with the dynamic attributes to obtain a joint representation of the target object; A portrait corresponding to the target object is determined based on the joint representation.
7. A device for determining a recommendation strategy based on data mining and knowledge graph, characterized in that: include: an acquisition module, configured to acquire original data corresponding to the operating area of a financial institution, wherein the original data includes first data corresponding to a target object within the operating area and second data reflecting environmental characteristics of the operating area; a determination module, configured to construct a knowledge graph corresponding to the financial institution based on the original data, and determine a profile corresponding to the target object based on the knowledge graph; a mapping module for mapping the portrait to a sub-region corresponding to the business area, wherein the business area corresponds to at least one sub-region, and each sub-region corresponds to different economic characteristics; dividing the business area into at least one first sub-region within a first administrative region, comprising: determining a plurality of target indicators corresponding to the original data, wherein the target indicators are used to quantitatively represent industry attributes and economic activity within the business area; converting the original data into data points on a map of the business area based on the target indicators, wherein the data points include the target indicators and geographic location information; performing cluster analysis on the data points to obtain the first sub-region; converting the original data into data points on a map of the business area based on the target indicators, comprising: obtaining actual geographic location information corresponding to the original data; when the original data does not have geographic location information, determining a connection between the non-geographic location information of the original data and the geographic location information, and determining virtual location information corresponding to the original data without geographic location information based on the connection; determining the data point based on the actual geographic location information and the virtual location information; The clustering module is used to perform cluster analysis on the portraits in the sub-areas and determine a recommendation strategy corresponding to the clustering results.
8. An electronic device, characterized in that: include: A memory and a processor, the memory being used to store program instructions; the processor being connected to the memory and being used to execute the method for determining a recommendation strategy based on data mining and knowledge graphs as described in any one of claims 1 to 6.
9. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the method for determining a recommendation strategy based on data mining and knowledge graph as described in any one of claims 1 to 6 by running the computer program.
10. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by the processor, the method for determining the recommendation strategy based on data mining and knowledge graph described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
High-level talent recommendation method and system based on multi-dimensional portraits and medium
CN116756409A
User portrait generation query method based on knowledge graph
CN119149755A
Client portrait construction method and device, storage medium and electronic equipment
CN119336977A