A county digital management map data analysis method and system
Patent Information
- Application Number
- CN202610727356.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-09-15
AI Technical Summary
本发明解决了县域金融多源异构数据整合难、客群经营缺乏抓手及管理决策缺支撑的问题,实现了数字化经营地图的精准分析、商机自动分配与全流程闭环管理
[0017]This invention discloses a county-level digital business map data analysis method and system. It generates standardized unified subject data by performing conflict detection and normalization on multi-source heterogeneous data based on primary key identification and priority coverage rules. It automatically allocates business opportunities by standardizing address information through mapping and spatial coordinate transformation, and calculating weighted matching degrees based on administrative division affiliation and service radius weights. It generates a panoramic portrait of the industrial chain by identifying industry characteristics and extracting indicators through a keyword rule engine and distributed parallel computing. It generates risk warning signals by collecting real-time capital change data and comparing it with dynamic thresholds, and then performs a closed-loop circulation and visualization of business opportunity allocation results, industrial chain portraits, and risk warnings. This method effectively solves the problems of difficult integration of multi-source financial data in counties, lack of key tools for customer group management, and lack of data support for management decisions. It achieves a fully automated closed loop from data collection, fusion, and calculation to monitoring, allocation, and display, significantly improving the digital business capabilities and marketing management efficiency of county-level institutions, and enhancing the accuracy and intelligence of financial services reaching lower levels.
Smart Images

Figure CN122760118A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of financial digital technology, and in particular to a method and system for analyzing county-level digital business map data. Background Technology
[0002] To implement the requirements of platform empowerment and comprehensive services for financial institutions, and to continuously improve the digital operation capabilities of county-level institutions, it is urgent to address the following core operational pain points currently faced by county-level financial institutions. First, multi-dimensional data on the economy, industry, and customers are scattered across multiple independent business systems (such as the credit system, core banking system, and customer relationship management system). Some data even relies on manual offline collection and integration, lacking a unified data governance and standardized processing mechanism. This results in an inability to comprehensively and accurately grasp regional customer resources, making it difficult to form effective customer profiles and operational views. Second, customer group management lacks digital tools. Lists of characteristic industry enterprises are usually managed in the form of manual ledgers or fragmented tables, leading to delayed information updates. The progress of customer managers' marketing expansion, visit records, and results cannot be tracked and quantitatively evaluated in real time. Performance evaluation and supervision lack data support, affecting operational efficiency and accuracy. Third, key business processes such as customer expansion, industry support, and funding lack systematic data-driven decision support. Management struggles to analyze business opportunity distribution, industry chain characteristics, and capital flow trends from a holistic perspective, leading to blind resource allocation and management decisions, weakening overall management efficiency, and failing to meet the refined and intelligent requirements of financial services reaching county-level areas. In summary, there is an urgent need to develop an integrated and automated digital business map data analysis method to solve the problems of difficulty in integrating multi-source heterogeneous data, lack of leverage in customer group management, and lack of support for management decisions, and to provide county-level financial institutions with a closed-loop digital tool for the entire process. Summary of the Invention
[0003] The present invention aims to at least partially solve one of the technical problems in the related art.
[0004] To address this, this invention discloses a method for analyzing county-level digital business map data. By acquiring multi-source, heterogeneous county-level financial business data, conflict detection and normalization are performed based on primary key identification and priority coverage rules to generate standardized unified subject data. Address information in the unified subject data is standardized and mapped using spatial coordinate transformation. Based on administrative division affiliation and service radius weighting calculations and a weighted matching degree with preset service outlets, business opportunity allocation results are generated. Industry characteristic identification and indicator extraction are performed on the unified subject data using a keyword rule engine. Distributed parallel computing is used for cleaning and normalization to generate a panoramic portrait of the industrial chain. Real-time collection of account fund change data, combined with customer ratings and historical transaction characteristics, dynamically generates anomaly identification thresholds. Change data is compared with these thresholds to generate risk warning signals. The business opportunity allocation results, the panoramic portrait of the industrial chain, and the risk warning signals are then circulated in a closed loop and visualized. This invention solves the problems of difficult integration of multi-source, heterogeneous county-level financial data, lack of tools for customer group management, and lack of support for management decisions, achieving accurate analysis of digital business maps, automatic allocation of business opportunities, and closed-loop management throughout the entire process.
[0005] Another objective of this invention is to propose a county-level digital business map data analysis system.
[0006] To achieve the above objectives, this invention proposes a method for analyzing county-level digital business map data, comprising:
[0007] Acquire multi-source heterogeneous county-level financial operation data, perform conflict detection and normalization processing based on primary key identification rules and priority coverage rules, and generate standardized unified subject data; The address information in the unified subject data is standardized and mapped and spatial coordinates are transformed. Based on the weighted matching degree of administrative division affiliation and service radius with preset service outlets, business opportunity allocation results are generated. Based on the keyword rule engine, the unified subject data is used to identify industry characteristics and extract indicators. Distributed parallel computing is used for cleaning and normalization to generate a panoramic portrait of the industrial chain. Real-time collection of account fund change data, combined with customer rating and historical transaction characteristics to dynamically generate anomaly identification thresholds, comparison of change data with thresholds to generate risk warning signals, and closed-loop circulation and visualization of business opportunity allocation results, industry chain panoramic profile data and risk warning signals.
[0008] In one embodiment of the present invention, the step of acquiring multi-source heterogeneous county-level financial operation data, performing conflict detection and normalization processing based on primary key identification rules and priority coverage rules, and generating standardized unified subject data includes: A full extraction is performed at a fixed time every day. Multi-source snapshot data is integrated through union operation, and a hash algorithm is used to perform a modulo operation on the unified social credit code to generate a unique identifier key. When there is no unified social credit code, the triplet of "enterprise name + registered address + legal representative" is used as the basis for unique identification. Based on the priority of government data > business data > internal data > external data, conflicting fields are automatically processed by having higher priority fields override lower priority fields, while retaining the data status of the latest valid timestamp, so as to obtain standardized and unified main data after deduplication, verification and overwriting.
[0009] In one embodiment of the present invention, the step of standardizing and mapping the address information in the unified subject data and transforming its spatial coordinates, and generating a business opportunity allocation result based on the weighted matching degree of administrative division affiliation and service radius with preset service outlets, includes: The address information is standardized and mapped to the three-level administrative divisions of province, city, district and county, and the misspellings, abbreviations and incorrect road names are automatically corrected based on the county address database; The corrected standardized address is converted into longitude and latitude in the WGS84 coordinate system by calling the map API interface, thus obtaining the spatial coordinate data of the enterprise and the spatial coordinate data of the bank service outlets. Based on two types of spatial coordinate data, the weighted Euclidean distance is calculated by combining the service radius weight of the outlet, the administrative division weight, and the historical marketing matching weight, and the business opportunity allocation result is obtained.
[0010] In one embodiment of the present invention, calculating the weighted Euclidean distance includes: Based on the weight of administrative division affiliation, a set of service outlets belonging to the same administrative region as the enterprise is selected; For each point in the set, calculate the basic spatial distance between the enterprise coordinates and the point coordinates: d = √[(x2 - x1)² + (y2 - y1)²] Where x and y are the enterprise coordinates, and x2 and y2 are the network coordinates; The weighted matching degree is obtained by multiplying the base distance by the service radius weight of the outlet, the administrative division weight, and the historical marketing matching weight, respectively, by 0.6, 0.3, and 0.1. The network outlets are sorted in ascending order of weighted matching degree, and the outlet with the highest ranking is determined as the optimal service outlet to obtain the business opportunity allocation result.
[0011] In one embodiment of the present invention, the step of performing industry feature identification and indicator extraction on the unified subject data based on a keyword rule engine, and using distributed parallel computing for cleaning and normalization processing to generate a panoramic portrait data of the industrial chain includes: A distributed parallel computing framework is used to split a large amount of unified main data into multiple subtasks with a fixed partition size, and then schedules them concurrently to different computing nodes to perform keyword rule matching. Based on regular expression matching and keyword rule engine, the system automatically identifies industry tags, scale level and operating status from enterprise name, business scope, revenue and tax fields to form a standardized indicator set; Missing value imputation and outlier truncation are performed on the original indicator values in the standardized indicator set, and the normalization formula is used to map them to the [0,1] interval to obtain the panoramic portrait data of the industrial chain.
[0012] In one embodiment of the present invention, the normalization process includes: Obtain the original value x of a certain indicator in the current industry chain, and calculate the minimum value min and the maximum value max of this indicator among all enterprise data in the current industry chain; Substituting x, min, and max into the normalization formula, we obtain the standard index value x': x' = (x - min) / (max - min); All x' are aggregated according to the hierarchical relationship of core enterprises, supporting enterprises and upstream and downstream enterprises, and output panoramic portrait data of the industrial chain containing multi-level relationships.
[0013] In one embodiment of the present invention, real-time acquisition and dynamic threshold comparison includes: A real-time channel is established through a Kafka message queue configured with the number of partitions and replicas to receive account transaction records and balance change data at a millisecond-level transmission rate. Dynamic thresholds are automatically generated based on customer ratings, industry characteristics, historical trading habits, and seasonal fluctuations, and real-time fund change data is continuously compared with the dynamic thresholds. When a change in funds is detected to exceed a dynamic threshold, a real-time alert is immediately triggered to obtain a risk warning signal.
[0014] In one embodiment of the present invention, a real-time channel is established through a Kafka message queue configured with a number of partitions and replicas to receive account transaction logs and balance change data at a millisecond-level transmission rate, including: Set up a Kafka message queue cluster with 3 partitions and 2 replicas, configure JDBC or ODBC read-only permissions to connect to the account system and perform data anonymization; The anonymized account funds data is written to the Kafka cluster using a dual-mode approach of T+1 batch synchronization and real-time incremental synchronization. The anomaly detection engine subscribes to real-time data streams in the Kafka cluster, pulling account transaction records and balance change data at millisecond levels as input sources for comparison.
[0015] In one embodiment of the present invention, closed-loop circulation includes: The standardized enterprise information, spatial coordinates, and entity IDs in the generated business opportunity allocation results are transferred to the industry feature identification stage as input data for generating the panoramic portrait data of the industrial chain. The generated industry chain panorama data, including the industry chain list, industry tags, and enterprise hierarchical relationships, is transferred to the capital monitoring stage to serve as a targeted monitoring list for generating risk warning signals. The abnormal fund tags, large inflow markers, and high-value customer tags from the generated risk warning signals are sent back to the map visualization module, and the corresponding enterprises are highlighted on the business map, forming a fully automated closed loop of "collection → fusion → calculation → monitoring → allocation → display → feedback".
[0016] This invention also proposes a county-level digital management map data analysis system, comprising: The normalization module is used to acquire multi-source heterogeneous county-level financial operation data, and performs conflict detection and normalization processing based on primary key identification rules and priority coverage rules to generate standardized unified subject data. The business opportunity module is used to standardize the address information in the unified entity data and perform spatial coordinate transformation. Based on the weighted matching degree of administrative division affiliation and service radius, and preset service outlets, it generates business opportunity allocation results. The profiling module is used to identify industry characteristics and extract indicators from the unified subject data based on the keyword rule engine, and to perform cleaning and normalization processing using distributed parallel computing to generate panoramic profiling data of the industrial chain. The early warning module is used to collect account fund change data in real time, dynamically generate anomaly identification thresholds by combining customer ratings and historical transaction characteristics, compare the change data with the thresholds to generate risk warning signals, and perform closed-loop circulation and visualization of business opportunity allocation results, industry chain panoramic profile data and risk warning signals.
[0017] This invention discloses a county-level digital business map data analysis method and system. It generates standardized unified subject data by performing conflict detection and normalization on multi-source heterogeneous data based on primary key identification and priority coverage rules. It automatically allocates business opportunities by standardizing address information through mapping and spatial coordinate transformation, and calculating weighted matching degrees based on administrative division affiliation and service radius weights. It generates a panoramic portrait of the industrial chain by identifying industry characteristics and extracting indicators through a keyword rule engine and distributed parallel computing. It generates risk warning signals by collecting real-time capital change data and comparing it with dynamic thresholds, and then performs a closed-loop circulation and visualization of business opportunity allocation results, industrial chain portraits, and risk warnings. This method effectively solves the problems of difficult integration of multi-source financial data in counties, lack of key tools for customer group management, and lack of data support for management decisions. It achieves a fully automated closed loop from data collection, fusion, and calculation to monitoring, allocation, and display, significantly improving the digital business capabilities and marketing management efficiency of county-level institutions, and enhancing the accuracy and intelligence of financial services reaching lower levels.
[0018] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0019] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart of a county-level digital business map data analysis method according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the structure of a county-level digital business map data analysis system according to an embodiment of the present invention. Detailed Implementation
[0020] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0021] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0022] The following description, with reference to the accompanying drawings, describes a method and system for analyzing county-level digital business map data according to an embodiment of the present invention.
[0023] Figure 1 This is a flowchart of a generator stator grounding fault location method considering neutral point displacement voltage according to an embodiment of the present invention.
[0024] like Figure 1 As shown, a method for analyzing county-level digital business map data includes the following steps: S1. Obtain multi-source heterogeneous county-level financial operation data, perform conflict detection and normalization processing based on primary key identification rules and priority coverage rules, and generate standardized unified subject data. S2, standardize the address information in the unified main data and perform spatial coordinate transformation, calculate the weighted matching degree of administrative division affiliation and service radius weight and preset service outlets, and generate business opportunity allocation results; S3. Based on the keyword rule engine, the unified subject data is used to identify industry characteristics and extract indicators. Distributed parallel computing is used for cleaning and normalization to generate a panoramic portrait of the industrial chain. S4 collects account fund change data in real time, dynamically generates anomaly identification thresholds by combining customer ratings and historical transaction characteristics, compares the change data with the thresholds to generate risk warning signals, and performs closed-loop circulation and visualization of business opportunity allocation results, industry chain panoramic profile data and risk warning signals.
[0025] This invention uses map data components as the core technology carrier to construct a complete technical architecture consisting of three major business modules, a unified data platform, and visualization output. Through end-to-end automation of data collection, fusion, computation, monitoring, allocation, and display, it achieves unified data source, unified map management, and a closed-loop process for county-level operational data. The overall technical solution adopts a microservice architecture, streaming computing, and spatial algorithms, supporting high concurrency, low latency, and large-scale data processing. Specifically, it includes: 1. Digital Sand Table Technology Module.
[0026] Core functions: automatic identification of business opportunities, latitude and longitude positioning, matching of marketing agencies based on proximity, map visualization and automatic allocation.
[0027] ① Multi-source data fusion technology.
[0028] This invention employs a data acquisition mechanism combining ETL full extraction and CDC incremental synchronization. Full extraction is performed at a fixed daily time T=05:00, using the following algorithm: FullExtract(T) = Union{Source1(T), Source2(T),…,Source n (T)}; Simultaneously, a unique identifier is provided using a hash algorithm: KeyHash(E) = Hash(Unified Social Credit Code) MOD N.
[0029] During the ETL extraction, transformation, and loading transformation stage, the system automatically executes data conflict resolution rules to deduplicate, verify, overwrite, and normalize data extracted from multiple sources, ensuring that the data entering the data platform is unique, accurate, and consistent.
[0030] Synchronization failure handling: Automatically log the error → enter the circuit breaker queue → trigger SMS + system alarm, maximum number of retries 3, with an interval of 1 minute.
[0031] Data conflict resolution rules (embedded in the ETL transformation process): Primary key rule: The unified social credit code is used as the unique primary key. If there is no credit code, the "company name + registered address + legal representative" triplet is used for verification to ensure that each record can be uniquely identified during the ETL process; Priority rule: The priority is set as government data > business data > internal data > external data. Low priority fields are automatically overwritten during the ETL transformation stage to ensure data authority; Time rule: The latest valid timestamp is used as the standard. The latest status data is retained before ETL loading to avoid conflicts between old and new data.
[0032] Through the above mechanism, the ETL full extraction and conflict resolution rules are deeply coupled and the process is integrated, so as to realize the unified storage of multi-source data after automatic cleaning, deduplication, and calibration.
[0033] ② Address standardization and spatial positioning technology.
[0034] Address standardization rules: Standardized mapping of administrative divisions at the provincial, municipal, and county levels; Address error correction: Correcting typos, abbreviations, and incorrect road names based on the county address database; Anomaly handling: Address cannot be standardized → Marked as an address to be verified → Triggers manual verification → Latitude and longitude conversion is performed after verification is passed; Latitude and longitude conversion: Calling the map API to convert the standardized address to the WGS84 coordinate system (Lon, Lat) to complete the spatial mapping.
[0035] ②Weighted Euclidean distance nearest matching algorithm.
[0036] Basic formula: d = √[(x2-x1)² + (y2-y1)²]; Parameter description: d: Spatial distance between enterprise coordinates and branch coordinates; x1, y1: Longitude and latitude of the enterprise in WGS84 coordinate system after standardized address conversion; x2, y2: Longitude and latitude of the bank service branch in WGS84 coordinate system.
[0037] In conjunction with county-level financial service scenarios, the weighting coefficients for basic distances have been optimized: Service radius of branches: 0.6; Administrative division affiliation: 0.3; Historical marketing matching: 0.1.
[0038] Calculation logic: First, filter the branch outlets belonging to the administrative division, then calculate the weighted distance, and automatically allocate business opportunities to the best service outlets and account managers according to the principle of minimum distance, so as to achieve accurate matching and efficient reach of customers in the county.
[0039] Furthermore, by defining a universal primary key identification logic, key feature combinations that uniquely identify business entities are extracted from multi-source data, serving as the basis for data association and deduplication. Building upon this, a configurable priority coverage strategy is introduced, assigning weight levels based on the authority and timeliness of different data sources. When multi-source data conflicts are detected for the same entity, field-level validation, coverage, or normalization operations are automatically performed according to preset rules, thereby eliminating data redundancy and contradictions and ensuring the consistency and accuracy of the output data. This process encompasses fully automated processing from data extraction and transformation to loading, aiming to solve common problems such as conflicts, duplication, and untimely synchronization during multi-source data fusion. As a specific implementation method, the system can adopt a collection mechanism that combines full extraction and incremental synchronization, using the unified social credit code as the core of the primary key identification rule, or using a triple consisting of enterprise name, registered address and legal representative for verification when missing; in conflict resolution, it follows the priority order of government data over business data, business data over internal data and external data, and combines the latest valid timestamp rule to automatically complete the overwriting of low priority fields and the retention of the latest status data, ultimately realizing the automatic cleaning and unified storage of multi-source data.
[0040] Specifically, the process begins by standardizing and correcting the address information in the unified entity data. This is achieved by establishing administrative division hierarchical mapping relationships and an address database verification mechanism to correct typos, abbreviations, or incorrect street names, transforming heterogeneous address text into standardized address strings. Subsequently, spatial positioning services are used to convert the standardized address information into spatial coordinate data, enabling precise mapping of business entities in geographic space. Then, based on a pre-defined service network layout, and considering both administrative division constraints and service radius coverage, a weighted matching model is constructed to calculate the weighted matching degree between the spatial coordinate data and each pre-defined service network. Administrative division affiliation is used to filter candidate network areas, while service radius weights quantify spatial accessibility. Finally, based on the optimal principle of weighted matching degree, business opportunity allocation results are generated, automatically associating business entities with the optimal service network. As one implementation method, the standardized address can be converted to longitude and latitude in the WGS84 coordinate system, and a weighted Euclidean distance algorithm can be used for calculation. The basic distance formula is... The allocation is then comprehensively weighted by combining the service radius weight of the outlets, the administrative division weight, and the historical marketing matching weight, and is completed according to the principle of minimum distance.
[0041] 2. Industry Sketching Technology Module.
[0042] Core functions: batch processing of data from distinctive industrial chains, automatic extraction of indicators, cleaning, processing and integration, and standardized output.
[0043] Batch computing service interface construction: Adopting a distributed parallel computing framework, it achieves task splitting, concurrent execution, queuing and rate limiting, and failure retries through thread pool scheduling, supporting high-throughput processing of tens of thousands of enterprise lists, and providing computing power support for batch operation of keyword rule engine; Automatic extraction of characteristic indicators: Based on regular expression matching + keyword rule engine, it automatically identifies industry tags, scale level, and operating status from fields such as enterprise name, business scope, revenue, and tax payment, forming a standardized indicator set; Distributed computing and rule engine collaboration mechanism: The distributed framework automatically splits large amounts of enterprise data into multiple sub-tasks according to fixed shard sizes, and concurrently schedules them to different computing nodes to execute keyword rule matching and indicator extraction. After a single node completes processing, the results are uniformly aggregated, realizing "large data volume splitting, parallel execution of rule engine, and unified result aggregation".
[0044] Data cleaning, processing, and integration: Missing value imputation, outlier truncation, and dimension normalization are employed. The normalization formula is: x' = (x min) / (max min); x: Original indicator value (such as revenue, tax payment, number of employees); min: Minimum value of the indicator in the current industry chain; max: Maximum value of the indicator in the current industry chain; x': Normalized standard indicator value (range [0,1]).
[0045] The system aggregates enterprises at different levels, including core enterprises, supporting enterprises, and upstream and downstream enterprises, to generate a panoramic portrait of the county's distinctive industries.
[0046] Output results: Data is sent back to the front end via file interfaces such as Excel / CSV / JSON, and supports one-click export of standardized industry chain lists.
[0047] Furthermore, firstly, relying on a pre-set keyword rule engine, the system performs deep scanning and semantic matching on multi-dimensional fields such as enterprise name, business scope, revenue, and tax payment contained in the standardized unified entity data. This automatically identifies feature tags, scale levels, and operating status representing industry attributes, forming a preliminary standardized indicator set. Subsequently, to cope with the high-throughput processing requirements of a large-scale enterprise list, a distributed parallel computing architecture is introduced. The massive indicator data to be processed is automatically split into multiple sub-tasks according to a fixed sharding strategy and concurrently scheduled to different computing nodes for subsequent processing. During the calculation process, the system performs cleaning operations on the extracted indicator data, including missing value imputation and outlier truncation, and further implements dimension normalization processing to eliminate numerical differences between indicators of different dimensions, ensuring that the data is comparable under the same standard scale. As a specific implementation method, a minimum-maximum normalization algorithm can be used, based on the formula... Map the original indicator values to a preset range, where These are the original indicator values. and These represent the minimum and maximum values of this indicator within the current industry chain. These are the normalized standard indicator values. Finally, the system aggregates the processed data according to the hierarchical relationship between core enterprises, supporting enterprises, and upstream and downstream enterprises, generating a panoramic portrait of the industrial chain that reflects the overall picture of the county's characteristic industries.
[0048] 3. Funds monitoring technology module.
[0049] Core functions: real-time collection, dynamic tracking, anomaly identification, and risk warning of account funds.
[0050] Real-time channel: Employs Kafka message queues to achieve millisecond-level transmission of account fund data streams. With 3 partitions and 2 replicas, it ensures high availability, low latency, and no data loss, providing real-time data stream support for dynamic threshold calculation and anomaly detection. Data integration: Uses JDBC / ODBC read-only access to integrate with the account system, and performs data anonymization. Employs both T+1 batch synchronization and real-time incremental synchronization modes to ensure data security, integrity, and timeliness.
[0051] Kafka's working mechanism in conjunction with dynamic thresholds: Kafka receives real-time fund data such as account transaction history and balance changes, pushing it to the anomaly detection engine in milliseconds. The system automatically generates dynamic thresholds based on customer ratings, industry characteristics, historical trading habits, and seasonal fluctuations, continuously comparing and judging real-time inflow data. When fund changes exceed the dynamic thresholds, a real-time alert is immediately triggered, forming a closed-loop chain of real-time data collection, real-time calculation, and real-time alerts. Data capture and identification: Key indicators such as fund inflows, outflows, large transactions, and sudden balance changes are captured in real-time according to preset rules. Combined with customer ratings and industry characteristics, a comprehensive judgment is made, and fund tracking reports and alert lists are automatically generated.
[0052] Furthermore, firstly, a high-throughput message queue channel is used to capture real-time changes in account funds, ensuring low latency and integrity of data flow. Then, based on multi-dimensional customer characteristic parameters and historical behavior patterns, an anomaly detection threshold model is adaptively constructed. This model can dynamically adjust according to factors such as customer rating, industry attributes, and seasonal fluctuations, thus overcoming the poor adaptability of fixed thresholds in complex financial scenarios. Next, real-time inflow fund change data is continuously compared with the dynamically generated anomaly detection threshold. Once data deviates from the preset safety range, a risk warning signal is immediately triggered. Finally, this risk warning signal is logically linked and circulated in a closed loop with the business opportunity allocation results and industry chain panoramic profile data generated in the preceding steps, and integrated and displayed in a visual interface, achieving simultaneous presentation and coordinated decision-making regarding operational risks and marketing opportunities. As one implementation method, the system can use a Kafka message queue as a real-time channel to receive transaction flow and balance change data, and automatically generate dynamic thresholds based on customer ratings, industry characteristics, and historical transaction habits. When the fund change exceeds the threshold, a real-time alert is immediately triggered. At the same time, risk signals including abnormal fund tags, large inflow markers, and high-value customer tags are sent back to the map module for highlighting.
[0053] 4. Supporting technical technology and data closed loop.
[0054] This invention adopts a microservice layered architecture, decoupling the three core capabilities of digital sand table, industry quick mapping, and capital monitoring into independent and scalable service units. Through a unified data platform and standard API interfaces, it achieves interconnection and interoperability between modules, and deep collaboration with the data closed loop, forming a fully automated operation system of "collection → fusion → calculation → monitoring → allocation → application → feedback".
[0055] The microservice layered architecture divides the system into a data access layer, a data middle platform layer, a business service layer, and an application presentation layer. Each layer is deployed independently and scales elastically: Data Access Layer: Provides full ETL extraction, incremental CDC synchronization, and a real-time Kafka channel, providing stable, real-time, and standardized data input for the closed loop; Data Middle Platform Layer: Provides a unified subject library, indicator library, tag library, and spatial library, achieving single data source and global reuse, serving as the data hub for the closed loop; Business Service Layer: Digital sandbox, industry quick mapping, and financial monitoring run as independent microservices, exchanging data through standard API calls, without interference and capable of independent upgrades; Application Presentation Layer: Outputs results with map visualization, industry profiling, financial dashboards, and operational reports, supporting decision-making and marketing implementation.
[0056] The process for each stage of the data closed loop is as follows: (1) Digital Sand Table → Industry Quick Sketch
[0057] Transfer data: standardized enterprise information, spatial coordinates, and entity ID.
[0058] Core algorithms: address standardization algorithm, hash deduplication algorithm, and weighted Euclidean distance matching algorithm, automatically feed spatial enterprise data into the industry processing module, realize positioning first and then profiling, avoid duplicate input, and improve the accuracy of industry tags.
[0059] (2) Industry quick sketch → Capital monitoring.
[0060] Circulation data: supply chain list, industry tags, enterprise level, core / supporting / upstream and downstream relationships.
[0061] Core algorithms: keyword rule engine, distributed parallel computing, and data normalization algorithm. The monitoring list is pushed to the industry chain in a targeted manner to realize full-chain fund tracking, upgrading from "single-household monitoring" to "circle chain group monitoring".
[0062] (3) Funds monitoring → Digital sand table. Flow data: abnormal fund tags, large inflow markers, fund outflow warnings, and high-value customer tags.
[0063] Core algorithms: Kafka real-time stream computing and dynamic threshold anomaly recognition algorithm, which transmit risk and business opportunity signals back to the map module in real time, highlighting and highlighting them on the sandbox and providing tiered alerts, enabling "one map to see business opportunities and risks".
[0064] (4) Three major modules → Business reports → Marketing allocation → Data feedback.
[0065] Circulation data: operating indicators, business opportunity conversion rate, capital accumulation, and institution ranking.
[0066] Core algorithms: A comprehensive evaluation weighted algorithm and a task intelligent scheduling and retry algorithm form an automatic closed loop of decision-making, execution, feedback and optimization, enabling data to be collected once, reused throughout the process and continuously iterated.
[0067] The method of this invention effectively solves the problems of difficulty in integrating multi-source heterogeneous financial data in county-level areas, lack of tools for customer group management, and lack of data support for management decisions. It realizes a fully automated closed loop from collection, integration, and calculation to monitoring, distribution, and display, significantly improving the digital operation capabilities and marketing management efficiency of county-level institutions, and enhancing the accuracy and intelligence of financial services in rural areas.
[0068] In another example, taking a county's distinctive tea industry chain: the digital sandbox automatically integrated 32 newly registered tea enterprises, mapped them all using latitude and longitude positioning, and allocated them to 3 township outlets according to the nearest matching algorithm, reducing the time for business opportunity allocation from 2 days to 5 minutes; the industry sketching module imported the tea enterprise list in batches, automatically extracted four categories of tags: planting, processing, sales, and packaging, completed data cleaning and aggregation, and generated a standard industry chain profile, improving manual processing efficiency by 90%; the capital monitoring module monitored the accounts of 128 tea enterprises in real time, identified 8 abnormal large inflows and 5 capital outflow warnings, improving the warning response time from T+1 to real time; unified reports were output as a county-level tea industry operation dashboard, allowing management to grasp customer coverage, business opportunity conversion, and capital accumulation in real time, resulting in a 42% increase in tea enterprise loan disbursement and a 28% increase in capital retention rate that month.
[0069] The application scenarios of the method in this invention, through technology integration and process reengineering, comprehensively realize the visualization of county-level financial operations, marketing automation, management datafication, and real-time risk monitoring, and have significant technological innovation and business application value.
[0070] To implement the method of the above embodiments, such as Figure 2 As shown, this invention proposes a county-level digital business map data analysis system 10, comprising: The normalization module 100 is used to acquire multi-source heterogeneous county-level financial operation data, perform conflict detection and normalization processing based on primary key identification rules and priority coverage rules, and generate standardized unified subject data.
[0071] The business opportunity module 200 is used to standardize the address information in the unified subject data and perform spatial coordinate transformation. Based on the administrative division affiliation and service radius weight calculation and the weighted matching degree of the preset service outlets, it generates business opportunity allocation results.
[0072] The profiling module 300 is used to identify industry characteristics and extract indicators from the unified subject data based on the keyword rule engine, and to perform cleaning and normalization processing using distributed parallel computing to generate panoramic profiling data of the industrial chain.
[0073] The early warning module 400 is used to collect account fund change data in real time, dynamically generate anomaly identification thresholds by combining customer ratings and historical transaction characteristics, compare the change data with the thresholds to generate risk warning signals, and perform closed-loop circulation and visualization of business opportunity allocation results, industry chain panoramic profile data and risk warning signals.
[0074] The system of this invention effectively solves the problems of difficulty in integrating multi-source financial data in county-level financial institutions, lack of tools for customer group management, and lack of data support for management decisions. It realizes a fully automated closed loop from collection, integration, and calculation to monitoring, distribution, and display, significantly improving the digital operation capabilities and marketing management efficiency of county-level institutions.
[0075] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0076] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
Claims
1. A method for analyzing county-level digital business map data, characterized in that, include: Obtain multi-source heterogeneous county-level financial operation data, perform conflict detection and normalization processing based on primary key identification rules and priority coverage rules, and generate standardized unified subject data; The address information in the unified subject data is standardized and mapped and spatial coordinates are transformed. Based on the weighted matching degree of administrative division affiliation and service radius with preset service outlets, business opportunity allocation results are generated. Based on the keyword rule engine, the unified subject data is used to identify industry characteristics and extract indicators. Distributed parallel computing is used for cleaning and normalization to generate a panoramic portrait of the industrial chain. Real-time collection of account fund change data, combined with customer rating and historical transaction characteristics to dynamically generate anomaly identification thresholds, comparison of change data with thresholds to generate risk warning signals, and closed-loop circulation and visualization of business opportunity allocation results, industry chain panoramic profile data and risk warning signals.
2. The method as described in claim 1, characterized in that, The process involves acquiring multi-source, heterogeneous county-level financial operation data, performing conflict detection and normalization based on primary key identification rules and priority coverage rules, and generating standardized, unified entity data, including: A full extraction is performed at a fixed time every day. Multi-source snapshot data is integrated through union operation, and a hash algorithm is used to perform a modulo operation on the unified social credit code to generate a unique identifier key. When there is no unified social credit code, the triplet of "enterprise name + registered address + legal representative" is used as the basis for unique identification. Based on the priority of government data > business data > internal data > external data, conflicting fields are automatically processed by having higher priority fields override lower priority fields, while retaining the data status of the latest valid timestamp, so as to obtain standardized and unified main data after deduplication, verification and overwriting.
3. The method as described in claim 1, characterized in that, The process of standardizing and mapping the address information in the unified entity data and transforming its spatial coordinates, calculating the weighted matching degree between administrative division affiliation and service radius and preset service outlets, and generating business opportunity allocation results includes: The address information is standardized and mapped to the three-level administrative divisions of province, city, district and county, and the misspellings, abbreviations and incorrect road names are automatically corrected based on the county address database; The corrected standardized address is converted into longitude and latitude in the WGS84 coordinate system by calling the map API interface, thus obtaining the spatial coordinate data of the enterprise and the spatial coordinate data of the bank service outlets. Based on two types of spatial coordinate data, the weighted Euclidean distance is calculated by combining the service radius weight of the outlet, the administrative division weight, and the historical marketing matching weight, and the business opportunity allocation result is obtained.
4. The method as described in claim 3, characterized in that, Calculating the weighted Euclidean distance includes: Based on the weight of administrative division affiliation, a set of service outlets belonging to the same administrative region as the enterprise is selected; For each point in the set, calculate the basic spatial distance between the enterprise coordinates and the point coordinates: d = √[(x2 - x1)² + (y2 - y1)²] Where x and y are the enterprise coordinates, and x2 and y2 are the network coordinates; The weighted matching degree is obtained by multiplying the base distance by the service radius weight of the outlet, the administrative division weight, and the historical marketing matching weight, respectively, by 0.6, 0.3, and 0.
1. The network outlets are sorted in ascending order of weighted matching degree, and the outlet with the highest ranking is determined as the optimal service outlet to obtain the business opportunity allocation result.
5. The method as described in claim 1, characterized in that, The unified subject data is analyzed and its industry characteristics are identified and indicators are extracted using a keyword rule engine. Distributed parallel computing is then used for cleaning and normalization to generate a comprehensive picture of the industry chain, including: A distributed parallel computing framework is used to split a large amount of unified main data into multiple subtasks with a fixed partition size, and then schedules them concurrently to different computing nodes to perform keyword rule matching. Based on regular expression matching and keyword rule engine, the system automatically identifies industry tags, scale level and operating status from enterprise name, business scope, revenue and tax fields to form a standardized indicator set; Missing value imputation and outlier truncation are performed on the original indicator values in the standardized indicator set, and the normalization formula is used to map them to the [0,1] interval to obtain the panoramic portrait data of the industrial chain.
6. The method as described in claim 5, characterized in that, Normalization processing includes: Obtain the original value x of a certain indicator in the current industry chain, and calculate the minimum value min and the maximum value max of this indicator among all enterprise data in the current industry chain; Substituting x, min, and max into the normalization formula, we obtain the standard index value x': x' = (x - min) / (max - min); All x' are aggregated according to the hierarchical relationship of core enterprises, supporting enterprises and upstream and downstream enterprises, and output panoramic portrait data of the industrial chain containing multi-level relationships.
7. The method as described in claim 1, characterized in that, Real-time data acquisition and dynamic threshold comparison, including: A real-time channel is established through a Kafka message queue configured with the number of partitions and replicas to receive account transaction records and balance change data at a millisecond-level transmission rate. Dynamic thresholds are automatically generated based on customer ratings, industry characteristics, historical trading habits, and seasonal fluctuations, and real-time fund change data is continuously compared with the dynamic thresholds. When a change in funds is detected to exceed a dynamic threshold, a real-time alert is immediately triggered to obtain a risk warning signal.
8. The method as described in claim 7, establishing a real-time channel through a Kafka message queue configured with a number of partitions and replicas, and receiving account transaction logs and balance change data at a millisecond-level transmission rate, includes: Set up a Kafka message queue cluster with 3 partitions and 2 replicas, configure JDBC or ODBC read-only permissions to connect to the account system and perform data anonymization; The anonymized account funds data is written to the Kafka cluster using a dual-mode approach of T+1 batch synchronization and real-time incremental synchronization. The anomaly detection engine subscribes to real-time data streams in the Kafka cluster, pulling account transaction records and balance change data at millisecond levels as input sources for comparison.
9. The method according to claim 1, characterized in that, Closed-loop circulation includes: The standardized enterprise information, spatial coordinates, and entity IDs in the generated business opportunity allocation results are transferred to the industry feature identification stage as input data for generating the panoramic portrait data of the industrial chain. The generated industry chain panorama data, including the industry chain list, industry tags, and enterprise hierarchy, is transferred to the capital monitoring stage to serve as a targeted monitoring list for generating risk warning signals. The abnormal fund tags, large inflow markers, and high-value customer tags from the generated risk warning signals are sent back to the map visualization module, and the corresponding enterprises are highlighted on the business map, forming a fully automated closed loop of "collection → fusion → calculation → monitoring → allocation → display → feedback".
10. A county-level digital business map data analysis system, characterized in that, include: The normalization module is used to acquire multi-source heterogeneous county-level financial operation data, perform conflict detection and normalization processing based on primary key identification rules and priority coverage rules, and generate standardized unified subject data. The business opportunity module is used to standardize the address information in the unified entity data and perform spatial coordinate transformation. Based on the weighted matching degree of administrative division affiliation and service radius with preset service outlets, it generates business opportunity allocation results. The profiling module is used to identify industry characteristics and extract indicators from the unified subject data based on the keyword rule engine, and to perform cleaning and normalization processing using distributed parallel computing to generate panoramic profiling data of the industrial chain. The early warning module is used to collect account fund change data in real time, dynamically generate anomaly identification thresholds by combining customer ratings and historical transaction characteristics, compare the change data with the thresholds to generate risk warning signals, and perform closed-loop circulation and visualization of business opportunity allocation results, industry chain panoramic profile data and risk warning signals.