System and method for dynamic switching of data source registration
Through the hierarchical aggregation model and intelligent matching engine, dynamic registration and real-time scheduling of data sources are realized, solving the problem of insufficient flexibility and intelligence in traditional methods, and improving the stability and efficiency of data access.
Patent Information
- Application Number
- CN202510824702.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-19
AI Technical Summary
Traditional data source registration methods are difficult to cope with the dynamic addition and deletion of business changes or routing adjustment requirements, resulting in inefficient call efficiency and waste of resources, and the inability to achieve flexibility and intelligence.
The hierarchical aggregation model and an intelligent matching engine are adopted to realize dynamic registration and real-time scheduling of data sources through small collection registration, metadata annotation, time window clustering and dynamic weight adjustment, combined with an exception circuit breaking mechanism.
It improves the flexibility and scalability of data management, ensures the stability and efficiency of data access, and improves data call efficiency and system performance.
Smart Images

Figure CN120336293A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital data processing, and particularly relates to a system and method for dynamic switching of data source registration. Background Art
[0002] Dynamic switching of data source registration is a technology for dynamically managing multi-data source connections and flexibly switching data access paths during application operation; In traditional switching methods, data source configurations are fixed during the system startup phase, making it difficult to meet the requirements for dynamically adding, deleting, or routing and adjusting data sources due to business changes. For example, in a multi-tenant scenario, it is impossible to allocate independent data sources to new tenants in real time, lacking an automatic clustering and matching mechanism based on data characteristics (such as business tags, performance metrics), and manual coding is required to implement data source routing, resulting in low call efficiency and the inability to dynamically optimize paths based on real-time performance data, easily causing high latency or resource waste.
[0003] Therefore, the present invention proposes a system and method for dynamic switching of data source registration. Through a hierarchical aggregation model (small collection → large collection) and an intelligent matching engine, dynamic registration, clustering, and real-time scheduling of data sources are achieved, and the stability and efficiency of data access are ensured through weight dynamic adjustment and exception fusing mechanisms, solving the core problems of traditional solutions in terms of flexibility, intelligence, and reliability. Summary of the Invention
[0004] Technical problems to be solved: Traditional switching methods are difficult to meet the requirements for dynamically adding, deleting, or routing and adjusting data sources due to business changes.
[0005] In view of the deficiencies of the prior art, the present invention provides a system and method for dynamic switching of data source registration, thereby solving the technical problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for dynamic switching of data source registration, the method comprising the following steps: Step 1, data storage step; Step 2, data calling step; Step 2.1, business request parsing, request feature extraction, extracting key information from the business request to provide a basis for subsequent data source matching; determining the weight range and performance index requirements of the required data source according to the urgency and performance requirements of the business request; Step 2.2, Large collection matching: Use the rule engine to screen out the eligible large collections according to the request features; sort the screened large collections in descending order of weight, and preferentially select the large collections with high weights for subsequent small collection positioning; check the metrics such as the call duration and error rate of the data sources within the large collection. If they exceed the set thresholds, reduce or exclude the weight of this large collection. Step 2.3, Small collection positioning: In the selected large collection, screen out the small collections that meet the request requirements according to the business tags and performance tags; sort the screened small collections in descending order of weight, and preferentially select the small collections with high weights; check the metrics such as the call duration and error rate of the small collections again. If they exceed the set thresholds, skip this small collection and select the next eligible small collection. Step 2.4, Data source call and feedback: According to the located small collection, call the corresponding data source to perform data operations; monitor the metrics such as the call duration and error rate of the data source in real time, and feedback the monitoring data to the intelligent analysis module for dynamically adjusting the weights of the data sources; the intelligent analysis module uses the sliding window algorithm and the exponentially weighted moving average method to dynamically adjust the weights of the small collections and large collections according to the monitoring data; if an exception occurs during the call, trigger the automatic fusing mechanism, pause the call of this data source, switch to the backup data source, and notify the operation and maintenance personnel to handle it.
[0007] In a possible implementation manner, the data storage step further includes the following steps: Step 1.1, Small collection registration and metadata construction: Register each data source as an independent small collection, and assign a unique identifier to each data source small collection for accurate identification during subsequent data storage, call, and management processes, ensuring that there is no confusion between different data sources. Step 1.2, Large collection clustering and weight calculation: Cluster the small collections with similar usage characteristics into large collections, and determine the clustering relationship by analyzing the call frequency and call time distribution data of the data sources in each time window; adopt the DBSCAN density clustering algorithm, and use the usage frequency and time distribution of the data sources in the time window as the feature vectors for clustering. Step 1.3, Data storage and protection: Select the Elasticsearch + PostgreSQL combination as the storage medium to store metadata; ensure the data consistency and integrity of the collection through data backup and version control.
[0008] In a possible implementation manner, the reduction or exclusion of the large collection weight is based on the set threshold judgment rule. If the call duration of more than a certain proportion (such as 30%) of the data sources within the large collection exceeds the threshold or the error rate exceeds the threshold, the weight of this large collection is reduced; the weight reduction can adopt the formula , that is, the original weight multiplied by 0.8.
[0009] In a possible implementation, the small collection screening adopts a string matching algorithm to match the requested service tags and performance tags with the corresponding tags of the small collection one by one.
[0010] In a possible implementation, the small collection weight sorting uses the quicksort algorithm to sort the small collection weights.
[0011] In a possible implementation, the weight dynamic adjustment uses the sliding window algorithm and the exponential weighted moving average method to dynamically adjust the weights of the small collection and the large collection; set the time window size, calculate the average call duration and error rate metrics of the data source within this window, and adjust the weights according to the metric changes; the formula for the exponential weighted moving average method is , where is the current weight, is the smoothing coefficient, takes values from 0 to 1, is the metric score within the current time window, is the weight at the previous moment.
[0012] In a possible implementation, a data source registration dynamic switching system for executing the above-mentioned method for data source registration dynamic switching includes the following modules: The data source registration module is responsible for the entry of the data source, metadata annotation, and dynamic weight initialization to form independent small collection units; The data aggregation and management module realizes the clustering of small collections into large collections, dynamically calculates the weights of the large collections, and completes data storage and protection; The intelligent matching and invocation module parses the service requests, dynamically matches the data source collections and executes the invocations, and at the same time feeds back the monitoring data; The visualization management interface module provides a graphical interface for system operation and maintenance and configuration, reducing the operation threshold; The system interaction process module is responsible for coordinating the automated execution of the entire process of data source registration, clustering, matching, and invocation.
[0013] Beneficial effects compared with the prior art: 1. In this solution, through the small collection registration and metadata annotation of the data source, combined with time window clustering and weight calculation to form a large collection, the hierarchical dynamic management of the data source is realized. The data source can be flexibly aggregated and split according to business requirements, improving the system's organizational ability for multi-type and multi-domain data sources, facilitating the quick positioning and invocation of the required data source, and enhancing the flexibility and scalability of data management; 2. In this solution, during data invocation, through processes such as business request parsing, large collection matching, small collection positioning, and dynamic weight adjustment, intelligent matching and efficient switching of data sources are achieved. Based on real-time monitoring data, the weights are dynamically optimized, combined with risk assessment and exception handling mechanisms, ensuring the stability and reliability of data source invocation, and improving data access efficiency and system performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The above description is only an overview of the technical solution of the present invention. In order to understand the technical means of the present invention more clearly and be able to implement it according to the content of the specification, the following describes the preferred embodiments of the present invention in detail with reference to the accompanying drawings.
[0015] Figure 1 is the flowchart of the method steps of the present invention; Figure 2 is the schematic diagram of the system framework of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] The preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. However, the present invention can be implemented in various different forms, so the present invention is not limited to the embodiments described below; The technical solutions in the embodiments of the present application are to solve the problems in the above-mentioned background technology, and the general idea is as follows: Embodiment 1: Please refer to Figure 1 As shown, this embodiment introduces a method for dynamic switching of data source registration, including the following steps: Step 1. Data storage step Step 1.1. Small collection registration and metadata construction Step 1.1.1. Data source information entry The data source is the basic unit for data storage and invocation. Each data source (such as a MySQL database connection, a RESTful API interface, a file storage path, etc.) is registered as an independent small collection; when entering, the basic connection information needs to be completely recorded, such as the URL of the database (including IP, port, database name), username, password, driver type; the address of the API interface, request method (GET / POST, etc.), request parameter format, response data format, etc. Example: Register a small collection of MySQL databases, and the entered information is as follows: Generation of small collection ID: Assign a unique identifier to each data source small collection for accurate identification during subsequent data storage, invocation, and management processes, ensuring that there is no confusion between different data sources; for example, a format similar to "db_001" can be used, where "db" represents the database type and "001" is the sequential number; Connection Information Record: Completely record the connection-related information of the data source; for data sources of database type, record the URL in detail. The URL should include the IP address, port number, and database name so that the program can accurately locate the database location; at the same time, record the username and password as the identity credentials for accessing the database; also record the driver type to clarify which database driver program is used to establish the connection; taking the MySQL database as an example, the URL is in the form of "jdbc:mysql: / / 192.168.1.100:3306 / user_db", where "192.168.1.100" is the IP address, "3306" is the port number, "user_db" is the database name, and "jdbc:mysql:" indicates using the JDBC driver protocol of MySQL; Information Integration: Integrate the generated small collection ID, recorded connection information, etc. to form a complete registration information unit for the small collection of data sources. This unit contains the key information required for subsequent data operations, laying a foundation for data storage and invocation; Step 1.1.2, Metadata Annotation (1) Business Label. The business label is used to identify the business area to which the data source belongs, facilitating the subsequent quick location of the data source according to business requirements; during annotation, it is necessary to combine the actual business situation to ensure that the label accurately reflects the purpose of the data source; Example: If the data source is used to store user registration and login-related data, then annotate the business label as ; if it is used to store order creation and payment data, then annotate it as ; (2) Performance Label. The performance label is annotated based on the performance of the data source and can be determined through performance monitoring data (such as average response time, throughput, etc.) over a period of time; Example: If the data source can still maintain a low-latency response (average response time < 100ms) under high concurrency, annotate the performance label as ; if the throughput is high (number of requests processed per second > 1000), annotate it as ; (3) Security Label. The security label is used to mark the data security attributes of the data source and is annotated according to whether the data is encrypted during storage and transmission and whether data masking is performed, etc.; Example: If the user password in the data source is stored using SHA-256 encryption, annotate the security label as ; if sensitive information such as the user's mobile phone number is masked in the query results, annotate it as ; Step 1.1.3, Dynamic Weight Initialization The dynamic weight is used to measure the priority of data sources during data invocation. The initial weight can be set according to historical usage data (such as usage frequency, stability, etc.) or initial configuration; If the initial weight is set based on historical usage frequency, the formula is adopted, where is the initial weight, is the historical usage times of this data source, is the total historical usage times of all data sources; Example: Suppose there are 3 data sources with historical usage times of , , respectively. Then the initial weight of data source 1 , the initial weight of data source 2 , and the initial weight of data source 3 ; Step 1.2, Large Aggregation Clustering and Weight Calculation Step 1.2.1, Time Window Clustering. Time window clustering is to cluster small aggregations with similar usage characteristics into large aggregations according to the usage patterns of data sources in different time periods (such as daily, weekly, monthly); by analyzing data such as the invocation frequency and invocation time distribution of data sources within each time window, the clustering relationship is determined; The DBSCAN density clustering algorithm is adopted to perform clustering with the usage frequency and time distribution of data sources within the time window as the feature vectors; set the neighborhood radius and the minimum number of samples . If the number of samples of a certain data source within its neighborhood is greater than or equal to , then it is taken as a core point and forms a clustering cluster with other points within the neighborhood; Example: In a certain e-commerce system, analyze the daily usage frequency of each data source; it is found that data sources A, B, and C have a higher invocation frequency from 9 to 11 am and 3 to 5 pm on weekdays, while other data sources do not have this pattern; set (representing the usage frequency difference threshold), , and through the DBSCAN algorithm, data sources A, B, and C can be clustered into the "High-Frequency Usage on Weekdays Large Aggregation"; Step 1.2.2, Weight Aggregation. The weight of the large aggregation is calculated by weighted calculation of the weights of the internal small aggregations. Different weight coefficients can be assigned according to the importance or performance of the small aggregations during weighted calculation; The weighted average algorithm is adopted, and the formula is , where is the weight of the large aggregation, is the weight of the th small aggregation, is the Weight coefficients of small collections ( ), where is the number of small collections in the large collection; Example: Suppose there are 3 small collections in the "High-frequency Usage Large Collection on Weekdays", and their weights are , , , and the weight coefficients are , , , then the weight of this large collection is ; Step 1.2.3, Resource Limitation and Domain Locking. According to business requirements, clearly define the resource usage scope for each large collection, lock the main domain direction, avoid the mixing of irrelevant data sources, and improve the efficiency and accuracy of data invocation; Example: For the "Financial Transaction Large Collection", it is clearly stipulated that it only includes data sources related to financial transactions, such as transaction databases, payment interfaces, etc.; when registering a new data source, if its business label does not contain content related to "financial transactions", it is not allowed to be added to this large collection; Step 1.3, Data Storage and Protection Step 1.3.1, Selection of Storage Media. Select the Elasticsearch + PostgreSQL combination to store metadata; Elasticsearch is used to implement full-text search for convenient and fast retrieval of metadata information; PostgreSQL is used to store structured metadata, supporting complex queries and transaction processing; Example: Establish an index in Elasticsearch to store the metadata information of small collections and large collections. Each document contains the detailed metadata of small collections or large collections; create a table structure in PostgreSQL to store the structured information of metadata, such as fields like small collection ID, large collection ID, tags, etc.; Step 1.3.2, Collection Protection Mechanism. Ensure the data consistency and integrity of the collection through data backup and version control; regularly take snapshot backups of metadata, and when data anomalies occur, it can be restored to the backup state; record the version of metadata modification operations for convenient traceability and rollback; Example: Use the crontab scheduled task to take snapshot backups of the Elasticsearch index and PostgreSQL database at 2 am every day; add a version number field to the metadata table. Each time the metadata is modified, the version number increases, and the modification time and content are recorded; Step 2, Data Invocation Steps Step 2.1, Business Request Parsing Step 2.1.1, Request Feature Extraction: Extract key information from the business request, including business type (such as user query, order creation, etc.), data format (JSON, XML, etc.), permission level (ordinary user, administrator, etc.), request urgency (high, medium, low), etc., to provide a basis for subsequent data source matching; Example: For a request from a user to query an order, the extracted request features are: business type = "order query", data format = "JSON", permission level = "ordinary user", request urgency = "medium"; Step 2.1.2, Weight Requirement Analysis: Determine the weight range and performance index requirements of the required data sources according to the urgency and performance requirements of the business request; for example, requests with high urgency require data sources with high weights and low latency; ordinary requests can choose data sources with relatively lower weights; Example: If the request urgency is "high", it is required that the data source weight is not less than 0.7 and the average response time < 200ms; if the request urgency is "low", the data source weight is not less than 0.3 and the average response time < 500ms; Step 2.2, Big Collection Matching Step 2.2.1, Rule Screening: Use a rule engine (such as Drools) to screen out the eligible big collections according to the request features; the rules can be set based on business tags, performance tags, security tags, etc.; The rule engine matches the request features with the metadata tags of the big collections through a pattern matching algorithm; for example, if the request business type is "real-time transaction query", then screen out the big collections whose business tags contain "real-time transaction"; Example: Set the rule "If the request business type is'real-time transaction query' and the request urgency is 'high', then screen out the big collections whose business tags contain'real-time transaction' and performance tags contain 'low latency'"; 2.2.2, Weight Sorting: Sort the screened big collections in descending order of weight, and preferentially select the big collections with high weights for subsequent small collection positioning; Use the quicksort algorithm to sort the weights of the big collections, and the time complexity is ; Example: Suppose there are 3 screened big collections with weights of , , , and the sorted order is big collection 2, big collection 1, big collection 3; Step 2.2.3, Risk Assessment: Check the call duration, error rate and other indicators of the data sources within the big collection. If they exceed the set thresholds (such as call duration exceeding 500ms, error rate exceeding 5%), reduce or exclude the weight of this big collection; Set the threshold judgment rule. If the data source call time of more than a certain proportion (such as 30%) in the large collection exceeds the threshold or the error rate exceeds the threshold, then reduce the weight of the large collection; the weight reduction can be achieved by using the formula (that is, the original weight multiplied by 0.8); Example: If there are 5 data sources in a "certain large collection", and the call time of 2 data sources exceeds 500ms, and the exceeding proportion is 40%, then the weight of this large collection will be reduced from to ; Step 2.3, Small collection positioning Step 2.3.1, Label matching. In the selected large collection, according to the business label and performance label, screen out the small collections that meet the request requirements; ensure that the labels of the small collections match the request characteristics; Adopt the string matching algorithm to match the business label and performance label of the request with the corresponding labels of the small collections one by one; Example: If the requested business label is , and the performance label is , screen out the small collections in the selected large collection whose business label contains and whose performance label contains ; Step 2.3.2, Weight screening. For the screened small collections, sort them in descending order of weight, and give priority to selecting small collections with high weights; Also use the quicksort algorithm to sort the weights of the small collections; Example: Suppose there are 4 screened small collections with weights of , , , , and the sorted order is small collection 4, small collection 2, small collection 1, small collection 3; Step 2.3.3, Secondary risk assessment. Check the call time, error rate and other indicators of the small collection again. If they exceed the set threshold, skip this small collection and select the next eligible small collection; Similar to the large collection risk assessment algorithm, set the threshold judgment rule. If the call time or error rate of the small collection exceeds the threshold, then do not select this small collection; Example: If the call time of small collection 1 is 600ms, exceeding the set threshold of 500ms, then skip small collection 1 and select small collection 2 for data source call; Step 2.4, Data source call and feedback Step 2.4.1, Data source call. According to the located small collection, call the corresponding data source and perform data operations (such as query, insert, update, etc.); Example: If the small collection is located as "db_001" (a small collection of MySQL databases), a database connection pool is used to obtain a connection, and SQL statements are executed for data query operations; Step 2.4.2, Monitoring and Feedback: Real-time monitor metrics such as the call duration and error rate of the data source, and feedback the monitoring data to the intelligent analysis module for dynamically adjusting the data source weight; Example: Use Prometheus to monitor the call duration and error rate of the data source, collect data every 10 seconds, and perform visual display through Grafana; at the same time, send the data to the intelligent analysis module; Step 2.4.3, Dynamic Weight Adjustment: The intelligent analysis module uses the sliding window algorithm and the exponentially weighted moving average method to dynamically adjust the weights of the small collection and the large collection according to the monitoring data; Set the time window size (such as 5 minutes), calculate metrics such as the average call duration and error rate of the data source within this window, and adjust the weight according to the metric changes; The formula for the exponentially weighted moving average method is , where is the current weight, is the smoothing coefficient (taking values from 0 to 1, such as 0.2), is the metric score within the current time window, is the weight at the previous moment; Example: If the current weight of the data source , within the new time window, due to the increase in call duration, the new metric score , , then the adjusted weight ; Step 2.4.4, Exception Handling: If an exception occurs during the call (such as connection failure, data reading error), trigger the automatic circuit breaker mechanism, suspend the call to this data source, and switch to the backup data source. At the same time, notify the operation and maintenance personnel for handling; Example: Use Sentinel to implement the automatic circuit breaker mechanism. When the data source fails to be called continuously for 3 times, trigger the circuit breaker, and do not call this data source within the next 1 minute. Instead, call the backup data source and notify the operation and maintenance personnel by email or text message.
[0017] Example 2: Please refer to Figure 2 As shown, based on the method for dynamic switching of data source registration in Example 1, in this example, a system for dynamic switching of data source registration aims to achieve dynamic management and intelligent switching of data sources through modular design. The system is specifically as follows: 1. Data Source Registration Module: Responsible for the entry of data sources, metadata annotation, and dynamic weight initialization, forming independent small collection units; 1.1. Information Input Unit: Provide a visual interface or API interface to support the input of basic information (such as URL, username, permissions, etc.) from multiple types of data sources, including databases, APIs, file storage, etc., and generate a unique small collection ID (such as "db_001"); 1.2. Metadata Annotation Unit: Automatically or manually annotate the data source with business tags (such as "user management"), performance tags (such as "low latency"), and security tags (such as "data encryption") according to business rules; 1.3. Weight Calculation Unit: Generate an initial weight for each small collection based on historical usage frequency (formula: ) or custom rules; 2. Data Aggregation and Management Module: Implement the clustering of small collections into large collections, dynamically calculate the weights of large collections, and complete data storage and protection; 2.1. Clustering Engine Unit: Adopt the DBSCAN algorithm to cluster according to the data source usage patterns (such as call frequency, time distribution) within a time window, and generate large collections (such as "high-frequency usage large collection on weekdays"); 2.2. Weight Aggregation Unit: Calculate the weights of large collections through a weighted average algorithm (formula: ), and support dynamic adjustment of weight coefficients (such as giving priority to high-performance data sources); 2.3. Resource Limitation Unit: Lock the data source range of large collections according to the business domain (such as "financial transactions") and reject the addition of irrelevant data sources; 2.4. Metadata Storage: Use the combination of Elasticsearch + PostgreSQL to store the metadata of small collections and large collections, and support full-text search and structured queries (such as retrieving data sources by tags); 2.5. Data Protection Unit: Ensure data consistency through regular snapshot backups (such as at 2 am every day) and version control (recording the metadata modification history); 3. Intelligent Matching and Invocation Module: Parse business requests, dynamically match data source collections and execute invocations, and at the same time feedback monitoring data; 3.1. Request Parser Unit: Extract the characteristics of business requests (such as business type "order query", urgency "high"), and generate matching conditions (such as requiring the data source weight ≥ 0.7, tags containing "order processing"); 3.2. Rule Engine Unit: Implement rule matching based on Drools to filter out large collections that meet the conditions (such as "business tags contain'real-time transactions' and performance tags contain 'low latency'"); 3.3. Weight Sorting and Risk Assessment Unit: Use the quicksort algorithm to sort large collections / small collections in descending order of weight, and exclude abnormal data sources through threshold verification (such as triggering weight degradation when the call time exceeds 500ms); 3.4. Data source connector unit: According to the positioning result (such as the small collection "db_001"), perform data operations through a database connection pool or an HTTP client, supporting multiple types of interfaces such as MySQL and RESTful API; 3.5. Monitoring and feedback unit: Real-time collect metrics such as call duration and error rate (e.g., Prometheus samples every 10 seconds), dynamically adjust the data source weight through the exponentially weighted moving average method (formula: ), and trigger abnormal fusing (e.g., suspend the call if Sentinel fails continuously 3 times); 4. Visualization management interface module, providing a graphical interface for system operation and maintenance and configuration, reducing the operation threshold; 4.1. Data source management unit: View / edit the details of the small collection, and manually adjust the tags or weights; 4.2. Large collection clustering monitoring unit: Visually display the members of the large collection, the trend of weight changes, and the clustering rules; 4.3. Call log analysis unit: Query the data source matching path, duration statistics, and exception records of historical requests; 4.4. Rule configuration unit: Dynamically modify the clustering rules, weight calculation formula, and threshold parameters (e.g., adjust the "high latency" threshold to 300 ms); 5. System interaction process module 5.1. Registration stage: The user enters the data source information through the interface or API, and the system generates a small collection and annotates the metadata, completing the initial weight calculation and storage; 5.2. Clustering stage: The clustering engine regularly scans the data used by the small collection, generates a large collection through the DBSCAN algorithm, and the resource limiting unit filters out irrelevant data sources; 5.3. Call stage: The business request triggers the parser to generate matching conditions, the rule engine filters the large collection, locates the high-priority small collection after weight sorting, and the connector executes the call and feeds back the monitoring data; 5.4. Optimization stage: According to the real-time monitoring data, the system automatically adjusts the weights of the small collection / large collection, and the abnormal data source triggers fusing and notifies the operation and maintenance.
[0018] Finally, it should be noted that: Obviously, the above embodiments are only examples for clearly illustrating the present invention, rather than limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the implementation manners here. And the obvious changes or modifications derived therefrom are still within the protection scope of the present invention.
Claims
1. A method for dynamically switching data source registration, characterized in that The method includes the following steps: Step 1, data storage step; Step 2, data invocation step; Step 2.1, business request parsing and request feature extraction. Extract key information from the business request to provide a basis for subsequent data source matching; determine the weight range and performance index requirements of the required data source according to the urgency and performance requirements of the business request; Step 2.2, large collection matching. Use the rule engine to screen out the qualified large collections according to the request features; sort the screened large collections in descending order of weight, and preferentially select the large collections with high weight for subsequent small collection positioning; check the call time consumption and error rate indicators of the data sources in the large collection. If they exceed the set threshold, reduce or exclude the weight of the large collection; Step 2.3, small collection positioning. In the selected large collection, screen out the small collections that meet the request requirements according to the business label and performance label; sort the screened small collections in descending order of weight, and preferentially select the small collections with high weight; check the call time consumption and error rate indicators of the small collection again. If they exceed the set threshold, skip the small collection and select the next qualified small collection; Step 2.4, data source invocation and feedback. According to the located small collection, invoke the corresponding data source and execute data operations; monitor the call time consumption and error rate indicators of the data source in real time, and feedback the monitoring data to the intelligent analysis module for dynamically adjusting the data source weight; the intelligent analysis module uses the sliding window algorithm and the exponentially weighted moving average method according to the monitoring data to dynamically adjust the weights of the small collection and the large collection; if an exception occurs during the invocation, trigger the automatic fusing mechanism, suspend the invocation of the data source, switch to the standby data source, and notify the operation and maintenance personnel to handle it.
2. The method for dynamically switching data source registration according to claim 1, wherein The data storage step further includes the following steps: Step 1.1, small collection registration and metadata construction. Register each data source as an independent small collection, and assign a unique identifier to each data source small collection for accurate identification during subsequent data storage, invocation, and management processes to ensure that there is no confusion between different data sources; Step 1.2, large collection clustering and weight calculation. Cluster the small collections with similar usage characteristics into large collections, and determine the clustering relationship by analyzing the call frequency and call time distribution data of the data sources in each time window; adopt the DBSCAN density clustering algorithm to cluster with the usage frequency and time distribution of the data source in the time window as the feature vector; Step 1.3, data storage and protection. Select Elasticsearch + PostgreSQL combination as the storage medium to store metadata; ensure the data consistency and integrity of the collection through data backup and version control.
3. The method for dynamically switching data source registration according to claim 1, characterized in that The weight of the large collection is reduced or excluded based on the set threshold judgment rule. If the data source call time in the large collection exceeds a certain proportion (such as 30%) and exceeds the threshold or the error rate exceeds the threshold, the weight of the large collection is reduced; the weight reduction can be achieved by using the formula , that is, the original weight is multiplied by 0.
8.
4. A method for dynamic switching of data source registration as claimed in claim 1, characterized in that, The string matching algorithm is used to screen the small collections, and the business label and performance label of the request are matched one by one with the corresponding labels of the small collection.
5. The method for dynamically switching data source registration according to claim 1, characterized in that, The quicksort algorithm is used to sort the weights of the small collections.
6. The method for dynamically switching data source registration according to claim 1, characterized in that The weight dynamic adjustment uses the sliding window algorithm and the exponentially weighted moving average method to dynamically adjust the weights of the small collection and the large collection; set the time window size, calculate the average call time and error rate metrics of the data source within this window, and adjust the weights according to the metric changes; the formula for the exponentially weighted moving average method is , where is the current weight, is the smoothing coefficient, takes values from 0 to 1, is the metric score within the current time window, is the weight at the previous moment.
7. A data source registration dynamic switching system for executing a method for data source registration dynamic switching according to any one of claims 1-6, characterized in that, The system includes the following modules: The data source registration module is responsible for the entry of the data source, metadata annotation, and dynamic weight initialization, and forms an independent small collection unit; The data aggregation and management module realizes the clustering of small aggregations into large aggregations, dynamically calculates the weights of the large aggregations, and completes data storage and protection; The intelligent matching and invocation module parses business requests, dynamically matches data source aggregations and executes invocations, and at the same time feeds back monitoring data; The visualization management interface module provides a graphical interface for system operation and maintenance and configuration, reducing the operation threshold; The system interaction process module is responsible for coordinating the automated execution of the entire process of data source registration, clustering, matching, and invocation.
Citation Information
Patent Citations
Metadata management system and method oriented to heterogeneous data sources
CN118796903A
Data source switching method and device, medium and electronic equipment
CN119917592A
Cloud Computing-Based Adaptive Storage Layering System and Method
US20240028604A1