A comprehensive data management system for multi-source databases

CN122822248APending Publication Date: 2026-09-25HANGZHOU YIRUIXIN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610687703.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

这种多源异构的数据管理模式导致数据孤岛现象严重,数据之间的关联性弱,难以实现跨业务模块的综合查询与全局统计分析

Benefits of technology

1、通过设置数据接入与特征提取模块以及关联映射模块,能够自动从多个异构业务数据库中抽取并标准化关键特征字段,进而识别并绑定跨业务数据间的隐含关联关系,生成多维关联映射表。有效地打破了不同业务系统间的数据孤岛,实现了项目信息、财务信息、人员信息与物资信息的全链路动态整合,显著提升了跨部门、跨业务的数据协同效率与全局数据一致性,解决了现有技术中数据整合能力差、业务流程联动性不足的根本性技术问题;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122822248A_ABST
    Figure CN122822248A_ABST
Patent Text Reader

Abstract

The application discloses a kind of comprehensive data management systems of multi-source database, it is related to computer processing technical field, comprising: data feature extraction module, it connects at least two heterogeneous service databases, respectively from each service database extraction original data table, and according to the preset service characteristic rule, from the original data table extraction key feature field, form service characteristic vector set, the service characteristic vector set at least includes project characteristic vector, personnel characteristic vector, financial characteristic vector and material characteristic vector;Association link interactive module is used to obtain the service characteristic vector set.The application realizes the full-link dynamic integration of multi-source heterogeneous data by data access and association mapping, combines fine-grained permission filtering and field-level desensitization to strengthen security access control, and automatically triggers cross-database cascade update and consistency check with the aid of state tracking and distributed transaction coordination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer processing technology, and in particular to a comprehensive data management system for multi-source databases. Background Technology

[0002] As enterprises continue to deepen their IT infrastructure development, especially medical institutions and project management companies, they face the challenge of managing multiple types of data in their daily operations.

[0003] In existing technologies, separate database systems are typically used to manage different types of business data, such as project contract information, patient records, financial billing information, inventory information (e.g., invoices), and staff scheduling information, all stored in isolated application systems. This multi-source, heterogeneous data management model leads to severe data silos, weak correlation between data, and difficulty in achieving comprehensive queries and global statistical analysis across business modules. For example, in project management, it is impossible to link project contract payment progress with the allocation of implementation personnel and the payment status of outsourced equipment in real time; in medical institution management, patient registration, hospitalization, billing, and medical insurance settlement information are scattered, resulting in inefficient and error-prone daily financial settlement and report compilation.

[0004] Therefore, existing systems generally suffer from technical problems such as poor data integration capabilities, insufficient business process linkage, crude access control, and inability to provide a one-stop global monitoring view. Summary of the Invention

[0005] To address the aforementioned technical problems, the present invention provides a comprehensive data management system for multi-source databases, comprising: The data feature extraction module connects to at least two heterogeneous business databases, extracts raw data tables from each business database, and extracts key feature fields from the raw data tables according to preset business feature rules to form a business feature vector set. The business feature vector set includes at least project feature vectors, personnel feature vectors, financial feature vectors, and material feature vectors. The association link interaction module is used to obtain the business feature vector set and create link derivation between different business feature vectors according to predefined cross-business association rules to generate a multi-dimensional association mapping table. The multi-dimensional association mapping table records the dynamic link relationship between project information and payment plan, implementation personnel and external procurement contract, and patient file and expense bill. The user configuration module binds users to preset roles based on the RBAC model and assigns data access permissions and function operation permissions to each role based on different business feature vectors. The data access permissions include at least row-level data filtering conditions and column-level field visibility rules. The dynamic configuration module is configured to monitor status change events in each business database in real time, and perform cascading updates on changes in related data based on the multidimensional association mapping table. At the same time, the updated business feature vectors are displayed in the form of a graphical dashboard.

[0006] Preferably, the data feature extraction module further includes: establishing an independent incremental data capture pipeline for each heterogeneous business database, wherein the incremental data capture pipeline captures change data records that have occurred since the last capture time periodically or in real time based on timestamps or log sequence numbers; After cleaning, deduplicating, and formatting the captured change data records, the key feature fields are then extracted. The key feature fields are extracted using a multimodal feature fusion algorithm based on an attention mechanism. Differentiated weights are assigned to different types of business data sources, and a dynamic weight adjustment mechanism is used to balance the feature contribution of multiple data sources in order to optimize the quality of the feature vector set.

[0007] Preferably, generating the multidimensional association mapping table includes the following steps: S100. Traverse the contract number field in the project feature vector, retrieve the payment records with the same contract number in the financial feature vector, and establish the first-level association. S101. Using the external supplier field in the project feature vector as the key, retrieve the corresponding purchase contract and payment plan in the material feature vector to establish a second-level association; S102. Using the task assignment identifier in the personnel feature vector as the key, match it with the implementation stage field in the project feature vector to establish a third-level association; merge the first-level association, the second-level association, and the third-level association to form the multi-dimensional association mapping table. S103. Before merging the first-level association, the second-level association and the third-level association, conflict detection is performed on each association based on the consistent hashing algorithm. When an association conflict is detected, a manual intervention process is triggered to ensure the data consistency of the multi-dimensional association mapping table. S104. The generated multidimensional association mapping table is stored in a directed acyclic graph data structure. The nodes in the graph are the key fields in each business feature vector, and the edges are the association relationships between each feature vector. The directed acyclic graph also serves as the topological sorting basis when the dynamic configuration module performs state cascading updates, ensuring that the execution order of cascading updates follows the edge dependency relationships in the directed acyclic graph.

[0008] Preferably, the user configuration module includes: The conditional filtering unit is used to dynamically add a WHERE clause to the SQL statement layer when executing a data query based on the row-level data filtering conditions, so as to limit the returned data rows to only the department or project scope of the current user. The field filtering unit is used to perform masking, truncation or encryption processing on sensitive fields in the query results according to the column-level field visibility rules before returning them to the front end for display. The sensitive fields include the patient's name, ID number, and contract amount; The field filtering unit uses a dynamic desensitization strategy based on data lineage to track the transmission path of sensitive fields from the source database to the query result set in real time, and adaptively adjusts the desensitization intensity according to the changes in the sensitivity level of the fields in the transmission path.

[0009] Preferably, the dynamic desensitization strategy for data lineage is linked with the multidimensional association mapping table generated by the association link interaction module: When a sensitive field is associated with multiple business feature vectors in a multidimensional association mapping table, the field filtering unit calculates a comprehensive desensitization rule based on the sensitivity level of each association path and performs differentiated desensitization levels for each association path. When a sensitive field is associated with only a single business feature vector, the single desensitization rule preset by that business feature vector shall be followed. Before performing dynamic desensitization, the field filtering unit first verifies whether the current user's role has access permissions for the corresponding associated path in the multidimensional association mapping table, and only desensitizes the authorized path before outputting it.

[0010] Preferably, the dynamic configuration module further includes automatically retrieving the associated payment plan from the financial feature vector of the second business database when the project status in the first business database changes from "under implementation" to "accepted", and updating the corresponding stage payment status in the payment plan to "pending payment". Simultaneously, the system retrieves relevant external procurement contracts from the material feature vectors in the third business database and generates payment reminder notifications. A hybrid consistency guarantee mechanism based on a two-phase commit protocol and TCC compensation transactions is used: two-phase commit is used to ensure strong consistency for mandatory synchronous updates across databases, and TCC compensation transactions are used to ensure eventual consistency for notification operations that are not subject to mandatory synchronization. The mechanism automatically switches between the two mechanisms dynamically through a preset fault threshold.

[0011] Preferably, an invoice management module is also included, which is used to receive invoice entry requests from the financial feature vector and store the invoice code and invoice number as unique keys in the material feature vector. When an invoice is issued or a registration fee is charged, the system automatically retrieves the available invoice number from the material feature vector and establishes a binding relationship between the invoice and the expense bill. When an invoice is voided or recycled, the binding relationship is unbound and the invoice status is reset to unused; The invoice management submodule is further configured with a blockchain-based evidence storage unit, which is used to write the entire lifecycle operation records of invoices, including warehousing, issuance, binding, and cancellation, into the blockchain to form an immutable audit log.

[0012] Preferably, the heterogeneous business database includes at least: A project management database used to store project contracts and implementation schedules; The HIS database is used to store patient records, outpatient registration, and inpatient registration. And a financial management database for storing expense items, billing records, and daily summaries; The system also includes a database connection pool manager, which maintains independent connection pools for different types of business databases and dynamically adjusts the size of the connection pools based on the historical query load data of each business database to avoid resource contention when accessing cross databases.

[0013] The present invention has at least the following beneficial effects: 1. By setting up data access and feature extraction modules as well as association mapping modules, key feature fields can be automatically extracted and standardized from multiple heterogeneous business databases. This allows for the identification and binding of implicit relationships between cross-business data, generating a multi-dimensional association mapping table. This effectively breaks down data silos between different business systems, achieving dynamic integration of project information, financial information, personnel information, and material information across the entire chain. It significantly improves cross-departmental and cross-business data collaboration efficiency and global data consistency, solving the fundamental technical problems of poor data integration capabilities and insufficient business process linkage in existing technologies. 2. By setting up a permission and role management module, combined with fine-grained permission filters and field-level data maskers, fine-grained access control for multi-source data fusion is achieved. It not only supports role-based row-level data filtering but also supports dynamic data masking for sensitive fields. This multi-layered, fine-grained permission management mechanism, while ensuring full data sharing and flow, greatly enhances the system's data security, meeting the high compliance requirements of medical institutions and enterprise management regarding financial data and patient privacy protection. 3. By setting up a status tracking and visualization module and introducing a distributed transaction coordinator, it is possible to monitor data change events in any business database in real time and automatically trigger cross-database cascading updates and consistency checks based on the multidimensional association mapping table. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a flowchart of a comprehensive data management system for multi-source databases provided in Embodiment 1 of the present invention; Figure 2 This is a flowchart of a multidimensional association mapping table provided in Embodiment 1 of the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0018] Example 1

[0019] This embodiment provides a comprehensive data management system for multi-source databases. The method includes the following steps: Figure 1 As shown: The data feature extraction module connects to at least two heterogeneous business databases, extracts raw data tables from each business database, and extracts key feature fields from the raw data tables according to preset business feature rules to form a business feature vector set. The business feature vector set includes at least project feature vectors, personnel feature vectors, financial feature vectors, and material feature vectors. The above embodiments also include: establishing an independent incremental data capture pipeline for each heterogeneous business database. The incremental data capture pipeline captures change data records that have occurred since the last capture time periodically or in real time based on timestamps or log sequence numbers. After cleaning, deduplicating, and formatting the captured change data records, key feature fields are extracted. Among them, the extraction of key feature fields is performed by a multimodal feature fusion algorithm based on the attention mechanism. Differentiated weights are assigned to different types of business data sources, and the feature contribution between multiple data sources is balanced through a dynamic weight adjustment mechanism to optimize the quality of the feature vector set.

[0020] Specifically, the system deployment process requires connecting to at least two heterogeneous business databases—for example, a project management database (Oracle) storing project contracts and implementation progress, a HIS database (SQL Server) storing patient records and registration information, and a financial management database (MySQL) storing expense items and billing records. After the connection is established, the module independently establishes an incremental data capture pipeline for each heterogeneous business database. This pipeline captures data records that have changed since the last capture, based on the database's timestamp field (e.g., last_update_time) or transaction log sequence number (e.g., binlog position in MySQL, SCN in Oracle), periodically (e.g., every 5 minutes) or in real-time (by listening to the log stream). For example, in the project management database, when the implementation status of a project changes from "under implementation" to "accepted," the update_time field of that record is updated to the current timestamp, and the incremental pipeline can capture this change record in the next capture cycle.

[0021] The captured raw change records typically contain redundant, incomplete, or inconsistently formatted data. Therefore, the module further cleans, deduplicates, and converts the format of the captured data. For example, patient names captured from the HIS database may contain leading or trailing spaces or special characters, which need to be removed uniformly; synchronization delays between multiple databases may cause the same contract record to be captured repeatedly, requiring deduplication based on the contract number; different databases store date formats differently (e.g., "2025-03-15" vs. "15-Mar-2025"), requiring conversion to a unified internal timestamp format. After preprocessing, the module enters the crucial feature extraction stage.

[0022] Feature extraction employs a multimodal feature fusion algorithm based on an attention mechanism. The algorithm operates as follows: First, preprocessed business data is treated as different "modalities," each corresponding to a data type. Examples include structured tabular data in a project management database, text-based diagnostic descriptions in an HIS database, and numerical expense details in a financial management database. For each modality, the module uses a corresponding encoder (e.g., a fully connected network for numerical features, BERT for text descriptions) to extract initial feature representations. Then, an attention mechanism is introduced to assign differentiated weights to features from different modalities. The core idea of ​​the attention mechanism is that not all business data is equally important to the current management objective (e.g., payment collection prediction, resource scheduling); the system should learn to "focus" on more relevant data sources. In practice, the module calculates the similarity score between each modal feature and the current task context (e.g., the project risk level to be predicted), and normalizes the score into weights using a Softmax function, giving higher attention to key modalities (e.g., in the project acceptance phase, the weight of financial feature vectors should be higher than that of material feature vectors).

[0023] To further optimize the balance of feature contributions among multi-source data, the module also incorporates a dynamic weight adjustment mechanism. This mechanism operates based on a closed-loop feedback principle: after each feature extraction, the system evaluates the performance of downstream tasks (such as the matching accuracy of the association mapping module or the success rate of state updates in the dynamic configuration module), generating a quality feedback score. This score is backpropagated to the attention mechanism layer to update the confidence coefficients of each data source. For example, if frequent missing or erroneous payment data is found in the financial feature vector (low data quality score), the system automatically reduces the weight of the financial data source while relatively increasing the weight of the project feature vector, thus achieving adaptive optimization in the next iteration. This closed-loop feedback feature extraction cycle allows the system to continuously adapt to fluctuations in data source quality and changes in business focus.

[0024] The association link interaction module is used to obtain the business feature vector set and create link derivation between different business feature vectors according to predefined cross-business association rules, generating a multi-dimensional association mapping table. The multi-dimensional association mapping table records the dynamic link relationship between project information and payment plan, implementation personnel and external procurement contract, and patient records and expense bills. The above embodiments, generating a multidimensional association mapping table includes the following steps (such as...). Figure 2 (as shown) S100. Traverse the contract number field in the project feature vector, retrieve the payment records with the same contract number in the financial feature vector, and establish the first-level association. S101. Using the external supplier field in the project feature vector as the key, retrieve the corresponding purchase contract and payment plan in the material feature vector to establish a second-level association; S102. Using the task assignment identifier in the personnel feature vector as the key, match it with the implementation stage field in the project feature vector to establish a third-level association; merge the first-level association, second-level association and third-level association to form a multi-dimensional association mapping table; S103. Before merging the first-level association, the second-level association and the third-level association, perform conflict detection on each association based on the consistent hashing algorithm. When an association conflict is detected, trigger the manual intervention process to ensure the data consistency of the multi-dimensional association mapping table. S104. Store the generated multidimensional association mapping table in a directed acyclic graph data structure. The nodes in the graph are the key fields in each business feature vector, and the edges are the association relationships between each feature vector. The directed acyclic graph also serves as the topological sorting basis when the dynamic configuration module performs state cascading updates, ensuring that the execution order of cascading updates follows the edge dependency relationships in the directed acyclic graph.

[0025] Specifically, while maintaining the physical independence of each business database, a multi-dimensional association mapping table is constructed by logically linking feature vectors from different data sources through predefined cross-business association rules. This dynamically reflects the implicit dependencies between business entities. The module's input receives a standardized set of business feature vectors (including project, personnel, financial, and material feature vectors) output by the data feature extraction module, while the output generates a multi-dimensional association mapping table with a directed acyclic graph (DAG) as its data structure. The following section provides a detailed explanation of the implementation, operating principles, and examples.

[0026] The generation of the multidimensional association mapping table, as described above, follows the logic of three-level association and hierarchical derivation. Each level of association is based on specific key fields for cross-vector retrieval and matching.

[0027] First-level association: The module first iterates through the "Contract Number" field in the project feature vector, using this number as the search key to find payment records with the same contract number in the financial feature vector. Once a match is found, a first-level dynamic link is established between the project information and its corresponding payment plan.

[0028] The second level of association: The module further uses the "external supplier field" in the project feature vector as the key to retrieve the corresponding purchase contracts and payment plans from the material feature vector. The core value of this level of association lies in connecting the external procurement chain during project execution.

[0029] The third level of association: The module uses the "assigned task identifier" in the personnel feature vector as the key and matches it with the "implementation phase field" in the project feature vector. Here, the "assigned task identifier" can be an employee ID, job code, or task batch number, while the "implementation phase field" identifies the current phase of the project (such as requirements analysis, equipment installation, system debugging, acceptance and delivery, etc.). Through this level of association, the module dynamically binds specific implementers to the project phases they are responsible for.

[0030] Before merging the three levels of associations mentioned above, the module introduces a conflict detection mechanism based on the consistent hashing algorithm. The consistent hashing ring maps each association edge (such as the "project-payment" edge in the first-level association and the "project-purchase contract" edge in the second-level association) to a specific position on the hash ring. When two different association edges collide on the hash ring (i.e., they have the same hash value or point to the same node), the system determines that there is an association conflict.

[0031] After completing the three-level association merging and resolving all conflicts, the module persistently stores the multidimensional association mapping table as a directed acyclic graph (DAG) data structure. In this DAG, each node represents a key field in a business feature vector (such as "contract number", "payment amount", "outsourced supplier", "installation personnel", etc.), and each directed edge represents the association relationship from one field to another, with the direction of the edge determined by the dependency direction of the association.

[0032] The DAG structure provides a natural topological ordering basis for the cascading state updates of the dynamic configuration module. Specifically, when the state in a business database changes (e.g., a project status changes from "under implementation" to "accepted"), the dynamic configuration module needs to determine the execution order of the cascading updates: should the payment status in the financial feature vector be updated first, or should the material payment reminder be generated first? The DAG's topological ordering algorithm can automatically calculate the linear execution sequence of all related edges, ensuring that all dependencies are satisfied—that is, the update of the dependent node must occur before the dependent node.

[0033] The user configuration module binds users to preset roles based on the RBAC model and assigns data access permissions and function operation permissions to each role based on different business feature vectors. The data access permissions include at least row-level data filtering conditions and column-level field visibility rules. Furthermore, the aforementioned user configuration module includes: The conditional filtering unit is used to dynamically add a WHERE clause to the SQL statement layer when executing a data query based on row-level data filtering conditions, so as to limit the returned data rows to only the department or project scope of the current user. The field filtering unit is used to perform masking, truncation, or encryption on sensitive fields in the query results based on column-level field visibility rules before returning them to the front end for display. Sensitive fields include patient name, ID number, and contract amount; The field filtering unit uses a dynamic desensitization strategy based on data lineage to track the transmission path of sensitive fields from the source database to the query result set in real time, and adaptively adjusts the desensitization intensity according to the changes in the sensitivity level of the fields in the transmission path.

[0034] Furthermore, the dynamic desensitization strategy for data lineage is linked with the multi-dimensional association mapping table generated by the association link interaction module: When a sensitive field is associated with multiple business feature vectors in a multidimensional association mapping table, the field filtering unit calculates a comprehensive desensitization rule based on the sensitivity level of each association path and performs differentiated desensitization levels for each association path. When a sensitive field is associated with only a single business feature vector, the single desensitization rule preset by that business feature vector shall be followed. Before performing dynamic desensitization, the field filtering unit first verifies whether the current user's role has the access permissions for the corresponding associated path in the multidimensional association mapping table, and only desensitizes the authorized path before outputting it.

[0035] Specifically, the first step is to establish a mapping relationship between users and roles, with each role corresponding to a set of data access permissions and functional operation permissions. Unlike traditional solutions, this module's permission allocation is based on "business feature vectors," which are project feature vectors, personnel feature vectors, financial feature vectors, and material feature vectors generated by the data feature extraction module. Each role can be granted read, modify, and export permissions for one or more feature vectors. Furthermore, row-level data filtering conditions and column-level field visibility rules can be set individually for each feature vector.

[0036] Then, the conditional filtering unit is responsible for converting the row-level data filtering conditions in the RBAC configuration into the SQL WHERE clause for the actual database query. Specifically, this unit intercepts all query requests at the system data access layer (DAO layer), parses the currently logged-in user's role and its bound row-level filtering conditions, dynamically generates filtering predicates, and appends them to the original SQL statement. This process is completely transparent to the upper-layer business logic and requires no modification to any business code.

[0037] To achieve compatibility across heterogeneous databases, the conditional filtering unit incorporates a built-in SQL dialect adapter, which can automatically generate WHERE clauses with appropriate syntax for project management databases (such as MySQL), HIS databases (such as Oracle), and financial management databases (such as SQL Server). Furthermore, the filtering conditions support multiple condition combinations (nested AND / OR) and dynamic parameter binding (such as current user ID, current department ID, current time range, etc.). The field filtering unit implements two layers of functionality: the first layer is static filtering based on column-level field visibility rules, which directly hides fields that the role does not have permission to view; the second layer is a dynamic desensitization strategy based on data lineage, which covers, truncates, or encrypts fields that the user has permission to view but contain sensitive information.

[0038] Static filtering is relatively simple: the system maintains a whitelist of visible fields for each role. Before the query result set is returned to the front end, the field filtering unit traverses the column names of the result set and removes columns that are not in the whitelist.

[0039] During the data access phase (i.e., when the data feature extraction module is running), the system records the complete transmission path of each sensitive field from the original table and column in the source database to the field in the business feature vector, and finally to the final query result set. This path is stored in the system's metadata database in the form of a directed graph, with each edge marking the sensitivity level of that step (e.g., the sensitivity level of "patient ID number" in the source database is 10, but after encryption and mapping to "patient identifier" in the business feature vector, the sensitivity level is reduced to 5). When the field filtering unit processes a result set containing a sensitive field, it queries the lineage path of that field in real time to obtain the sensitivity level of the current step, and then selects the corresponding desensitization strength according to the preset desensitization strategy table (e.g., sensitivity level 10 performs "the first 6 digits and the last 4 digits are kept with asterisks" masking, sensitivity level 5 performs "only the last 4 digits are displayed" masking, sensitivity level 2 only performs format validation without desensitization, etc.).

[0040] After receiving the query result set, the field filtering unit, for each sensitive field, not only tracks its individual lineage path but also queries the multidimensional association mapping table to find all related edges originating from or passing through that field, forming a set of association paths. Each association path is associated with one or more business feature vectors, and each business feature vector has a preset business sensitivity level under that path. The field filtering unit calculates the weighted average level according to the formula: Comprehensive Desensitization Strength = Σ(Path Weight × Path Sensitivity Level) / ΣPath Weight, where the path weight is determined by the business importance coefficient of the association path (e.g., the association path weight for contract amount is higher than the association path weight for patient name). Then, the corresponding desensitization rule is selected based on the comprehensive desensitization strength. Simultaneously, before performing desensitization, the system verifies whether the current user's role has permission to traverse these association paths. If the user's role is not authorized to access the business feature vector corresponding to a certain association path, the sensitivity level of that path will not participate in the weighted calculation, and the related fields in the desensitized data will be additionally masked.

[0041] The dynamic configuration module is configured to monitor state change events in each business database in real time, and perform cascading updates on changes in related data based on the multi-dimensional association mapping table. At the same time, the updated business feature vectors are displayed in the form of a graphical dashboard.

[0042] Furthermore, the dynamic configuration module in the above embodiment also includes automatically retrieving the associated repayment plan from the financial feature vector of the second business database when the project status in the first business database changes from "under implementation" to "accepted", and updating the corresponding stage repayment status in the repayment plan to "pending repayment". Simultaneously, the system retrieves relevant external procurement contracts from the material feature vectors in the third business database and generates payment reminder notifications. A hybrid consistency guarantee mechanism based on a two-phase commit protocol and TCC compensation transactions is used: two-phase commit is used to ensure strong consistency for mandatory synchronous updates across databases, and TCC compensation transactions are used to ensure eventual consistency for notification operations that are not subject to mandatory synchronization. The mechanism automatically switches between the two mechanisms dynamically through a preset fault threshold.

[0043] Specifically, by deploying a lightweight listening agent or a change data capture (CDC) component based on database log parsing on each heterogeneous business database (such as project management database, HIS database, financial management database), the system can capture state change events (such as INSERT, UPDATE, DELETE operations) of key tables in each database in real time, and push the captured events to a message queue (such as Kafka or RocketMQ) in a unified message format. When the status field of a project in the first business database (project management database) changes from "Under Implementation" to "Accepted," the dynamic configuration module consumes the event from the message queue and immediately performs path retrieval based on the multi-dimensional association mapping table generated by the association link interaction module (this mapping table stores the dependencies between the project and entities such as payment plans and procurement contracts in a directed acyclic graph (DAG) structure): First, it locates the payment plan record matching the project contract number in the financial feature vector of the second business database (finance database) and atomically updates the corresponding stage payment status to "Pending Payment." Simultaneously, it retrieves the associated procurement contract in the material feature vector of the third business database (materials database), calls the notification service to generate a payment reminder, and pushes it to the procurement manager. To ensure the reliability of cross-database update operations, the dynamic configuration module incorporates a hybrid consistency guarantee mechanism. Specifically, it adopts a strategy combining two-phase commit (2PC) and TCC (Try-Confirm-Cancel) compensation transactions. For mandatory synchronization update operations (such as successfully modifying the payment plan status after a status change), the module initiates the 2PC protocol: first, a Prepare request is sent to all involved database nodes, and after all nodes respond Ready, a Commit command is sent, thereby achieving strong consistency. For non-mandatory synchronization notification operations (such as generating payment reminders), TCC compensation transactions are used, that is, a Try phase is executed first (reserving resources for sending reminders), and Confirm is executed after confirming the success of the main update (actually sending the notification). If the main update fails, Cancel is executed (canceling the reservation), thereby ensuring eventual consistency. More importantly, the module also has dynamic fault threshold switching logic: the system counts the cumulative number of abnormal forced synchronization updates per unit time in real time (such as the number of timeouts in the Prepare phase or the number of node crashes). When the cumulative number of abnormalities is lower than the preset threshold (e.g., less than 3 times per minute), the 2PC protocol is used first to ensure zero error in critical data. Once the cumulative number of abnormalities reaches or exceeds the threshold, the module automatically switches the forced synchronization update operation in the current and subsequent window to TCC compensation transaction mode for execution. That is, it first attempts to update, and if the update fails, it records the compensation log and retryes asynchronously, while generating a fault site report for operation and maintenance intervention.For example, in the Yiruixin project management system of a top-tier hospital, after the "Smart Ward Construction Project" passes the implementation acceptance test, the dynamic configuration module detects the change and automatically updates the second-phase payment status of the project from "not received" to "pending payment," while simultaneously pushing a "Payment Reminder for ECG Monitor Outsourcing Contract" to the equipment department. If the financial database experiences three consecutive timeouts in the 2PC Prepare phase due to network jitter (reaching the threshold), the module automatically switches to TCC mode, first completing the final consistent update of the payment status, and then retrying multiple times through compensation tasks until the reminder is successfully sent. This ensures business integrity while preventing system collapse due to frequent retries. Finally, all updated business feature vectors (including changed project status, payment status, and reminder records) are pushed to the front-end visualization engine in real time, dynamically refreshing in the form of graphical dashboards (such as bar charts comparing planned and actual payments, and project progress heatmaps), providing managers with a real-time and consistent decision-making view.

[0044] Example 2

[0045] Based on the above embodiment one, this embodiment also includes an invoice management module, which is used to receive invoice entry requests from the financial feature vector and store the invoice code and invoice number as unique keys in the material feature vector; When an invoice is issued or a registration fee is charged, the system automatically retrieves the available invoice number from the material feature vector and establishes a binding relationship between the invoice and the expense bill. When an invoice is voided or recycled, the binding relationship is unbound and the invoice status is reset to unused; The invoice management submodule is further configured with a blockchain-based evidence storage unit, which is used to write the entire lifecycle operation records of invoices, including warehousing, issuance, binding, and cancellation, into the blockchain to form an immutable audit log.

[0046] Specifically, the invoice management module receives invoice entry requests from the financial feature vector. The financial feature vector is typically generated from invoice purchase records in the financial management database, data imported from the tax system, or manually entered information. It contains key fields such as invoice code, invoice number, invoice type (VAT general invoice, special invoice, electronic invoice, etc.), invoice date, and amount. The module uses "invoice code + invoice number" as a unique key and stores this information in the material feature vector. The material feature vector is a dynamic data structure specifically for managing physical or virtual materials (including invoices, consumables, equipment, etc.), supporting fast retrieval by unique key, status updates, and lifecycle tracking.

[0047] When an invoice is issued (e.g., by finance personnel using blank invoices) or a registration and payment operation occurs (e.g., when an outpatient needs an invoice after registration), the module automatically retrieves an available invoice number with a "not used" status from the material feature vector. This retrieval process follows a first-in, first-out (FIFO) or batch-first scheduling algorithm to ensure the effective utilization of invoice number ranges. After retrieving the invoice number, the module establishes a binding relationship between the invoice and the expense bill. This binding relationship is stored in a multidimensional association mapping table, forming a link derivation of "invoice-expense bill-patient / item".

[0048] When an invoice is issued (e.g., by finance personnel using blank invoices) or a registration and payment operation occurs (e.g., when an outpatient needs an invoice after registration), the module automatically retrieves an available invoice number with a "not used" status from the material feature vector. This retrieval process follows a first-in, first-out (FIFO) or batch-first scheduling algorithm to ensure the effective utilization of invoice number ranges. After retrieving the invoice number, the module establishes a binding relationship between the invoice and the expense bill. This binding relationship is stored in a multidimensional association mapping table, forming a link derivation of "invoice-expense bill-patient / item".

[0049] The permissioned blockchain (Hyperledger Fabric or enterprise-grade Ethereum private blockchain) is jointly maintained by multiple trusted nodes within the system (such as financial servers, audit servers, and third-party regulatory nodes). The evidence storage unit generates an immutable evidence record for every status change operation of invoices, including their entry, exit, binding, cancellation, and recycling.

[0050] Example 3

[0051] Based on the above embodiment one, the heterogeneous business database of this embodiment includes at least: A project management database used to store project contracts and implementation schedules; The HIS database is used to store patient records, outpatient registration, and inpatient registration. And a financial management database for storing expense items, billing records, and daily summaries; It also includes a database connection pool manager, which maintains independent connection pools for different types of business databases and dynamically adjusts the size of the connection pools based on the historical query load data of each business database to avoid resource contention when accessing cross databases.

[0052] Specifically, in the multi-source database integrated data management system described in this embodiment, the heterogeneous business database includes at least three types: project management database, HIS database (Hospital Information System database), and financial management database. The project management database is mainly used to store core information of project contracts and progress data for each stage of project implementation, such as project number, contract amount, contract signing date, planned payment time, actual implementation progress percentage, project leader and implementation team allocation records, etc. The HIS database focuses on the medical business field, storing patient files (including name, ID number, medical card number, medical history summary), outpatient registration records (registration department, registration doctor, consultation time, registration fee), and inpatient registration information (admission date, bed number, attending physician, ward number), etc. The financial management database is responsible for storing expense categories (such as consultation fees, drug fees, examination fees, project payments, external procurement payments, etc.), detailed records of each charge (charge time, charge item, amount, payment method, invoice number), and daily and monthly summary data (total revenue, revenue percentage of each department, prepayment balance, etc.). These three types of databases are physically independent of each other, may be deployed on different servers, and may even use different database management systems (for example, project management databases use MySQL, HIS databases use Oracle, and financial management databases use SQL Server). Their data formats, access interfaces, and concurrency control strategies are also different.

[0053] Instead of mixing all database connections into a single public pool, a separate connection pool is maintained for each type of business database (project management, HIS, finance). Each connection pool manages a set of pre-created database connection objects, which can be reused by various modules of the system, avoiding the overhead of frequently creating and destroying connections. More importantly, the connection pool manager continuously collects historical query load data for each business database.

[0054] Example 4

[0055] This invention provides a non-transitory computer-readable storage medium storing at least one instruction or at least one program segment, which is loaded and executed by a processor to implement the following steps: Connect at least two heterogeneous business databases, extract raw data tables from each business database, and extract key feature fields from the raw data tables according to preset business feature rules to form a business feature vector set; the business feature vector set includes at least project feature vectors, personnel feature vectors, financial feature vectors, and material feature vectors; Obtain a set of business feature vectors and, based on predefined cross-business association rules, create links between different business feature vectors to generate a multidimensional association mapping table. The multidimensional association mapping table records the dynamic link relationships between project information and payment plans, implementation personnel and external procurement contracts, and patient records and expense bills. Users are bound to preset roles based on the RBAC model, and each role is assigned data access permissions and function operation permissions based on different business feature vectors; data access permissions include at least row-level data filtering conditions and column-level field visibility rules; It monitors status change events in each business database in real time, and performs cascading updates on changes to related data based on the multidimensional association mapping table. At the same time, it displays the updated business feature vectors in the form of a graphical dashboard.

[0056] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0057] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0058] Example 5

[0059] This invention provides an electronic device, including a processor and a memory, wherein the memory stores at least one instruction or at least one program segment, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement the following steps: Connect at least two heterogeneous business databases, extract raw data tables from each business database, and extract key feature fields from the raw data tables according to preset business feature rules to form a business feature vector set; the business feature vector set includes at least project feature vectors, personnel feature vectors, financial feature vectors, and material feature vectors; Obtain a set of business feature vectors and, based on predefined cross-business association rules, create links between different business feature vectors to generate a multidimensional association mapping table. The multidimensional association mapping table records the dynamic link relationships between project information and payment plans, implementation personnel and external procurement contracts, and patient records and expense bills. Users are bound to preset roles based on the RBAC model, and each role is assigned data access permissions and function operation permissions based on different business feature vectors; data access permissions include at least row-level data filtering conditions and column-level field visibility rules; It monitors status change events in each business database in real time, and performs cascading updates on changes to related data based on the multidimensional association mapping table. At the same time, it displays the updated business feature vectors in the form of a graphical dashboard.

[0060] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A comprehensive data management system for multi-source databases, characterized in that, include: The data feature extraction module connects to at least two heterogeneous business databases, extracts raw data tables from each business database, and extracts key feature fields from the raw data tables according to preset business feature rules to form a business feature vector set. The business feature vector set includes at least project feature vectors, personnel feature vectors, financial feature vectors, and material feature vectors. The association link interaction module is used to obtain the business feature vector set and create link derivation between different business feature vectors according to predefined cross-business association rules to generate a multi-dimensional association mapping table. The multi-dimensional association mapping table records the dynamic link relationship between project information and payment plan, implementation personnel and external procurement contract, and patient file and expense bill. The user configuration module binds users to preset roles based on the RBAC model and assigns data access permissions and function operation permissions to each role based on different business feature vectors. The data access permissions include at least row-level data filtering conditions and column-level field visibility rules. The dynamic configuration module is configured to monitor status change events in each business database in real time, and perform cascading updates on changes in related data based on the multidimensional association mapping table. At the same time, the updated business feature vectors are displayed in the form of a graphical dashboard.

2. The integrated data management system for multi-source databases according to claim 1, characterized in that, The data feature extraction module also includes: establishing an independent incremental data capture pipeline for each heterogeneous business database, wherein the incremental data capture pipeline captures change data records that have occurred since the last capture time periodically or in real time based on timestamps or log sequence numbers; After cleaning, deduplicating, and formatting the captured change data records, the key feature fields are then extracted. The key feature fields are extracted using a multimodal feature fusion algorithm based on an attention mechanism. Differentiated weights are assigned to different types of business data sources, and a dynamic weight adjustment mechanism is used to balance the feature contribution of multiple data sources in order to optimize the quality of the feature vector set.

3. The integrated data management system for multi-source databases according to claim 1, characterized in that, Generating the multidimensional association mapping table includes the following steps: S100. Traverse the contract number field in the project feature vector, retrieve the payment records with the same contract number in the financial feature vector, and establish the first-level association. S101. Using the external supplier field in the project feature vector as the key, retrieve the corresponding purchase contract and payment plan in the material feature vector to establish a second-level association; S102. Using the task assignment identifier in the personnel feature vector as the key, match it with the implementation stage field in the project feature vector to establish a third-level association; merge the first-level association, the second-level association, and the third-level association to form the multi-dimensional association mapping table. S103. Before merging the first-level association, the second-level association and the third-level association, conflict detection is performed on each association based on the consistent hashing algorithm. When an association conflict is detected, a manual intervention process is triggered to ensure the data consistency of the multi-dimensional association mapping table. S104. The generated multidimensional association mapping table is stored in a directed acyclic graph data structure. The nodes in the graph are the key fields in each business feature vector, and the edges are the association relationships between each feature vector. The directed acyclic graph also serves as the topological sorting basis when the dynamic configuration module performs state cascading updates, ensuring that the execution order of cascading updates follows the edge dependency relationships in the directed acyclic graph.

4. The integrated data management system for multi-source databases according to claim 1, characterized in that, The user configuration module includes: The conditional filtering unit is used to dynamically add a WHERE clause to the SQL statement layer when executing a data query based on the row-level data filtering conditions, so as to limit the returned data rows to only the department or project scope of the current user. The field filtering unit is used to perform masking, truncation or encryption processing on sensitive fields in the query results according to the column-level field visibility rules before returning them to the front end for display. The sensitive fields include the patient's name, ID number, and contract amount; The field filtering unit uses a dynamic desensitization strategy based on data lineage to track the transmission path of sensitive fields from the source database to the query result set in real time, and adaptively adjusts the desensitization intensity according to the changes in the sensitivity level of the fields in the transmission path.

5. A comprehensive data management system for multi-source databases according to claim 4, characterized in that, The dynamic desensitization strategy for data lineage is linked with the multidimensional association mapping table generated by the association link interaction module: When a sensitive field is associated with multiple business feature vectors in a multidimensional association mapping table, the field filtering unit calculates a comprehensive desensitization rule based on the sensitivity level of each association path and performs differentiated desensitization levels for each association path. When a sensitive field is associated with only a single business feature vector, the single desensitization rule preset by that business feature vector shall be followed. Before performing dynamic desensitization, the field filtering unit first verifies whether the current user's role has access permissions for the corresponding associated path in the multidimensional association mapping table, and only desensitizes the authorized path before outputting it.

6. A comprehensive data management system for multi-source databases according to claim 1, characterized in that, The dynamic configuration module also includes automatically retrieving the associated payment plan from the financial feature vector of the second business database when the project status in the first business database changes from "under implementation" to "accepted", and updating the corresponding stage payment status in the payment plan to "pending payment". Simultaneously, the system retrieves relevant external procurement contracts from the material feature vectors in the third business database and generates payment reminder notifications. A hybrid consistency guarantee mechanism based on a two-phase commit protocol and TCC compensation transactions is used: two-phase commit is used to ensure strong consistency for mandatory synchronous updates across databases, and TCC compensation transactions are used to ensure eventual consistency for notification operations that are not subject to mandatory synchronization. The mechanism automatically switches between the two mechanisms dynamically through a preset fault threshold.

7. A comprehensive data management system for multi-source databases according to claim 1, characterized in that, It also includes an invoice management module, which is used to receive invoice entry requests from the financial feature vector and store the invoice code and invoice number as unique keys in the material feature vector; When an invoice is issued or a registration fee is charged, the system automatically retrieves the available invoice number from the material feature vector and establishes a binding relationship between the invoice and the expense bill. When an invoice is voided or recycled, the binding relationship is unbound and the invoice status is reset to unused; The invoice management submodule is further configured with a blockchain-based evidence storage unit, which is used to write the entire lifecycle operation records of invoices, including warehousing, issuance, binding, and cancellation, into the blockchain to form an immutable audit log.

8. A comprehensive data management system for multi-source databases according to claim 1, characterized in that, The heterogeneous business database includes at least: A project management database used to store project contracts and implementation schedules; The HIS database is used to store patient records, outpatient registration, and inpatient registration. And a financial management database for storing expense items, billing records, and daily summaries; The system also includes a database connection pool manager, which maintains independent connection pools for different types of business databases and dynamically adjusts the size of the connection pools based on the historical query load data of each business database to avoid resource contention when accessing cross databases.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the integrated data management system for the multi-source database as described in any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the integrated data management system for the multi-source database as described in any one of claims 1 to 8.