Metadata intelligent driving-based data integration system
By using a data integration system driven by metadata intelligence, the parameters of multi-dimensional meta-features and deep learning models are dynamically adjusted, solving the problem of insufficient data integration stability and achieving more efficient data integration and consistent collection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING JUHUI RONGSHENG INTERNET TECH CO LTD
- Filing Date
- 2025-12-15
- Publication Date
- 2026-04-28
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The lack of a dynamic adaptation mechanism in existing technologies leads to insufficient data integration stability, task timeouts, and incomplete data collection when dynamic issues such as link congestion during peak hours and fluctuations in edge device load occur.
A data integration system based on metadata intelligence is adopted. Through data processing module, integration module, confidence threshold adjustment module, step size adjustment module, and trigger threshold adjustment module, the system dynamically adjusts the confidence threshold for missing features of multi-dimensional meta-features, the iteration step size of the deep learning model, and the trigger threshold for redundant features of multi-dimensional meta-features, thereby optimizing the data integration process.
It improves the stability of financial metadata integration, reduces invalid matching problems caused by feature completion errors, alleviates data accumulation and loss, and enhances the timeliness and consistency of data integration.
Smart Images

Figure CN121935481A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a data integration method and system based on metadata-driven intelligence. Background Technology
[0002] In the current era of explosive growth in multi-source heterogeneous data, data integration has become a core element in unlocking data value. However, existing technologies suffer from several drawbacks. Firstly, metadata application is superficial, often serving only as a format specification guideline without exploring its business relevance and link characteristics. This leads to a significant decrease in integration value when core metadata is missing. Secondly, dynamic adaptability is weak; in scenarios such as peak link congestion and terminal hardware fluctuations, integration strategies relying on fixed rules are prone to timing failures and feature deviations. Thirdly, quality control is fragmented, lacking comprehensive closed-loop management. These issues result in insufficient consistency of integrated data, low timing synchronization rates, and an inability to meet the stability requirements of data integration.
[0003] Chinese Patent Publication No. CN119046356A discloses a system and method for integrating and managing metadata based on Airflow. The method includes the following steps: Business Platform: Metadata information in the metadata management platform can be obtained through API, without considering the database type; Metadata Management Platform: Provides a unified external API and defines a unified entity model for metadata; Configures connection information and configuration information for various relational data sources, and issues integration tasks to the integration service; Integration Service: Used to verify whether the integration tasks issued by the metadata management platform are correct; Integration Task: Executes the integration task and retrieves the corresponding information from the corresponding database; Database: Integrates metadata from multiple relational databases. It is evident that the aforementioned system and method for integrating and managing metadata based on Airflow suffers from a lack of dynamic adaptation mechanisms, resulting in weak ability to cope with complex scenarios. When faced with dynamic issues such as link congestion during peak hours and fluctuations in edge device load, it cannot adjust the task execution rhythm and transmission strategy, leading to task timeouts, incomplete data collection, and insufficient data integration stability. Summary of the Invention
[0004] To address this, the present invention provides a data integration method and system based on metadata-driven intelligence, which overcomes the problems in the prior art where the lack of a dynamic adaptation mechanism results in weak ability to cope with complex scenarios. When faced with dynamic issues such as link congestion during peak hours and fluctuations in edge device load, the system is unable to adjust the task execution rhythm and transmission strategy, leading to task timeouts, incomplete data collection, and insufficient data integration stability.
[0005] To achieve the above objectives, the present invention provides a data integration system based on metadata-driven intelligence, comprising: The data processing module includes a data acquisition unit for collecting financial metadata from several data sources and a preprocessing unit connected to the data acquisition unit for preprocessing the financial metadata to output multi-dimensional metadata features. An integration module, which is connected to the data processing module, includes a model training unit for training an initial model based on the multi-dimensional meta-features to obtain a deep learning model, and an integration unit connected to the model training unit for integrating financial metadata based on the deep learning model to obtain an integration scheme. A confidence threshold adjustment module, which is connected to the integration module, is used to determine the confidence threshold for missing features of multi-dimensional meta-features based on the invalid matching rate of financial metadata integration per unit time. The step size adjustment module is connected to the data processing module and the confidence threshold adjustment module respectively, and is used to determine the iteration step size of the deep learning model based on the backlog rate of the financial metadata queue to be processed within a unit time. The trigger threshold adjustment module is connected to the data processing module and the step size adjustment module respectively, and is used to determine the trigger threshold for multi-dimensional meta-feature redundancy verification based on the loss rate of financial metadata per unit time.
[0006] Furthermore, the confidence threshold adjustment module determines that the integration stability of the financial metadata does not meet the requirements when the invalid matching rate of the financial metadata integration within a unit time is greater than the preset first matching rate.
[0007] Furthermore, the confidence threshold adjustment module responds to the fact that the invalid matching rate of the financial metadata integration within the unit time is greater than the preset first matching rate and less than the preset second matching rate, and initially determines that the processing effectiveness of the financial metadata does not meet the requirements. It then determines whether the processing effectiveness of the financial metadata meets the requirements based on the backlog rate of the pending data queue of the financial metadata within the unit time.
[0008] Furthermore, the confidence threshold adjustment module increases the confidence threshold for missing features of multi-dimensional meta-features in response to the invalid matching rate of financial metadata integration within the unit time being greater than or equal to the preset second matching rate. The increase in the confidence threshold for missing feature completion of the multi-dimensional meta-features is determined by the difference between the invalid matching rate of the financial metadata integration per unit time and the preset second matching rate.
[0009] Furthermore, the step size adjustment module determines that the processing effectiveness of the financial metadata does not meet the requirements when the backlog rate of the pending data queue of financial metadata within a unit time is greater than a preset first backlog rate.
[0010] Furthermore, the step size adjustment module increases the iteration step size of the deep learning model in response to the fact that the backlog rate of the pending data queue of financial metadata within the unit time is greater than the preset first backlog rate and less than the preset second backlog rate. The step size adjustment module responds when the backlog rate of the pending data queue of financial metadata within a unit time is greater than or equal to the preset second backlog rate, initially determines that the consistency of financial metadata collection does not meet the requirements, and determines whether the consistency of financial metadata collection meets the requirements based on the loss rate of financial metadata within a unit time.
[0011] Furthermore, the increase in the iteration step size of the deep learning model is determined by the difference between the backlog rate of the unprocessed data queue of financial metadata per unit time and the preset first backlog rate.
[0012] Furthermore, the trigger threshold adjustment module responds to the fact that the loss rate of financial metadata per unit time is greater than the preset loss rate, determines that the consistency of financial metadata collection does not meet the requirements, and reduces the trigger threshold for multi-dimensional meta-feature redundant feature verification.
[0013] Furthermore, the reduction in the multi-dimensional meta-feature redundancy feature verification trigger threshold is determined by the difference between the loss rate of financial metadata per unit time and the preset loss rate.
[0014] This invention also provides a data integration method based on metadata-driven intelligence, comprising: Collect financial metadata from several data sources and preprocess the financial metadata to output multi-dimensional metadata features; The initial model is trained based on the multi-dimensional meta-features to obtain a deep learning model, and the deep learning model is used to integrate financial metadata to obtain an integration scheme. Obtain the invalid match rate of financial metadata integration within a unit of time, and determine whether the integration stability of financial metadata meets the requirements based on the invalid match rate of financial metadata integration within the unit of time. If the integration stability of the financial metadata does not meet the requirements, it is determined whether it is necessary to increase the confidence threshold for missing feature completion of multi-dimensional meta-features. If it is not necessary to increase the confidence threshold for missing feature completion of multi-dimensional meta-features, then obtain the backlog rate of the unprocessed data queue of financial metadata per unit time to determine whether the processing effectiveness of financial metadata meets the requirements. If the processing effectiveness of the financial metadata does not meet the requirements, then determine whether it is necessary to increase the iteration step size of the deep learning model; If it is not necessary to increase the iteration step size of the deep learning model, the threshold for triggering redundant feature verification of multi-dimensional meta-features can be determined based on the loss rate of financial metadata per unit time.
[0015] Compared with existing technologies, the beneficial effects of this invention are as follows: The system of this invention, through a data processing module, an integration module, a confidence threshold adjustment module, a step size adjustment module, and a trigger threshold adjustment module, adjusts the confidence threshold for missing feature completion of multi-dimensional meta-features based on the invalid matching rate of financial metadata integration within a unit of time. Due to the significant heterogeneity of multi-source data storage formats and differences in metadata description specifications among different data sources, the identification and extraction of core metadata during the integration process are hindered. By increasing the confidence threshold for missing feature completion of multi-dimensional meta-features, highly reliable data with low deviation between the completion results and the true features can be screened out, while speculative completion data with low confidence can be eliminated, reducing invalid matching problems caused by feature completion errors from the source. Furthermore, the iteration step size of the deep learning model is adjusted based on the backlog rate of the pending data queue of financial metadata within a unit of time. Since financial metadata processing requires operations such as format verification, and in multi-dimensional meta-feature integration... In the dimensional feature extraction stage, the increased time consumption of feature noise filtering for low-quality data leads to a longer processing cycle for a single piece of metadata, resulting in decreased processing efficiency and queue backlog. By increasing the iteration step size of the deep learning model, the model can be accelerated to prioritize the allocation of computing power to process core data, making the feature filtering logic for low-quality data more efficient, reducing the time consumption of invalid feature calculation, and effectively alleviating the problem of data backlog. The trigger threshold for redundant feature verification of multi-dimensional features is adjusted according to the loss rate of financial metadata per unit time. Due to hardware buffer overflow and insufficient data bus bandwidth in the acquisition terminal, some key information of financial metadata is lost. By reducing the trigger threshold for redundant feature verification of multi-dimensional features, the redundancy verification mechanism can be reduced. Even if a small feature is detected, the system can still trigger the redundancy verification process, reducing the transmission load of the hardware bus, repairing the loss of key information, preventing the accumulation of lost data, and improving the integration stability of financial metadata.
[0016] Furthermore, the system of the present invention adjusts the dynamic weight coefficient of metadata feature matching by setting a preset first matching rate and a preset second matching rate. Due to the prominent heterogeneity of multi-source data storage formats and the differences in metadata description specifications of different data sources, the identification and extraction of core metadata are hindered during the integration process. By increasing the confidence threshold for missing feature completion of multi-dimensional meta-features, highly reliable data with low deviation between the completion results and the true features can be screened out, and speculative completion data with low confidence can be eliminated. This reduces invalid matching problems caused by feature completion errors from the source and further improves the integration stability of financial metadata.
[0017] Furthermore, the system of the present invention adjusts the iteration step size of the deep learning model by setting a preset first stacking rate and a preset second stacking rate. Since financial metadata processing requires operations such as format verification, and the feature noise filtering of low-quality data in the multi-dimensional meta-feature extraction stage increases the time consumption, the processing cycle of a single piece of metadata is extended, the processing efficiency decreases, and queue stacking occurs. By increasing the iteration step size of the deep learning model, the model can be accelerated to prioritize the allocation of computing power to process core data, the feature screening logic for low-quality data is more efficient, the time consumption of invalid feature calculation is reduced, the problem of data stacking to be processed is effectively alleviated, and the integration stability of financial metadata is further improved.
[0018] Furthermore, the system described in this invention adjusts the trigger threshold for redundant feature verification of multi-dimensional meta-features by setting a preset loss rate. Since the acquisition terminal suffers from hardware buffer overflow and insufficient data bus bandwidth, resulting in partial loss of key information in financial metadata, reducing the trigger threshold for redundant feature verification of multi-dimensional meta-features can reduce the redundancy verification mechanism. Even if a minor feature loss is detected, the system can still trigger the redundancy verification process, reduce the transmission load of the hardware bus, repair the loss of key information, prevent the accumulation of lost data, and further improve the integration stability of financial metadata. Attached Figure Description
[0019] Figure 1 This is a block diagram of the overall structure of the data integration system based on metadata-driven intelligence, as described in an embodiment of the present invention. Figure 2 This is a flowchart illustrating the process of determining the confidence threshold for filling in missing features of multi-dimensional meta-features in a data integration system based on metadata intelligence, as described in this embodiment of the invention. Figure 3 This is a flowchart illustrating the process of determining the iterative step size of a deep learning model in a data integration system based on metadata intelligence, as described in an embodiment of the present invention. Figure 4 This is an overall flowchart of the data integration method based on metadata-driven intelligence, as described in an embodiment of the present invention. Detailed Implementation
[0020] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0021] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0022] Please see Figure 1As shown, it is an overall structural block diagram of the data integration system based on metadata intelligent drive according to an embodiment of the present invention.
[0023] This invention provides a data integration system based on metadata-driven intelligence, comprising: The data processing module includes a data acquisition unit for collecting financial metadata from several data sources and a preprocessing unit connected to the data acquisition unit for preprocessing the financial metadata to output multi-dimensional metadata features. An integration module, which is connected to the data processing module, includes a model training unit for training an initial model based on the multi-dimensional meta-features to obtain a deep learning model, and an integration unit connected to the model training unit for integrating financial metadata based on the deep learning model to obtain an integration scheme. A confidence threshold adjustment module, which is connected to the integration module, is used to determine the confidence threshold for missing features of multi-dimensional meta-features based on the invalid matching rate of financial metadata integration per unit time. The step size adjustment module is connected to the data processing module and the confidence threshold adjustment module respectively, and is used to determine the iteration step size of the deep learning model based on the backlog rate of the financial metadata queue to be processed within a unit time. The trigger threshold adjustment module is connected to the data processing module and the step size adjustment module respectively, and is used to determine the trigger threshold for multi-dimensional meta-feature redundancy verification based on the loss rate of financial metadata per unit time.
[0024] Specifically, the data integration system based on metadata-driven intelligence also includes an execution module connected to the integration module to execute the integration scheme.
[0025] Specifically, data sources include financial regulatory systems, payment and settlement systems, and accounting systems.
[0026] Specifically, financial metadata includes clearing timestamps, transaction serial numbers, and exchange rate conversion parameters.
[0027] Specifically, preprocessing includes data cleaning, data completion, data normalization, and feature extraction.
[0028] Specifically, standard metadata includes the clearing timestamp after data cleaning, the transaction serial number after data completion, and the exchange rate conversion parameters after data normalization.
[0029] Specifically, multi-dimensional meta-features include transaction frequency, the proportion of transaction types, and the customer subscription rate of financial products.
[0030] Specifically, the initial model is an initial framework with multi-dimensional meta-feature fusion capabilities and suitable for intelligent regulation of data integration status.
[0031] Specifically, the deep learning model can be a temporal convolutional network, a long short-term memory network model, or a graph neural network model, with the preferred embodiment being a long short-term memory network model.
[0032] Specifically, the process of training an initial model based on multi-dimensional meta-features to obtain a deep learning model involves dividing the multi-dimensional meta-features into a training set, a validation set, and a test set. The training set is used to allow the model to learn the mapping relationship between features and states. The validation set and the test set are used to optimize the model to obtain a deep learning model.
[0033] Specifically, the integration solution includes a multi-source financial metadata priority scheduling scheme, a financial metadata quality dynamic governance scheme, and a financial metadata integration compliance verification scheme.
[0034] Specifically, the process of integrating financial metadata into an integration solution based on a deep learning model involves inputting the financial metadata into the deep learning model, which then captures the multi-dimensional meta-features of the financial metadata through multi-dimensional feature modeling, dynamically updates the feature fusion weights and strategy decision parameters, adapts to financial business rules and the timeliness requirements of data integration, and outputs the integration solution.
[0035] Specifically, the invalid match rate of financial metadata integration per unit time is the ratio of the number of invalid matches of financial metadata in the feature matching process to the total number of matches per unit time.
[0036] Specifically, the feature matching process involves preprocessing financial metadata to generate multi-dimensional meta-features. The feature matching process then verifies the validity of the multi-dimensional meta-features based on the multi-source data relationships established by the multi-dimensional meta-features in the deep learning model.
[0037] Specifically, invalidity is caused by missing multi-dimensional meta-features, unreliable completion, or conflicting matching logic, resulting in the metadata failing to meet the integration standard.
[0038] Specifically, the confidence threshold for missing feature completion of multi-dimensional meta-features is a threshold standard for evaluating the reliability of the matching between the model's completion results of missing features of financial metadata and the true features.
[0039] Specifically, the backlog rate of unprocessed data queues in financial metadata per unit time is the ratio of the amount of unprocessed data queues in financial metadata to the total amount of financial metadata per unit time.
[0040] Specifically, the iteration step size of a deep learning model is the step size of the ensemble strategy optimization that extends the parameter update step size during the model training phase to the inference phase.
[0041] Specifically, the loss rate of financial metadata per unit time is the ratio of the amount of financial metadata lost per unit time to the total amount of financial metadata collected.
[0042] Specifically, the multi-dimensional meta-feature redundancy feature verification trigger threshold is a core indicator used to quantify whether to start the redundant feature filtering process and control the verification sensitivity. It is a critical value set to detect whether there is information redundancy among multiple features.
[0043] In implementation, the system of this invention adjusts the confidence threshold for missing feature completion of multi-dimensional meta-features based on the invalid matching rate of financial metadata integration per unit time through a data processing module, an integration module, a confidence threshold adjustment module, a step size adjustment module, and a trigger threshold adjustment module. Due to the significant heterogeneity of multi-source data storage formats and differences in metadata description specifications across different data sources, the identification and extraction of core metadata during integration is hindered. By increasing the confidence threshold for missing feature completion of multi-dimensional meta-features, highly reliable data with low deviation between the completion results and the true features can be screened out, while speculative completion data with low confidence can be eliminated, reducing invalid matching problems caused by feature completion errors from the source. The iteration step size of the deep learning model is adjusted based on the backlog rate of the pending data queue of financial metadata per unit time. Since financial metadata processing requires format verification and other operations, and multi-dimensional meta-feature extraction... In the process, the increased time consumption of feature noise filtering for low-quality data leads to a longer processing cycle for a single piece of metadata, resulting in decreased processing efficiency and queue backlog. By increasing the iteration step size of the deep learning model, the model can be accelerated to prioritize the allocation of computing power to process core data, making the feature filtering logic for low-quality data more efficient, reducing the time spent on invalid feature calculations, and effectively alleviating the problem of data backlog. The trigger threshold for redundant feature verification of multi-dimensional meta-features is adjusted according to the loss rate of financial metadata per unit time. Due to hardware buffer overflow and insufficient data bus bandwidth in the acquisition terminal, some key information of financial metadata is lost. By reducing the trigger threshold for redundant feature verification of multi-dimensional meta-features, the redundancy verification mechanism can be reduced. Even if a minor feature loss is detected, the system can still trigger the redundancy verification process, reducing the transmission load of the hardware bus, repairing the loss of key information, preventing the accumulation of lost data, and improving the integration stability of financial metadata.
[0044] Please continue reading. Figure 2 As shown, it is a logical flowchart of the process of determining the confidence threshold for filling missing features of multi-dimensional meta-features in a data integration system based on metadata intelligence according to an embodiment of the present invention.
[0045] Specifically, the confidence threshold adjustment module determines that the integration stability of the financial metadata meets the requirements when the invalid matching rate of the financial metadata integration is less than or equal to the preset first matching rate within a unit of time. The confidence threshold adjustment module determines that the integration stability of the financial metadata does not meet the requirements when the invalid matching rate of the financial metadata integration within a unit time is greater than the preset first matching rate.
[0046] Specifically, the confidence threshold adjustment module responds when the invalid matching rate of the financial metadata integration within a unit time is greater than the preset first matching rate and less than the preset second matching rate, preliminarily determines that the processing effectiveness of the financial metadata does not meet the requirements, and determines whether the processing effectiveness of the financial metadata meets the requirements based on the backlog rate of the pending data queue of the financial metadata within a unit time.
[0047] It is understandable that the preset first matching rate is lower than the preset second matching rate. The three intervals divided by the preset first matching rate and the preset second matching rate correspond to three different scenarios: The first interval is when the invalid matching rate of financial metadata integration per unit time is less than or equal to the preset first matching rate, which corresponds to the situation where the integration stability of financial metadata meets the requirements. The second interval is when the invalid matching rate of financial metadata integration per unit time is greater than the preset first matching rate and less than the preset second matching rate. The corresponding situation is: because financial metadata processing requires the execution of operations such as format verification, and in the multi-dimensional meta-feature extraction stage, the feature noise filtering of low-quality data takes longer, resulting in a longer processing cycle for a single piece of metadata, a decrease in processing efficiency, and queue accumulation. The third interval is when the invalid matching rate of financial metadata integration per unit time is greater than or equal to the preset second matching rate. The corresponding situation is that due to the prominent heterogeneity of multi-source data storage formats, there are differences in the metadata description specifications of different data sources, which leads to the obstruction of core metadata identification and extraction during the integration process.
[0048] Understandably, using preset first and second matching rates to characterize the stability of metadata integration is fundamentally about leveraging invalid match rate thresholds to quantify and grade integration quality. This avoids the limitations of a single threshold judgment and aligns with the dual control requirements of stability and accuracy in financial scenarios. The preset first matching rate is set as the system's acceptable threshold for normal invalid matches, ensuring that invalid matches caused by the inherent heterogeneity of financial data and slight format differences between interfaces of different business systems are not over-reacted to, thus maintaining the operational efficiency of the integration system. The preset second matching rate is set as the system's tolerable upper limit for invalid matches, ensuring that when the invalid match rate exceeds this limit, the threshold adjustment mechanism can be triggered in a timely manner, preventing the continuous deterioration of integration quality from affecting the accuracy of subsequent financial transactions. This dual-threshold setting ensures efficient system operation under steady-state conditions while accurately identifying integration mismatches of different causes and degrees, achieving progressive diagnosis and optimization. The preset first and second matching rates can be set according to actual operating conditions. The setting of the preset first and second matching rates aims to ensure the stability and usability of financial metadata integration. Optionally, the preset first matching rate and preset second matching rate are determined through a limited number of trials by evaluating the integration effect of different invalid matching rates on financial metadata. The determined preset first matching rate and preset second matching rate should be neither too small nor have too great an impact on the financial metadata integration process. For example, the preset first matching rate is generally selected in the range of [1.5%, 3%], and the preset second matching rate is generally selected in the range of [4%, 6%].
[0049] Preferably, the first matching rate is 2% in a preferred embodiment, and the second matching rate is 5% in a preferred embodiment.
[0050] Specifically, the confidence threshold adjustment module increases the confidence threshold for missing features of multi-dimensional meta-features in response to the invalid matching rate of financial metadata integration within the unit time being greater than or equal to the preset second matching rate. The increase in the confidence threshold for missing feature completion of the multi-dimensional meta-features is determined by the difference between the invalid matching rate of the financial metadata integration per unit time and the preset second matching rate.
[0051] Specifically, when the difference between the invalid matching rate of financial metadata integration and the preset second matching rate within a unit time is within 1%, the confidence threshold for missing feature completion of multi-dimensional meta-features is increased to 1.15 times the original value. When the difference between the invalid matching rate of financial metadata integration and the preset second matching rate within a unit time exceeds 1%, in addition to increasing to 1.15 times the original value, for every 0.5% exceeding 1%, the confidence threshold for missing feature completion of multi-dimensional meta-features increases by 0.08. For example, if the difference between the invalid matching rate of financial metadata integration and the preset second matching rate within a unit time is 2%, and the current confidence threshold for missing feature completion of multi-dimensional meta-features is 0.6, the increased confidence threshold for missing feature completion of multi-dimensional meta-features is 0.6 × 1.15 + 0.08 × 2 = 0.85.
[0052] In implementation, the system of the present invention adjusts the dynamic weight coefficient of metadata feature matching by setting a preset first matching rate and a preset second matching rate. Due to the prominent heterogeneity of multi-source data storage formats and the differences in metadata description specifications of different data sources, the identification and extraction of core metadata are hindered during the integration process. By increasing the confidence threshold of missing feature completion for multi-dimensional meta-features, highly reliable data with low deviation between the completion results and the true features can be screened out, and speculative completion data with low confidence can be eliminated. This reduces invalid matching problems caused by feature completion errors from the source and further improves the integration stability of financial metadata.
[0053] Please continue reading. Figure 3 As shown, it is a logical flowchart of the process of determining the iterative step size of a deep learning model in a data integration system based on metadata intelligence according to an embodiment of the present invention.
[0054] Specifically, the step size adjustment module determines that the processing effectiveness of the financial metadata meets the requirements when the backlog rate of the pending data queue of financial metadata is less than or equal to a preset first backlog rate within a unit of time. The step size adjustment module determines that the processing effectiveness of the financial metadata does not meet the requirements when the backlog rate of the pending data queue of financial metadata within a unit time is greater than the preset first backlog rate.
[0055] Specifically, the step size adjustment module increases the iteration step size of the deep learning model in response to the fact that the backlog rate of the pending data queue of financial metadata within the unit time is greater than the preset first backlog rate and less than the preset second backlog rate. The step size adjustment module responds when the backlog rate of the pending data queue of financial metadata within a unit time is greater than or equal to the preset second backlog rate, initially determines that the consistency of financial metadata collection does not meet the requirements, and determines whether the consistency of financial metadata collection meets the requirements based on the loss rate of financial metadata within a unit time.
[0056] It is understandable that the preset first stacking ratio is less than the preset second stacking ratio, and the three intervals divided by the preset first stacking ratio and the preset second stacking ratio correspond to three different situations: The first interval is when the backlog rate of the pending data queue of financial metadata is less than or equal to the preset first backlog rate within a unit of time. The corresponding situation is that the processing effectiveness of the financial metadata meets the requirements. The second interval is when the queue backlog rate of financial metadata to be processed is greater than the preset first backlog rate and less than the preset second backlog rate within a unit of time. The corresponding situation is: because financial metadata processing requires the execution of operations such as format verification, and in the multi-dimensional meta-feature extraction stage, the feature noise filtering of low-quality data takes longer, resulting in a longer processing cycle for a single piece of metadata, a decrease in processing efficiency, and queue backlog. The third interval is when the backlog rate of the pending data queue of financial metadata is greater than or equal to the preset second backlog rate within a unit of time. The corresponding situation is that due to hardware cache overflow and insufficient data bus bandwidth of the acquisition terminal, some key information of financial metadata is lost.
[0057] Understandably, using preset first and second backlog rates to characterize the synergy of financial metadata collection and integration is fundamentally about leveraging backlog rate thresholds to quantify and classify data flow pressure. This avoids the limitations of a single threshold and aligns with the high-reliability control requirements for the timeliness and consistency of metadata integration in the fintech field. The preset first backlog rate is set as the system's acceptable normal backlog threshold, ensuring that temporary and minor queue backlogs caused by peak fluctuations in financial business and network latency are not over-responded to, maintaining the operational stability of the deep learning model. The preset second backlog rate is set as the system's tolerable upper backlog threshold, ensuring that when the backlog exceeds this limit, a deep diagnostic process for collection consistency can be triggered promptly, avoiding risks such as metadata loss and distorted integration results due to escalating backlog. This dual-threshold setting ensures system integration efficiency under steady-state conditions while accurately identifying different degrees of flow imbalance, achieving progressive early warning and control. The preset first and second backlog rates can be set according to actual operating conditions. The setting of the preset first and second backlog rates aims to ensure the stability and usability of financial metadata integration. Optionally, the preset first stacking rate and preset second stacking rate are determined through a limited number of trials by evaluating the integration effect of different stacking rates on financial metadata. The determined preset first stacking rate and preset second stacking rate should meet the requirement that they are neither too small nor have too great an impact on the financial metadata integration process. For example, the preset first stacking rate is generally selected in the range of [1%, 5%], and the preset second stacking rate is generally selected in the range of [6%, 9%].
[0058] Preferably, the first stacking ratio is 3% in a preferred embodiment, and the second stacking ratio is 7% in a preferred embodiment.
[0059] Specifically, the increase in the iteration step size of the deep learning model is determined by the difference between the backlog rate of the unprocessed data queue of financial metadata per unit time and a preset first backlog rate.
[0060] Specifically, when the difference between the backlog rate of the financial metadata queue and the preset first backlog rate is within 2% per unit time, the iteration step size of the deep learning model increases to 1.2 times the original size. When the difference exceeds 2%, in addition to increasing to 1.2 times the original size, the iteration step size of the deep learning model increases by 0.1 for every 1% exceeding the preset first backlog rate. For example, if the difference between the backlog rate of the financial metadata queue and the preset first backlog rate is 4% per unit time, and the current iteration step size of the deep learning model is 0.05, then the increased iteration step size of the deep learning model will be 0.05 × 1.2 + 0.1 × 2 = 0.26.
[0061] In practice, the system described in this invention adjusts the iteration step size of the deep learning model by setting a preset first stacking rate and a preset second stacking rate. Since financial metadata processing requires operations such as format verification, and the feature noise filtering of low-quality data in the multi-dimensional meta-feature extraction stage increases the time consumption, the processing cycle of a single piece of metadata is extended, the processing efficiency decreases, and queue stacking occurs. By increasing the iteration step size of the deep learning model, the model can be accelerated to prioritize the allocation of computing power to process core data, the feature screening logic for low-quality data is more efficient, the calculation time of invalid features is reduced, the problem of data stacking to be processed is effectively alleviated, and the integration stability of financial metadata is further improved.
[0062] Specifically, the trigger threshold adjustment module responds to the fact that the loss rate of financial metadata per unit time is less than or equal to the preset loss rate, and determines that the consistency of the collection of financial metadata meets the requirements. The trigger threshold adjustment module responds when the loss rate of financial metadata per unit time is greater than the preset loss rate, determines that the consistency of financial metadata collection does not meet the requirements, and reduces the trigger threshold for multi-dimensional meta-feature redundant feature verification.
[0063] It is understandable that the two intervals defined by the preset loss rate correspond to two different scenarios: If the loss rate of financial metadata per unit time in the first interval is less than or equal to the preset loss rate, the corresponding situation is: the consistency of the collection of financial metadata meets the requirements. The second interval is when the loss rate of financial metadata per unit time is greater than the preset loss rate. The corresponding situation is that due to hardware buffer overflow and insufficient data bus bandwidth of the acquisition terminal, some key information of financial metadata is lost.
[0064] It is understandable that using a preset loss rate to characterize the health status of financial metadata collection consistency is primarily based on its core properties of quantifying the degree of data integrity deviation, accurately defining the threshold for redundancy verification, and enabling dynamic allocation of computing resources. The preset loss rate is a threshold determined during the system design phase through simulation testing of peak financial business scenarios, verification of cross-institutional data transmission stability, and analysis of core metadata integrity requirements. Its core function is to define the critical point between the acceptable range of data loss and the range requiring intervention and repair; enabling the system to accurately judge the operational health of the collection link, dynamically adjust the triggering timing of redundancy feature verification, and ultimately ensure the consistency and integrity of financial metadata collection, providing high-quality input data for the intelligent integration strategy model. The preset loss rate can be set according to actual working conditions. The setting of the preset loss rate aims to ensure the stability and usability of financial metadata integration. Optionally, the preset loss rate is determined through a limited number of experiments by evaluating the integration effect of different loss rates on financial metadata. The determined preset loss rate should be neither too low nor have an excessive impact on the financial metadata integration process.
[0065] For example, the preset loss rate is typically selected in the range of [0.8%, 1.5%].
[0066] Preferably, the preset loss rate is 1% in this preferred embodiment.
[0067] Specifically, the reduction in the multi-dimensional meta-feature redundancy feature verification trigger threshold is determined by the difference between the loss rate of financial metadata per unit time and the preset loss rate.
[0068] Specifically, when the difference between the loss rate of financial metadata per unit time and the preset loss rate is within 0.5%, the trigger threshold for multi-dimensional meta-feature redundancy verification is reduced to 0.8 times the original value. When the difference exceeds 0.5%, in addition to reducing it to 0.8 times the original value, for every 0.3% exceeding the original value, the trigger threshold for multi-dimensional meta-feature redundancy verification is reduced by 0.08. For example, if the difference between the loss rate of financial metadata per unit time and the preset loss rate is 1.1%, and the current trigger threshold for multi-dimensional meta-feature redundancy verification is 0.6, the reduced trigger threshold would be 0.6 × 0.8 - 0.08 × 2 = 0.32.
[0069] In practice, the system described in this invention adjusts the trigger threshold for redundant feature verification of multi-dimensional meta-features by setting a preset loss rate. Due to hardware buffer overflow and insufficient data bus bandwidth in the acquisition terminal, some key information of financial metadata is lost. By reducing the trigger threshold for redundant feature verification of multi-dimensional meta-features, the redundancy verification mechanism can be reduced. Even if a minor feature is detected, the system can still trigger the redundancy verification process, reduce the transmission load of the hardware bus, repair the loss of key information, prevent the accumulation of lost data, and further improve the integration stability of financial metadata.
[0070] Please continue reading. Figure 4 The diagram shown is an overall flowchart of the data integration method based on metadata-driven intelligence according to an embodiment of the present invention.
[0071] This invention also provides a data integration method based on metadata-driven intelligence, comprising: Step S1: Collect financial metadata from several data sources and preprocess the financial metadata to output multi-dimensional metadata features; Step S2: Train the initial model based on the multi-dimensional meta-features to obtain a deep learning model, and use the deep learning model to integrate financial metadata to obtain an integration scheme; Step S3: Obtain the invalid matching rate of financial metadata integration within a unit time period, and determine whether the integration stability of financial metadata meets the requirements based on the invalid matching rate of financial metadata integration within the unit time period. Step S4: If the integration stability of the financial metadata does not meet the requirements, determine whether it is necessary to increase the confidence threshold for filling missing features of multi-dimensional meta-features. Step S5: If it is not necessary to increase the confidence threshold for missing feature completion of multi-dimensional meta-features, then obtain the data queue backlog rate of financial metadata to be processed per unit time to determine whether the processing effectiveness of financial metadata meets the requirements. Step S6: If the processing effectiveness of the financial metadata does not meet the requirements, determine whether it is necessary to increase the iteration step size of the deep learning model. Step S7: If it is not necessary to increase the iteration step size of the deep learning model, determine the multi-dimensional meta-feature redundancy feature verification trigger threshold based on the loss rate of financial metadata per unit time.
[0072] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
Claims
1. A data integration system based on metadata-driven intelligence, characterized in that, include: The data processing module includes a data acquisition unit for collecting financial metadata from several data sources and a preprocessing unit connected to the data acquisition unit for preprocessing the financial metadata to output multi-dimensional metadata features. An integration module, which is connected to the data processing module, includes a model training unit for training an initial model based on the multi-dimensional meta-features to obtain a deep learning model, and an integration unit connected to the model training unit for integrating financial metadata based on the deep learning model to obtain an integration scheme. A confidence threshold adjustment module, which is connected to the integration module, is used to determine the confidence threshold for missing features of multi-dimensional meta-features based on the invalid matching rate of financial metadata integration per unit time. The step size adjustment module is connected to the data processing module and the confidence threshold adjustment module respectively, and is used to determine the iteration step size of the deep learning model based on the backlog rate of the unprocessed data queue of financial metadata per unit time. The trigger threshold adjustment module is connected to the data processing module and the step size adjustment module respectively, and is used to determine the trigger threshold for multi-dimensional meta-feature redundancy verification based on the loss rate of financial metadata per unit time.
2. The data integration system based on metadata-driven intelligence according to claim 1, characterized in that, The confidence threshold adjustment module determines that the integration stability of financial metadata does not meet the requirements when the invalid matching rate of financial metadata integration within a unit time is greater than the preset first matching rate.
3. The data integration system based on metadata-driven intelligence according to claim 2, characterized in that, The confidence threshold adjustment module responds when the invalid matching rate of the financial metadata integration within a unit time is greater than the preset first matching rate and less than the preset second matching rate, and initially determines that the processing effectiveness of the financial metadata does not meet the requirements. It then determines whether the processing effectiveness of the financial metadata meets the requirements based on the backlog rate of the pending data queue of the financial metadata within a unit time.
4. The data integration system based on metadata-driven intelligence according to claim 3, characterized in that, The confidence threshold adjustment module increases the confidence threshold for missing features of multi-dimensional meta-features in response to the invalid matching rate of financial metadata integration within the unit time being greater than or equal to the preset second matching rate. The increase in the confidence threshold for missing feature completion of the multi-dimensional meta-features is determined by the difference between the invalid matching rate of the financial metadata integration per unit time and the preset second matching rate.
5. The data integration system based on metadata-driven intelligence according to claim 4, characterized in that, The step size adjustment module determines that the processing effectiveness of the financial metadata does not meet the requirements when the backlog rate of the pending data queue of financial metadata within a unit time is greater than the preset first backlog rate.
6. The data integration system based on metadata-driven intelligence according to claim 5, characterized in that, The step size adjustment module increases the iteration step size of the deep learning model in response to the fact that the backlog rate of the pending data queue of financial metadata within the unit time is greater than the preset first backlog rate and less than the preset second backlog rate. The step size adjustment module responds when the backlog rate of the pending data queue of financial metadata within a unit time is greater than or equal to the preset second backlog rate, initially determines that the consistency of financial metadata collection does not meet the requirements, and determines whether the consistency of financial metadata collection meets the requirements based on the loss rate of financial metadata within a unit time.
7. The data integration system based on metadata-driven intelligence according to claim 6, characterized in that, The increase in the iteration step size of the deep learning model is determined by the difference between the backlog rate of the unprocessed data queue of financial metadata per unit time and the preset first backlog rate.
8. The data integration system based on metadata-driven intelligence according to claim 7, characterized in that, The trigger threshold adjustment module responds when the loss rate of financial metadata per unit time is greater than the preset loss rate, determines that the consistency of financial metadata collection does not meet the requirements, and reduces the trigger threshold for multi-dimensional meta-feature redundant feature verification.
9. The data integration system based on metadata-driven intelligence according to claim 8, characterized in that, The reduction in the multi-dimensional meta-feature redundancy feature verification trigger threshold is determined by the difference between the loss rate of financial metadata per unit time and the preset loss rate.
10. A data integration method applied to the metadata-driven intelligent data integration system according to any one of claims 1-9, characterized in that, include: Collect financial metadata from several data sources and preprocess the financial metadata to output multi-dimensional metadata features; The initial model is trained based on the multi-dimensional meta-features to obtain a deep learning model, and the deep learning model is used to integrate financial metadata to obtain an integration scheme. Obtain the invalid match rate of financial metadata integration within a unit of time, and determine whether the integration stability of financial metadata meets the requirements based on the invalid match rate of financial metadata integration within the unit of time. If the integration stability of the financial metadata does not meet the requirements, it is determined whether it is necessary to increase the confidence threshold for missing feature completion of multi-dimensional meta-features. If it is not necessary to increase the confidence threshold for missing feature completion of multi-dimensional meta-features, then obtain the backlog rate of the unprocessed data queue of financial metadata per unit time to determine whether the processing effectiveness of financial metadata meets the requirements. If the processing effectiveness of the financial metadata does not meet the requirements, then determine whether it is necessary to increase the iteration step size of the deep learning model; If it is not necessary to increase the iteration step size of the deep learning model, the threshold for triggering redundant feature verification of multi-dimensional meta-features can be determined based on the loss rate of financial metadata per unit time.
Citation Information
Patent Citations
Airflow-based metadata integration and management system and working method
CN119046356A