Report automatic generation method and system based on multi-source data integration and medium
By deploying a lightweight verification rule engine and hierarchical tag rules to integrate data in a multi-source data system, the problems of low storage accuracy and high false alarm rate caused by data system isolation are solved, and efficient and accurate data retrieval and report generation are achieved.
Patent Information
- Application Number
- CN202511555083.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-29
- Publication Date
- 2026-03-03
- Estimated Expiration
- 2045-10-29
AI Technical Summary
In existing technologies, isolated data systems rely on manual operation for multi-source data acquisition and processing, resulting in low accuracy of mixed data storage and a lack of conflict resolution mechanisms. This leads to data redundancy, delays in cross-system data retrieval response, and high report error rates.
By deploying a lightweight validation rule engine, local preprocessing is performed on multiple source business terminals to output standardized data streams. Then, cross-business terminal redundant validation and integration of hierarchical tag rules are performed on the data integration cloud to generate a unified resource pool. Real-time associated tag combinations are obtained by parsing user queries and generating permission filtering reports.
It improves the data storage accuracy of multi-source heterogeneous data hybrid storage, and realizes millisecond-level data retrieval and efficient report visualization generation with low error rate.
Smart Images

Figure CN121029865B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and specifically to a method, system, and medium for automated report generation based on multi-source data integration. Background Technology
[0002] Currently, there are significant technical bottlenecks in the field of multi-source heterogeneous data integration: isolated independent business systems using heterogeneous database architectures force operators to manually perform data collection and processing. Specifically, operators need to repeatedly enter the same entity information into different system interfaces, manually clean up format conflicts using offline spreadsheet tools, and manually correct outliers. This operating mode directly leads to severe degradation in the accuracy of hybrid data storage, and the lack of a conflict resolution engine results in redundant copies being stored independently in different systems.
[0003] The aforementioned defects lead to a systemic performance collapse at the application layer: when cross-system related data needs to be retrieved, operators must manually log into multiple systems to perform serial searches, and the generation of comprehensive reports has a high false alarm rate due to reliance on manual merging operations.
[0004] In summary, existing technologies suffer from the technical problems of requiring manual operation for multi-source data collection and processing in isolated data systems. This results in low accuracy of mixed data storage and a lack of conflict resolution mechanisms, leading to data redundancy. Consequently, in practical applications, this causes delays in cross-system data retrieval and a high rate of report errors. Summary of the Invention
[0005] This invention provides a method, system, and medium for automated report generation based on multi-source data integration. It addresses the technical problems in existing technologies where manual operation is used to collect and process multi-source data from isolated data systems, resulting in low accuracy of mixed data storage and a lack of conflict resolution mechanisms, leading to data redundancy. Consequently, in practical applications, this results in delayed response times for cross-system data retrieval and high report error rates.
[0006] In view of the above problems, the present invention provides a method, system and medium for automated report generation based on multi-source data integration.
[0007] The first aspect of this invention provides a method for automatically generating reports based on multi-source data integration. The method includes: preprocessing real-time monitoring business data captured by multiple source business terminals using multiple lightweight verification rule engines deployed on multiple source business terminals, outputting multiple standardized data streams; receiving data in the cloud and performing cross-business terminal redundancy verification and integration of the multiple standardized data streams based on hierarchical tag rules, outputting a unified resource pool, wherein each data record in the unified resource pool carries a composite tag identifier; parsing a user-input form-based query to obtain real-time associated tag combinations, dynamically extracting matching data from the unified resource pool based on the real-time associated tag combinations, generating a temporary dataset; and after retrieving viewing permissions based on the user ID, converting the temporary dataset into a permission-filtered report using the viewing permissions as a field filtering strategy.
[0008] A second aspect of the present invention provides an automated report generation system based on multi-source data integration. The system includes: a local data processing unit, configured to perform local preprocessing of real-time monitoring business data captured by multiple source business terminals using multiple lightweight verification rule engines deployed at multiple source business terminals, and output multiple standardized data streams; a data integration execution unit, configured to integrate data received from the cloud and perform cross-business terminal redundancy verification and integration of the multiple standardized data streams based on hierarchical tag rules, and output a unified resource pool, wherein each data record in the unified resource pool carries a composite tag identifier; a data dynamic extraction unit, configured to parse user-input form-based queries to obtain real-time associated tag combinations, and then dynamically extract matching data from the unified resource pool based on the real-time associated tag combinations to generate a temporary dataset; and a report generation execution unit, configured to retrieve viewing permissions based on user IDs, and then convert the temporary dataset into a permission-filtered report using the viewing permissions as a field filtering strategy.
[0009] A third aspect of the present invention provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the report automation generation method based on multi-source data integration in the first aspect.
[0010] One or more technical solutions provided in this invention have at least the following technical effects or advantages:
[0011] The method provided in this invention utilizes multiple lightweight verification rule engines deployed at multiple source business terminals to perform local preprocessing of real-time monitoring business data captured by these terminals, outputting multiple standardized data streams. The cloud-based data integration system receives the data and performs cross-business terminal redundancy verification and integration based on hierarchical tag rules, outputting a unified resource pool. Each data record in the unified resource pool carries a composite tag identifier. After parsing the user-input form-based query to obtain real-time associated tag combinations, matching data is dynamically extracted from the unified resource pool based on these combinations to generate a temporary dataset. After retrieving viewing permissions based on the user ID, the temporary dataset is converted into a permission-filtered report using the viewing permissions as a field filtering strategy. This achieves the technical effect of improving data storage accuracy in multi-source heterogeneous data hybrid storage while enabling millisecond-level data retrieval based on user needs and high-efficiency, low-error-report visualization generation based on a hierarchical tag system.
[0012] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0013] Figure 1 This is a schematic diagram of the automated report generation method based on multi-source data integration provided by the present invention.
[0014] Figure 2 This is a schematic diagram of the structure of the report automation generation system based on multi-source data integration provided by the present invention.
[0015] Explanation of reference numerals in the attached figures: Local data processing unit 11, data integration execution unit 12, dynamic data extraction unit 13, report generation execution unit 14. Detailed Implementation
[0016] This invention provides a method, system, and medium for automated report generation based on multi-source data integration. It addresses the technical problems of existing technologies that rely on manual operation for multi-source data acquisition and processing in isolated data systems. These methods result in low accuracy of mixed data storage, a lack of conflict resolution mechanisms leading to data redundancy, and consequently, delayed response times for cross-system data retrieval and high report error rates in practical applications. The invention achieves improved data storage accuracy for mixed storage of heterogeneous multi-source data while simultaneously enabling millisecond-level data retrieval based on a hierarchical tagging system and efficient, low-error-prone report visualization generation.
[0017] The technical solutions of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. It should be understood that the present invention is not limited to the exemplary embodiments described herein. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention. It should also be noted that, for ease of description, only the parts related to the present invention are shown in the accompanying drawings, not all of them.
[0018] Example 1, as Figure 1 As shown, this invention provides a method for automatically generating reports based on multi-source data integration, the method comprising:
[0019] A100: By deploying multiple lightweight verification rule engines on multiple source business terminals, it performs local preprocessing of the real-time monitoring business data captured by the multiple source business terminals and outputs multiple standardized data streams.
[0020] Furthermore, the method provided by the present invention also includes:
[0021] A110: Locally invoke multiple data update cycles of the multiple source business terminals.
[0022] A120: Based on the business type priority association attributes of the multiple source business terminals, the multiple data update cycles are modified to obtain multiple dynamic scheduling cycles.
[0023] A130: Using the multiple dynamic scheduling cycles as task trigger constraints, drive the multiple lightweight verification rule engines to perform periodic data processing and uploading on the multiple source business terminals.
[0024] Specifically, the lightweight verification rule engine described in this embodiment is used for basic data verification and standardization. In this embodiment, the lightweight verification rule engine is deployed on the multiple source business terminals that serve as edge data upload nodes.
[0025] Multiple lightweight validation rule engines deployed at various source business endpoints capture raw business data streams through a real-time monitoring mechanism and perform standardized preprocessing operations on the local edge. Preprocessing includes data format conversion, such as XML / JSON normalization; basic logic validation, such as field integrity checks; and outlier interception, such as numerical range filtering, ultimately outputting a standardized data stream that conforms to cloud processing specifications.
[0026] This step transforms scattered, heterogeneous business data into structured data units, providing a unified input source for cloud integration and effectively solving the problem of fragmented multi-source data.
[0027] To further ensure the timeliness of data uploads, this embodiment obtains the original data update cycle parameters defined within each business terminal, thereby obtaining the multiple data update cycles that reflect the inherent generation patterns of the business data from the multiple source business terminals, providing a benchmark for subsequent dynamic scheduling.
[0028] Based on the priority association attributes of the business types of the multiple source business terminals, such as the rule that people's livelihood and medical data takes precedence over agricultural statistics data, the multiple data update cycles are modified to obtain multiple dynamic scheduling cycles.
[0029] The specific correction methods are not strictly limited in this embodiment. Preferably, a weighted algorithm is used to integrate the urgency of the business, the sensitivity of the data, and the system load status to generate a dynamic scheduling cycle that adapts to the actual scenario. For example, the low-frequency business data cycle is extended to twice the original value.
[0030] The dynamic scheduling cycle is transformed into task triggering constraints. Specifically, the multiple dynamic scheduling cycles are used as task triggering constraints to drive the multiple lightweight verification rule engines to perform data processing and uploading according to the optimized time-series strategy.
[0031] Each lightweight verification rule engine is automatically activated according to periodic instructions to complete data cleaning, standardized encapsulation, and cloud transmission, enabling collaborative scheduling across business systems.
[0032] The periodic data upload control method in this embodiment avoids network congestion caused by concentrated requests, and reduces resource consumption at the edge through peak-shaving processing, thus ensuring the efficiency and stability of the data acquisition process.
[0033] A200: Data integration cloud receives and integrates the multiple standardized data streams across business terminals based on hierarchical tag rules, and outputs a unified resource pool, wherein each data record in the unified resource pool is marked with a composite tag.
[0034] Furthermore, the data integration cloud receives and performs cross-business redundancy verification and integration of the multiple standardized data streams based on hierarchical tagging rules, outputting a unified resource pool. The method step A200 provided by this invention includes:
[0035] A210: Based on the benchmark classification rules, static labeling of the multiple standardized data streams is performed in the data integration cloud to obtain multiple basic label data units.
[0036] A220: After performing cross-node redundancy check correction on the multiple basic tag data units, the data of the multiple basic tag data units is fused based on the primary key driven fusion mechanism to output a basic resource pool, wherein the basic resource pool stores multiple discrete data records, and each discrete data record is marked with a basic tag identifier.
[0037] A230: Based on cross-node correlation retrieval, extended data correlation is used to construct composite label classification rules, wherein the baseline classification rules and composite label classification rules constitute the hierarchical label rules.
[0038] A240: Based on the composite label classification rules, dynamically add labels to the basic resource pool and output the unified resource pool, wherein each discrete data record in the unified resource pool is labeled with a composite label.
[0039] Furthermore, after performing cross-node redundancy check correction on the multiple basic tag data units, data fusion of the multiple basic tag data units is performed based on the primary key-driven fusion mechanism to output a basic resource pool. The method step A220 provided by this invention includes:
[0040] A221: Based on the business-side confidence priority, the multiple basic tag data units are cross-polled to perform multi-round field conflict verification and calibration for duplicate records of the same entity, resulting in multiple conflict resolution data units and multiple redundant copies. The multiple redundant copies are mapped and backed up to the local storage of the multiple source business units.
[0041] A222: Set multiple business primary key identifiers for the multiple conflict resolution data units, and perform resource pool existence judgment based on the multiple business primary key identifiers.
[0042] A223: Based on the judgment result, the fields of the multiple conflict resolution data units are adaptively written to construct the data consistency of the basic resource pool.
[0043] Furthermore, the method provided by the present invention also includes:
[0044] A221-1: Interact to obtain the multiple data update time distributions of the multiple basic tag data units.
[0045] A221-2: Locally retrieve the authority ratings of multiple data sources from the multiple source business terminals.
[0046] A221-3: Based on the multiple data update time distributions and multiple data source authority ratings, perform dynamic weighted calculation of confidence priorities and output multiple business-side confidence scores.
[0047] A221-4: Arrange the multiple source business terminals in descending order according to their confidence scores, and output the confidence priority of each business terminal.
[0048] Furthermore, the method provided by the present invention also includes:
[0049] A223-1: If the first business primary key identifier of the first conflict resolution data unit exists in the historical resource pool, then the updatable field is retrieved from the first conflict resolution data unit according to the field mapping table to perform incremental update of the basic resource pool.
[0050] A223-2: If the first business primary key identifier of the first conflict resolution data unit does not exist in the historical resource pool, then the data record of the first conflict resolution data unit is completely inserted into the basic resource pool.
[0051] Furthermore, the method provided by the present invention also includes:
[0052] A241: Perform cross-node correlation retrieval based on business primary key identifiers to obtain multiple historical multi-source data records of multiple sample business entities.
[0053] A242: Based on the predefined business logic rules, associate the basic tag numbers of the multiple historical multi-source data records to construct multiple composite tag combination rules.
[0054] A243: Store the multiple composite label combination rules into the classification rule library to obtain the composite label classification rules.
[0055] A244: Based on the composite tag classification rules, traverse the basic resource pool, add dynamic tag identifiers to data records that meet the tag association conditions, and output the unified resource pool.
[0056] Specifically, the data integration cloud is used to receive standardized data streams from the edge preprocessing stage, and then performs cross-business system data integration operations based on a preset hierarchical tagging rule system to form an integrated data asset that supports multi-dimensional retrieval, namely the unified resource pool. This completely breaks down data barriers between business systems, realizes the transformation from raw data to reusable knowledge assets, and lays the data foundation for subsequent dynamic report generation.
[0057] The specific technical implementation of the data integration across multiple edge devices in the cloud to generate the unified resource pool is as follows:
[0058] A predefined baseline classification rule is provided, which includes a dual core mechanism: a business attribute classification mechanism and a business primary key identifier generation rule.
[0059] For each data record, it is divided into basic categories according to its business attributes, such as population information being classified into the 100 series and economic indicators into the 200 series. A unique basic label number is assigned to each type of data record, and multiple basic label data units are output. Each unit corresponds to a data record and carries a basic label number, retaining the original business characteristics and forming the atomic input for subsequent data fusion.
[0060] In this embodiment, the static identification process establishes a basic semantic framework for data management, ensuring that subsequent operations have a traceable classification basis.
[0061] For data units with completed basic tagging, cross-system redundancy verification and correction are performed. First, multiple source records of the same entity are associated based on primary key identifiers, and field-level conflicts are detected. Then, a primary key-driven fusion mechanism is used to control data writing behavior through existence checks: if the primary key exists, incremental updates are performed; otherwise, complete insertion is performed. The output is a basic resource pool that eliminates redundancy. Each discrete record in the pool carries a basic tag identifier and satisfies consistency constraints. This step solves the problem of duplicate data entry from multiple sources and ensures the accuracy and integrity of core data through primary key uniqueness.
[0062] The specific technical implementation of cross-system redundancy verification and correction is as follows:
[0063] First, the confidence priority of the business side is dynamically set. Specifically, the data update time distribution of each basic tag data unit is dynamically obtained, including the last update timestamp, update frequency and historical change trajectory. For example, the low-income data of the civil affairs system is updated daily, while the cultivated land data is statistically analyzed weekly.
[0064] The data update time distribution quantifies data freshness, provides objective time-series evidence for conflict resolution, and ensures that the latest version of data is used in calibration.
[0065] Authoritative rating data is retrieved from the source business. The rating criteria include the level of the data source institution, the importance of the business, and historical accuracy. For example, the authority score for data from the health system is set at 90 points, while data collected from third parties is set at 60 points. This rating establishes a quantitative benchmark for data credibility, supporting the objectivity of priority calculation.
[0066] By integrating the update time distribution and authority rating, a business confidence score is calculated using a dynamic weighted model. Here, the dynamic weighted model assigns a differentiation coefficient to the timeliness factor and the authority factor to generate the multiple business confidence scores.
[0067] The preferred built-in formula of the dynamic weighted model is: Score = Timeliness weight × Standardized freshness + Authority weight × Rating value. The output is a sortable quantitative indicator. For example, a certain medical data receives a score of 0.95 because it is updated in real time and comes from an authoritative source.
[0068] The business-side sequences are arranged in descending order of confidence scores to form a business-side confidence priority for conflict resolution. Business-side sequences with higher scores have decision-making priority in field conflict scenarios, such as health data with a score of 0.95 covering community data with a score of 0.75. This sequence dynamically adapts to changes in business scenarios to ensure the dominant position of key data in the integration process.
[0069] Based on the business-side confidence priority, multiple rounds of iterative field conflict verification and calibration are performed on cross-node data records carrying the same basic tag identifier.
[0070] Specifically, when duplicate data describing the same business entity from multiple sources is detected, such as a resident's ID number being duplicated in the health record of the health system and the welfare record of the civil affairs system, a cross-polling mechanism is initiated:
[0071] The system iterates through the data units of each business unit in descending order of confidence score, comparing key attribute values such as address and marital status field by field. When a conflict is found, such as the health system registering the address as village A while the civil affairs system registers it as village B, the latest data version from the business unit with higher confidence is adopted as the authoritative value, and the conflicting data from the business unit with lower confidence is marked as a redundant copy.
[0072] These redundant copies are not simply discarded, but are sent back to the local storage nodes of the corresponding source business system through metadata mapping relationships, thus preserving the traceability of data while respecting the data sovereignty of the business side.
[0073] After multiple rounds of fine calibration, standardized data units with conflict-free output are output, providing unambiguous, high-quality input for subsequent primary key-driven fusion.
[0074] The cross-polling mechanism and business-side confidence priority mechanism in this embodiment ensure data consistency in the cloud while constructing a dual-track architecture of centralized management of authoritative data and distributed retention of raw data, achieving a technical effect of effectively balancing global integration needs and local data autonomy.
[0075] Each conflict-resolved data unit is assigned a globally unique business primary key identifier, and an existence check is performed in the historical resource pool based on the business primary key identifiers:
[0076] Specifically, when the primary key of the conflict resolution data is detected to exist in the historical resource pool, the subset of fields that need to be updated is extracted according to the preset field mapping table. For example, Field_Map defines "number of family members" as mapping to the "permanent resident population" field.
[0077] In this scenario, this embodiment only writes the changed fields to the resource pool, while keeping the non-mapped fields at their original values. This incremental update strategy minimizes I / O overhead and ensures the real-time performance of high-frequency data.
[0078] Specifically, when the primary key of conflict resolution data does not exist in the historical resource pool, the data record is inserted into the resource pool as a new entity. The insertion operation includes all fields, such as the population, land, and economic data of the new village, and automatically inherits its basic tag identifier.
[0079] Furthermore, by associating cross-node data records through business primary key identifiers, the business logic relationships between basic tags are extended, and composite tag classification rules are constructed. For example, associating "101-Population Tag" with "205-Medical Insurance Tag" yields the composite tag "Elderly Medical Insurance = 101 + 205".
[0080] Composite label classification rules and baseline classification rules together constitute a hierarchical label system. The former describes the cross-semantics of business scenarios, while the latter defines atomic data attributes, realizing the knowledge sublimation from basic classification to business scenarios.
[0081] Based on the business primary key identifier (such as ID card number) in the baseline classification rules, cross-node association retrieval is performed in historical multi-source data records. For example, all historical records of the same resident in the medical, social security, and civil affairs systems are retrieved, and their tag combination patterns are identified (such as often containing the "medical + social security" tag). This step uncovers potential business relationships between data, providing an empirical basis for the generation of composite rules.
[0082] It should be understood that the composite label classification rule in this embodiment is derived from the baseline classification rule. Using the baseline classification rule as the atomic-level classification framework, and based on cross-node correlation retrieval to extend data correlation, the specific implementation of the composite label classification rule is as follows:
[0083] Using business primary key identifiers such as resident ID numbers or village-level administrative division codes as globally unique indexes, cross-business node association retrieval operations are performed in a distributed data environment.
[0084] This process uses a primary key value matching mechanism to extract historical data records of the same business entity generated at different times from multiple heterogeneous business systems.
[0085] For example, when searching for a resident's ID number, their health records in the health system, welfare records in the civil affairs system, and insurance information in the social security system are obtained simultaneously, forming a complete vertical multi-source data chain for that entity.
[0086] This step breaks through the barriers of the business system, provides empirical evidence for subsequent tag combination analysis, and ensures that the construction of composite tag rules is based on real business behavior trajectories.
[0087] Basic tag parsing is performed on the retrieved multi-source historical records. For example, health records are marked as population attributes and medical insurance records are marked as medical insurance status. Based on the pre-set business logic rules, such as "elderly people must meet the conditions of being over 60 years old and participating in medical insurance", high-frequency co-occurring tag combination patterns are identified.
[0088] The tag numbers that meet the conditions are associated according to the rules to generate tag combination rules that describe complex business scenarios.
[0089] For example, when a resident's historical records show a repeated coexistence of population attributes and medical insurance labels, the "elderly medical insurance" composite rule is automatically abstracted.
[0090] The rule-building process integrates domain knowledge to ensure that composite tags have business interpretability and practical application value, ultimately constructing multiple composite tag combination rules.
[0091] The multiple composite label combination rules are persistently stored in the classification rule library to obtain composite label classification rules.
[0092] The composite tag classification rules adopt a structured storage design. Each rule includes a composite tag identifier, a basic tag combination expression, and a description of the applicable business scenario. These rules support efficient retrieval and dynamic expansion, providing a standardized basis for the dynamic addition of tags to the unified resource pool while ensuring the maintainability of the rules.
[0093] Traverse all data records in the basic resource pool, match each record against the tag association conditions in the composite tag classification rules, and add composite tag identifiers to records that meet the conditions to form a unified resource pool carrying multi-layer semantic tags.
[0094] For example, a resident record is dynamically marked as "elderly medical insurance" because it simultaneously possesses population attributes and medical insurance labels.
[0095] The output results achieve semantic enhancement from atomic data attributes to business scenarios, and support one-click data aggregation by compound tags, such as quickly counting all elderly people covered by medical insurance, significantly improving the efficiency of multidimensional analysis.
[0096] It should be understood that each discrete data record in the unified resource pool will eventually be identified by a composite label. The composite label is not a simple classification mark, but a scenario-based identifier formed by intelligently combining the basic label numbers based on predefined business logic rules. For example, population attribute labels and medical insurance status labels are merged into a composite label of "elderly medical insurance".
[0097] This identification mechanism enables previously isolated data records to acquire multi-dimensional business semantics, greatly improving multi-dimensional retrieval efficiency and decision support capabilities, while also providing rich semantic basis for automatic binding of subsequent reports.
[0098] For example, when a resident record carries both population attributes and medical insurance records, a corresponding composite identifier is automatically added, forming a two-layer semantic structure of basic attributes + business scenario.
[0099] A300: After parsing the user's form-based query to obtain the real-time associated tag combination, it dynamically extracts matching data from the unified resource pool based on the real-time associated tag combination to generate a temporary dataset.
[0100] Based on the form-based query requests entered by users on the front-end interface, such as selecting multiple conditions like "elderly medical insurance" and "farmland subsidies", the system analyzes the logical relationships of the tag combinations contained in the request in real time and transforms them into executable composite tag retrieval instructions.
[0101] Then, based on the instruction, a millisecond-level dynamic retrieval is performed in the unified resource pool. The composite tag index is used to accurately locate discrete data records that simultaneously meet all tag conditions. Finally, the full field data of the matching records is extracted as needed to generate a temporary memory dataset that only retains the data related to this query.
[0102] This process breaks through the static limitations of traditional pre-generated reports, enabling end-to-end real-time response from user needs to data results. It ensures that when workers query data in specific scenarios, they can instantly obtain the latest authoritative data across business domains, completely eliminating data lag and the time loss from manual screening.
[0103] A400: After retrieving the viewing permissions based on the user ID, the temporary dataset is converted into a permission filtering report using the viewing permissions as the field filtering strategy.
[0104] Furthermore, after retrieving viewing permissions based on the user ID, the temporary dataset is converted into a permission filtering report using the viewing permissions as a field filtering strategy. Step A400 of the method provided by this invention includes:
[0105] A410: Query the permission configuration table based on the user ID to obtain viewing permissions. The viewing permissions include a list of field-level access permissions and a list of business domain access permissions. The list of field-level access permissions includes the range of accessible fields and restrictions on sensitive data.
[0106] A420: Use the business domain access permission list to perform record-level filtering of the temporary dataset, and use the field-level access permission list to perform field-level filtering of the temporary dataset to obtain the filtered dataset.
[0107] A430: After converting the filtered dataset into the permission filtering report based on the standard report format, dynamic watermarks are added to the permission filtering report according to the user ID.
[0108] Specifically, in this embodiment, after obtaining the temporary dataset generated by the user query, the system queries the preset permission configuration rule base based on the user's identity identifier (ID) to extract the user's data access permission policy.
[0109] This strategy incorporates a dual control dimension: the field-level access permission list explicitly limits the range of specific data fields that users can view, while the business domain access permission list defines the range of business topics that users are allowed to access.
[0110] These permission rules serve as the mandatory enforcement standards for data filtering, performing fine-grained processing on temporary datasets, removing content that exceeds permissions, and finally outputting permission filtering reports that comply with security standards. This step ensures the output of data value while completely eliminating the risk of unauthorized access.
[0111] The specific technical implementation for filtering permission information is as follows:
[0112] The user ID is used to retrieve the associated permission configuration entry in the permission configuration table. The permission configuration table adopts a structured storage design, and each record contains the user ID, business domain permission code, field permission code, and sensitive data control flag.
[0113] The field-level access permission list explicitly enumerates the set of field names that users can operate on, such as allowing access to the "number of family members" and "area of cultivated land" fields, and marks the access restriction rules for sensitive fields, such as only displaying the last four digits of the "ID number" field for village-level users.
[0114] The business domain access permission list limits the boundaries of data topics that users can access through business codes.
[0115] This mechanism enables precise and dynamic management of permissions, ensuring that users with different roles only have access to the minimum dataset within their authorized scope.
[0116] First, the business domain access permission list is applied to perform record-level filtering on the temporary dataset. Specifically, the business tag identifier carried by each data record is scanned, and only business domain records permitted by the permission list are retained.
[0117] Then, the field-level access permission list is applied to perform field-level filtering. Specifically, the field metadata of each record is traversed, unauthorized fields are removed, and sensitive fields are de-identified.
[0118] After double filtering, a subset of data containing only authorized content is generated, which satisfies business needs while strictly adhering to data security red lines.
[0119] The filtered dataset is converted to a predefined report template, such as an HTML table or a PDF document, to generate a structured report file.
[0120] Then, an invisible watermark is dynamically embedded in the report based on the user ID. This watermark is deeply integrated with the report content and does not interfere with data reading.
[0121] In the event of a data breach, the responsible party can be quickly traced through watermark analysis. This design enhances data traceability while ensuring report readability, forming a complete security loop. The final output permission filtering report simultaneously meets the three core requirements of data availability, security, and traceability.
[0122] This embodiment achieves the technical effect of improving the data storage accuracy of multi-source heterogeneous data hybrid storage, while realizing millisecond-level data retrieval and high-efficiency, low-error report visualization generation based on a hierarchical tag system.
[0123] Example 2 is based on the same inventive concept as the report automation generation method based on multi-source data integration in the previous examples, such as... Figure 2 As shown, this invention provides an automated report generation system based on multi-source data integration, wherein the system includes:
[0124] The local data processing unit 11 is used to perform local preprocessing of the real-time monitoring business data captured by the multiple source business terminals through multiple lightweight verification rule engines deployed on multiple source business terminals, and output multiple standardized data streams.
[0125] The data integration execution unit 12 is used to integrate data received from the cloud and perform cross-business redundancy verification and integration of the multiple standardized data streams based on hierarchical tag rules, and output a unified resource pool, wherein each data record in the unified resource pool is marked with a composite tag.
[0126] The data dynamic extraction unit 13 is used to parse the form-type query input by the user to obtain the real-time associated tag combination, and then dynamically extract matching data from the unified resource pool according to the real-time associated tag combination to generate a temporary dataset.
[0127] The report generation execution unit 14 is used to retrieve the viewing permissions based on the user ID, and then convert the temporary dataset into a permission-filtered report using the viewing permissions as the field filtering strategy.
[0128] Furthermore, the data integration execution unit 12 is also used for:
[0129] In the data integration cloud, static labels are assigned to the multiple standardized data streams based on benchmark classification rules to obtain multiple basic label data units. After cross-node redundancy verification and correction are performed on the multiple basic label data units, data fusion is performed on the multiple basic label data units based on the primary key-driven fusion mechanism to output a basic resource pool. The basic resource pool stores multiple discrete data records, each with a basic label identifier. Based on cross-node correlation retrieval, extended data correlation is constructed to build composite label classification rules. The benchmark classification rules and composite label classification rules constitute the hierarchical label rules. Based on the composite label classification rules, dynamic label identifiers are appended to the basic resource pool to output a unified resource pool. Each discrete data record in the unified resource pool has a composite label identifier.
[0130] Furthermore, the data integration execution unit 12 is also used for:
[0131] Based on the business-side confidence priority, the multiple basic tag data units are cross-polled to perform multi-round field conflict verification and calibration for duplicate records of the same entity, resulting in multiple conflict resolution data units and multiple redundant copies. The multiple redundant copies are mapped and backed up to the local storage of the multiple source business units. Multiple business primary key identifiers are set for the multiple conflict resolution data units, and resource pool existence judgment is performed based on the multiple business primary key identifiers. Based on the judgment results, the fields of the multiple conflict resolution data units are adaptively written to build the data consistency of the basic resource pool.
[0132] Furthermore, the data integration execution unit 12 is also used for:
[0133] If the first business primary key identifier of the first conflict resolution data unit exists in the historical resource pool, then the updatable field is retrieved from the first conflict resolution data unit according to the field mapping table to perform incremental update of the basic resource pool; if the first business primary key identifier of the first conflict resolution data unit does not exist in the historical resource pool, then the data record of the first conflict resolution data unit is completely inserted into the basic resource pool.
[0134] Furthermore, the report generation and execution unit 14 is also used for:
[0135] The user ID is used to query the permission configuration table to obtain viewing permissions. These viewing permissions include a field-level access permission list and a business domain access permission list. The field-level access permission list includes the range of accessible fields and sensitive data restrictions. The business domain access permission list is used to perform record-level filtering of the temporary dataset, and the field-level access permission list is used to perform field-level filtering of the temporary dataset to obtain a filtered dataset. After converting the filtered dataset into the permission filtering report based on a standard report format, a dynamic watermark is added to the permission filtering report according to the user ID.
[0136] Furthermore, the data integration execution unit 12 is also used for:
[0137] The system interactively obtains multiple data update time distributions of the multiple basic tag data units; locally retrieves multiple data source authority ratings of the multiple source business terminals; performs dynamic weighted calculation of confidence priority based on the multiple data update time distributions and multiple data source authority ratings, and outputs multiple business terminal confidence scores; sorts the multiple source business terminals in descending order according to the multiple business terminal confidence scores, and outputs the business terminal confidence priority.
[0138] Furthermore, the local data processing unit 11 is also used for:
[0139] The system locally invokes multiple data update cycles of the multiple source business terminals; based on the business type priority association attributes of the multiple source business terminals, it modifies the multiple data update cycles to obtain multiple dynamic scheduling cycles; using the multiple dynamic scheduling cycles as task trigger constraints, it drives the multiple lightweight verification rule engines to perform periodic data processing and uploading on the multiple source business terminals.
[0140] Furthermore, the data integration execution unit 12 is also used for:
[0141] Cross-node association retrieval is performed based on the business primary key identifier to obtain multiple historical multi-source data records of multiple sample business entities; multiple composite label combination rules are constructed by associating the basic tag numbers of the multiple historical multi-source data records based on predefined business logic rules; the multiple composite label combination rules are stored in the classification rule library to obtain the composite label classification rules; the basic resource pool is traversed based on the composite label classification rules, and dynamic tag identifiers are appended to data records that meet the tag association conditions, and the unified resource pool is output.
[0142] Example 3 provides a storage medium on which a computer program is stored, which, when executed by a processor, implements any step of Example 1.
[0143] In summary, any of the methods or steps described above can be stored as computer instructions or programs in various types of computer memory, and the computer instructions or programs can be recognized by various types of computer processors to implement any of the above methods or steps.
[0144] Based on the above specific embodiments of the present invention, any improvements and modifications made to the present invention by those skilled in the art without departing from the principle of the present invention shall fall within the patent protection scope of the present invention.
Claims
1. A method for automatically generating reports based on multi-source data integration, characterized in that: The method includes: By deploying multiple lightweight verification rule engines on multiple source business terminals, local preprocessing is performed on the real-time monitoring business data captured by the multiple source business terminals, and multiple standardized data streams are output. The data integration cloud receives and performs cross-business redundancy verification and integration of the multiple standardized data streams based on hierarchical tag rules, and outputs a unified resource pool, wherein each data record in the unified resource pool is marked with a composite tag. After parsing the user's form-based query to obtain real-time associated tag combinations, matching data is dynamically extracted from the unified resource pool based on the real-time associated tag combinations to generate a temporary dataset; After retrieving the viewing permissions based on the user ID, the temporary dataset is converted into a permission filtering report using the viewing permissions as the field filtering strategy. The method involves receiving data from the cloud and performing cross-business-end redundancy verification and integration of multiple standardized data streams based on hierarchical labeling rules, outputting a unified resource pool. The method includes: Based on the benchmark classification rules, the data integration cloud performs static labeling on the multiple standardized data streams to obtain multiple basic label data units; After performing cross-node redundancy check correction on the multiple basic tag data units, the data of the multiple basic tag data units is fused based on the primary key driven fusion mechanism to output a basic resource pool, wherein the basic resource pool stores multiple discrete data records, and each discrete data record is marked with a basic tag identifier. Based on cross-node correlation retrieval, extended data correlation is constructed, and composite label classification rules are built, wherein the baseline classification rules and composite label classification rules constitute the hierarchical label rules; Based on the composite label classification rules, dynamic label identifiers are added to the basic resource pool, and the unified resource pool is output. Each discrete data record in the unified resource pool is labeled with a composite label. The method further includes: Perform cross-node correlation retrieval based on business primary key identifiers to obtain multiple historical multi-source data records of multiple sample business entities; Based on the predefined business logic rules, associate the basic tag numbers of the multiple historical multi-source data records, and construct multiple composite tag combination rules; The multiple composite label combination rules are stored in the classification rule library to obtain the composite label classification rules; Based on the composite label classification rules, the basic resource pool is traversed, dynamic label identifiers are added to data records that meet the label association conditions, and the unified resource pool is output.
2. The method for automatically generating reports based on multi-source data integration as described in claim 1, characterized in that, After performing cross-node redundancy check correction on the multiple basic tag data units, data fusion of the multiple basic tag data units is performed based on the primary key-driven fusion mechanism to output a basic resource pool. The method includes: Based on the business-side confidence priority, the multiple basic tag data units are cross-polled to perform multi-round field conflict verification and calibration for duplicate records of the same entity, resulting in multiple conflict resolution data units and multiple redundant copies. The multiple redundant copies are mapped and backed up to the local storage of the multiple source business units. Multiple business primary key identifiers are set for the multiple conflict resolution data units, and the existence of the resource pool is determined based on the multiple business primary key identifiers; Based on the judgment results, the fields of the multiple conflict resolution data units are adaptively written to build the data consistency of the basic resource pool.
3. The method for automatically generating reports based on multi-source data integration as described in claim 2, characterized in that, The method further includes: If the first business primary key identifier of the first conflict resolution data unit exists in the historical resource pool, then the updatable field is retrieved from the first conflict resolution data unit according to the field mapping table to perform incremental update of the basic resource pool. If the first business primary key identifier of the first conflict resolution data unit does not exist in the historical resource pool, then the data record of the first conflict resolution data unit is completely inserted into the basic resource pool.
4. The method for automatically generating reports based on multi-source data integration as described in claim 1, characterized in that, After retrieving view permissions based on the user ID, the temporary dataset is converted into a permission filtering report using the view permissions as the field filtering strategy. The method includes: The user ID is used to query the permission configuration table to obtain the viewing permission. The viewing permission includes a list of field-level access permissions and a list of business domain access permissions. The list of field-level access permissions includes the range of accessible fields and restrictions on sensitive data. The temporary dataset is filtered at the record level using the business domain access permission list and at the field level using the field-level access permission list to obtain the filtered dataset. After converting the filtered dataset into the permission filtering report based on the standard report format, a dynamic watermark is added to the permission filtering report according to the user ID.
5. The method for automatically generating reports based on multi-source data integration as described in claim 2, characterized in that, The method further includes: Interactively obtain multiple data update time distributions of the multiple basic tag data units; Locally retrieve the authority ratings of multiple data sources from the multiple source business terminals; Based on the multiple data update time distributions and multiple data source authority ratings, a dynamic weighted calculation of confidence priorities is performed to output multiple business-side confidence scores; The multiple source business terminals are sorted in descending order based on their confidence scores, and the confidence priority of each business terminal is output.
6. The method for automatically generating reports based on multi-source data integration as described in claim 1, characterized in that, The method further includes: Locally invoke multiple data update cycles from the multiple source business terminals; Based on the business type priority association attributes of the multiple source business terminals, the multiple data update cycles are corrected to obtain multiple dynamic scheduling cycles; Using the multiple dynamic scheduling cycles as task trigger constraints, the multiple lightweight verification rule engines drive the multiple source business terminals to perform periodic data processing and uploading.
7. An automated report generation system based on multi-source data integration, characterized in that: The steps for implementing the method according to any one of claims 1 to 6 include: The local data processing unit is used to perform local preprocessing of the real-time monitoring business data captured by the multiple source business terminals through multiple lightweight verification rule engines deployed on multiple source business terminals, and output multiple standardized data streams. The data integration execution unit is used to integrate data received from the cloud and perform cross-business redundancy verification and integration of the multiple standardized data streams based on hierarchical tag rules, and output a unified resource pool, wherein each data record in the unified resource pool is marked with a composite tag. The data dynamic extraction unit is used to parse the user's input form-type query to obtain real-time related tag combinations, and then dynamically extract matching data from the unified resource pool based on the real-time related tag combinations to generate a temporary dataset. The report generation execution unit is used to retrieve the viewing permissions based on the user ID, and then convert the temporary dataset into a permission-filtered report using the viewing permissions as the field filtering strategy.
8. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the report automation generation method based on multi-source data integration as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Financial statement generation method and system based on cloud computing
CN120764505A