Intelligent machine account report fusion method and system based on large model
Through the intelligent fusion method of ledger report based on large-model, the problem of format heterogeneity and semantic ambiguity in cross-department ledger data fusion is solved, and efficient and accurate data merging and traceability are achieved, adapting to dynamic business scenarios.
Patent Information
- Application Number
- CN202510528276.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-01
AI Technical Summary
It is difficult to efficiently integrate ledger data across departments and business systems. The existing technology has problems such as format heterogeneity, semantic ambiguity, and statistical caliber differences, resulting in the risk of redundant labor and data errors, and lack of automated end-to-end solutions.
The intelligent fusion method of ledger reports based on large models is adopted, including multi-modal data collection, dynamic Prompt construction and big model analysis, unified ledger generation, blood relationship recording and visualization, audit management and data traceability. Through large-model-driven semantic analysis and multi-dimensional similarity calculation, semantic similarity fields across reports are automatically merged.
It significantly improves data processing efficiency, reduces manual verification and rule configuration time, improves field merging accuracy and data credibility, and supports dynamic adjustments to adapt to business changes.
Smart Images

Figure CN120409439A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and specifically provides an intelligent fusion method and system for ledger reports based on large models. Background Art
[0002] In current grass-roots governance of government affairs, the need for data integration across departments and business systems is becoming increasingly urgent. However, multi-source reports have problems such as format heterogeneity, semantic ambiguity, and differences in statistical calibers, resulting in difficult efficient fusion of ledger data. There are a large number of common data in the ledgers required to be reported by superior units. However, due to subtle differences in the field formats, classification levels, etc. of the filling templates, grass-roots departments are forced to re-enter or manually adjust the data, resulting in redundant labor and risks of data errors. Traditional methods mainly rely on manual alignment of fields and writing conversion rules, with defects such as low efficiency, high cost, and insufficient standardization. There are differences in the naming rules, data units, and statistical logics of the same business entity among different departments. Traditional rule engines are difficult to cover dynamically changing business scenarios, resulting in low credibility and poor traceability of the fusion results.
[0003] In the prior art, some solutions attempt to achieve field alignment through keyword matching or simple clustering, but ignore the semantic associations in the context and cannot handle synonyms, abbreviations, and specific descriptions of business scenarios. For example, "the number of residents" and "the number of permanent residents" may not be directly merged due to different statistical calibers and require manual judgment. In addition, there is a lack of a complete blood relationship record and version management mechanism after the ledger is generated, making it difficult to meet the compliance requirements of data governance. Although large model technology has been applied to some data cleaning scenarios, its application in cross-modal data fusion, dynamic business adaptation, and conflict resolution is still immature. Especially in the government affairs field where industry norms and flexibility need to be considered, the prior art has not formed an end-to-end automated solution. Summary of the Invention
[0004] The purpose of the present invention is to provide an intelligent fusion method and system for ledger reports based on large models to solve the problems raised in the above background art.
[0005] To achieve the above purpose, the present invention provides the following technical solution: An intelligent fusion method for ledger reports based on large models, including the following steps:
[0006] Multi-modal data collection and preprocessing;
[0007] Dynamic Prompt construction and large model analysis;
[0008] Unified ledger generation and recommendation;
[0009] Blood relationship record and visualization;
[0010] Audit management and data traceability.
[0011] Preferably, the specific operations for multi-modal data collection and preprocessing include: collecting EXCEL and CSV report files uploaded by users, supporting OCR recognition of scanned documents, using the Tesseract 5.0 engine combined with an image enhancement algorithm to improve recognition accuracy; automatically parsing the file content, extracting the table headers, data rows, and metadata, and outputting structured tabular data; performing standardized preprocessing on data items, including duplicate removal, null value filling, and format verification; extracting the description content of each data item as the input for large model analysis.
[0012] Preferably, the specific operations for dynamic Prompt construction and large model analysis include: presetting Prompt templates to guide the large model to understand the business scenario, defining field parsing instructions, embedding government industry specifications, and guiding conflict resolution strategies; using a pre-trained large model as the base, which is based on the Transformer architecture and fine-tuned for a specific government grass-roots governance context to adapt to the ledger integration scenario, and the training data includes historical report samples, manually annotated merging rules, and government domain corpora; dynamically optimizing the Prompt template according to user feedback and business changes; achieving semantic alignment across reports through multi-dimensional analysis, and the specific dimensions include word meaning similarity, sentence meaning similarity, type matching degree, distribution similarity, and scenario association degree, performing multi-dimensional analysis on the file name, data item description, and content, and extracting keywords and context semantics; calculating the similarity of data items through vector representation and generating merging suggestions in combination with confidence scores.
[0013] Preferably, the specific operations for unified ledger generation and recommendation include: based on the output of the large model, recommending ledger names and ledger fields, including the name, data type, and unit of the fields, and supporting manual modification; automatically adding units and precisions according to business specifications, marking conflicts for data items with different statistical calibers and prompting administrators to intervene.
[0014] Preferably, the specific operations for lineage recording and visualization include: using a graph database to store the integrated lineage relationship, designing a compression storage algorithm to record the original fields, integrated fields, and conversion logic; enhancing visualization interaction: topological graph to display the field-level integration path; traceability tree to show the complete derivation chain from the original field to the integrated field.
[0015] The specific operations for audit management and data traceability include: administrators auditing the generated standard ledger, and the audit content includes the rationality of field naming, the accuracy of data mapping, the consistency of units, and semantic compliance checking by matching the business term library; after passing the audit, the ledger template is stored in the electronic ledger library for business systems to call; recording the ledger version history, supporting rollback to any version, and tracking the evolution of the ledger version with a timeline; security enhancement technology, adopting the SM4 national secret algorithm, with permission control based on the RBAC model, and auditing logs recording the full life cycle operations.
[0016] A system for the intelligent fusion method of ledger reports based on large models according to claim 5, comprising:
[0017] A data collection module, responsible for multi-modal data collection and preprocessing;
[0018] A large model analysis engine module, used for dynamic Prompt construction and large model analysis;
[0019] A ledger generation module, used for unified ledger generation and recommendation;
[0020] A lineage management module, used for lineage recording and visualization;
[0021] An audit management module, used for audit management and data traceability.
[0022] Preferably, the data collection module collects EXCEL and CSV report files uploaded by users, supports OCR recognition of scanned documents, uses the Tesseract 5.0 engine combined with an image enhancement algorithm to improve the recognition accuracy; automatically parses the file content, extracts the table headers, data rows and metadata, and outputs structured tabular data; performs standardized preprocessing on data items, including duplicate removal, null value filling, and format verification; extracts the description content of each data item as the input for large model analysis.
[0023] Preferably, the large model analysis engine module presets a Prompt template to guide the large model to understand the business scenario, defines field parsing instructions, embeds government industry specifications, and guides conflict resolution strategies; uses a pre-trained large model as the base, and this large model is based on the Transformer architecture and is fine-tuned for a specific government grass-roots governance context to adapt to the ledger fusion scenario. The training data includes historical report samples, manually marked merging rules, and government domain corpora; dynamically optimizes the Prompt template according to user feedback and business changes; realizes semantic alignment across reports through multi-dimensional analysis. The specific dimensions include word sense similarity, sentence sense similarity, type matching degree, distribution similarity, and scenario association degree. Perform multi-dimensional analysis on the file name, data item description, and content, extract keywords and context semantics; calculate the similarity of data items through vectorized representation, and generate merging suggestions in combination with confidence scores.
[0024] Preferably, the ledger generation module recommends ledger names and ledger fields based on the output of the large model, including the name, data type, and unit of the fields, and supports manual modification; automatically adds units and precisions according to business specifications, marks conflicts for data items with different statistical calibers and prompts the administrator to intervene.
[0025] Preferably, the blood relationship management module uses a graph database to store and integrate blood relationships, designs a compression storage algorithm, and records the original fields, integrated fields, and conversion logic; visual enhancement interaction: topological graph to display the field-level integration path; traceability tree to show the complete derivation chain from the original field to the integrated field;
[0026] The audit management module allows administrators to audit the generated standard ledger. The audit content includes the rationality of field naming, the accuracy of data mapping, the consistency of units, and semantic compliance checks by matching the business term library. After passing the audit, the ledger template is stored in the electronic ledger library for business systems to call; record the ledger version history, support rolling back to any version, and track the evolution of the ledger version with a timeline; security enhancement technology, adopt the SM4 national cryptography algorithm, the permission control is based on the RBAC model, and the audit log records the full life cycle operations.
[0027] Compared with the prior art, the beneficial effects of the present invention are:
[0028] The intelligent integration method and system for ledger reports based on a large model proposed by the present invention automatically merge semantically similar fields across reports through semantic analysis and multi-dimensional similarity calculation driven by the large model, reducing the time for manual verification and rule configuration, and significantly improving the data processing efficiency.
[0029] By comprehensively scoring in five dimensions of word meaning, sentence meaning, type, distribution, and scenario, it solves the problem of incorrect merging caused by the traditional method relying on single keyword matching, introduces a dynamic optimization mechanism, supports dynamically adjusting the Prompt template and large model parameters according to user feedback, continuously improves the semantic understanding accuracy, and improves the accuracy of field merging. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is the flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] In order to clearly and completely describe the objectives, technical solutions of the present invention, and make the advantages more clear, the following further details the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are part of the embodiments of the present invention, rather than all of the embodiments, and are only used to explain the embodiments of the present invention, not to limit the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0032] Embodiment 1, please refer to Figure 1 , the present invention provides a technical solution: an intelligent integration method for ledger reports based on a large model, including the following steps:
[0033] Step 1: Multimodal data collection and preprocessing
[0034] Collects Excel and CSV report files uploaded by users, supports OCR recognition of scanned documents (such as PDF scanned reports), and uses the Tesseract 5.0 engine combined with image enhancement algorithms to improve recognition accuracy; automatically parses file content, extracts table headers, data rows, and metadata (such as file name, creation time, etc.), and outputs structured table data;
[0035] Perform standardization preprocessing on data items, including deduplication, null value filling, and format verification (such as standardizing the date format to YYYY-MM-DD);
[0036] Extract the descriptive content of each data item (such as field comments and business meaning description) as input for large model analysis.
[0037] Step 2: Dynamic Prompt Construction and Large Model Analysis
[0038] The preset prompt template is used to guide the large model to understand business scenarios, define field parsing instructions, embed government industry standards, and guide conflict resolution strategies;
[0039] The pre-trained large model is used as the foundation. This large model is based on the Transformer architecture and is fine-tuned for the specific context of grassroots government governance to adapt to the ledger fusion scenario. The training data includes historical report samples, manually annotated merging rules, and government domain corpus.
[0040] Dynamically optimize Prompt templates based on user feedback and business changes;
[0041] Achieve semantic alignment across reports through multi-dimensional analysis, including word meaning similarity, sentence meaning similarity, type matching, distribution similarity, and scenario relevance. Multi-dimensional analysis is performed on file names, data item descriptions, and content to extract keywords and contextual semantics.
[0042] Calculate the similarity of data items through vectorized representation and generate merge suggestions based on confidence scores;
[0043] Step 3: Unified ledger generation and recommendation
[0044] Based on the large model output, the ledger name and ledger field are recommended, including the field name, data type, unit, etc., and manual modification is supported;
[0045] Automatically add units and precision according to business specifications, mark conflicts in data items with statistical caliber differences (such as amounts including tax and amounts excluding tax), and prompt administrators to intervene.
[0046] Step 4: Bloodline Recording and Visualization
[0047] Use a graph database to store the integrated lineage relationship, design a compression storage algorithm, and record the original fields, integrated fields, and conversion logic;
[0048] Visualization enhanced interaction:
[0049] Topological graph to display the field-level integration path;
[0050] Traceability tree to show the complete derivation chain from the original fields to the integrated fields.
[0051] Step 5: Audit management and data traceability
[0052] The administrator audits the generated standard ledger, and the audit content includes the rationality of field naming, the accuracy of data mapping, the consistency of units, etc., and performs semantic compliance checks by matching the business term library;
[0053] After the audit is passed, the ledger template is stored in the electronic ledger library for the business system to call;
[0054] Record the ledger version history, support rolling back to any version, and track the evolution of the ledger version with a timeline;
[0055] Security enhancement technology, adopt the SM4 national encryption algorithm, the permission control is based on the RBAC model, and the audit log records the full life cycle operations.
[0056] In addition to supporting the integration of user-reported reports, it also supports the integration of existing electronic ledgers. When the user creates a ledger, they input information such as data item names and business descriptions, and call the large model to recommend an integration solution.
[0057] The specific steps are as follows:
[0058] The user inputs the field information of the new ledger;
[0059] The system retrieves semantically similar fields from the ledger library and recommends a merging solution;
[0060] Automatically generate a cross-business ledger after integration.
[0061] Example 2, based on Example 1, proposes an intelligent integration system for ledger reports based on a large model, including:
[0062] The data acquisition module is responsible for multi-modal data acquisition and preprocessing, supports users to upload Excel, CSV files and scanned documents (such as PDF scanned reports), the system automatically parses the content, extracts the business description information of the data items and performs data cleaning and standardization.
[0063] The large model analysis engine module incorporates one or more pre-trained large models, a preset Prompt template library to guide the large model to understand the business scenario, fine-tunes according to the ledger report by integrating the business scenario, and optimizes the model based on historical report samples, manual annotation rules, and government domain corpora.
[0064] It also includes a dynamic feedback and learning mechanism that dynamically optimizes the Prompt content according to user feedback and business changes, continuously optimizing the performance and accuracy of the large model to improve the integration accuracy and user experience.
[0065] The semantic similarity calculation module achieves semantic alignment across reports through multi-dimensional analysis, calculates the comprehensive score by weighting, and generates field merging suggestions and confidence scores.
[0066] The ledger generation module generates a standardized ledger template based on the output of the large model, recommends field names and ledger names, and automatically executes conflict resolution strategies such as unit conversion and precision retention.
[0067] The lineage management module uses a graph database to record the mapping relationship and conversion logic between the original fields and the merged fields, and provides an interactive interface for users to view the integration path and confidence scores.
[0068] The audit management module is used to audit and release the ledger, manage versions, and trace operations, auditing the rationality of ledger field naming, the accuracy of data mapping, and the consistency of units to ensure data compliance and traceability.
[0069] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An intelligent fusion method for ledger reports based on large models, characterized in that: It includes the following steps: Multi-modal data collection and preprocessing; Dynamic Prompt construction and large model analysis; Unified ledger generation and recommendation; Lineage record and visualization; Audit management and data traceability.
2. The intelligent fusion method of ledger reports based on a large model according to claim 1, characterized in that: The specific operations of multi-modal data collection and preprocessing include: collecting EXCEL and CSV report files uploaded by users, supporting OCR recognition of scanned documents, using the Tesseract 5.0 engine combined with image enhancement algorithms to improve recognition accuracy; automatically parsing the file content, extracting the table headers, data rows and metadata, and outputting structured tabular data; performing standardized preprocessing on data items, including duplicate removal, null value filling, and format verification; extracting the description content of each data item as the input for large model analysis.
3. The intelligent fusion method of ledger reports based on a large model according to claim 2, characterized in that: The specific operations of dynamic Prompt construction and large model analysis include: presetting Prompt templates to guide the large model to understand the business scenario, defining field parsing instructions, embedding government industry specifications, and guiding conflict resolution strategies; using a pre-trained large model as the base, which is based on the Transformer architecture and is fine-tuned for specific government grass-roots governance contexts to adapt to the ledger integration scenario, and the training data includes historical report samples, manually annotated merging rules, and government domain corpora; dynamically optimizing the Prompt template according to user feedback and business changes; achieving semantic alignment across reports through multi-dimensional analysis, and the specific dimensions include word similarity, sentence similarity, type matching degree, distribution similarity, and scenario association degree, performing multi-dimensional analysis on the file name, data item description and content, and extracting keywords and context semantics; calculating the similarity of data items through vector representation and generating merging suggestions in combination with confidence scores.
4. The intelligent fusion method of ledger reports based on large models according to claim 3, characterized in that: The specific operations of unified ledger generation and recommendation include: recommending ledger names and ledger fields based on the large model output, including the name, data type, and unit of the fields, and supporting manual modification; automatically adding units and precisions according to business specifications, marking conflicts for data items with different statistical calibers and prompting the administrator to intervene.
5. The intelligent fusion method of ledger statements based on a large model according to claim 4, characterized in that: The specific operations of lineage record and visualization include: using a graph database to store the integrated lineage relationship, designing a compression storage algorithm to record the original fields, integrated fields and conversion logic; visual enhancement interaction: topology graph to display the field-level integration path; traceability tree to show the complete derivation chain from the original field to the integrated field. The specific operations of audit management and data traceability include: the administrator audits the generated standard ledger, and the audit content includes the rationality of field naming, the accuracy of data mapping, the consistency of units, and semantic compliance check by matching the business term library; after the audit is passed, the ledger template is stored in the electronic ledger library for the business system to call; recording the ledger version history, supporting rollback to any version, and tracking the evolution of the ledger version with a timeline; security enhancement technology, adopting the SM4 national encryption algorithm, and the permission control is based on the RBAC model, and the audit log records the full life cycle operations.
6. A system for the intelligent fusion method of ledger reports based on a large model according to claim 5, characterized in that: It includes: A data collection module responsible for multi-modal data collection and preprocessing; A large model analysis engine module for dynamic Prompt construction and large model analysis; A ledger generation module for unified ledger generation and recommendation; Blood relationship management module, used for blood relationship recording and visualization; Audit management module, used for audit management and data traceability.
7. A system according to claim 6, wherein: Data collection module, which collects EXCEL and CSV report files uploaded by users, supports OCR recognition of scanned documents, adopts the Tesseract 5.0 engine combined with image enhancement algorithms to improve recognition accuracy; automatically parses the file content, extracts the table header, data rows and metadata, and outputs structured tabular data; performs standardized preprocessing on data items, including duplicate removal, null value filling, and format verification; extracts the description content of each data item as the input for large model analysis.
8. A system according to claim 7, characterized in that: Large model analysis engine module, which presets Prompt templates to guide the large model to understand the business scenario, defines field parsing instructions, embeds government industry specifications, and guides conflict resolution strategies; uses a pre-trained large model as the base, which is based on the Transformer architecture and is fine-tuned for specific government grass-roots governance contexts to adapt to the ledger integration scenario. The training data includes historical report samples, manually annotated merging rules, and government domain corpora; dynamically optimizes the Prompt template according to user feedback and business changes; realizes semantic alignment across reports through multi-dimensional analysis. The specific dimensions include word similarity, sentence similarity, type matching degree, distribution similarity, and scenario association degree. Perform multi-dimensional analysis on the file name, data item description, and content, extract keywords and context semantics; calculate the similarity of data items through vector representation, and generate merging suggestions combined with confidence scores.
9. A system according to claim 8, wherein: Ledger generation module, based on the output of the large model, recommends ledger names and ledger fields, including the name, data type, and unit of the fields, and supports manual modification; automatically adds units and precisions according to business specifications, marks conflicts for data items with different statistical calibers, and prompts the administrator to intervene.
10. A system according to claim 9, wherein: Blood relationship management module, which uses a graph database to store the integrated blood relationship, designs a compression storage algorithm, and records the original fields, integrated fields, and conversion logic; Visualization enhanced interaction: topological graph, which shows the field-level integration path; Traceability tree, which shows the complete derivation chain from the original field to the integrated field; Audit management module, where the administrator audits the generated standard ledger. The audit content includes the rationality of field naming, the accuracy of data mapping, the consistency of units, and semantic compliance checks by matching the business term library; after the audit is passed, the ledger template is stored in the electronic ledger library for business systems to call; records the ledger version history, supports rolling back to any version, and tracks the evolution of the ledger version with a timeline; security enhancement technology, adopts the SM4 national secret algorithm, the permission control is based on the RBAC model, and the audit log records the full life cycle operations.
Citation Information
Cited By
Data integration and fusion method and system based on large model
CN120596563A
Prompt prompt template automatic construction method for intelligent reimbursement system
CN120832371A
Automatic data management method and system based on large model
CN121301805A
Multi-source data fusion analysis method and system based on large water affair model
CN122414935A