Accounting archive management method and system based on blockchain technology and large language model
By combining blockchain technology with large language models, the automated and secure management of accounting archive data has been achieved, solving the problems of insufficient data association and storage security, as well as low efficiency in tag management and retrieval in traditional accounting archive management, and realizing efficient data processing and secure storage.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-15
- Publication Date
- 2026-04-07
AI Technical Summary
Traditional accounting record management suffers from insufficient data association and storage security, and low efficiency in tag management and retrieval, failing to meet the needs for efficient processing and secure storage of massive amounts of data.
An accounting record management method based on blockchain technology and large language model is adopted. Semantic features are extracted through a semantic understanding and reasoning model of invoices. Combined with an accounting reconciliation relationship identification algorithm, the reconciliation relationship between invoices and accounting entries is automatically matched to generate multi-level classification labels. A blockchain transaction sequence encapsulation algorithm is used for distributed storage to achieve data security and traceability.
It has enabled automated and secure management of accounting archive data, solved the risks of data tampering and loss, improved retrieval efficiency, and met the needs of efficient management of massive accounting archives.
Smart Images

Figure CN121542223B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of accounting file management, and in particular to an accounting file management method and system based on blockchain technology and a large language model. BACKGROUND
[0002] With the acceleration of digital transformation of the accounting industry, the number of accounting files is growing explosively, including various data such as original invoices, accounting entries, and transaction records. Traditional management methods rely on manual classification and storage, which cannot meet the efficient processing needs of massive data. At the same time, accounting file management needs to meet the requirements of data authenticity, traceability, and long-term storage. The breakthrough of large language models in semantic understanding and the decentralized and tamper-proof features of blockchain technology provide a technical basis for solving the data processing and secure storage problems in accounting file management, promoting the industry to explore new management solutions that integrate the two technologies to achieve intelligent and secure management of the entire process of accounting file data collection and storage retrieval.
[0003] The existing accounting file management technology has two significant shortcomings: on the one hand, data correlation and storage security are insufficient. Traditional technologies mostly use centralized storage architecture, lack automated identification capabilities for the cross-checking relationship between invoice data and accounting entries, and rely on manual verification, which can easily lead to correlation errors. Centralized storage also faces the risk of data tampering and loss, and cannot guarantee the long-term integrity and traceability of file data. On the other hand, label management and retrieval efficiency are low. Existing label generation is mostly based on fixed rules, which cannot generate accurate multi-level labels by combining file formation time, business type, storage period, and other multi-dimensional parameters, resulting in low correlation between labels and file data. Subsequent retrieval requires a lot of time to filter, which cannot meet the actual needs of fast positioning and efficient utilization of accounting files. SUMMARY
[0004] In order to overcome the shortcomings and deficiencies of the existing technology, the present application provides an accounting file management method and system based on blockchain technology and a large language model.
[0005] The technical scheme adopted by the present application is an accounting archive management method based on blockchain technology and a large language model, comprising the following steps: S1, performing semantic feature extraction on original bill data associated with the accounting archive through a bill semantic understanding reasoning model, wherein the model calls a preset semantic dictionary library in the accounting field to perform multi-dimensional semantic mapping on the text field, numerical field and format field in the bill, and generates a semantic feature vector set comprising bill type identification, amount association identification and transaction subject identification; S2, calculating the reconciliation relationship between different bill data in the semantic feature vector set generated in S1 based on an accounting reconciliation relationship identification algorithm, and traversing and matching the correspondence between the bill semantic feature vector and the accounting entry data through the accounting subject association rule library built in the algorithm, and outputting a reconciliation relationship matching result set; S3, converting the reconciliation relationship matching result set output in S2 into a transaction sequence by using a blockchain transaction sequence packaging algorithm, arranging the accounting archive data that passes the matching in order according to the timestamp and transaction number, and generating a transaction data unit conforming to the block structure of the blockchain; S4, generating labels for the blockchain transaction data unit generated in S3 through an archive label intelligent matching engine, wherein the engine calls an accounting archive classification standard library, combines the archive formation time, business type and storage period parameters, and generates a label set comprising multi-level classification labels; S5, associating and binding the label set generated in S4 with the blockchain transaction data unit to form a labeled accounting archive blockchain data block, calculating the hash value of the data block, and generating a unique identifier of the data block; S6, storing the accounting archive blockchain data block with the unique identifier generated in S5 to a distributed blockchain node based on the storage parameters and access permission parameters of the accounting archive management, and updating the archive index library of the blockchain node to perform distributed storage and index updating of the accounting archive data.
[0006] Further, the expression of the bill semantic understanding reasoning model is: wherein, is a bill semantic feature vector set, is the number of bill field types, is the semantic weight coefficient of the first class bill field, is the semantic mapping matrix of the first class bill field, is the text feature matrix of the first class bill field, is the numerical feature matrix of the first class bill field, is the format feature matrix of the first class bill field, is a feature matrix fusion operator, is an accounting field dictionary adjustment coefficient, is a semantic vector matrix of the accounting field semantic dictionary library, This represents the number of semantic reasoning layers.
[0007] Furthermore, the expression for the accounting reconciliation relationship identification algorithm is as follows: ,in, For the set of matching results of reconciliation relationships, The number of types of reconciliation relationships. For the first Matching weights for similar logical relationships For the first A rule matrix for reconciliation-like relationships. For matrix dot product operation, For the related adjustment factor of accounting items, This is the rule vector matrix of the accounting subject association rule base. Number of accounting entry types For the first Class-based logical relationships and the first The correlation coefficient of similar accounting entries, This is the feature vector of accounting entry data.
[0008] Furthermore, the expression for the blockchain transaction sequence encapsulation algorithm is as follows: ,in, For blockchain transaction data units, These are the conversion coefficients for the transaction sequence. A timestamp vector matrix, For matrix tensor product operations, For transaction number vectors, For ordered vectors, This is the block structure adjustment coefficient. For hash calculation matrix, This is the parameter matrix for the blockchain block structure.
[0009] Furthermore, the tag generation expression of the intelligent matching engine for file tags is: ,in, For a set of tags, For the number of tag levels, For the first The generation weight of hierarchical labels For the first The classification matrix of hierarchical labels, For the classification vector of the accounting record classification standard library, Adjust the coefficient for tag association. For the time vector of archive formation, For business type vectors, For the storage period parameter vector, This is a label dictionary matrix.
[0010] Furthermore, the unique identifier generation expression for the accounting record blockchain data block is: ,in, A unique identifier for the data block. Adjustment factors are calculated for hash calculation. Calculate the matrix for the hash values. Adjust the coefficients for the node index. A vector for numbering blockchain nodes. This is a vector of archive index parameters.
[0011] Further, S3 includes the following sub-steps: S31, calling the initialization module of the blockchain transaction sequence encapsulation algorithm, reading the blockchain block capacity parameters and transaction sequence sorting rule parameters preset by the accounting file management system, and importing the parameters into the parameter configuration layer of the algorithm; S32, cleaning the data of the reconciliation matching result set output by S2, removing abnormal data records in the matching result set, and initially sorting the valid data records according to the order of transaction occurrence time; S33, comparing the initially sorted valid data records with the blockchain block structure parameters to determine whether the number of data records meets the block capacity requirements. If it exceeds the requirements, the block is split; if it is insufficient, it is temporarily stored in the data pool to be supplemented; S34, assigning transaction numbers to the data records that meet the block capacity requirements, generating an ordered transaction data queue according to the assigned transaction number and timestamp, and converting it into blockchain transaction data units.
[0012] Further, S4 includes the following sub-steps: S41, activating the intelligent matching engine for archive tags, loading industry classification standards, business classification standards, and retention period classification standards from the accounting archive classification standard library, and establishing a standard parameter system for tag generation; S42, extracting the labeling information from the blockchain transaction data unit generated in S3, including the archive formation timestamp, business type code, and related accounting subject code, and converting the labeling information into feature parameters that the engine can recognize; S43, comparing the feature parameters with the standard parameter system in multiple dimensions to determine the primary classification label, secondary classification label, and tertiary classification label corresponding to the archive, forming an initial label group; S44, performing redundancy verification on the initial label group, deleting duplicate and invalid labels, and supplementing missing related labels according to the label association rules for accounting archive management to generate the final label set.
[0013] Further, step S5 includes the following sub-steps: S51, read the tag set generated in S4 and the blockchain transaction data unit generated in S3, establish an association mapping table between the two, and clarify the correspondence between each tag and the specific fields in the transaction data unit; S52, embed the association mapping table into the extended fields of the blockchain transaction data unit, and perform format conversion on the embedded transaction data unit to make it conform to the format requirements of the blockchain data block; S53, call the hash calculation module to perform two hash calculations on the format-converted transaction data unit. The first calculation obtains the primary hash value of the data unit, and the second calculation calculates the final hash value based on the primary hash value and the feature value of the tag set; S54, use the final hash value as the unique identifier of the data block and bind it with the format-converted transaction data unit to form a tagged accounting archive blockchain data block.
[0014] This accounting record management system, based on blockchain technology and a large language model, is applied to an accounting record management method based on blockchain technology and a large language model. It includes: a document semantic feature extraction unit, connected to the original accounting record data acquisition device, which extracts semantic features from the acquired original document data using a document semantic understanding and reasoning model, and outputs a semantic feature vector set to an accounting reconciliation relationship matching unit; an accounting reconciliation relationship matching unit, connected to both the document semantic feature extraction unit and the accounting subject rule database, which calls an accounting reconciliation relationship recognition algorithm to calculate the reconciliation relationships of the semantic feature vector set, and outputs a reconciliation relationship matching result set to a blockchain transaction sequence conversion unit; and a blockchain transaction sequence conversion unit, connected to both the accounting reconciliation relationship matching unit and the blockchain parameter configuration module, which uses a blockchain transaction sequence encapsulation algorithm to convert the matching result set into blockchain transaction data. The system consists of several modules: a data block generation unit and a data block distribution unit. The data block generation unit is connected to both the blockchain transaction sequence conversion unit and the accounting archive classification standard library. It generates a tag set using an intelligent tag matching engine, associates the tag set with blockchain transaction data units, and outputs it to the blockchain data block generation unit. The blockchain data block generation unit, connected to both the intelligent tag generation unit and the hash calculation module, performs hash calculations on the associated tag transaction data units to generate unique identifiers, forming accounting archive blockchain data blocks which are then output to the distributed storage unit. The distributed storage unit, connected to both the blockchain data block generation unit and the blockchain node network, stores the accounting archive blockchain data blocks on different blockchain nodes based on the storage parameters and access permission parameters for accounting archive management, and updates the archive index database of the nodes, enabling distributed management of accounting archive data.
[0015] Beneficial Effects: This invention proposes an accounting archive management method and system based on blockchain technology and a large language model. It extracts multi-dimensional semantic features from original invoice data using a semantic understanding and reasoning model, and automatically matches the reconciliation relationships between invoices and accounting entries using an accounting reconciliation relationship recognition algorithm, replacing manual verification and avoiding association errors. Simultaneously, a blockchain transaction sequence encapsulation algorithm converts matched archive data into blockchain transaction data units, and decentralized management is achieved through distributed storage, completely resolving the risks of data tampering and loss under traditional centralized storage, ensuring the long-term integrity and traceability of archive data. Through an intelligent matching engine for archive tags, it calls upon an accounting archive classification standard library, combining multiple dimensions such as archive formation time, business type, and retention period to generate accurate multi-level tags, solving the problem of low relevance in existing fixed-rule tag generation. Subsequently, archives can be quickly located based on tags, significantly improving retrieval efficiency. Furthermore, the entire process of technical operation automates the processing of accounting archives from semantic extraction and reconciliation matching to blockchain storage and tag retrieval, reducing manual intervention and meeting the needs of efficient management of massive accounting archives, achieving the goal of intelligent and secure management throughout the entire process. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the overall steps of the method of the present invention;
[0017] Figure 2 This is a flowchart of method step S3 of the present invention;
[0018] Figure 3 This is a flowchart of method step S4 of the present invention;
[0019] Figure 4 This is a flowchart of step S5 of the method of the present invention;
[0020] Figure 5 This is a diagram showing the system unit composition of the present invention. Detailed Implementation
[0021] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. The application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0022] like Figure 1 As shown, the accounting record management method based on blockchain technology and a large language model includes the following steps:
[0023] S1. The semantic features of the original invoice data associated with the accounting archives are extracted through the invoice semantic understanding reasoning model. The model calls the preset accounting domain semantic dictionary library to perform multi-dimensional semantic mapping on the text fields, numeric fields and format fields in the invoice, and generate a set of semantic feature vectors including invoice type identifier, amount association identifier and transaction entity identifier.
[0024] Specifically, when implementing step S1, the semantic understanding and reasoning model for bills is first used. This model is preloaded with a semantic dictionary library for the accounting field, which includes 1,200 core vocabulary words in the accounting field and 300 types of bill features. The semantic reasoning layer of the model is set to 5 layers, and each layer is set with different semantic weight coefficients, of which the weight coefficient for text fields is 0.4, the weight coefficient for numeric fields is 0.35, and the weight coefficient for format fields is 0.25. Subsequently, the original invoice data associated with the accounting records was imported into the model. The data included 20 common types of invoices such as invoices, receipts, and bank statements. The model performed word segmentation on the text fields (such as the names of the transacting parties and business descriptions) of each type of invoice, numerical mapping on the numeric fields (such as amount and quantity), and structural feature extraction on the format fields (such as invoice numbering rules and signature positions). Through a multi-dimensional semantic mapping algorithm, the features of the three types of fields were fused to generate a semantic feature vector set, including invoice type identifiers (such as "VAT special invoice" and "expense reimbursement form"), amount association identifiers (such as "associated with accounts receivable" and "associated with management expense"), and transaction entity identifiers (such as "Party A's enterprise code XXX" and "Party B's enterprise code XXX"). Each vector in the vector set was set to 256 dimensions to ensure that it could completely represent the semantic information of a single invoice, providing an accurate data foundation for subsequent reconciliation relationship identification.
[0025] S2 calculates the reconciliation relationship between different invoice data in the semantic feature vector set generated by S1 based on the accounting reconciliation relationship recognition algorithm. Through the built-in accounting subject association rule library of the algorithm, the correspondence between the semantic feature vector of the invoice and the accounting entry data is traversed and matched, and the reconciliation relationship matching result set is output.
[0026] Specifically, in step S2, processing is carried out based on the accounting reconciliation relationship identification algorithm. The algorithm has a built-in rule base containing 150 accounting subject association rules, including subject matching rules corresponding to the six major accounting elements such as assets, liabilities, and owner's equity. The algorithm sets the number of reconciliation relationship types to 8, including account balance reconciliation, report item reconciliation, and bill and journal entry reconciliation. The matching weight of each type of reconciliation relationship is set according to its importance, with account balance reconciliation having a weight of 0.2 and bill and journal entry reconciliation having a weight of 0.3. First, the semantic feature vector set generated in S1 is input into the algorithm. The algorithm calls the rules in the rule base to traverse and match the semantic features of each bill in the vector set with the accounting journal entry data. The number of traversals is set to 3 to ensure the comprehensiveness of the matching. During each traversal, the algorithm calculates the corresponding similarity between the bill semantic feature vector and the accounting journal entry data. The similarity threshold is set to 0.85. When the similarity is higher than the threshold, it is judged as a successful match; when it is lower than the threshold, it is marked as pending verification. After the matching is completed, the algorithm outputs a set of matching results, which includes the correspondence between the matched invoices and journal entries, the matching similarity value, the invoice identifiers that failed to match and the reasons for failure, providing compliant data for subsequent blockchain transaction sequence conversion.
[0027] S3 uses a blockchain transaction sequence encapsulation algorithm to convert the matching result set of the reconciliation relationship output by S2 into a transaction sequence, and arranges the matched accounting file data in an orderly manner according to timestamp and transaction number to generate transaction data units that conform to the blockchain block structure.
[0028] Specifically, in step S3, a blockchain transaction sequence encapsulation algorithm is used for conversion. The algorithm first reads the blockchain block capacity parameters (each block can store a maximum of 1000 transaction data records) and transaction sequence sorting rule parameters (sorted in ascending order by transaction timestamp, with timestamp precision accurate to milliseconds) preset by the accounting file management system, and imports these parameters into the algorithm's parameter configuration layer. Next, the data cleaning of the reconciliation matching result set output by S2 is performed, removing abnormal data records with a similarity lower than 0.85, and deleting redundant data that has been submitted repeatedly. After cleaning, the valid data records are initially sorted according to the order of transaction occurrence. During the sorting process, if two data records have the same timestamp, they are sorted by transaction number from smallest to largest. The initially sorted valid data records are then compared with the blockchain block structure parameters (the block header includes 10 fields such as version number, previous block hash value, and timestamp; the block body includes transaction data records). If the number of valid data records exceeds 1000, a block is created for every 1000 records. If there are fewer than 1000 records, they are temporarily stored in a supplementary data pool and will be encapsulated when the number of records reaches 1000 or when the block generation interval (one block is generated every 30 minutes) is reached. Finally, a 16-bit transaction number is assigned to each data record that meets the block capacity requirement. An ordered transaction data queue is generated based on the assigned transaction number and timestamp, and then converted into transaction data units that conform to the blockchain block structure. Each unit includes 20 fields such as transaction number, timestamp, ticket identifier, and entry information.
[0029] S4 generates tags for the blockchain transaction data units generated by S3 through the archive tag intelligent matching engine. The engine calls the accounting archive classification standard library and combines the parameters of archive formation time, business type, and retention period to generate a tag set including multi-level classification tags.
[0030] Specifically, during step S4, the intelligent matching engine for file tags is activated. The engine first loads the accounting file classification standard library, which includes industry classification standards (including 30 industries such as manufacturing, finance, and services), business classification standards (including 50 types of business such as procurement, sales, and expense reimbursement), and retention period classification standards (divided into three categories: permanent retention, 30-year retention, and 10-year retention). Based on these standards, a standard parameter system for tag generation is established. The system includes 10 categories for first-level tags, 50 categories for second-level tags, and 200 categories for third-level tags. Subsequently, key information is extracted from the blockchain transaction data units generated in S3, including the file formation timestamp (accurate to year, month, and day), business type code (using a 6-digit numerical code, with the first two digits representing the industry and the last four digits representing the specific business), and the relevant accounting subject code (using a 4-digit numerical code, following the enterprise accounting standards subject coding rules). This key information is converted into feature parameters that the engine can recognize. The feature parameters are represented numerically (e.g., 20240510 represents May 10, 2024, and 010001 represents manufacturing procurement business). Next, the feature parameters are compared with the standard parameter system from multiple dimensions. Based on the timestamp of formation, a retention period label is determined (e.g., accounting records formed in 2024 are labeled "permanent" for permanent retention and "30 years" for ordinary vouchers). Based on the business type code, a business category label is determined, and based on the accounting subject code, an account association label is determined, forming an initial label group. Finally, redundancy checks are performed on the initial label group, deleting duplicate labels (if the same record is labeled both "Purchasing Business" and "Purchasing-related Business") and invalid labels (e.g., blank labels without a corresponding business type). Based on the label association rules for accounting record management (e.g., the "Purchasing Business" label must be associated with the "Accounts Payable" account label), missing association labels are added, generating a final label set containing 3-5 labels.
[0031] S5 associates and binds the tag set generated in S4 with the blockchain transaction data unit to form a tagged accounting record blockchain data block. The hash value of the data block is calculated to generate a unique identifier for the data block.
[0032] Specifically, during step S5, the tag set generated in S4 and the blockchain transaction data unit generated in S3 are first read. A mapping table is established between the two through the data association module. The mapping table clearly defines the correspondence between each tag and a specific field in the transaction data unit (e.g., the "Procurement Business" tag corresponds to the "Business Type Code" field in the transaction data unit, and the "30-Year Storage" tag corresponds to the "Storage Period Identifier" field). The mapping table is stored in key-value pair format, with the tag name as the key and the field name and value of the transaction data unit as the value. Next, the mapping table is embedded into the extended field of the blockchain transaction data unit. The extended field is set to 1024 bytes to ensure that it can accommodate the mapping table data. Subsequently, the embedded transaction data unit is formatted to conform to the format requirements of blockchain data blocks (using JSON format, field order following blockchain data specifications, and character encoding in UTF-8). After conversion, the hash calculation module is invoked to perform two hash calculations on the converted transaction data units. The first calculation uses the SHA-256 algorithm to obtain the initial hash value of the data unit. The second calculation combines the initial hash value with the feature value of the tag set (obtained by summing the ASCII codes of the tag names) and again uses the SHA-256 algorithm to calculate the final hash value. Finally, the final hash value is used as the unique identifier (32 bytes long) of the data block and bound to the converted transaction data unit. During the binding process, a digital signature algorithm (using the RSA algorithm with a key length of 2048 bits) ensures that the association between the identifier and the data unit is tamper-proof, forming a tagged accounting record blockchain data block.
[0033] S6, based on the storage parameters and access permission parameters of accounting file management, stores the uniquely identified accounting file blockchain data blocks generated in S5 to the distributed blockchain nodes, and updates the file index library of the blockchain nodes to perform distributed storage and index update of accounting file data.
[0034] Specifically, during step S6, the storage parameters for accounting file management are first read (the number of distributed storage nodes is set to 10, each node stores the full amount of blockchain data, and the node storage capacity is no less than 10TB) and access permission parameters are also read (three roles are set: administrator, financial personnel, and auditor. Administrators have full access permissions, financial personnel have data viewing and uploading permissions, and auditors have data viewing and export permissions. The permissions adopt a role-based access control model). Then, the accounting file blockchain data block with a unique identifier generated in S5 is transmitted to the 10 distributed blockchain nodes via a P2P network. During transmission, SSL / TLS protocol is used for encryption to ensure data transmission security. After receiving the data block, each node first verifies the consistency between the unique identifier of the data block and the data content (by recalculating the hash value and comparing it with the identifier). After successful verification, the data block is stored in the node's local database (using LevelDB database, which supports efficient read and write operations). After storage is completed, the node sends storage confirmation information back to the system. Once the system receives storage confirmation information from at least 6 nodes, it triggers an update to the blockchain node's archive index. The index uses a B+ tree index structure. During the update, key information such as the unique identifier of the data block, tag information, and storage node address are written into the index. After the update is completed, the index can support fast queries by multiple dimensions such as tags, timestamps, and transaction numbers, ultimately realizing distributed storage and index updates of accounting archive data, ensuring data security and accessibility.
[0035] Preferably, the expression of the document semantic understanding reasoning model is: ,in, For the semantic feature vector set of the bill, The number of field types for the document. For the first The semantic weight coefficient of the document-like field. For the first The semantic mapping matrix of the fields of the bill class. For the first The text feature matrix of the document-like field. For the first Numerical feature matrix of invoice fields, For the first The format feature matrix of the fields of the document type. For the feature matrix fusion operator, Adjustment coefficients for the accounting field dictionary. This is the semantic vector matrix of a semantic dictionary for the accounting field. This represents the number of semantic reasoning layers.
[0036] Specifically, when implementing the semantic understanding and reasoning model for invoices, the number of invoice field types is first determined to be three, corresponding to text fields, numeric fields, and format fields. The semantic weight coefficient for the first type of text field is set to 0.4, the semantic weight coefficient for the second type of numeric field is set to 0.35, and the semantic weight coefficient for the third type of format field is set to 0.25. The semantic mapping matrix dimension for each field is set to 256×128 to ensure that the field features can be fully represented. The text feature matrix is composed of word vectors after word segmentation of the invoice text field, the numeric feature matrix is formed by standardizing the numerical values such as invoice amount and quantity, and the format feature matrix is generated by encoding format information such as invoice number structure and signature position. The feature matrix fusion operator adopts element-level addition operation to achieve effective fusion of the three types of features. The adjustment coefficient of the accounting domain dictionary is set to 0.15 to balance the impact of the semantic dictionary on semantic extraction. The semantic vector matrix dimension of the accounting domain semantic dictionary is set to 1200×256, including the semantic vectors of 1200 core accounting terms. The number of semantic reasoning layers is set to 5 to improve the accuracy of semantic understanding through multi-layer reasoning. During implementation, the model first calls the semantic dictionary library to load the semantic vector matrix, then inputs the feature matrices of three types of fields. After weight allocation, matrix fusion, and multi-layer inference, a 256-dimensional semantic feature vector set of bills is generated. This vector set can fully reflect the key information such as the type of bill, the amount association, and the transaction subject, providing accurate data support for subsequent identification of reconciliation relationships.
[0037] Preferably, the expression for the accounting reconciliation relationship identification algorithm is: ,in, For the set of matching results of reconciliation relationships, The number of types of reconciliation relationships. For the first Matching weights for similar logical relationships For the first A rule matrix for reconciliation-like relationships. For matrix dot product operation, For the related adjustment factor of accounting items, This is the rule vector matrix of the accounting subject association rule base. Number of accounting entry types For the first Class-based logical relationships and the first The correlation coefficient of similar accounting entries, This is the feature vector of accounting entry data.
[0038] Specifically, when implementing the accounting reconciliation relationship identification algorithm, the number of reconciliation relationship types is set to 8, including account balance reconciliation, report item reconciliation, and invoice and journal entry reconciliation. The matching weight for the first type of account balance reconciliation is set to 0.2, the matching weight for the second type of invoice and journal entry reconciliation is set to 0.3, and the matching weights for the remaining types of reconciliation relationships are set between 0.1 and 0.15 according to their importance. The rule matrix dimension for each type of reconciliation relationship is set to 150×256, and the matrix elements are generated by encoding 150 rules in the accounting account association rule base. The matrix dot product operation is used to calculate the matching degree between the semantic feature vector of the invoice and the rule matrix. The accounting account association adjustment coefficient is set to 0.2 to adjust the degree of influence of the rule base on the matching results. The rule vector matrix dimension of the accounting account association rule base is set to 150×128, storing the feature vectors of various account association rules. The number of accounting entry types is 12, including entry types corresponding to the six major accounting elements such as assets and liabilities. The correlation coefficient between the j-th type of reconciliation relationship and the k-th type of accounting entry ranges from 0 to 1, depending on the degree of their correlation. The feature vector dimension of the accounting entry data is set to 256 dimensions, generated from information such as the account code and amount of the accounting entry. During implementation, the algorithm first loads the rule matrix and rule vector matrix, inputs the semantic feature vector set of the invoice and the feature vector of the accounting entry, and after weight calculation, matrix dot product, and correlation adjustment, outputs a reconciliation relationship matching result set including matching results and similarity values, ensuring the accuracy of the correlation between invoices and accounting entries.
[0039] Preferably, the expression for the blockchain transaction sequence encapsulation algorithm is: ,in, For blockchain transaction data units, These are the conversion coefficients for the transaction sequence. A timestamp vector matrix, For matrix tensor product operations, For transaction number vectors, For ordered vectors, This is the block structure adjustment coefficient. For hash calculation matrix, This is the parameter matrix for the blockchain block structure.
[0040] Specifically, in the implementation of the blockchain transaction sequence encapsulation algorithm, the transaction sequence conversion coefficient is set to 0.3 to adjust the data scaling ratio during the transaction sequence conversion process. The timestamp vector matrix is set to a dimension of 1000×32, with each element storing the timestamp information of one transaction (accurate to milliseconds). The transaction number vector is set to a dimension of 1000×16, with each element being a unique 16-bit transaction number. Matrix tensor product operation is used to expand the dimensions of the timestamps and transaction numbers, forming a high-dimensional matrix that integrates both information. The ordered arrangement vector is set to a dimension of 1000×1, with element values arranged in ascending order of transaction timestamps to determine the arrangement order of transaction data. The block structure adjustment coefficient is set to 0.25 to balance the impact of block structure parameters on data unit generation. The hash calculation matrix is set to a dimension of 256×256, using the parameter matrix of the SHA-256 algorithm. The blockchain block structure parameter matrix is set to a dimension of 50×256, storing the feature vectors of structural parameters such as the block header and block body. The reconciliation matching result set and the block structure parameter matrix achieve feature fusion through element-level addition operations. During implementation, the algorithm first reads parameters such as timestamps and transaction numbers to generate corresponding matrices. After tensor product operation, ordered arrangement, and feature fusion, it combines the hash calculation matrix to generate transaction data units that conform to the blockchain block structure. Each unit includes 1,000 transaction data entries, ensuring that the data can be directly used for blockchain storage.
[0041] Preferably, the tag generation expression of the file tag intelligent matching engine is: ,in, For a set of tags, For the number of tag levels, For the first The generation weight of hierarchical labels For the first The classification matrix of hierarchical labels, For the classification vector of the accounting record classification standard library, Adjust the coefficient for tag association. For the time vector of archive formation, For business type vectors, For the storage period parameter vector, This is a label dictionary matrix.
[0042] Specifically, when implementing the intelligent matching engine for archive tags, the number of tag levels is set to 3, corresponding to first-level, second-level, and third-level classification tags. The generation weight of the first-level tag is set to 0.4, the second-level to 0.35, and the third-level to 0.25. The classification matrix dimensions for each level of tag are set to 10×256, 50×256, and 200×256, respectively, and the matrix elements are generated by the classification rules encoded in the accounting archive classification standard library. The classification vector dimension of the accounting archive classification standard library is set to 256 dimensions, including classification features such as industry, business, and retention period. The tag association adjustment coefficient is set to 0.2, used to adjust the impact of multi-dimensional parameters on tag generation. The archive formation time vector dimension is set to 32 dimensions, generated by encoding information such as the year, month, and day of archive formation. The business type vector dimension is set to 64 dimensions, including feature information of 50 types of businesses. The retention period parameter vector dimension is set to 16 dimensions, corresponding to three retention periods: permanent, 30 years, and 10 years. The tag dictionary matrix dimension is set to 200×128, storing the semantic vectors of 200 tags. The document creation time vector, business type vector, and retention period parameter vector are fused using element-wise addition. During implementation, the engine first loads the classification matrix and classification vector, inputs the multi-dimensional parameter vector, and after weight allocation, classification matching, and feature fusion, generates a tag set including 3-5 tags to ensure that the tags accurately reflect the classification information of the document.
[0043] Preferably, the unique identifier generation expression for the accounting record blockchain data block is: ,in, A unique identifier for the data block. Adjustment factors are calculated for hash calculation. Calculate the matrix for the hash values. Adjust the coefficients for the node index. A vector for numbering blockchain nodes. This is a vector of archive index parameters.
[0044] Specifically, during the implementation of unique identifier generation for accounting archive blockchain data blocks, the hash calculation adjustment coefficient is set to 0.4 to regulate the impact of hash calculation on unique identifier generation. The hash value calculation matrix is set to a dimension of 256×256, using the parameter matrix of the SHA-256 algorithm to ensure the security and uniqueness of the hash calculation. Blockchain transaction data units and tag sets are fused using element-level addition operations, resulting in a fused matrix of 256×256 dimensions, including complete feature information for both data units and tags. The node index adjustment coefficient is set to 0.25 to balance the impact of node index parameters on the identifier. The blockchain node number vector is set to a dimension of 10×16, storing the unique numbers of 10 distributed nodes. The archive index parameter vector is set to a dimension of 64, including key field features of the archive index. During implementation, the blockchain transaction data units and tag sets are first read, fused, and then input into the hash calculation matrix to obtain a preliminary hash value. This hash value is then combined with the node number vector and the archive index parameter vector, and weighted by the adjustment coefficient to generate a 32-byte unique identifier for the data block. Once this identifier is bound to a data block, it ensures the uniqueness of each blockchain data block, facilitating subsequent data retrieval and integrity verification, while also supporting data synchronization and consistency maintenance between distributed nodes.
[0045] Preferred, such as Figure 2 As shown, step S3 includes the following sub-steps: S31, calling the initialization module of the blockchain transaction sequence encapsulation algorithm, reading the blockchain block capacity parameters and transaction sequence sorting rule parameters preset by the accounting file management system, and importing the parameters into the parameter configuration layer of the algorithm; S32, cleaning the data of the reconciliation matching result set output by S2, removing abnormal data records in the matching result set, and performing preliminary sorting of valid data records according to the order of transaction occurrence time; S33, comparing the preliminary sorted valid data records with the blockchain block structure parameters to determine whether the number of data records meets the block capacity requirements. If it exceeds the requirements, the block is split; if it is insufficient, it is temporarily stored in the data pool to be supplemented; S34, assigning transaction numbers to the data records that meet the block capacity requirements, generating an ordered transaction data queue according to the assigned transaction number and timestamp, and converting it into blockchain transaction data units.
[0046] Specifically, step S3 includes four sub-steps. S31 first calls the initialization module of the blockchain transaction sequence encapsulation algorithm, reading the preset blockchain block capacity parameters (each block stores a maximum of 1000 transaction data records) and transaction sequence sorting rule parameters (sorted in ascending order by transaction timestamp, with timestamp precision accurate to milliseconds) from the accounting file management system. These parameters are then imported into the algorithm's parameter configuration layer, which uses a key-value pair storage structure to ensure efficient parameter calls. S32 performs data cleaning on the matching result set output by S2. During cleaning, a similarity threshold of 0.85 is set, removing abnormal data records with similarity below this threshold. Simultaneously, a duplicate data detection algorithm is used to delete redundant data submitted repeatedly. After cleaning, the valid data records are initially sorted according to the transaction occurrence time using a quicksort algorithm with a time complexity controlled within O(n log n). To ensure efficiency; S33 compares the initially sorted valid data records with the blockchain block structure parameters (the block header includes 10 fields such as version number and hash value of the previous block, and the block body includes transaction data records) to determine whether the number of data records meets the block capacity requirements. If it exceeds 1,000 records, it splits into a block of 1,000 records. If it is insufficient, it is temporarily stored in the supplementary data pool. The supplementary data pool is set to a maximum cache time of 30 minutes. After the timeout, a new block is automatically generated. S34 assigns a 16-digit transaction number (the numbering rule is "year + month + random 8-digit number") to the data records that meet the block capacity requirements. An ordered transaction data queue is generated according to the assigned transaction number and timestamp. The queue is converted into transaction data units that conform to the blockchain block structure through the format conversion module. The converted data units are stored in JSON format, and the field order strictly follows the blockchain data specifications.
[0047] Preferred, such as Figure 3 As shown, step S4 includes the following sub-steps: S41, activating the intelligent matching engine for archive tags, loading industry classification standards, business classification standards, and retention period classification standards from the accounting archive classification standard library, and establishing a standard parameter system for tag generation; S42, extracting the labeling information from the blockchain transaction data unit generated in S3, including the archive formation timestamp, business type code, and related accounting subject code, and converting the labeling information into feature parameters that the engine can recognize; S43, comparing the feature parameters with the standard parameter system in multiple dimensions to determine the primary classification label, secondary classification label, and tertiary classification label corresponding to the archive, forming an initial label group; S44, performing redundancy verification on the initial label group, deleting duplicate and invalid labels, and supplementing missing related labels according to the label association rules for accounting archive management to generate the final label set.
[0048] Specifically, step S4 includes four sub-steps. S41 first activates the intelligent matching engine for archive tags, loading industry classification standards (including 30 industries such as manufacturing and finance, each with a unique 6-digit code), business classification standards (including 50 business categories such as procurement and sales, each with a 4-digit code), and retention period classification standards (divided into three categories: permanent retention, 30-year retention, and 10-year retention, each with a 1-digit code) from the accounting archive classification standard library. Based on these standards, a standard parameter system for tag generation is established. The system adopts a hierarchical structure: first-level tags include 10 categories, second-level tags include 50 categories, and third-level tags include 200 categories. S42 extracts key information from the blockchain transaction data unit generated in S3, including the archive formation timestamp (accurate to year, month, and day, in the format "YYYYMMDD"), business type code (6-digit code), and relevant accounting subject code (4-digit code, following enterprise accounting standards). The code rules are used to convert these key information into numerical feature parameters that the engine can recognize through the data conversion module. During the conversion process, Min-Max standardization is used to map the parameter values to the 0-1 range. S43 compares the feature parameters with the standard parameter system in multiple dimensions. During the comparison, the cosine similarity algorithm is used to calculate the matching degree. The matching degree threshold is set to 0.9. If it is higher than the threshold, the first-level, second-level, and third-level classification labels corresponding to the file are determined to form an initial label group. Each label in the initial label group includes three attributes: label name, code, and matching degree. S44 performs redundancy verification on the initial label group. Duplicate labels are deleted by label name deduplication algorithm, and blank labels without corresponding business types are deleted by validity detection algorithm. Then, according to the label association rules of accounting file management (such as the "Purchasing Business" label must be associated with the "Accounts Payable" account label), the missing associated labels are supplemented. Finally, a label set including 3-5 labels is generated. The set is stored in XML format to ensure compatibility.
[0049] Preferred, such as Figure 4 As shown, step S5 includes the following sub-steps: S51, read the tag set generated in S4 and the blockchain transaction data unit generated in S3, establish an association mapping table between the two, and clarify the correspondence between each tag and the specific fields in the transaction data unit; S52, embed the association mapping table into the extended fields of the blockchain transaction data unit, and perform format conversion on the embedded transaction data unit to make it conform to the format requirements of the blockchain data block; S53, call the hash calculation module to perform two hash calculations on the format-converted transaction data unit. The first calculation obtains the primary hash value of the data unit, and the second calculation calculates the final hash value based on the primary hash value and the feature value of the tag set; S54, use the final hash value as the unique identifier of the data block and bind it with the format-converted transaction data unit to form a tagged accounting archive blockchain data block.
[0050] Specifically, step S5 includes four sub-steps. S51 first reads the tag set generated in S4 and the blockchain transaction data unit generated in S3, establishing a mapping table between them through the data association module. The mapping table clearly defines the correspondence between each tag and a specific field in the transaction data unit (e.g., the "Procurement Business" tag corresponds to the "Business Type Code" field, and the "30-Year Storage" tag corresponds to the "Storage Period Identifier" field). The mapping table uses a hash table storage structure, with the key being the tag name and the value being the field name and field value. The hash table's load factor is set to 0.75 to balance storage and query efficiency. S52 embeds the mapping table into the extended field of the blockchain transaction data unit. The extended field is set to 1024 bytes to ensure it can accommodate the mapping table data. Then, the format conversion module performs format conversion on the embedded transaction data unit to meet the format requirements of blockchain data blocks (using Protocol Buffers format, with a higher compression rate than JS). (ON format 30%), strict validation of field type and length is performed during the conversion process; S53 calls the hash calculation module to perform two hash calculations on the converted transaction data unit. The first calculation uses the SHA-256 algorithm to output a 256-bit primary hash value. The second calculation concatenates the primary hash value with the feature value of the tag set (obtained by summing the ASCII codes of the tag names) and then uses the SHA-256 algorithm again to calculate the final hash value. Salt values are set for both calculations to improve security; S54 uses the final hash value as the unique identifier of the data block (32 bytes long). The digital signature module uses the RSA algorithm (key length of 2048 bits) to bind the identifier to the transaction data unit. During the binding process, a digital signature is generated and appended to the end of the data block to form a tagged accounting archive blockchain data block. After the data block is generated, the data integrity is verified by the integrity verification algorithm. Only after the verification is passed can it enter the subsequent storage stage.
[0051] The semantic understanding and reasoning model for invoices in this invention is a semantic feature extraction technology model specifically designed for original invoices in accounting archives. It transforms unstructured invoice data such as invoices, receipts, and bank statements into structured semantic feature vectors that include key information, providing standardized input for subsequent data processing. The implementation process first completes the basic configuration, pre-loading an accounting semantic dictionary library containing 1200 core accounting terms (including terms related to the six major accounting elements such as assets, liabilities, and owner's equity) and 300 types of invoice features (including the format and content features of 20 common invoices such as VAT invoices and expense reimbursement forms). Next, the model parameters are set, with the number of semantic reasoning layers set to 5 to ensure the depth of semantic understanding. Semantic weight coefficients of 0.4, 0.35, and 0.25 are assigned to the text fields of the invoice (names of the transacting parties, business description), numeric fields (amount, quantity), and format fields (invoice numbering rules, signature position) respectively to ensure a reasonable proportion of different types of field features. Finally, through a multi-dimensional semantic mapping algorithm, the three types of field features are fused and calculated to generate a 256-dimensional semantic feature vector set, which clearly includes key information such as invoice type identifier, amount association identifier, and transaction entity identifier. This model provides a precise and structured data foundation for subsequent accounting reconciliation identification, completely solving the problems of low efficiency (processing a single invoice takes 3-5 minutes) and large errors (human identification error rate is about 5%) in traditional manual extraction of invoice information. It realizes full automation and intelligence of invoice information extraction, controlling the semantic extraction time of each type of invoice to within 0.5 seconds. It can efficiently process 20 common types of invoices in a single batch, significantly improving the initial data quality of accounting file processing, laying a solid data foundation for intelligent management of the entire accounting file process, while reducing labor costs and reducing the risk of information deviation caused by human operation.
[0052] The accounting reconciliation relationship identification algorithm is used to automatically determine the reasonableness of the relationship between invoice data and accounting entry data in accounting archives. It verifies the existence of a logical correspondence between the two based on accounting rules, avoiding data association errors. Its implementation first constructs a basic algorithm support system, including a built-in rule library containing 150 accounting subject association rules. These rules cover common correspondences such as asset accounts and purchase invoices, liability accounts and accounts payable, etc. Eight reconciliation relationship types are set, including account balance reconciliation, report item reconciliation, and invoice and entry reconciliation. Different matching weights are assigned according to the importance of each reconciliation relationship. Invoice and entry reconciliation directly affects the data correlation determination, with a weight set at 0.3. The weights of the other types range from 0.1 to 0.15. Subsequently, the algorithm performs calculations, generating a semantic understanding and reasoning model for invoices. The semantic feature vector set and accounting entry data (including account codes, amounts, debit / credit directions, etc.) are input into the algorithm. The similarity between the invoice feature vector and the accounting entry data is calculated one by one through a traversal matching method. A similarity threshold of 0.85 is set. When the calculated result is higher than the threshold, the reconciliation relationship between the two is determined to be matched. If it is lower than the threshold, it is marked as data to be verified. Finally, the reconciliation relationship matching result set is output. The result set includes the matching invoice-entry correspondence, the specific similarity value, the invoice identifier and the reason for failure (such as account mismatch or inconsistent amount). This algorithm replaces the traditional manual verification method, solving the problems of low efficiency (requiring hours or even days to complete massive amounts of data) and error-proneness (human verification error rate is about 3%). It realizes automated verification of the association between invoices and accounting entries, ensuring the accuracy of the association of accounting archive data. The matching accuracy rate can reach over 98%, avoiding financial risks caused by human operation errors (such as incorrect account entries or omissions of related invoices). At the same time, the algorithm takes no more than 2 seconds for a single traversal matching and supports 3 repeated traversal verifications, further improving the matching reliability. It can meet the rapid verification needs of massive accounting archives (the daily processing volume can reach more than 100,000 records), providing compliant and accurate data basis for the subsequent storage and management of accounting archives, and helping to standardize financial work.
[0053] The blockchain transaction sequence encapsulation algorithm is an algorithm that connects the identification of accounting reconciliation relationships with blockchain storage. It converts accounting file data that has passed the reconciliation relationship matching into transaction data units that conform to blockchain technical specifications, ensuring that the data can be successfully integrated into the blockchain system. Its implementation first reads preset system parameters, including blockchain block capacity parameters (each block can store a maximum of 1000 transaction data records to avoid excessively large blocks affecting transmission and storage efficiency) and transaction sequence sorting rules parameters (sorted in ascending order by transaction timestamp, with timestamp precision accurate to milliseconds to ensure data timeliness). Next, data preprocessing is performed, cleaning the data in the reconciliation relationship matching result set. An outlier detection algorithm removes abnormal data records with a similarity lower than 0.85, and a duplicate data identification algorithm deletes redundant data submitted repeatedly to ensure data quality. Then, data splitting and numbering are performed, and the cleaned valid data records are compared with the block capacity parameters. If the data volume exceeds 1... If there are 000 records, the data will be split into groups of 1000 records each. If there are fewer than 1000 records, they will be temporarily stored in a pool to be replenished (with a maximum cache time of 30 minutes, after which a new block will be automatically generated). Each data group that meets the block capacity requirement will be assigned a 16-digit transaction number (the numbering rule is "last 4 digits of the year + 2 digits of the month + 8 random digits" to ensure the uniqueness of the number). Finally, the format will be converted into blockchain transaction data units stored in JSON format. The field order of the data units will strictly follow the blockchain data specifications (including fields such as transaction number, timestamp, invoice identifier, accounting entry summary, and hash pre-value) to ensure compatibility with the blockchain system. This algorithm adapts accounting archive data to blockchain storage formats, solving the problem that traditional data formats cannot be directly accessed by the blockchain system. It provides standardized data units for distributed data storage, ensuring the effective migration of accounting archive data to the blockchain system. The algorithm can generate a complete block every 30 minutes, with a block data conversion success rate of 100%, ensuring that data can be seamlessly accessed by the blockchain network and avoiding data loss or storage failure due to format incompatibility. At the same time, the standardized data units facilitate data synchronization and consistency verification between blockchain nodes, providing a secure and traceable storage foundation for accounting archive data, helping to achieve decentralized management of accounting archives and improving data anti-tampering capabilities.
[0054] The intelligent matching engine for archive tags is a smart tool for generating accurate classification tags for blockchain-based accounting archive data. It can automatically generate multi-level tags by combining multi-dimensional archive parameters, providing efficient support for archive retrieval and management. Its implementation first establishes a basic tag generation system, loading an accounting archive classification standard library that includes 30 industries (such as manufacturing, finance, and services), 50 business categories (such as procurement, sales, and expense reimbursement), and 3 retention periods (permanent retention, 30-year retention, and 10-year retention). Based on this standard library, a three-level tag system is established: Level 1 tags include 10 categories (such as asset archives and liability archives); Level 2 tags include 50 categories (such as procurement asset archives and sales liability archives); and Level 3 tags include 200 categories (such as raw material procurement asset archives and product sales liability archives). Then, key archive parameters are extracted, including the archive formation timestamp (format "YYYYMMDD", accurate to the specific date), a 6-digit business type code (the first 2 digits represent the industry, and the last 4 digits represent the specific business), and a 4-digit accounting subject code (from the blockchain transaction data unit). Following the coding rules of enterprise accounting standards, these parameters are converted into numerical feature parameters in the range of 0-1 using the Min-Max standardization algorithm, facilitating matching calculations by the engine. Next, label matching and optimization are performed. A cosine similarity algorithm is used to compare the feature parameters with the standard parameter system in the classification standard library from multiple dimensions. A matching threshold of 0.9 is set. If the match is higher than the threshold, the corresponding first-level, second-level, and third-level initial label groups for the file are determined. Redundancy checks are performed on the initial label groups (duplicate labels are deleted; for example, if the same file is labeled with both "Purchase Business File" and "Purchase-related Business File," the former is retained) and associated labels are supplemented (according to accounting file management rules, such as "Purchase Business File" needing to be associated with the "Accounts Payable Account File" label). Finally, a final label set is generated, with each file generating a set of 3-5 labels, stored in XML format to ensure compatibility and readability. This engine generates accurate and comprehensive classification tags for accounting records, solving the problems of low tag-to-record correlation (less than 60%) and poor retrieval efficiency (response time exceeding 5 seconds) under traditional fixed-rule tag generation methods. It significantly improves the utilization efficiency of accounting records, with tag generation accuracy exceeding 95%. The response time for tag-based record retrieval is reduced to less than 0.3 seconds, allowing users to quickly locate target records and reduce retrieval time costs. At the same time, accurate tag classification facilitates the classification management and expiration destruction of records, realizing the transformation of accounting record management from "simple storage" to "efficient utilization and precise management," and helping to improve the overall efficiency of financial work.
[0055] like Figure 5As shown, an accounting record management system based on blockchain technology and a large language model is applied to an accounting record management method based on blockchain technology and a large language model. The system includes: a document semantic feature extraction unit, which is connected to the original accounting record data acquisition device. This unit extracts semantic features from the acquired original document data using a document semantic understanding and reasoning model, and outputs a semantic feature vector set to an accounting reconciliation relationship matching unit; an accounting reconciliation relationship matching unit, which is connected to both the document semantic feature extraction unit and the accounting subject rule database. This unit calls an accounting reconciliation relationship recognition algorithm to calculate the reconciliation relationships of the semantic feature vector set, and outputs a reconciliation relationship matching result set to a blockchain transaction sequence conversion unit; and a blockchain transaction sequence conversion unit, which is connected to both the accounting reconciliation relationship matching unit and the blockchain parameter configuration module. This unit uses a blockchain transaction sequence encapsulation algorithm to convert the matching result set into blockchain transactions. The data unit is output to the intelligent document tag generation unit. This unit is connected to the blockchain transaction sequence conversion unit and the accounting document classification standard library. It generates a tag set through an intelligent document tag matching engine, associates the tag set with the blockchain transaction data unit, and outputs it to the blockchain data block generation unit. The blockchain data block generation unit is connected to both the intelligent document tag generation unit and the hash calculation module. It performs hash calculations on the transaction data units after associating tags to generate unique identifiers, forming accounting document blockchain data blocks which are then output to the distributed storage unit. The distributed storage unit is connected to both the blockchain data block generation unit and the blockchain node network. Based on the storage parameters and access permission parameters for accounting document management, it stores the accounting document blockchain data blocks on different blockchain nodes and updates the node's document index library, enabling distributed management of accounting document data.
[0056] The accounting record management method and system based on blockchain technology and large language models extract multi-dimensional semantic features from the text, number, and format fields of the original invoice data through a semantic understanding and reasoning model. This is then combined with an accounting reconciliation relationship recognition algorithm that calls upon an accounting subject association rule base to achieve automated matching of reconciliation relationships between invoices and accounting entries. This replaces traditional manual verification, significantly reducing the probability of association errors and improving the accuracy of data association. Simultaneously, a blockchain transaction sequence encapsulation algorithm converts matched record data into blockchain transaction data units based on timestamps and transaction numbers. Combined with distributed storage, this replaces the traditional centralized storage architecture, fundamentally avoiding the risks of data tampering and loss, ensuring the long-term integrity and traceability of accounting record data, and meeting the core data security requirements of record management.
[0057] This method and system utilize an intelligent matching engine for archive tags to access an accounting archive classification standard library. By combining multiple dimensions such as archive creation time, business type, and retention period, it generates accurate multi-level tags, breaking the limitations of traditional fixed-rule tag generation and solving the problem of low correlation between tags and archive data. Based on the generated multi-level tags, target archives can be quickly located during subsequent searches without extensive manual screening, significantly improving search efficiency. Furthermore, the entire process, from semantic extraction and reconciliation matching of invoices to blockchain data conversion, tag generation, and distributed storage, is automated, reducing manual intervention. This enables efficient handling of massive accounting archive management needs and promotes the transformation of accounting archive management from traditional models to intelligent and efficient ones.
[0058] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," "link," and "fix" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0059] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An accounting record management method based on blockchain technology and a large language model, characterized in that: Includes the following steps: S1. Semantic features are extracted from the original invoice data associated with accounting records using a semantic understanding and reasoning model for invoices. This model calls upon a pre-defined semantic dictionary in the accounting domain to perform multi-dimensional semantic mapping on text fields, numeric fields, and format fields in the invoices, generating a set of semantic feature vectors including invoice type identifiers, amount association identifiers, and transaction entity identifiers. S2. Based on an accounting reconciliation relationship recognition algorithm, the reconciliation relationships between different invoice data in the semantic feature vector set generated in S1 are calculated. Using the algorithm's built-in accounting subject association rule library, the correspondence between the invoice semantic feature vectors and accounting entry data is traversed and matched, outputting a set of reconciliation relationship matching results. S3. A blockchain transaction sequence is used. The encapsulation algorithm transforms the matching result set of the correlation relationship output by S2 into a transaction sequence, and arranges the matched accounting archive data in an orderly manner according to timestamp and transaction number to generate transaction data units that conform to the blockchain block structure; S4, the archive tag intelligent matching engine generates tags for the blockchain transaction data units generated by S3. The engine calls the accounting archive classification standard library and combines the parameters of archive formation time, business type, and retention period to generate a tag set including multi-level classification tags; S5, the tag set generated by S4 is associated and bound with the blockchain transaction data units to form a tagged accounting archive blockchain data block. The hash value of the data block is calculated to generate a unique identifier for the data block; S6, based on the storage parameters and access permission parameters of accounting file management, stores the uniquely identified accounting file blockchain data block generated by S5 to the distributed blockchain node, and updates the file index library of the blockchain node to perform distributed storage and index update of accounting file data. The expression for the semantic understanding and reasoning model of the bill is: ,in, For the semantic feature vector set of the bill, The number of field types for the document. For the first The semantic weight coefficient of the document-like field. For the first The semantic mapping matrix of the fields of the bill class. For the first The text feature matrix of the document-like field. For the first Numerical feature matrix of invoice fields, For the first The format feature matrix of the fields of the document type. For the feature matrix fusion operator, Adjustment coefficients for the accounting field dictionary. This is the semantic vector matrix of a semantic dictionary for the accounting field. This refers to the number of semantic reasoning layers; The expression for the accounting reconciliation relationship identification algorithm is: ,in, For the set of matching results of reconciliation relationships, The number of types of reconciliation relationships. For the first Matching weights for similar logical relationships For the first A rule matrix for reconciliation-like relationships. For matrix dot product operation, For the related adjustment factor of accounting items, This is the rule vector matrix of the accounting subject association rule base. Number of accounting entry types For the first Class-based logical relationships and the first The correlation coefficient of similar accounting entries, This is the feature vector of accounting entry data; The expression for the blockchain transaction sequence encapsulation algorithm is: ,in, For blockchain transaction data units, These are the conversion coefficients for the transaction sequence. A timestamp vector matrix, For matrix tensor product operations, For transaction number vectors, For ordered vectors, This is the block structure adjustment coefficient. For hash calculation matrix, This is a matrix of blockchain block structure parameters; The tag generation expression of the intelligent matching engine for file tags is: ,in, For a set of tags, For the number of tag levels, For the first The generation weight of hierarchical labels For the first The classification matrix of hierarchical labels, For the classification vector of the accounting record classification standard library, Adjust the coefficient for tag association. For the time vector of archive formation, For business type vectors, For the storage period parameter vector, A matrix of label dictionaries; The unique identifier generation expression for the accounting record blockchain data block is: ,in, A unique identifier for the data block. Adjustment factors are calculated for hash calculation. Calculate the matrix for the hash values. Adjust the coefficients for the node index. A vector for numbering blockchain nodes. This is a vector of archive index parameters.
2. The accounting record management method based on blockchain technology and a large language model according to claim 1, characterized in that, S3 includes the following sub-steps: S31, calling the initialization module of the blockchain transaction sequence encapsulation algorithm, reading the blockchain block capacity parameters and transaction sequence sorting rule parameters preset by the accounting file management system, and importing the parameters into the parameter configuration layer of the algorithm; S32, cleaning the data of the reconciliation matching result set output by S2, removing abnormal data records in the matching result set, and initially sorting the valid data records according to the order of transaction occurrence time; S33, comparing the initially sorted valid data records with the blockchain block structure parameters to determine whether the number of data records meets the block capacity requirements. If it exceeds the requirements, the block is split; if it is insufficient, it is temporarily stored in the data pool to be supplemented; S34, assigning transaction numbers to the data records that meet the block capacity requirements, generating an ordered transaction data queue according to the assigned transaction number and timestamp, and converting it into blockchain transaction data units.
3. The accounting record management method based on blockchain technology and a large language model according to claim 1, characterized in that, S4 includes the following steps: S41, start the intelligent matching engine for archive tags, load the industry classification standards, business classification standards, and retention period classification standards in the accounting archive classification standard library, and establish a standard parameter system for tag generation; S42, extract the labeling information from the blockchain transaction data unit generated in S3, including the archive formation timestamp, business type code, and relevant accounting subject code, and convert the labeling information into feature parameters that the engine can recognize. S43. Compare the feature parameters with the standard parameter system in multiple dimensions to determine the primary, secondary, and tertiary classification labels corresponding to the archives, forming an initial label group; S44. Perform redundancy checks on the initial label group, delete duplicate and invalid labels, and supplement missing related labels according to the label association rules of accounting archive management to generate the final label set.
4. The accounting record management method based on blockchain technology and a large language model according to claim 1, characterized in that, S5 includes the following sub-steps: S51, read the tag set generated in S4 and the blockchain transaction data unit generated in S3, establish an association mapping table between the two, and clarify the correspondence between each tag and the specific fields in the transaction data unit; S52, embed the association mapping table into the extended fields of the blockchain transaction data unit, and perform format conversion on the embedded transaction data unit to make it conform to the format requirements of the blockchain data block; S53, call the hash calculation module to perform two hash calculations on the format-converted transaction data unit. The first calculation obtains the primary hash value of the data unit, and the second calculation calculates the final hash value based on the primary hash value and the feature value of the tag set; S54, use the final hash value as the unique identifier of the data block and bind it with the format-converted transaction data unit to form a tagged accounting archive blockchain data block.
5. An accounting record management system based on blockchain technology and a large language model, characterized in that: This system is applied to the accounting archive management method based on blockchain technology and a large language model as described in claim 1, comprising: a document semantic feature extraction unit, which is connected to the original data acquisition device for accounting archives, extracts semantic features from the acquired original document data through a document semantic understanding and reasoning model, and outputs a semantic feature vector set to an accounting reconciliation relationship matching unit; an accounting reconciliation relationship matching unit, which is connected to the document semantic feature extraction unit and the accounting subject rule database, respectively, calls an accounting reconciliation relationship recognition algorithm to calculate the reconciliation relationship of the semantic feature vector set, and outputs a reconciliation relationship matching result set to a blockchain transaction sequence conversion unit; and a blockchain transaction sequence conversion unit, which is connected to the accounting reconciliation relationship matching unit and the blockchain parameter configuration module, uses a blockchain transaction sequence encapsulation algorithm to convert the matching result set into a blockchain transaction data unit, and outputs it to the archive. The system comprises: a smart tag generation unit; a blockchain transaction sequence conversion unit; an accounting archive classification standard library; a tag set generated by a smart tag matching engine; a blockchain data block generation unit; a hash calculation module connected to the smart tag generation unit; a distributed storage unit connected to the blockchain data block generation unit and a blockchain node network; and a distributed storage unit that stores accounting archive blockchain data blocks on different blockchain nodes based on storage and access parameters for accounting archive management, and updates the archive index database of the nodes for distributed management of accounting archive data.
Citation Information
Patent Citations
Financial data compliance examination method and device based on block chain
CN115952240A
Intelligent archive opening identification method based on large model
CN120611077A