Enterprise financial data security storage management system and method based on internet of things
By constructing a set of trusted authentication device data and a minimum trust path, and by collecting and structuring invoice data, as well as encrypting and controlling its traceability, the problem of secure management of enterprise financial data in the Internet of Things environment is solved, and full-process secure management and efficient data archiving are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2026-03-31
AI Technical Summary
Traditional enterprise financial data security management methods lack full lifecycle security supervision capabilities in the Internet of Things environment, making them vulnerable to man-in-the-middle attacks, data hijacking, and data leakage, especially due to the lack of end-to-end encryption measures during wireless communication protocol transmission.
By acquiring financial data transmission logs based on the Internet of Things, building trusted authentication device data, generating a minimum set of trusted paths, collecting and structuring invoice data, performing homomorphic structured encryption, and archiving through multi-tag traceability control, the security of data during transmission and storage is ensured.
It enables full-process secure management of enterprise financial data in the Internet of Things environment, enhances the credibility of data collection, ensures basic security of transmission links, improves the field identifiability and reproducibility of data encryption and indexing processes, and improves system stability and business implementation capabilities.
Smart Images

Figure CN121143715B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data storage technology, and in particular to an Internet of Things-based enterprise financial data secure storage management system and method. Background Technology
[0002] Driven by the continuous maturation of "Internet+" and Internet of Things (IoT) technologies, various types of corporate financial information are being automatically collected, remotely transmitted, and centrally processed through a variety of smart terminal devices. In this process, the transmission path of financial data exhibits characteristics such as multi-source heterogeneity, distributed transmission, and edge intelligent processing. The data content involves core business information such as corporate revenue, expenses, taxes, labor costs, and assets and liabilities, and possesses high sensitivity and confidentiality.
[0003] Traditional enterprise financial data security management methods largely rely on database encryption, access control policies, and physical isolation within a local area network (LAN) environment for information protection. Specific technical measures include setting up firewalls, antivirus software, permission-based access control mechanisms, and symmetric or asymmetric encryption storage of the database. However, these methods have the following technical drawbacks:
[0004] Traditional systems generally lack the capability to provide full lifecycle security oversight of data accessed by IoT devices. During the data transmission process from edge sensing nodes to central processing nodes, they are vulnerable to security threats such as man-in-the-middle attacks, data hijacking, and forged command injection. This is especially true when using wireless communication protocols (such as Wi-Fi, LoRa, and NB-IoT) for data transmission, where the lack of end-to-end encryption measures can easily lead to data leaks. Summary of the Invention
[0005] Therefore, it is necessary for the present invention to provide an Internet of Things-based enterprise financial data secure storage management system and method to solve at least one of the above-mentioned technical problems.
[0006] To achieve the above objectives, an IoT-based method for secure storage and management of enterprise financial data includes the following steps:
[0007] Step S1: Obtain enterprise financial data transmission logs based on the Internet of Things, and perform data cache chain node trust anchoring based on the enterprise financial data transmission logs to obtain trusted authentication device data;
[0008] Step S2: Construct a minimum trust path set based on the trust authentication device data, and use the minimum trust path set to analyze the jump point trust transmission nodes of enterprise financial data transmission information to obtain the jump point trust transmission map;
[0009] Step S3: Using the data from the trusted authentication device, perform multi-channel parallel financial document data collection on the Internet of Things to obtain trusted financial document data; perform multi-category field structure enhancement on the trusted financial document data to obtain document data to be encrypted;
[0010] Step S4: Perform homomorphic structured encryption on the data to be encrypted, and sort the encryption results by data clustering time to obtain a set of ciphertext sliding window data blocks; encapsulate the set of ciphertext sliding window data blocks with chain-side security consensus to obtain verifiable consensus blocks;
[0011] Step S5: Perform multi-label traceability control archiving on verifiable consensus blocks to obtain traceability label archives; use the jump point trust transmission graph to execute jump point transmission of the traceability label archives to the enterprise private database to perform data archiving tasks.
[0012] Optionally, step S1 specifically includes:
[0013] Step S11: Collect enterprise financial data transmission logs through IoT sensing devices;
[0014] Step S12: Standardize the data packet format of the enterprise financial data transmission log, and remove duplicate data packets, broadcast test packets and format-forged data from the standardization results to obtain standardized transmission data;
[0015] Step S13: Construct a dynamic trust assessment model for device nodes based on the device identification code information extracted from standardized transmission data, historical communication stability records, and edge authentication corresponding latency, and then use the dynamic trust assessment model for device nodes to assess the trust score of device nodes.
[0016] Step S14: Perform trust score matching between the standardized transmission data and the corresponding device node trust score, filter the data packets corresponding to devices with a trust score higher than 0.75, and obtain the certified financial data packets;
[0017] Step S15: Cluster the certified financial data packets according to timestamp and business number, write them into the cache linked list structure of the edge computing node in the Internet of Things, and assign a digest root hash to each batch of cluster sorting results to obtain the trusted certified device data.
[0018] Optionally, step S15 specifically includes:
[0019] Step S151: Extract the timestamp field and business number field information from the certified financial data package, construct a time-business two-dimensional index structure, and thus obtain the time-business index dataset;
[0020] Step S152: Perform adaptive density clustering based on the time business index dataset, divide the certified financial data packets into business clusters within each time period, and generate a cluster ranking candidate set;
[0021] Step S153: Reorder the data packet set within the cluster sorting candidate set by indexing the data packets and write it into the cache linked list structure of the edge computing node in the Internet of Things to generate the edge linked list cache body;
[0022] Step S154: Perform node-by-node hash digest calculation on the data group within each edge list cache body, and iteratively construct a binary Merkle tree structure to generate a set of cluster digest roots;
[0023] Step S155: Bind each digest root in the cluster digest root set to the first address of the linked list structure corresponding to the cluster, the number of the business cluster group, and the time period information to generate trust authentication device data.
[0024] Optionally, step S2 specifically includes:
[0025] Step S21: Extract the device communication path and device identification number from the trusted authentication device data to construct the initial communication graph dataset;
[0026] Step S22: Based on the initial communication graph dataset, calculate the average hop count and response latency difference between each device node, and construct the trust hop weight matrix;
[0027] Step S23: Using the trust hop point weight matrix as the path cost input and the device node trust score as the node confidence weight, dynamically calculate the minimum trust transmission path between nodes and generate a set of minimum trust transmission paths.
[0028] Step S24: Calculate the relay frequency, average link stability and cumulative trust score of each hop node in the minimum trust transmission path set, filter the high-frequency steady-state node set, and obtain the hop steady-state trust set.
[0029] Step S25: Perform multi-factor hierarchical clustering on the steady-state trust set of the jump point, select path relay points, and generate a set of trust nodes for the jump point;
[0030] Step S26: Using the set of trusted hop nodes as the core nodes, and combining the path connectivity and node bandwidth load in the initial communication graph dataset, construct the hop path graph structure, thereby outputting the hop trusted transmission graph.
[0031] Optionally, step S24 specifically includes:
[0032] Step S241: Deconstruct each path in the minimum trust transmission path set, count the number of times each node appears in all paths, and use this number of appearances as the node relay frequency to construct a relay frequency vector table;
[0033] Step S242: Calculate the historical packet loss rate, average RTT, and number of communication failure retries for each hop node in the initial communication graph dataset, and calculate the average link stability index of the hop node to generate a stability score table.
[0034] Step S243: Backtrack the number of times each hop node in the initial communication graph dataset has participated in the communication process in the trusted authentication device data, and calculate the cumulative trust score by combining the device node trust score, thereby constructing a trust score matrix;
[0035] Step S244: Normalize and integrate the relay frequency vector table, stability score table and trust integral matrix to generate a jump point node steady-state score vector, and set the steady-state score threshold to 0.8 to filter the high-frequency steady-state node set.
[0036] Step S245: Perform topology consistency verification on the corresponding physical locations of the high-frequency steady-state node set in the minimum trust transmission path set, remove nodes with high concentration, and obtain the jump point steady-state trust set.
[0037] Optionally, the structure enhancement of the multi-category fields of the invoice in step S3 specifically includes:
[0038] The trusted financial invoice data is classified and pre-parsed to obtain the original invoice field metadata set;
[0039] Based on the image file header structure, field label confidence level, and field storage format in the original invoice field metadata set, the original invoice field metadata set is divided into paper image invoice data, OCR-recognized invoice data, and electronic invoice data.
[0040] The paper image ticket data is structured into paper fields to obtain an image structured field table;
[0041] Perform field type re-identification and semantic correction on OCR-recognized invoice data to generate a semantically enhanced field dataset;
[0042] The fields in the image structured field table, semantically enhanced field dataset, and electronic invoice data are aligned and mapped to construct a general invoice field standard template table. Based on the invoice category label, redundant field clipping and structural standardization are performed to obtain the structured invoice field template.
[0043] Based on the structured invoice field template, the original invoice field metadata set is processed by field replacement, format completion, and value verification to obtain the invoice data to be encrypted.
[0044] Optionally, the data clustering time sorting in step S4 is specifically as follows:
[0045] Batch tags are extracted from the homomorphic structured encrypted invoice ciphertext data, and the corresponding invoice generation timestamp, business identifier and source device ID are extracted to construct an initial encrypted invoice index dataset;
[0046] The initial encrypted ticket index dataset is divided into time windows based on the timestamp field. The sliding window length is set to 10 minutes and the sliding step is 5 minutes. The ticket data is then grouped into windowed ciphertext cluster data groups.
[0047] Hash values are calculated based on the business identifiers in the windowed ciphertext cluster data group, and cluster sorting is performed based on the hash value calculation results to obtain sorted ciphertext cluster data.
[0048] Based on the bill amount field, source device ID and bill generation timestamp in the sorted ciphertext cluster data, perform cluster digest extraction operation, bind the corresponding ciphertext cluster, and generate time series indexing structure;
[0049] Each time series indexing structure is encapsulated into a sliding window data block, generating a set of encrypted sliding window data blocks.
[0050] Optionally, the chain-side security consensus encapsulation in step S4 specifically involves:
[0051] Extract the sliding window number, time range, number of tickets in the cluster, cluster digest hash value, and data source device ID from each encrypted sliding window data block to construct an encrypted sliding window information index table;
[0052] The encrypted sliding window information index table is used to perform encapsulation preprocessing on each encrypted sliding window data block to obtain a structured encapsulated draft body.
[0053] The chain-side edge computing nodes are screened from the jump point trust node set, and the screened chain-side edge computing nodes are used as consensus participants. The content of the structured encapsulated draft is verified to obtain the verified encapsulated draft.
[0054] The Byzantine voting consensus is executed on the encapsulated draft body that passes the verification, and the voting results, the signatures of the participating nodes and the consensus round parameters are encapsulated into a consensus encapsulation body.
[0055] For each consensus encapsulation, bind content verification and Byzantine consensus anchoring parameters to obtain verifiable consensus blocks;
[0056] Verifiable consensus blocks are written into the edge chain structure in the Internet of Things, and the corresponding chain-side block index table in the edge chain structure is updated.
[0057] Optionally, step S5 specifically includes:
[0058] Step S51: Extract the sliding window number, timestamp, service identifier, and device ID from the verifiable consensus block, and construct a consensus block service mapping index table;
[0059] Step S52: Based on the field characteristics of the consensus block business mapping index table, divide the field labels to obtain a multi-label identifier set, and embed the multi-label identifier set into the structure of the verifiable consensus block to obtain the traceability label archive.
[0060] Step S53: Select highly stable transmission relay nodes from the hop point trust transmission graph to construct the enterprise private link path structure;
[0061] Step S54: The traceability tag archive is distributed and transmitted to the enterprise private database according to the enterprise private link path structure to perform the data archiving task.
[0062] This invention addresses the characteristics of enterprise financial data transmission in IoT scenarios, such as diverse data transmission paths, heterogeneous source devices, complex communication protocols, and dynamically changing node topologies. It designs a comprehensive security management mechanism covering the entire data transmission chain from source data acquisition to final archiving and storage. Key settings and parameter configurations in its steps demonstrate significant technical effectiveness. In the data access phase, a dynamic trust assessment matrix is constructed to comprehensively score different types of IoT terminal devices within the enterprise based on factors such as access frequency, historical data quality, firmware version, registration source, and communication status. A trust threshold of 0.75 is set; devices below this threshold require re-registration and identity verification, thus blocking potential data injection from abnormal devices at the source. This method enhances the reliability of device-level data acquisition, ensuring basic security for inbound data in the transmission link. In the data standardization and structuring phase, addressing issues such as diverse fields, inconsistent naming, and variable positions in invoice image data, the system uses 31 main invoice types from a pre-set template library and their corresponding field layout features. Combined with logical table matrices extracted from the image-text structure segmentation map and character positioning tags, a field remapping operation is performed to form a unified set of field tags. The field label set is divided using a multi-level namespace, encompassing three dimensions: field type, business semantics, and data unit. This enhances the field recognizability and reproducibility during subsequent data encryption and indexing processes. In the encrypted document clustering stage, a hash function is introduced to generate a fixed-length 256-bit hash value for hashing and classifying the business identifier field in the encrypted data, avoiding grouping biases of similar documents due to differences in business naming. During clustering, a sliding window length of 10 minutes and a sliding step of 5 minutes are set to maintain the rationality of data clustering in terms of temporal continuity. The number of documents within each window is limited to between 200 and 800, ensuring the clusters are processable and representative, and effectively controlling the processing load caused by single-batch data volumes. In the sorting and cluster summary extraction stage, a three-dimensional feature structure, namely <timestamp, source device ID, amount field>, is designed to construct the intra-cluster index list. When timestamps are duplicated, they are prioritized for ascending order by device ID, thus ensuring the sequentiality and retrieval of documents within the cluster. The summary feature employs a hierarchical summary strategy. First, it aggregates floating ranges based on the amount field, dividing the ranges into [0–1000], [1000–5000], [5000–20000], and [20000+]. Then, it combines the geographic identifier of the source device ID to construct a cluster-level summary index, facilitating subsequent business traceability analysis. During the consensus encapsulation phase, based on the pre-consensus structured encapsulation draft, the system selects trusted nodes in the edge chain that meet the following criteria: stability score ≥ 0.87, online rate ≥ 95%, and average response latency ≤ 200ms, to participate in the verification process. The consensus adopts a 5-round Byzantine signature voting mechanism. The initial node pool consists of 15 nodes, and at least 9 nodes must achieve signature consistency in each round to generate the consensus encapsulation.Each encapsulation contains a consensus timestamp, a list of participating node signatures, a verification digest hash, and a consensus round number, ensuring the encapsulation process is verifiable, traceable, and auditable. During the archiving phase, the system constructs verifiable consensus blocks and selects highly stable relay nodes with an average packet loss rate ≤2% and an average bandwidth ≥20Mbps from the hop-point trust transmission graph to form private link paths. The link hop count is limited to no more than 5 hops to reduce data forwarding latency, and the optimal path is selected for hop-point transmission using a bandwidth-weighted path cost function. The final verifiable consensus block embeds multi-label business fields, such as sliding window number, time range, device identifier, and hash digest index, before being transmitted to the chain. It also interfaces with a field retrieval mapping table in the enterprise's local private database, achieving highly consistent and efficient data archiving and deployment. In summary, through systematic design across multiple technical dimensions such as structural design, parameter configuration, and path planning, this invention not only achieves a secure encapsulation and storage mechanism for end-to-end financial data in IoT scenarios, but also improves data adaptability and fault tolerance between processing stages through fine-grained data structure construction, significantly enhancing the system's stability, security, and business implementation capabilities.
[0063] Optionally, this specification also provides an IoT-based enterprise financial data secure storage management system for executing the IoT-based enterprise financial data secure storage management method described above. The IoT-based enterprise financial data secure storage management system includes:
[0064] The cache chain node anchoring module is used to acquire enterprise financial data transmission logs based on the Internet of Things, and to perform data cache chain node trust anchoring based on the enterprise financial data transmission logs to obtain trusted authentication device data.
[0065] The jump point path analysis module is used to construct a minimum set of trust paths based on the data from the trust authentication device, and to analyze the jump point trust transmission nodes of the enterprise's financial data transmission information using the minimum set of trust paths to obtain a jump point trust transmission map.
[0066] The bill field enhancement module is used to collect financial bill data from multiple parallel channels via trusted authentication devices to obtain trusted financial bill data; and to enhance the structure of multiple bill category fields in the trusted financial bill data to obtain the bill data to be encrypted.
[0067] The structured encryption module is used to perform homomorphic structured encryption on the data to be encrypted, and to sort the encryption results by data clustering and time to obtain a set of ciphertext sliding window data blocks; the set of ciphertext sliding window data blocks is then encapsulated with chain-side security consensus to obtain a verifiable consensus block.
[0068] The traceability archiving module is used to perform multi-label traceability control archiving on verifiable consensus blocks to obtain traceability label archives; the traceability label archives are then transmitted to the enterprise's private database using the jump point trust transmission graph to perform data archiving tasks.
[0069] The present invention relates to an IoT-based enterprise financial data secure storage management system. This system can implement any of the IoT-based enterprise financial data secure storage management methods of the present invention. It serves as a medium for the operation and signal transmission between various modules to complete the IoT-based enterprise financial data secure storage management method. The modules within the system cooperate with each other, thereby improving the data adaptability and fault tolerance between processing stages, and significantly enhancing the system's stability, security, and business implementation capabilities. Attached Figure Description
[0070] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0071] Figure 1 This is a flowchart illustrating the steps of the IoT-based enterprise financial data secure storage and management method of the present invention.
[0072] Figure 2 This is a detailed flowchart of step S1 in the present invention;
[0073] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0074] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0075] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.
[0076] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0077] To achieve the above objectives, please refer to Figures 1 to 2 This invention provides a method for secure storage and management of enterprise financial data based on the Internet of Things, the method comprising the following steps:
[0078] Step S1: Obtain enterprise financial data transmission logs based on the Internet of Things, and perform data cache chain node trust anchoring based on the enterprise financial data transmission logs to obtain trusted authentication device data;
[0079] In this embodiment, within an enterprise financial data IoT environment, communication log data between various smart sensing devices and financial terminals in the network is periodically collected. A log snapshot is generated every 5 minutes. The log format includes the sending device ID, receiving device ID, timestamp, data load type, and CRC check information. These logs are mapped to a five-dimensional structure matrix L[i][j][t][p][c], where i is the sending node number, j is the receiving node number, t is the time window, p is the data load parameter, and c is the communication stability index. Next, a trust score is generated based on log frequency, data consistency, and communication history stability, forming a trust vector T[i]. When a device node's trust score threshold is higher than 0.87, it is designated as a trusted authentication device, and a structured dataset D_trust={ID,Cert,TS,L_history} is generated, where Cert represents the device's digital authentication information, TS is the latest communication timestamp, and L_history is the device's 10 most recent valid communication records. This dataset serves as the foundation for subsequent trust path construction.
[0080] Step S2: Construct a minimum trust path set based on the trust authentication device data, and use the minimum trust path set to analyze the jump point trust transmission nodes of enterprise financial data transmission information to obtain the jump point trust transmission map;
[0081] In this embodiment, based on the obtained D_trust data, a device relationship graph structure G=(V,E,W) is first constructed, where V represents all certified device nodes, E represents the communication connection edges between them, and W represents the trust score weight of each edge. To establish the minimum trust path set, the minimum weight transmission structure extraction method is adopted to calculate the path cost matrix C[i][j]=1 / W[i][j] between any two certified devices; where W[i][j] represents the trust weight value between node i and node j. This weight is a positive number, and the larger the value, the more reliable and stable the communication between node i and j; C[i][j] represents the transmission cost value between node i and node j. It is the reciprocal of the trust weight and is used to reflect the path cost or "risk". Thus, the minimum transmission cost search is performed between all devices to obtain a set of hop transmission paths P_min={P1,P2,...,P...} that satisfy the path trust score greater than 0.85. n Each path is represented by a chain of device IDs, such as P1={D1→D5→D9}. Further, all paths are merged to construct a jump point trust graph structure H=(N,E',M), where N is the set of jump point devices, E' is the jump point connection edge, and M is the trust path index mapping table, supporting fast path structure lookup and improving the efficiency of subsequent data jump point scheduling.
[0082] Step S3: Using the data from the trusted authentication device, perform multi-channel parallel financial document data collection on the Internet of Things to obtain trusted financial document data; perform multi-category field structure enhancement on the trusted financial document data to obtain document data to be encrypted;
[0083] In this embodiment, based on the aforementioned trust authentication device set D_trust and trust graph H, a parallel data acquisition process is initiated in the Internet of Things (IoT). A data acquisition port based on the MQTT protocol is simultaneously activated for each trust authentication device to monitor the real-time output of the local invoice printing terminal and scanning device. Each data segment contains fields such as invoice number, invoice amount, taxpayer identification number, invoice time, and product name, initially in JSON format. These fields are then parsed into vector form. A data table B_total for invoices was constructed. To address the differences in invoice types, structural enhancement was used to strengthen the identification of the invoice type, amount type, and tax identification type fields, adding a new category marker field C. i And restructure the data structure as Data augmentation also includes anomaly field repair and field redundancy compression. The final dataset B_enc_ready={B_ext[i]}_{i=1}^{k} of the tickets to be encrypted is generated for the next step of structured encryption processing.
[0084] Step S4: Perform homomorphic structured encryption on the data to be encrypted, and sort the encryption results by data clustering time to obtain a set of ciphertext sliding window data blocks; encapsulate the set of ciphertext sliding window data blocks with chain-side security consensus to obtain verifiable consensus blocks;
[0085] In this embodiment, the generated B_enc_ready dataset is fed into the structured encryption module, and each record is grouped by field and homomorphic encryption is performed. Let the vector B_ext[i]=[f1,f2,...,f6], and the Paillier encryption method is applied to each field to generate the ciphertext vector C_enc[i]=[E(f1),E(f2),...,E(f6)]. All ciphertext vectors are combined into a ciphertext matrix C_total[m][n], where m is the number of tickets and n=6 is the number of fields. The ciphertext matrix is clustered and sorted with a 10-minute sliding window, the sliding window unit is set to 600 seconds, and the maximum sliding offset is 30 seconds, generating an ordered set of ciphertext sliding window data blocks S_win={C_t1,C_t2,...,C_tn}. Each data block is accompanied by a time index and a sliding window label. Then, each data block is encapsulated in an inter-block hash chain structure to construct a data linked list structure L_block={Hash_prev,C_block,Hash_next}. Finally, all L_blocks are encapsulated with chain-side consensus, and a verifiable consensus block set Q_consensus={Q1,Q2,...,Q_m} is generated through signature verification and multi-party cross-verification for final archiving and transmission.
[0086] Step S5: Perform multi-label traceability control archiving on verifiable consensus blocks to obtain traceability label archives; use the jump point trust transmission graph to execute jump point transmission of the traceability label archives to the enterprise private database to perform data archiving tasks.
[0087] In this embodiment, the generated Q_consensus consensus block set is archived. First, the timestamp, field digest, and signature information in the header of each block are parsed to generate a unique traceability tag structure Tag_i={Q_ID,TS,Sign,Field_Hash} for each block. Each block and its corresponding tag are then merged to form an archive body Archive_i={Tag_i,Q_i}, resulting in a multi-tag traceability archive body set A_set={Archive1,Archive2,...,Archive...}. kThe process then invokes the hop-point trust graph H constructed in step S2, selects the shortest trusted path to the target enterprise's private database DB_priv, and generates a transmission route chain R_path={Node1→Node4→Node7→DB_priv} through the hop-point scheduling controller. During each hop, data integrity verification and tag signature secondary confirmation mechanisms are executed to ensure the archive is not tampered with during transmission, ultimately delivering the complete archive to the private database. Data is stored hierarchically in the private database according to tag time index and field summary, supporting subsequent auditing and verification, forming a complete closed-loop secure archiving system.
[0088] Optionally, step S1 specifically includes:
[0089] Step S11: Collect enterprise financial data transmission logs through IoT sensing devices;
[0090] In this embodiment, 23 IoT sensing devices are deployed in a financial information management system. These devices are embedded in the enterprise's intranet, including gateway nodes, invoice scanners, financial printing terminals, approval machines, and virtual expense reimbursement interfaces. Every 20 seconds, these sensing devices capture enterprise financial data transmission logs from the TCP / UDP communication layer and upload them to the local edge computing service via the MQTT protocol. The collected information includes source device ID, target device ID, packet size, transmission timestamp, checksum, protocol type, and link layer handshake information. Each transmission log is approximately 1.5KB in size, with approximately 120,000 logs collected daily. The data collection task employs a FIFO buffering mechanism to avoid packet loss and uses physical network segment address verification to exclude data generated by non-enterprise nodes.
[0091] Step S12: Standardize the data packet format of the enterprise financial data transmission log, and remove duplicate data packets, broadcast test packets and format-forged data from the standardization results to obtain standardized transmission data;
[0092] In this embodiment, the collected financial data transmission logs are standardized according to a custom data packet structure specification, uniformly converted into a seven-tuple format: {DevSrc,DevDst,Time,Size,Checksum,Protocol,UUID}. Duplicate packets (based on comparison with the same UUID and timestamp), packets with all zeros broadcast during network testing, and illegal data packets with abnormal header signature information (e.g., forged HTTP GET packet format without valid header fields) are removed. The cleaning process uses 1000 log entries as a processing unit, performing content comparison within a sliding window, and batch updating every 10 minutes. After standardization, the number of log entries is reduced by an average of 15%, and the average dataset after cleaning is approximately 102,000 entries per day.
[0093] Step S13: Construct a dynamic trust assessment model for device nodes based on the device identification code information extracted from standardized transmission data, historical communication stability records, and edge authentication corresponding latency, and then use the dynamic trust assessment model for device nodes to assess the trust score of device nodes.
[0094] In this embodiment, a three-factor dynamic trust evaluation framework for devices is constructed based on the device identification code, communication records, and authentication response time extracted from standardized transmission data. This framework uses a scoring formula of the form T[i] = α × S[i] + β × R[i] + γ × D[i], where: S[i] represents the success rate of device i communicating with other devices (e.g., 0.98), R[i] represents the average authentication response delay in the last 10 communications (e.g., 120ms), D[i] represents the number of times the device has been downgraded due to anomalies in the past week (e.g., 2 times), and the weights of α, β, and γ are set to 0.5, 0.3, and 0.2 respectively. Substituting these parameters, a trust scoring matrix is output. (e.g., T=[0.91,0.83,0.67,...]). The scoring results are updated daily and cached on the edge nodes for local storage, ready for subsequent filtering.
[0095] Step S14: Perform trust score matching between the standardized transmission data and the corresponding device node trust score, filter the data packets corresponding to devices with a trust score higher than 0.75, and obtain the certified financial data packets;
[0096] In this embodiment, based on the results in the aforementioned scoring matrix T, all standardized data records are associated with the trust scores of the device nodes they generate. A scoring threshold of 0.75 is used as a filtering boundary to select records with a device trust score ≥ 0.75; the corresponding data packets are then retained as authentication financial data packets. For example, in a single batch processing session, a total of 40,000 data packets are processed, of which approximately 28,000 meet the filtering criteria and are retained for subsequent clustering and caching operations. The matching operation is completed in the SQL in-memory database of the edge nodes. The matching fields include the device identification code and UUID, and index optimization controls the query matching to be completed within 100 milliseconds.
[0097] Step S15: Cluster the certified financial data packets according to timestamp and business number, write them into the cache linked list structure of the edge computing node in the Internet of Things, and assign a digest root hash to each batch of cluster sorting results to obtain the trusted certified device data.
[0098] In this embodiment, the aforementioned certified financial data packets are initially sorted in ascending order according to timestamps and then grouped and clustered based on the business number field. Within each cluster, the data is sorted in a two-level order of "business number + timestamp", and each group constitutes a data batch. Subsequently, each batch is written into a cache linked list structure in the edge computing node, whose structure is defined as struct PacketNode{charUUID
[32] ;longts;charpayload
[2048] ;PacketNode*next;}, and each batch is managed using a head node. After sorting, a Merkle hash tree digest is generated for the UUID field in each batch, and the root hash value is extracted as the digest authentication code for that batch. The final generated data structure includes a linked list cache area and a digest code set, realizing batch ordered caching and fast verification of trusted authentication device data.
[0099] Optionally, step S15 specifically includes:
[0100] Step S151: Extract the timestamp field and business number field information from the certified financial data package, construct a time-business two-dimensional index structure, and thus obtain the time-business index dataset;
[0101] In this embodiment, the "timestamp field (ts)" and "business number field (bn)" are selected as key index dimensions in the authenticated financial data packet to construct a two-dimensional sparse index matrix structure. The index structure is in the form of IndexMap[T][B]={PktID1,PktID2,...,PktIDn}, where T represents the time period (divided in 5-minute increments, such as 14:00–14:05, 14:05–14:10), and B represents the business number (such as INV-2041, EXP-0039). Each cell stores the set of packet IDs corresponding to the business number within that time period. The index structure is implemented using a dictionary-nested hash table: the outermost layer is segmented by timestamp, and the inner layer is grouped by business number. A single index construction can process approximately 50,000 data packets, with a construction time of no more than 1.2 seconds, and can support on-demand rapid location of business traffic within a specific time window.
[0102] Step S152: Perform adaptive density clustering based on the time business index dataset, divide the certified financial data packets into business clusters within each time period, and generate a cluster ranking candidate set;
[0103] In this embodiment, adaptive density clustering is applied to the data distribution features extracted from the time-based service index dataset. First, the data packet density corresponding to each service number within each time period is calculated, using the formula ρ_{T,B}=Count(Pkts) / ΔT, where ΔT is 5 minutes and Count(Pkts) is the number of data packets within that time period. If the density exceeds a set threshold (e.g., more than 15 packets per second), the time-based service pair is marked as a high-density cluster core and expands towards adjacent time periods and the same service number, forming a service cluster group. For example, service INV-2041 has 510 packets between 14:00 and 14:05, thus being identified as a cluster core. The system continues to detect related data between 14:05 and 14:10 and between 14:10 and 14:15. If the density remains consistently high, a time-based cluster is formed. This process ultimately outputs a cluster ranking candidate set, where each cluster record includes the cluster group number, time window range, service number, and number of packets.
[0104] Step S153: Reorder the data packet set within the cluster sorting candidate set by indexing the data packets and write it into the cache linked list structure of the edge computing node in the Internet of Things to generate the edge linked list cache body;
[0105] In this embodiment, the data packets in each group of the cluster sorting candidate set are numbered and rearranged. The rearrangement method is based on the rule of "cluster number priority + timestamp order". That is, within each cluster, the packet IDs are sorted in ascending order of timestamp, and at the same time, the data is sorted into the chain in ascending order of cluster number between different clusters. The sorted data is written into a dedicated linked list structure cache in the memory of the edge computing node. The linked list structure is defined as: struct ClusNode{charPktID
[32] ;longts;intorder;ClusNode*next;}, and each cluster corresponds to a linked list head pointer, and the cache is built in order. Each linked list construction process does not exceed 20 clusters, each containing about 200-500 packets, and the memory overhead can be controlled within 30MB. It supports batch-level retrieval and serial reading. Finally, a complete edge linked list cache body set is generated in the node for use in summary generation.
[0106] Step S154: Perform node-by-node hash digest calculation on the data group within each edge list cache body, and iteratively construct a binary Merkle tree structure to generate a set of cluster digest roots;
[0107] In this embodiment, for each data node in the edge linked list cache, SHA-256 hash digests are generated sequentially for the PktID+ts+order fields. Each node generates a digest value in the form H_i=SHA256(PktID||ts||order). Based on the node order, a combination hash is performed on every two node hash digests: H_pair=SHA256(H_i||H_j), iteratively building the upper-level hash, ultimately completing the bottom-up binary Merkle tree structure. If the number of linked list nodes is n, the tree depth is ⌈log2n⌉. During the digest generation process, each cluster generates a Merkle tree, with the root node being the digest root of that cluster. Taking three samples as an example, clusters A, B, and C generate root hashes: Root_A=0xabc..., Root_B=0xdef..., Root_C=0x987..., forming a cluster digest root set, which is subsequently used for authentication binding.
[0108] Step S155: Bind each digest root in the cluster digest root set to the first address of the linked list structure corresponding to the cluster, the number of the business cluster group, and the time period information to generate trust authentication device data.
[0109] In this embodiment, each digest root hash value is structurally bound to the first address of the linked list of its cluster (in the form of a memory pointer, such as &ClusA_Head), the business number (such as EXP-0042), and the time range (such as 14:05–14:15) to construct the final trust authentication device data structure. The structure is defined as: struct TrustClusMeta{charRootHash
[64] ;void*ListHeadAddr;charBizID
[16] ;charTimeRange
[32] ;}; The bound structure is stored in the trust cache pool of the edge node, generating about 120–200 cluster digest records per day. The trust cache pool of the edge node is a highly available, limited-access structured memory block set deployed inside the IoT edge computing device, specifically used for temporarily storing intermediate result data of the completed trust authentication process. The cache pool adopts a multi-level memory partition structure, with each partition corresponding to a time window (such as 15 minutes). Within each time partition, logical sub-blocks are established according to the business number (such as INV-1042). Each sub-block contains several TrustClusMeta structure instances (see supplementary structure binding content below). Edge nodes are configured with a hardware-encrypted authentication mechanism, allowing access to cache pool memory data only through pre-registered device identity credentials (such as TPM certificates or pre-configured HMACs), preventing external tampering or unauthorized reading. The digest hash of all trust authentication data is registered as a Merkle tree root when written to the cache pool, and the digest is periodically uploaded to the upper-layer master node as an anchor point for integrity verification. The cache pool supports storing authentication data for up to the past 12 hours; content exceeding this period is automatically marked as reclaimable and triggers archiving. This structure supports both rapid data integrity verification via hash checking and query indexes based on business number or time range, ensuring data verifiability and cluster consistency. All binding information is synchronously backed up to the master node's solid-state encrypted storage, forming a complete trust authentication data graph.
[0110] Optionally, step S2 specifically includes:
[0111] Step S21: Extract the device communication path and device identification number from the trusted authentication device data to construct the initial communication graph dataset;
[0112] In this embodiment, the communication paths and device identification numbers recorded in the trusted authentication device data are extracted, and the source and target device pairs for each hop in the communication path are parsed out. For example, in a communication record "DevA→DevB→DevC", the edge set {(DevA,DevB),(DevB,DevC)} is extracted, and combined with the unique identification number of each device (such as a 16-bit hash format ID: e.g., A3F2:9B7E:...) to establish an initial communication graph. This graph is represented by a directed graph G=(V,E), where V is the set of devices and E is the set of communication path edges. Each edge is accompanied by basic communication frequency and communication result status code information. The entire graph data is stored in the edge side graph database, serving as the basis for subsequent path deduction and hop selection.
[0113] Step S22: Based on the initial communication graph dataset, calculate the average hop count and response latency difference between each device node, and construct the trust hop weight matrix;
[0114] In this embodiment, the edge communication frequency and status code stored in the initial communication graph are combined with the response time logs returned by the edge nodes to analyze all edges (e... ij Calculate the average hop count H[i][j] and latency difference ΔT[i][j]. Combining these two factors, the trust hop weight matrix W is defined as follows: W[i][j] = α·H[i][j] + β·ΔT[i][j], where α = 0.6 and β = 0.4 are weighting factors that measure path stability and response efficiency. This matrix is stored in a two-dimensional symmetric structure with a size of N×N (N is the number of device nodes), representing the communication cost between devices, providing a quantitative basis for subsequent trust path calculation.
[0115] Step S23: Using the trust hop point weight matrix as the path cost input and the device node trust score as the node confidence weight, dynamically calculate the minimum trust transmission path between nodes and generate a set of minimum trust transmission paths.
[0116] In this embodiment, the trust hop weight matrix W is used as the path cost input, and the trust score matrix S of each device (in the form S=[0.92,0.78,0.95,...], corresponding one-to-one with the device number) is also referenced, using the device node trust level as the node confidence factor. In the transmission path calculation, all possible paths are traversed, prioritizing combinations with low path weight costs and high node trust scores, and filtering layer by layer to find the minimum trust transmission path between devices. The final output path set is stored in a triplet structure: (start point, end point, [relay sequence]), forming the minimum trust transmission path set P. trust .
[0117] Step S24: Calculate the relay frequency, average link stability and cumulative trust score of each hop node in the minimum trust transmission path set, filter the high-frequency steady-state node set, and obtain the hop steady-state trust set.
[0118] In this embodiment, the statistical path set P is... trust The relay frequency N of each hop node relay Furthermore, considering the communication stability S between this hop point and the preceding and following nodes... link (Calculated as the percentage of stable and successful communications), and the cumulative trust score ΣS of the jump point itself. node (The calculation method is Rsum = w1·Rs + w2·Rh + w3·Ri; where w1 = 0.4, w2 = 0.3, w3 = 0.3; Rs is the communication success rate; Rh is the historical trust score; Ri is the stable relay index, which is whether it appears continuously in multiple different paths and whether its relay position in the path is relatively core, composed of path coverage and relay concentration). A jump point performance table is constructed based on the statistical indicators, with each record as follows: {NodeID:A3F2...,N relay :25,S link :0.91,ΣS node :4.27}. By setting a threshold N relay ≥10, S link ≥0.85、ΣS node ≥3.5 Filter out the set of high-frequency stable jump point nodes, denoted as T stable .
[0119] Step S25: Perform multi-factor hierarchical clustering on the steady-state trust set of the jump point, select path relay points, and generate a set of trust nodes for the jump point;
[0120] In this embodiment, T stable The device nodes in the cluster undergo multi-factor hierarchical clustering. The three indicators mentioned above are standardized and projected onto a three-dimensional space to form a vector point set. Grouping is performed based on a hierarchical method with a cluster radius of 0.15, and hop nodes with high centrality are marked as core trusted relay nodes. Within each group, the device node closest to other nodes is selected as the final hop node, generating a hop node trusted node set T. core This is used to ensure the stability and reliability of subsequent data transmission.
[0121] Step S26: Using the set of trusted hop nodes as the core nodes, and combining the path connectivity and node bandwidth load in the initial communication graph dataset, construct the hop path graph structure, thereby outputting the hop trusted transmission graph.
[0122] In this embodiment, T coreAs the core of constructing the jump point trust graph, the initial edge set recorded in the initial communication graph dataset is used to select edges with high connectivity (average communication frequency greater than 5) and available bandwidth (greater than 2MB / s) to the core nodes to construct the jump point path graph. This graph is stored as an adjacency matrix A[i][j], where A[i][j]=1 indicates that there is a trust path with sufficient bandwidth and frequency from device i to j. The final output is the jump point trust transmission graph structure G. trust =(T core E filtered This serves as the structural basis for subsequent data hop point transmission and security path selection.
[0123] Optionally, step S24 specifically includes:
[0124] Step S241: Deconstruct each path in the minimum trust transmission path set, count the number of times each node appears in all paths, and use this number of appearances as the node relay frequency to construct a relay frequency vector table;
[0125] In this embodiment, by traversing the set of minimum trusted transmission paths, the device nodes of each path are decomposed into a sequentially numbered node queue, and the cumulative number of times each node appears in all paths is counted to form a relay frequency vector table. Taking 45 device nodes participating in path forwarding in the enterprise financial communication process as an example, numbered N1–N45, if nodes N7, N12, and N30 appear in 38, 42, and 26 paths respectively, their corresponding values in the relay frequency vector are 38, 42, and 26 respectively. This vector data is stored in a two-dimensional array RelayFreq[i], where i represents the node number index.
[0126] Step S242: Calculate the historical packet loss rate, average RTT, and number of communication failure retries for each hop node in the initial communication graph dataset, and calculate the average link stability index of the hop node to generate a stability score table.
[0127] In this embodiment, to calculate the average link stability of hop nodes, the communication link information associated with each node in the initial communication graph dataset is called, and its historical packet loss rate (PktLoss), average round-trip time (RTT), and number of communication failure retries (RetryTimes) are extracted. Specifically, the associated link records for node N12 show an average packet loss rate of 1.8%, an average RTT of 65ms, and an average number of retries of 1.4 times per communication. The stability score is calculated using the formula StabScore[i] = (1 - PktLoss[i]) × (1 / (1 + RTT[i])) × (1 / (1 + RetryTimes[i])). In the example above, the stability score for N12 is approximately 0.497, and the result is written into the stability score table StabScore[i].
[0128] Step S243: Backtrack the number of times each hop node in the initial communication graph dataset has participated in the communication process in the trusted authentication device data, and calculate the cumulative trust score by combining the device node trust score, thereby constructing a trust score matrix;
[0129] In this embodiment, for each hop node, the total number of data packets it participated in communication is retrieved from the trust authentication device data, and its historical trust score is combined to accumulate points. Taking node N7 as an example, it appears 67 times in the authentication log, with an average trust score of 0.83. Therefore, the accumulated trust score is 67 × 0.83 = 55.61, which is used as the trust score input value for node N7 to construct the trust score matrix T[i], where i is the hop node index, and T[i] represents the trust score of the i-th node.
[0130] Step S244: Normalize and integrate the relay frequency vector table, stability score table and trust integral matrix to generate a jump point node steady-state score vector, and set the steady-state score threshold to 0.8 to filter the high-frequency steady-state node set.
[0131] In this embodiment, the relay frequency vector RelayFreq[i], stability score StabScore[i], and trust integral matrix T[i] obtained above are uniformly normalized. After being mapped to the [0,1] interval using a min-max normalization method, they are weighted and merged according to proportional weights (e.g., frequency weight is 0.4, stability weight is 0.3, and integral weight is 0.3) to generate the steady-state score vector StabilityVector[i] of the jump point node. For example, the normalized frequency of node N12 is 0.89, stability is 0.52, and trust integral is 0.94, so its steady-state score is 0.4×0.89+0.3×0.52+0.3×0.94=0.786. The steady-state score threshold is set to 0.8, and all nodes are filtered to select those with scores higher than the threshold to form a high-frequency steady-state node set.
[0132] Step S245: Perform topology consistency verification on the corresponding physical locations of the high-frequency steady-state node set in the minimum trust transmission path set, remove nodes with high concentration, and obtain the jump point steady-state trust set.
[0133] In this embodiment, a topology consistency check is further performed on the formed high-frequency steady-state node set. This involves calculating the physical concentration of nodes based on their positions within the minimum trust transmission path set. If a node, such as N12, is located in a densely populated position in the middle of multiple paths, with a probability significantly higher than the network average density standard (e.g., exceeding three times the network average topology density), this node is removed to prevent relay aggregation from causing single-point bottlenecks or communication link congestion. After removing nodes with excessively high concentration, the resulting node set is the hop-point steady-state trust set, which can be used as the core node set for constructing the trust path graph.
[0134] Optionally, the structure enhancement of the multi-category fields of the invoice in step S3 specifically includes:
[0135] The trusted financial invoice data is classified and pre-parsed to obtain the original invoice field metadata set;
[0136] Based on the image file header structure, field label confidence level, and field storage format in the original invoice field metadata set, the original invoice field metadata set is divided into paper image invoice data, OCR-recognized invoice data, and electronic invoice data.
[0137] In this embodiment, the raw document field metadata set is read from the trusted financial document data. Each data record includes a file identifier, document source path, field location coordinates, field confidence score (0–1 floating-point value), field label, original field value, and the document type to which the field belongs. During the pre-parsing stage, the metadata set is traversed to analyze the field extraction method and source structure of each record. Specifically, if a field record contains TIFF, JPEG, or PNG format image header information, and the field location is represented by coordinates but lacks OCR annotation information, it is classified as paper image document data; if the field annotation confidence score is less than 0.7 and the source is OCR engine output, it is classified as OCR-recognized document data; and fields with an original structure of structured XML or JSON format and field values with timestamps and field identifier labels are classified as electronic document data. This classification result is written to three data cache queues for subsequent processing.
[0138] The paper image ticket data is structured into paper fields to obtain an image structured field table;
[0139] In this embodiment, for paper image document data, a preset field location and contour extraction process is invoked, and a fixed pixel interval scanning mode is used to read the field frame region within the image area. Combining image grayscale gradient changes and coordinate difference information, a field position relationship graph is constructed based on 3×3 neighborhood blocks. Each graph node contains the coordinates of the upper left corner of the image region where the field is located, the region width and height, pixel density, and the average grayscale distribution. Subsequently, graph structure matching is performed on the position and arrangement relationships between adjacent nodes to extract key field sets such as table header, table body, and signature area, and output an image structured field table. Each field item in the table is presented in the form of a five-tuple: <field name, position coordinates, pixel frame value, normalized image content summary, recognition level>, which is used for subsequent field mapping.
[0140] Perform field type re-identification and semantic correction on OCR-recognized invoice data to generate a semantically enhanced field dataset;
[0141] In this embodiment, for OCR recognition of invoice data, the original recognition result, language type, and field context window are first extracted from each field record. A field type recognition table is constructed, employing a language rule matching structure with contextual collaborative recognition capabilities. Field types are corrected through field context label matching and content semantic analysis. The language rule matching structure is established through statistical analysis and pattern mining of a large amount of historical invoice text data. First, an OCR recognition text dataset containing various invoice types is collected. Each text contains field content and its contextual information, such as the preceding and following field names, field position order, typical vocabulary, and grammatical features. Based on this data, a contextual collaborative relationship table is constructed using "field context pairs" as the basic unit. This relationship table is a multi-dimensional matrix, where rows represent candidate field names, columns represent possible contextual field labels, and matrix elements are the co-occurrence frequency and conditional probability weights between fields. The weight values range from 0 to 1, reflecting the matching strength of field names in a specific context. The matching structure consists of three parts: 1) A field dictionary set: containing all standard field names and their common synonyms, aliases, and spelling variations, such as "amount," "total," and "total amount"; 2) A context adjacency matrix: a two-dimensional matrix representing the probability of a field collaborating with its preceding and following neighboring fields. The matrix dimension is N×N, where N is the size of the field dictionary, and the element value represents the conditional probability of field i in the context of field j, such as P(field i|field j); 3) A semantic rule base: stored in the form of condition-action rules, such as "If field A appears in the first two rows of field B and the content of field A conforms to the amount format, then field A is classified as the 'payment amount' field." The rules include regular expression matching and positional constraints on the field content. For example, fields like "Zhang San" and "Li Si" can be re-identified as "payee" type fields if they are located in the first five rows of the "amount" field; "20XX year XX month" can be uniformly corrected to the "ticket issuance date" field. After processing, the semantically enhanced field dataset is output, with the data items in the following format: <Standard field name, corrected field type, original value, semantic label, confidence level>.
[0142] The fields in the image structured field table, semantically enhanced field dataset, and electronic invoice data are aligned and mapped to construct a general invoice field standard template table. Based on the invoice category label, redundant field clipping and structural standardization are performed to obtain the structured invoice field template.
[0143] In this embodiment, field alignment processing is performed on the image structured field table, the semantic enhancement field data set, and the electronic bill data. First, a field reference set is established according to the bill type tags to which the fields belong, and there is a standard field layout table for each bill type. Taking the standard field names as rows and the three data sources as columns, an alignment mapping matrix is constructed, where each element represents whether a certain field exists in the three types of data, whether the content is consistent, and the source confidence level. After alignment, field redundancy trimming is performed on the alignment matrix, removing duplicate field items with a content consistency lower than 70%, retaining field data with consistent field labels and from two or more sources, and unifying the field names, formats, and structures. The processing result is a structured bill field template, and the field structure is represented by a nested table, which includes a main field group, an extended field group, and an audit field marking area.
[0144] Based on the structured bill field template, field replacement, format completion, and value verification processing are performed on the original bill field metadata set to obtain the bill data to be encrypted.
[0145] In this embodiment, taking the structured bill field template as a reference, field replacement, format completion, and numerical verification processing are performed on each record in the original bill field metadata set. In the replacement step, according to the corresponding standard names and values of each field in the field alignment mapping table, the original field names and values are batch-replaced with standard field items; in the format completion process, rule filling is performed on the missing currency unit fields and date format fields (such as adding the prefix "¥" to the amount field and completing the date field in the format of "YYYY年MM月DD日"); in the numerical verification step, by performing numerical range limitation and decimal point precision verification on the amount fields, it is ensured that all amount data is within a legal range (such as the value of the amount field shall not be less than 0 and shall not exceed the maximum limit set for the bill type). Finally, the output bill data to be encrypted is a structured bill data table with unified format, standard fields, and compliant numerical values, and each bill data is stored in a JSON structure and carries a source label, a processing timestamp, and an integrity identification code.
[0146] Especially importantly, the structuring of paper fields specifically includes:
[0147] Perform image type recognition and resolution analysis on the paper image bill data, screen clear bill images, and generate an effective bill image set;
[0148] In this embodiment, when performing image type recognition and resolution analysis on paper image ticket data, multi-dimensional parameters are first extracted for each ticket image, including the image width, height, color depth, resolution (DPI), and image format (such as JPEG, PNG, etc.). All these parameters are stored in an N×5 matrix, where N is the number of ticket images, and each row represents the width (pixels), height (pixels), number of color channels, resolution (DPI), and format encoding of the corresponding image. Based on this matrix, ticket images with acceptable clarity are selected by threshold judgment. The threshold is set to a resolution of not less than 300 DPI and a width and height of not less than 1024×768 pixels to ensure that the image quality meets the requirements of subsequent processing. Blurry images are removed by calculating the Laplacian variance of the image (a variance less than 50 is considered blurry). Finally, the set of valid ticket images after filtering is stored in a multi-dimensional array, and the valid image index is updated synchronously to support subsequent batch processing.
[0149] Perform image and text region partitioning on the valid invoice image set, identify the image and text boundaries and field blocks of the invoice, and obtain an image and text structure segmentation map;
[0150] In this embodiment, image-text region partitioning is performed on the valid ticket image set. This process uses the grayscale matrix of image pixels as a basis and employs multi-scale edge detection and local texture analysis techniques to generate an image-text structure segmentation map. First, the input ticket image is preprocessed and converted into a grayscale image G, where G is a two-dimensional matrix and the element G(i,j) represents the grayscale value of pixel (i,j), ranging from 0 to 255. To enhance edge information, multi-scale Gaussian filtering is used to construct a Gaussian pyramid. Specifically, a set of Gaussian kernel functions {K_σ} is generated, where σ is the scale parameter, selected as σ={1.0,2.0,4.0}, and convolved with the grayscale image G to generate a multi-scale blurred image set {G_σ}. This process can be represented by matrix convolution: G_σ=G*K_σ, where * represents the convolution operation. Based on the multi-scale image set, edge detection is performed on the image at each scale. Edge detection employs an improved Canny edge detection technique. The specific process includes: calculating the image gradient magnitude matrix M_σ and the direction matrix Θ_σ. , ,in The image is defined by its horizontal and vertical gradients. Then, edges are refined using non-maximum suppression (NMS), suppressing points with gradient magnitudes below a threshold T_1 (typically T_1 is set to 20% of the maximum gradient value), preserving edge points to form an edge matrix E_σ. The edge matrix E_σ is then fused across different scales. The fusion method involves maximizing the edge values at corresponding positions across all scales: E_fused(i,j)=max{E_σ(i,j)|σ∈{1.0,2.0,4.0}}. This operation ensures complete preservation of multi-scale edge information, enhancing the unified expression of details and coarse boundaries. Next, combined with the edge information, local texture analysis is performed to distinguish text regions from the background. The Gray-Level Co-occurrence Matrix (GLCM) analysis method is used to calculate texture feature parameters such as contrast, correlation, energy, and homogeneity within a sliding window of the image. Specifically, the image G is divided into blocks, each 16×16 pixels in size. The GLCM matrix P within each block is calculated, and the four texture feature vectors T=[contrast,correlation,energy,homogeneity] are extracted. Based on empirical thresholds, text regions typically exhibit high contrast and low energy. Classification is achieved by setting threshold ranges. Texture analysis results generate a texture mask matrix W, where W(i,j)=1 indicates the location belongs to a text texture region, and 0 indicates a non-text region. The fused edge matrix E_fused is combined with the texture mask W, and a logical AND operation is used to generate the final image-text structure segmentation image S: S(i,j)=E_fused(i,j)∧W(i,j). Pixels with a value of 1 in matrix S are marked as text or image-text boundary regions, while 0 represents the background. After performing image-text region segmentation on the ticket image, a binary image result is obtained, where the foreground part represents possible text or table regions, and the background part represents blank areas. To further identify each independent text or field region, connected component extraction technology is used for structured processing. First, in the preprocessed binary image, each pixel is scanned sequentially. When a foreground pixel (pixel value is black or "1") is encountered, it is determined whether the pixel values in the eight adjacent directions around this pixel also belong to the foreground. If so, these pixels are determined to belong to the same image-text region. All connected pixel blocks are labeled with numbers to ensure that each connected region has a unique number. Then, two scan processes are performed. The first scan traverses the image from the top left to the bottom right, numbering each foreground pixel and recording their connections. The second scan, based on the connections found in the first scan, unifies regions with the same number, merging all connected numbers into a single region number to form the final connected region partitioning map. After labeling, the spatial structure attributes of each connected region are extracted. These attributes include the pixel coordinates of the top left and bottom right corners of the region, the number of pixels occupied by the region (i.e., its area), the width and height of the region, and the density ratio of foreground pixels in the region.Regions that are too small or have too low pixel density will be filtered out to avoid processing invalid content such as blemishes, noise, or stamp fragments. The filtering criteria are preset by the user, such as a minimum of 30 pixels and a pixel density exceeding 0.1. In the remaining regions, they are numbered according to their order of appearance in the image, arranged from top to bottom and left to right, to ensure accurate matching of subsequent field labels. For each connected region, its spatial extent and sequence number are recorded and output as a candidate block for a field, thus obtaining the image-text structure segmentation map.
[0151] The table lines of the bill table in the image-text structure segmentation diagram are reconstructed to obtain a logical table region matrix.
[0152] In this embodiment, the table structure of the invoice is reconstructed based on the table line information identified in the image-text structure segmentation map. The table line information comes from the horizontal and vertical lines in the image-text structure matrix, and the table boundaries are extracted using matrix scanning technology. The extracted line information is represented by two sets of vectors: a set of horizontal lines L_h and a set of vertical lines L_v. Each line is defined by pixel start and end coordinates, such as L_h={(x_start_i,x_end_i,y_i)|i=1,...,m}, L_v={(y_start_j,y_end_j,x_j)|j=1,...,n}. Combining these two sets of vectors, a logical table region matrix M∈{0,1}^{m×n} is constructed, where M(i,j)=1 indicates that there is a valid data area within the cell, and M(i,j)=0 indicates that it is blank. This matrix is generated by scanning the logical table boundary points, ensuring that the boundary of each cell does not overlap with adjacent cells, and supporting multi-level nested table structures.
[0153] Based on the logical table area matrix, perform character area positioning to obtain character positioning coordinates and character area labels;
[0154] In this embodiment, after constructing the logical table region matrix, the character regions within the table cells are precisely located. Character location employs a combination of local pixel density peak detection and connected component analysis. The location result is represented by a character coordinate matrix C, where C∈R^{k×4}, and k is the number of detected characters. Each row contains the x and y coordinates of the top-left corner of the character, as well as the character's width and height. Character location also includes a confidence score ranging from [0,1], with a threshold set to 0.7 to filter out low-confidence recognitions. Corresponding to the coordinate matrix, the character region label set L contains identifiers for the corresponding field types, such as "invoice code," "date," and "amount," achieving precise mapping from characters to fields.
[0155] The system maps character positioning coordinates to character region labels, and performs field label matching based on a preset ticket category template to generate a structured field table for the image.
[0156] In this embodiment, the character positioning coordinate matrix C is combined with the field label set L, and field matching is performed based on a predefined standard template for ticket category fields. The field template is stored in the form of a multi-dimensional matrix T, which includes the field category, the corresponding coordinate range, and the field name. It supports multiple ticket types and can be flexibly expanded. During matching, the field attribution is calculated based on the overlap between the character positioning coordinates and the template coordinate range. The overlap threshold is set to 0.8 to ensure accurate field matching. Through this mapping, a complete image structured field table F is generated and presented in tabular form. Each row in the table contains the field label, character coordinates, field content, and matching confidence. This field table supports subsequent data standardization, verification, and encryption preparation.
[0157] Optionally, the data clustering time sorting in step S4 is specifically as follows:
[0158] Batch tags are extracted from the homomorphic structured encrypted invoice ciphertext data, and the corresponding invoice generation timestamp, business identifier and source device ID are extracted to construct an initial encrypted invoice index dataset;
[0159] In this embodiment, the label field is pre-decompressed and extracted for each encrypted invoice ciphertext data. Each invoice ciphertext contains a homomorphic encrypted field and a plaintext index field. The plaintext index field is stored in an independent encrypted carrier in a fixed-length structured data block format, including a generation timestamp (13-bit UNIX time format), a business identifier (8-bit business code), and a source device ID (fixed-length 32-bit device hash identifier). By sequentially reading the index area data block, the above three fields are parsed and stored in the initial encrypted invoice index dataset in the format of <ciphertext ID, timestamp, business identifier, device ID>. The index dataset is stored in a structured table format, providing a data foundation for subsequent sliding window segmentation and indexing.
[0160] The initial encrypted ticket index dataset is divided into time windows based on the timestamp field. The sliding window length is set to 10 minutes and the sliding step is 5 minutes. The ticket data is then grouped into windowed ciphertext cluster data groups.
[0161] In this embodiment, time period division processing is performed according to the timestamp field in the index dataset. To improve the continuity and coverage accuracy of the time series, the sliding window length is set to 10 minutes and the step size is set to 5 minutes, and continuous sliding window slicing processing is performed. In specific operations, a time window is generated every 5 minutes starting from the earliest timestamp, with a range from the current start time to the start time + 10 minutes. The window contains all ticket records whose timestamps fall within this interval. In this way, multiple cross tickets can be included in the same time slice, increasing the ticket aggregation density within the window. Each window processing result is encapsulated as a set of ciphertext ticket record sets, forming a windowed ciphertext cluster data group, and the data structure format is <window ID, ciphertext ID list, window start and end time>.
[0162] Calculate the hash value according to the business identifier in the windowed ciphertext cluster data group, and perform clustering sorting according to the hash value calculation result to obtain sorted ciphertext cluster group data;
[0163] In this embodiment, the window is used as the processing unit, and the hash value calculation is performed on the business identifier field of the ticket data within the window. A preset high-uniformity hash function is used to map the business identifier of each record to a fixed-length hash value, and the ticket records within the window are clustered and sorted with this hash value as the primary key. The sorting rule is executed according to the numerical order of the hash value, and the records clustered with the same hash value are arranged in ascending order of the timestamp from early to late, ensuring business semantic consistency and time continuity within each cluster of records. Each sorting result is output as a ciphertext cluster group data block, and the structure form is <cluster ID, hash value, ticket ID list, timestamp sequence>. [[ID=I7]]
[0164] Perform cluster summary extraction operations based on the ticket amount field, source device ID, and ticket generation timestamp in the sorted ciphertext cluster group data, and bind the corresponding ciphertext clusters to generate a time series indexing structure;
[0165] In this embodiment, the summary information of each cluster is further extracted for the construction of the subsequent sliding window data indexing structure. The specific operations include extracting the encrypted ticket amount field (field name: enc_amount), device ID, and generation timestamp from the original ciphertext record, and performing summary encapsulation while maintaining the encrypted state. The encapsulation result is a sequence summary set formed within each cluster, and the format is a set of <enc_amount, device ID, timestamp> triple sequences, which is bound to the original cluster structure. The final output structure is <cluster ID, summary sequence, bound ticket ciphertext ID list>, constructing a time series indexing structure to prepare for the sliding window encapsulation of ciphertext data.
[0166] Encapsulate each time series indexing structure into a sliding window data block respectively to generate a set of ciphertext sliding window data blocks.
[0167] In this embodiment, each time-series indexing structure is encapsulated using a sliding window method to construct a set of encrypted sliding window data blocks. Each sliding window data block is encapsulated with a fixed structure, including the window start and end times, cluster hash value, digest information, and a list of encrypted IDs, and is encapsulated as a self-describing data object. The packaging uses a binary structure for storage, where each field is written sequentially, and a CRC checksum and sliding window number are appended to the data block header. This set of encrypted sliding window data blocks serves as the basic data unit for subsequent data synchronization, off-site archiving, and blockchain traceability, realizing the structural encapsulation of distributed encrypted invoice data.
[0168] Most importantly, the cluster sorting is as follows:
[0169] The business identifier in the windowed encrypted data group is calculated using a preset hash function to generate a fixed-length hash value. The encrypted data is then classified according to the fixed-length hash value to obtain the initial hash cluster group set.
[0170] In this embodiment, when performing classification processing on windowed encrypted cluster data groups, a family of hash functions is preset. The business identifier field uses a 32-bit integer hash mapping function H(x) = (a·x+b)modp, where a and b are fixed multiplication and addition factors, and p is a prime number (e.g., 104729) to improve hash uniformity. The business identifier field of each ticket record is converted into a set of fixed-length hash values (e.g., 8 bytes) using this function, and records with the same hash value are grouped into the same initial hash cluster. Each hash cluster is constructed as a structure object, including a cluster ID (i.e., hash value), a list of contained encrypted ticket IDs, a cache of the original index field, and an identifier for the time period to which the window belongs. During the classification process, a distributed hash mapping table is used for cluster affiliation registration. After processing, a complete set of initial hash cluster groups is generated. The number of encrypted records in each cluster of this set is maintained between 100 and 500, and data balance is ensured through a dynamic cluster adjustment mechanism.
[0171] A local sorted index is built based on the timestamp field of the tickets generated in each hash cluster in the initial hash cluster grouping set. If there are duplicate timestamps in the local sorted index, the business identifier is introduced as a secondary sorting key to generate an ordered list of tickets within the cluster.
[0172] In this embodiment, a time-order structure is established for each cluster of data in the initial hash cluster grouping set based on the timestamp field. Specifically, the generation timestamp field (13-bit integer, milliseconds) of each ticket record is read, and a time-series index structure is built within the same hash cluster, using a min-heap to maintain the sorting result. If there are records with duplicate timestamps within the same cluster (i.e., multiple records have the exact same generation time), a business identifier field (8-bit encoded string) is introduced as a secondary sorting key, and lexicographical comparison is performed using ASCII code sequences. The final result is an ordered list of tickets within the cluster, a structured sequence of <ticket ID, timestamp, business identifier>, with all fields securely encapsulated in their original encrypted or plaintext state. To enhance data integrity, a content digest is calculated for each ordered list and bound to the metadata structure of that cluster.
[0173] The ordered list of tickets within each hash cluster is encapsulated into an ordered ciphertext cluster block, and all ordered ciphertext cluster blocks are combined in ascending order according to their corresponding hash cluster IDs to obtain sorted ciphertext cluster group data.
[0174] In this embodiment, after all the tickets within a cluster are constructed sequentially, the ordered list of tickets within each hash cluster is encapsulated to form ordered ciphertext cluster blocks. Each cluster block contains a hash cluster ID, an ordered list of tickets, an encapsulation time identifier, and a content digest. A segmented data structure is used for data layout, with the segment header containing the cluster block size, number of records, and checksum. Subsequently, all ordered ciphertext cluster blocks are arranged in ascending order according to their hash cluster IDs, outputting the final sorted ciphertext cluster group data. This data can be used for subsequent sliding window digest extraction, chained structure binding, or encrypted archiving. Its structural stability and clustering consistency are effectively guaranteed through standard hash mapping and time sorting mechanisms, ensuring that tickets from different sources have controllable organizational granularity in both logical time and business identifier dimensions.
[0175] Optionally, the chain-side security consensus encapsulation in step S4 specifically involves:
[0176] Extract the sliding window number, time range, number of tickets in the cluster, cluster digest hash value, and data source device ID from each encrypted sliding window data block to construct an encrypted sliding window information index table;
[0177] In this embodiment, when extracting the structure of the homomorphically encrypted sliding window ciphertext data block, the metadata fields attached to each data block are first deconstructed and analyzed, including: sliding window number (Window_ID, in the form of w0001~wN), window start and end timestamps (Start_Ts-End_Ts, accurate to the second), bill entry count (Bill_Count), the digest hash value of the corresponding cluster (Cluster_Hash, fixed-length 64-bit SHA-256 digest), and the source device identifier (Device_ID, in 64-bit Hex format). All field data are organized in key-value pairs and use a five-tuple structure:<Window_ID, Time_Range, Bill_Count, Cluster_Hash, Device_ID> The encrypted sliding window information index entries are constructed and written into the index table one by one, ultimately generating an encrypted sliding window information index table with an N×5 structure.
[0178] The encrypted sliding window information index table is used to perform encapsulation preprocessing on each encrypted sliding window data block to obtain a structured encapsulated draft body.
[0179] In this embodiment, during the preprocessing stage, index entries are read one by one according to the encrypted sliding window information index table, and the corresponding target sliding window data blocks are extracted. The data block structure is then reorganized into a packaged draft containing field mapping descriptions, content verification summaries, and time series labels. The packaged structure adopts a JSON-Like structure. Each packaged draft contains a field mapping area (containing a standardized field name mapping table), a cluster hash label area (referencing the Cluster_Hash field), a source verification label area (including the original Device_ID and its signature digest), and a data validity period field (derived from Time_Range). During preprocessing, redundant fields are temporarily retained, and the proportion of structured fields is set to be no less than 80% of the total number of data fields.
[0180] The chain-side edge computing nodes are screened from the jump point trust node set, and the screened chain-side edge computing nodes are used as consensus participants. The content of the structured encapsulated draft is verified to obtain the verified encapsulated draft.
[0181] In this embodiment, based on the previously constructed set of trusted nodes (node IDs and their trust scores), a screening operation is performed on the edge computing nodes accessing the chain. The limiting conditions include that the node's computing power is no less than 1.5 times the standard computing power index, there are no abnormal communication records in the last 10 minutes, and the trust score is no less than 0.82. After receiving the structured encapsulated draft, the screened node performs a fast content verification of the ticket structure consistency and hash consistency based on the field standard mapping table and the Cluster_Hash value. If the verification passes, a verification signature is attached, and a verified encapsulated draft is generated; otherwise, the verification fails, and the abnormal field type and source device ID are recorded.
[0182] The Byzantine voting consensus is executed on the encapsulated draft body that passes the verification, and the voting results, the signatures of the participating nodes and the consensus round parameters are encapsulated into a consensus encapsulation body.
[0183] In this embodiment, before entering the voting stage, data is synchronized to all consensus nodes according to the encapsulated draft content broadcast mechanism. Nodes initiate consensus requests using preset round rules, generate voting responses and sign verification results within each round. The voting data includes the round number (Round_ID), voting node ID (Node_ID), signature (generated using ECDSA), and local verification status (PASS / FAIL). When more than 2 / 3 of the votes are in agreement in a round, the final voting consensus result is generated, and the result, along with all signatures, round numbers, and signature times, is encapsulated into a structure:<Round_ID, Vote_Summary,Signatures[], Final_Result> This is the consensus encapsulation body.
[0184] For each consensus encapsulation, bind content verification and Byzantine consensus anchoring parameters to obtain verifiable consensus blocks;
[0185] In this embodiment, after each consensus encapsulation body is generated, its content verification record and voting signature record are extracted, and the corresponding hash anchor value is calculated (for example, by aggregating all signature values based on a Merkle tree to obtain the anchor hash). This anchor hash and the Cluster_Hash in the encapsulation draft body form the main anchor pair, which is embedded in the index header of the consensus block. The final generated verifiable consensus block structure includes: a block header (containing a timestamp, signature root anchor, and round identifier), a block body (encapsulating the draft body content), and a consensus digest record (recording the node voting ratio and verification process summary for each round). The entire structure is reversibly verifiable, has tamper-resistant capabilities, and supports cross-node consensus verification.
[0186] Verifiable consensus blocks are written into the edge chain structure in the Internet of Things, and the corresponding chain-side block index table in the edge chain structure is updated.
[0187] In this embodiment, the edge chain structure in the Internet of Things (IoT) adopts a lightweight distributed chain structure, writing each consensus block into a chained data area maintained by the edge computing nodes. The chain writing operation is performed on a block-by-block basis, with each block containing a hash pointer at the block header and a hash concatenation entry with the previous block, thus achieving chained data association. After writing is complete, the chain-side block index table is dynamically updated, inserting the timestamp index, device ID index, and Cluster_Hash reverse lookup table of the new block, facilitating efficient block retrieval by time window or device dimension. The index table uses an inverted index structure, with the following index entry format:<Key_Field, Value_List> The Key_Field can be a time window number, device ID, or cluster digest hash.
[0188] Optionally, step S5 specifically includes:
[0189] Step S51: Extract the sliding window number, timestamp, service identifier, and device ID from the verifiable consensus block, and construct a consensus block service mapping index table;
[0190] In this embodiment, verifiable consensus blocks that have been written are read from the edge chain in the Internet of Things (IoT). The data structure of these blocks includes three parts: a block header (Block_Header), a consensus digest record, and a block body (Block_Body). The block body encapsulates a set of structured ciphertext documents. Each document item includes a window ID field (Window_ID), a generation timestamp field (Gen_Ts), a business identifier field (Biz_Tag), and a source device field (Device_ID). Each document item is structurally deconstructed, and four key fields are extracted and uniformly formatted into quadruplets:<Window_ID,Gen_Ts,Biz_Tag,Device_ID> The business mapping index table is organized and constructed in ascending order by window number and timestamp. The index table adopts a two-dimensional matrix structure, with each row of the matrix corresponding to one ticket record. It has a total of 4 columns and an M×4 dimension, supporting subsequent label mapping and business traceability operations.
[0191] Step S52: Based on the field characteristics of the consensus block business mapping index table, divide the field labels to obtain a multi-label identifier set, and embed the multi-label identifier set into the structure of the verifiable consensus block to obtain the traceability label archive.
[0192] In this embodiment, during the tag partitioning process, the distribution characteristics of fields in the business mapping index table are analyzed, and a multi-tag identifier set is generated according to the set tag partitioning rules. The partitioning criteria include: time-series tag partitioning of Window_ID (e.g., forming a tag W_HH by hour), calendar-classification tagging of the Gen_Ts field (e.g., "weekdays" and "holidays"), business logic-based classification of the Biz_Tag field (e.g., "payment instructions", "tax verification", and "contract verification"), and device affiliation tagging of Device_ID by registration region (e.g., "East China node" and "Southwest node"). Each tag identifier uses the format <tag category::tag value> to form a multi-tag set, such as ["Window segment::W_13","Time category::weekdays","Business category::tax verification","Device category::East China node"]. Finally, it is embedded in the "Trace_Tags" field under each invoice data structure as a JSON object structure to form the expanded block body. The updated consensus block is called the "traceability tag archive body", which has a fine-grained business tag structure, making it easy for the private chain to trace the source.
[0193] Step S53: Select highly stable transmission relay nodes from the hop point trust transmission graph to construct the enterprise private link path structure;
[0194] In this embodiment, the hop-point trust transmission graph consists of a chain-like communication topology composed of multiple edge nodes. Its core structure is a directed graph G=(V,E,W), where V is the set of all reachable edge nodes, E is the set of historical communication paths between nodes, and W is the stability score matrix of the corresponding path. The stability score is composed of communication latency, packet loss rate, and response consistency between nodes, and the score value range is set to [0,1]. Taking the current enterprise identifier as the target, the target node set V_t⊂V with registration communication permissions with the enterprise's private database is selected as the destination node set. With the condition that the number of hops does not exceed 5, a path filtering operation is performed in the graph, prioritizing the selection of path sequences with a stability score of not less than 0.87 for all nodes, forming a stable hop-point path group. Finally, the enterprise's private link path structure is constructed as a matrix P[N][K], where N is the number of selectable hop-point paths, K is the node sequence number contained in each path, and the path node sequence will be used as the template for subsequent hop-point distribution order.
[0195] Step S54: The traceability tag archive is distributed and transmitted to the enterprise private database according to the enterprise private link path structure to perform the data archiving task.
[0196] In this embodiment, during the data distribution phase, the path P[0] with the highest score is selected as the primary hop path based on the constructed enterprise private link path structure matrix P[N][K]. After the traceability tag archive is initially signed and bound at the starting edge node, it is distributed hop-by-hop to the next hop according to the path node sequence. The distribution packet adopts a structured encapsulation format, including five parts: current path identifier (Path_ID), hop number (Hop_Index), signature chain record, block header digest, and archive content segment. Each hop relay node performs signature confirmation on the data and adds a local network timestamp. The entire transmission process adopts a symmetric encryption encapsulation method. The key is issued by the enterprise side, and only the target private database node has the decapsulation permission. After the final private database node receives the data, it performs signature verification, tag field parsing, and business field index confirmation, completes data parsing and database entry, and builds an inverted index for tag-type fields to realize the subsequent archive structure for tag retrieval. The entire hop transmission delay is controlled within 800ms, ensuring high reachability and low loss archive transmission performance.
[0197] Optionally, this specification also provides an IoT-based enterprise financial data secure storage management system for executing the IoT-based enterprise financial data secure storage management method described above. The IoT-based enterprise financial data secure storage management system includes:
[0198] The cache chain node anchoring module is used to acquire enterprise financial data transmission logs based on the Internet of Things, and to perform data cache chain node trust anchoring based on the enterprise financial data transmission logs to obtain trusted authentication device data.
[0199] The jump point path analysis module is used to construct a minimum set of trust paths based on the data from the trust authentication device, and to analyze the jump point trust transmission nodes of the enterprise's financial data transmission information using the minimum set of trust paths to obtain a jump point trust transmission map.
[0200] The bill field enhancement module is used to collect financial bill data from multiple parallel channels via trusted authentication devices to obtain trusted financial bill data; and to enhance the structure of multiple bill category fields in the trusted financial bill data to obtain the bill data to be encrypted.
[0201] The structured encryption module is used to perform homomorphic structured encryption on the data to be encrypted, and to sort the encryption results by data clustering and time to obtain a set of ciphertext sliding window data blocks; the set of ciphertext sliding window data blocks is then encapsulated with chain-side security consensus to obtain a verifiable consensus block.
[0202] The traceability archiving module is used to perform multi-label traceability control archiving on verifiable consensus blocks to obtain traceability label archives; the traceability label archives are then transmitted to the enterprise's private database using the jump point trust transmission graph to perform data archiving tasks.
[0203] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.
[0204] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. A method for managing the security storage of enterprise financial data based on the Internet of Things, characterized in that, The method comprises the following steps: Step S1: obtaining enterprise financial data transmission logs based on the Internet of Things, and performing data cache chain node trust anchoring according to the enterprise financial data transmission logs to obtain trust authentication device data; Step S2: constructing a minimum trust path set according to the trust authentication device data, and analyzing a jump point trust transmission node of enterprise financial data transmission information by using the minimum trust path set to obtain a jump point trust transmission graph; Step S3: collecting multi-path parallel financial bill data of the Internet of Things by using the trust authentication device data to obtain trust financial bill data; Performing bill multi-category field structured enhancement on the trust financial bill data to obtain encrypted bill data; Step S4: homomorphic structured encryption is performed on the encrypted bill data, and the encryption result is subjected to data clustering time sorting to obtain a ciphertext sliding window data block set; The ciphertext sliding window data block set is subjected to chain side security consensus encapsulation to obtain a verifiable consensus block; Step S5: performing multi-label source tracing control archiving on the verifiable consensus block to obtain a source tracing label archive, and performing jump point transmission of the source tracing label archive to an enterprise private database by using the jump point trust transmission graph to perform a data archiving task. 2.The Internet of Things based enterprise financial data security storage management method according to claim 1, characterized in that, Step S1 specifically comprises: Step S11: collecting enterprise financial data transmission logs through Internet of Things sensing devices; Step S12: standardizing the data packet format of the enterprise financial data transmission logs, and removing duplicate data packets, broadcast test packets and format fake data from the standardized results to obtain standardized transmission data; Step S13: constructing a device node dynamic trust evaluation model according to the device identification code information, historical communication stability records and edge authentication corresponding time delay extracted from the standardized transmission data, so as to evaluate the device node trust score by using the device node dynamic trust evaluation model; Step S14: performing trust share matching of the standardized transmission data and the corresponding device node trust score, screening data packets corresponding to devices with a trust score higher than 0.75 to obtain authenticated financial data packets; Step S15: clustering and sorting the authenticated financial data packets according to the time stamp and business number, writing them into the cache linked list structure of the edge computing node in the Internet of Things, and assigning an abstract root hash to each batch of clustering and sorting results, so as to obtain trust authentication device data. 3.The IoT-based enterprise financial data security storage management method of claim 2, wherein, Step S15 specifically comprises: Step S151: extracting the time stamp field and business number field information in the authenticated financial data packets, constructing a time-business two-dimensional index structure, and obtaining a time-business index data set; Step S152: performing adaptive density clustering based on the time-business index data set, dividing the authenticated financial data packets into business cluster groups in each time period to generate a clustering and sorting candidate set; Step S153: performing data packet index rearrangement on the data packet set in the clustering and sorting candidate set, and writing it into the cache linked list structure of the edge computing node in the Internet of Things to generate an edge linked list cache body; Step S154: performing node-by-node hash digest calculation on the data groups in each edge linked list cache body, and iteratively constructing a binary Merkle tree structure to generate a cluster abstract root set; Step S155: binding each abstract root in the cluster abstract root set with the chain table structure head address corresponding to the cluster, the number of the service cluster group and the time period information to which it belongs, to generate the trust authentication device data. 4.The IoT-based enterprise financial data security storage management method of claim 1, wherein, Step S2 is specifically: Step S21: extracting the device communication path and the device identification number in the trust authentication device data, thereby constructing an initial communication graph data set; Step S22: based on the initial communication graph data set, calculating the average hop number and the response time difference between each device node, and constructing a trust hop point weight matrix; Step S23: taking the trust hop point weight matrix as the path cost input and taking the device node trust score as the node confidence weight, dynamically calculating the minimum transmission trust path between nodes to generate a minimum trust transmission path set; Step S24: counting the relay frequency, average link stability and trust cumulative score of each hop point node in the minimum trust transmission path set, screening a high-frequency stable node set to obtain a hop point stable trust set; Step S25: performing multi-factor hierarchical clustering on the hop point stable trust set and selecting a path relay point to generate a hop point trust node set; Step S26: taking the hop point trust node set as the core node, combining the path connectivity and node bandwidth load in the initial communication graph data set to construct a hop point path graph structure, and thereby outputting a hop point trust transmission graph. 5.The Internet-of-Things based enterprise financial data security storage management method according to claim 1, characterized in that, Step S24 is specifically: Step S241: deconstructing each path in the minimum trust transmission path set, counting the number of occurrences of each node in all paths, and taking the number of occurrences as the node relay frequency to construct a relay frequency vector table; Step S242: counting the historical packet loss rate, average RTT and communication failure retry number of each hop point node associated link in the initial communication graph data set, calculating the average link stability index of the hop point node, and thereby generating a stability score table; Step S243: counting the number of times each hop point node in the initial communication graph data set accumulates in the trust authentication device data Participate in the communication process, and combine the device node trust score to calculate the trust cumulative score, thereby constructing a trust score matrix; Step S244: normalizing and integrating the stable score of the relay frequency vector table, the stability score table and the trust score matrix to generate a hop point node stable score vector, and setting the stable score threshold to 0.8 to screen a high-frequency stable node set; Step S245: topological consistency verification of the high-frequency stable node set in the corresponding physical position of the minimum trust transmission path set, and removing the nodes with high concentration to obtain a hop point stable trust set. 6.The Internet-of-Things based enterprise financial data security storage management method according to claim 1, characterized in that, The step S3 ticket multi-category field structure enhancement is specifically: Classifying and pre-analyzing the trust financial ticket data to obtain an original ticket field metadata set; According to the image file header structure, field label confidence and field storage format in the original ticket field metadata set, the original ticket field metadata set is divided into paper image ticket data, OCR recognized ticket data and electronic ticket data; Paper field structuring is performed on the paper image ticket data to obtain an image structured field table; The field type re-identification and semantic correction are performed on the OCR recognized ticket data to generate a semantic enhanced field data set; Align and map the image structured field table, semantic enhanced field dataset and fields in the electronic ticket data, construct a general ticket field standard template table, and perform field redundancy pruning and structure standardization processing based on the ticket category label to obtain a structured ticket field template; Based on the structured ticket field template, perform field replacement, format completion and value verification processing on the original ticket field metadata set to obtain encrypted ticket data. 7.The IoT-based enterprise financial data security storage management method of claim 1, wherein, The data clustering time sorting in step S4 is specifically: Extract the batch label from the ticket ciphertext data generated by the homomorphic structured encryption, extract the corresponding ticket generation timestamp, business identifier and source device ID, and construct an initial encrypted ticket index dataset; Divide the initial encrypted ticket index dataset into time windows based on the timestamp field, set the sliding window length to 10 minutes and the sliding step to 5 minutes, group the ticket data into windows, and generate windowed ciphertext cluster data groups; According to the business identifier in the windowed ciphertext cluster data group, calculate the hash value, and perform clustering sorting according to the hash value calculation result to obtain sorted ciphertext clustering group data; Based on the ticket amount field, source device ID and ticket generation timestamp in the sorted ciphertext clustering group data, perform clustering summary extraction operation, and bind the corresponding ciphertext cluster to generate time sequence index structure; Encapsulate each time sequence index structure into a sliding window data block to generate a ciphertext sliding window data block set. 8.The IoT-based enterprise financial data security storage management method of claim 1, wherein, The chain side security consensus encapsulation in step S4 is specifically: Extract the sliding window number, time range, number of tickets in the cluster, clustering summary hash value and data source device ID features from each ciphertext sliding window data block to construct a ciphertext sliding window information index table; Use the ciphertext sliding window information index table to perform encapsulation preprocessing on each ciphertext sliding window data block to obtain a structured encapsulation draft; Select chain side edge computing nodes from the hop trust node set, and use the selected chain side edge computing nodes as consensus participants to perform content verification on the structured encapsulation draft to obtain a verified encapsulation draft; Perform Byzantine voting consensus on the verified encapsulation draft, and encapsulate the voting result, participant signature and consensus round parameters into a consensus encapsulation body; Bind the content verification and Byzantine consensus anchor parameters to each consensus encapsulation body to obtain a verifiable consensus block; Write the verifiable consensus block into the edge chain structure in the Internet of Things, and update the corresponding chain side block index table in the edge chain structure. 9.The Internet-of-Things based enterprise financial data security storage management method according to claim 1, wherein, Step S5 is specifically: Step S51: Extract the sliding window number, timestamp, business identifier and device ID from the verifiable consensus block to construct a consensus block business mapping index table; Step S52: Divide the fields according to the field characteristics of the consensus block business mapping index table to obtain a multi-label identification set, and embed the multi-label identification set into the structure of the verifiable consensus block to obtain a traceability label archive; Step S53: Select high-stability transmission relay nodes from the hop trust transmission graph to construct an enterprise private link path structure; Step S54: Distribute and transmit the traceability label archive to the enterprise private database according to the enterprise private link path structure to perform data archiving tasks.
10. An Internet of Things-based enterprise financial data security storage management system, characterized in that, The application discloses an Internet of Things (IoT)-based enterprise financial data security storage management method and system. The cache chain node anchoring module is used for acquiring enterprise financial data transmission logs based on the IoT, performing data cache chain node trust anchoring according to the enterprise financial data transmission logs, and obtaining trust authentication device data. The hop path analysis module is used for constructing a minimum trust path set according to the trust authentication device data, and analyzing hop trust transmission nodes of enterprise financial data transmission information by using the minimum trust path set, to obtain a hop trust transmission graph. The bill field strengthening module is used for collecting multi-path parallel financial bill data of the IoT by using the trust authentication device data, to obtain trust financial bill data; performing multi-category field structured enhancement on the trust financial bill data, to obtain to-be-encrypted bill data; and performing homomorphic structured encryption on the to-be-encrypted bill data, and performing data clustering time sorting on an encryption result, to obtain a ciphertext sliding window data block set; and performing chain side security consensus encapsulation on the ciphertext sliding window data block set, to obtain a verifiable consensus block. The traceability archiving module is used for performing multi-label traceability control archiving on the verifiable consensus block, to obtain a traceability label archive; and performing hop transmission of the traceability label archive to an enterprise private database by using the hop trust transmission graph, to perform a data archiving task.
Citation Information
Patent Citations
Automatic financial information processing method based on AI
CN120494993A
Blockchain-based authentication system and method for authenticating electric vehicles or drones in a smart city
US20250233743A1