A data exchange system with a nested tag structure
By comparing the MIME type and pattern matching algorithm to locate file header features, combining DOM tree analysis and depth-first traversal, using the Hungarian algorithm to match field identification, construct a dynamic field mapping table and perform logical operations, the problems of high time cost and insufficient semantic boundary recognition in existing nested label analysis are solved, and efficient cross-platform data exchange and real-time response are achieved.
Patent Information
- Application Number
- CN202510639089.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The existing nested label analysis technology has high resolution time cost when dealing with multi-layer nested structures, lacks dynamic recognition of semantic boundaries, resulting in data attribution logic conflicts, static field mapping cannot adapt to database encoding updates, and policy lag or instruction redundancy in real-time data processing scenarios, affecting the accuracy and efficiency of data exchange.
The data carrier adaptation module is used to compare the MIME type and pattern matching algorithm to locate the file header features, combine DOM tree analysis and depth priority traversal to generate node sequences, and use the Hungarian algorithm to match field identification and database encoding to construct a dynamic field mapping table and perform logical operations to generate write strategy instructions, and reduce semantic ambiguity through multiple algorithms.
It improves the accuracy and adaptability of data carrier identification, optimizes the analytical efficiency of complex label structures, realizes dynamic field mapping, enhances the robustness and real-time response capabilities of cross-platform data exchange, and ensures the timeliness of data transmission in fields such as industrial Internet of Things.
Smart Images

Figure CN120162466B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data exchange, and in particular to a data exchange system with a nested tag structure. Background Art
[0002] The technical field of data exchange includes various transmission methods and conversion mechanisms for different communication scenarios and data formats. Its core content lies in realizing data interconnection and semantic consistency processing between heterogeneous systems. This technical field mainly involves communication protocol parsing, data structure mapping, message format standardization, and cross-platform data intercommunication mechanisms. In applications such as network communication, industrial Internet of Things, and information integration systems, different systems or devices often adopt their own defined data formats and protocols. Therefore, it is necessary to complete format conversion and content reconstruction through data exchange technology to ensure the correct transmission and recognition of information between multiple systems. This technical field further covers structured data processing, protocol encapsulation conversion, and semantic recognition based on tags or fields, supporting the requirements for data compatibility and adaptability in various transmission media and application scenarios.
[0003] Among them, a data exchange system with a nested tag structure refers to when processing structured tag data with a hierarchical relationship, by constructing an exchange mechanism that identifies the nested levels and parses their semantic boundaries, to complete data reception and reconstruction. For data packets containing multiple levels of nested elements, during the parsing process, tag nodes are extracted through preset tag recognition rules, and data content recombination is completed according to the tag path and nested hierarchical relationship. The system uses a path matching method to identify the tag nested structure, performs tag expansion and data splicing in sequence, and determines the data attribution relationship through tag sequence position calculation and nested depth judgment. Finally, a target data format with a consistent logical structure is formed for data transmission and processing between heterogeneous systems.
[0004] The nested tag parsing of the prior art relies on path matching and tag sequence position calculation. When processing a multi-level nested structure, the number of path traversals increases linearly with the increase in the hierarchical depth, resulting in too high a parsing time cost. The existing hierarchical relationship judgment only depends on the nested depth and position, lacking dynamic recognition of semantic boundaries, and easily causing data attribution logic conflicts between heterogeneous systems. The existing field mapping uses a static configuration table, which cannot adapt to the dynamic update of the target database field encoding, requires frequent manual maintenance, and there is a risk of mapping failure. The existing writing strategy is based on fixed rules and cannot be adjusted in real time according to the data content or logical state, and it is easy to have strategy lag or instruction redundancy in real-time data processing scenarios. For example, when parsing multi-level nested data packets of industrial equipment, the prior art is difficult to quickly identify key semantic nodes, resulting in the monitoring system being unable to accurately extract equipment operation parameters and affecting the accuracy of real-time warning. Summary of the Invention
[0005] The purpose of the present invention is to solve the shortcomings in the prior art and to propose a data exchange system with a nested tag structure.
[0006] In order to achieve the above object, the present invention adopts the following technical solution: A data exchange system with a nested tag structure comprises:
[0007] The data carrier adaptation module is used to compare the MIME type identifier of the input file with the preset type correspondence table to generate a carrier type identifier, call the KMP pattern matching algorithm to locate the binary sequence features of the file header, generate an extraction strategy parameter set, and pass the carrier type identifier and the extraction strategy parameter set to the structure analysis module;
[0008] A structure parsing module, used to call the DOM tree parsing algorithm through the carrier type identifier to process the metadata parameters in the extraction strategy parameter set, perform depth-first traversal on the file hierarchy path to generate a node sequence set, calculate the hierarchical nesting relationship and output a nested tag list and a field identifier set, and pass them to the field mapping module;
[0009] A field mapping module, used to match the field identification set with the target database field code through the Hungarian algorithm, generate a field mapping table, and transmit it to the logic driving module synchronously with the nested tag chain table;
[0010] The logic driving module is used to construct a decision matrix through the field mapping table, perform logical operations on the field values in the nested tag list to generate a status bit identifier, generate a write strategy instruction set after triggering a threshold, and transmit it to the script generation module synchronously with the nested tag list.
[0011] As a further solution of the present invention, the carrier type identifier is specifically PDF, XML, and JSON; the extraction strategy parameter set includes an offset, a delimiter, and a block size; the nested tag list includes a parent-child relationship, a hierarchical depth, and a closed state; the field identifier set includes a tag name, an attribute hash, and a path index; the field mapping table specifically refers to an encoding matching pair, a type conversion rule, and a constraint condition; and the write strategy instruction set includes a batch operation flag, a transaction isolation level, and a rollback condition.
[0012] As a further solution of the present invention, the data carrier adaptation module includes:
[0013] The type identification comparison submodule obtains the MIME type identifier of the input file, compares the character sequence of the identifier with the character sequence of the extension in the preset type correspondence table item by item, determines the mapping relationship priority between the identifier and the extension, and completes the type matching according to the weight value in the preset table to generate a carrier type identifier;
[0014] The feature localization sub-module constructs a partial match table based on the KMP pattern matching algorithm, slides a window over the binary stream of the input file in byte order to compare it with the preset file header feature sequence, calculates the index position where the feature sequence first appears in the binary stream, and generates a feature offset;
[0015] The parameter generation sub-module determines the starting position of the data block according to the segmentation length constraint condition corresponding to the carrier type identifier, combines the feature offset to split the binary stream by byte step and filters out non-continuous regions, and generates a set of extraction strategy parameters.
[0016] As a further solution of the present invention, the structure parsing module includes:
[0017] The metadata parsing sub-module matches the label separation rules in the DOM tree parsing algorithm based on the carrier type identifier, maps the metadata parameters in the set of extraction strategy parameters according to the label name length and the attribute key-value hash value, constructs a linked list of parent-child relationships between nodes, and generates a node tree structure;
[0018] The traversal control sub-module visits the child nodes layer by layer according to the depth-first traversal rule based on the number of child nodes of the root node of the node tree structure, records the difference between the node access order and the hierarchical number, and generates a traversal index sequence;
[0019] The nesting relationship sub-module extracts the difference in hierarchical numbers between adjacent nodes in the traversal index sequence, counts the number of parent node branches and the difference in nesting depth, calculates the label nesting level weight and the field identifier hash value, and generates a nesting level coefficient and a set of field identifiers.
[0020] As a further solution of the present invention, the field mapping module includes:
[0021] The weight calculation sub-module extracts the character hash value of multiple field names in the set of field identifiers and the character length of the target database field encoding, compares the coincidence degree of the character sequences of the field names and the encoding, and calculates the weight value according to the semantic similarity algorithm to generate a field weight matrix;
[0022] The matching execution sub-module splits the field weight matrix into row indexes and column indexes based on the Hungarian algorithm, traverses all nodes to construct a bipartite graph edge weight relationship, iteratively calculates the augmenting path and updates the matching status, and generates an initial matching pair;
[0023] The table optimization sub-module filters out the mapping items with weight values lower than the field conflict threshold in the initial matching pair according to the uniqueness constraint condition of the target database field encoding, reassigns the matching priorities of the conflicting fields, and generates a field mapping table.
[0024] As a further solution of the present invention, the logic driving module includes:
[0025] The matrix construction sub-module calls the field encoding in the field mapping table and the type encoding of the field values in the nested tag linked list, extracts the length parameter and priority weight of the field values, arranges the mapping relationship between the field encoding and the type encoding according to the row and column index rules, and generates a decision matrix;
[0026] The logical operation sub-module, based on the Boolean operation rules of the field values in the decision matrix, compares the matching degree between the field values and the truth table of the preset logical conditions, counts the number of fields and the triggering times that meet the conditions, and generates a status bit flag;
[0027] The instruction generation sub-module filters the fields in the status bit flag whose triggering times exceed the threshold according to the write policy trigger threshold, arranges the operation instructions according to the field encoding order and the triggering times priority, and generates a write policy instruction set.
[0028] As a further solution of the present invention, the system further includes:
[0029] A script generation module, which is used to call the XSD schema verification algorithm to detect the closure of the nested tag linked list, recursively verify the node levels based on the write policy instruction set, output a standard script file after correcting the errors, and write the standard script file into the target system transaction queue;
[0030] The standard script file includes a syntax tree structure, a verification log, and a transaction handle.
[0031] As a further solution of the present invention, the script generation module includes:
[0032] The closure detection sub-module calls the XSD schema verification algorithm to parse the tag closure rules of the nested tag linked list, extracts the sequence matching degree of the tag start symbol and the end symbol, counts the number of unclosed tags and the position offset, and generates a closure check value;
[0033] The level verification sub-module traverses the node levels of the nested tag linked list based on the field encoding order of the write policy instruction set, recursively compares the difference in the level depth between the parent node and the child node, calculates the legality weight of the node nested path, and generates the node level depth;
[0034] The script correction sub-module corrects the position of the end symbol of the unclosed tags according to the closure check value, adjusts the nesting order of the illegal paths in the node level depth, reorganizes the tag linked list structure according to the XSD schema rules, generates a standard script file and writes it into the target system transaction queue, and outputs a write index number.
[0035] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0036] In the present invention, by comparing the MIME type identifier of the input file to generate a carrier type identifier, and combining with a pattern matching algorithm to locate the binary sequence features of the file header, the accuracy and adaptability of data carrier recognition are improved. The file hierarchy path is traversed in a depth-first manner and the hierarchical nesting relationship is calculated, and combined with DOM tree parsing to generate a node sequence, optimizing the parsing efficiency of complex tag structures. An algorithm is used to match the field identifier with the target database encoding to generate a mapping table, realizing the construction of a dynamic field mapping relationship and reducing the dependence on manual configuration. Based on the field mapping table, a decision matrix is constructed, logical operations are performed to generate a status bit identifier and trigger a threshold generation write policy instruction set, enhancing the dynamic adjustment ability of the data conversion strategy. The above process reduces the semantic ambiguity of heterogeneous data conversion through multi-algorithm collaboration and a dynamic decision-making mechanism, improves the robustness of cross-platform data exchange, and at the same time supports rapid response to data state changes in real-time scenarios, ensuring the timeliness of data transmission in fields such as industrial Internet of Things. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 is the system flow chart of the present invention;
[0038] Figure 2 is the flow chart of the data carrier adaptation module of the present invention;
[0039] Figure 3 is the flow chart of the structure parsing module of the present invention;
[0040] Figure 4 is the flow chart of the field mapping module of the present invention;
[0041] Figure 5 is the flow chart of the logic drive module of the present invention;
[0042] Figure 6 is the flow chart of the script generation module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0044] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. are based on the orientation or positional relationships shown in the drawings. These are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, in the description of the present invention, the meaning of "plurality" is two or more, unless otherwise specifically defined.
[0045] Embodiment 1
[0046] Please refer to Figure 1 , a data exchange system with a nested tag structure includes:
[0047] A data carrier adaptation module, which is used to compare the MIME type identifier of the input file with a preset type correspondence table to generate a carrier type identifier, call the KMP pattern matching algorithm to locate the binary sequence features of the file header, generate a set of extraction strategy parameters, and transfer the carrier type identifier and the set of extraction strategy parameters to the structure parsing module;
[0048] The KMP pattern matching algorithm is an efficient string matching algorithm based on the longest common prefix and suffix, which is used to quickly locate the starting position of a specific pattern string in the text;
[0049] A structure parsing module, which is used to call the DOM tree parsing algorithm through the carrier type identifier to process the metadata parameters in the set of extraction strategy parameters, perform a depth-first traversal on the file hierarchical path to generate a set of node sequences, calculate the hierarchical nesting relationship and output a nested tag linked list and a field identifier set, and transfer them to the field mapping module;
[0050] Depth-first traversal is a graph traversal algorithm strategy that preferentially accesses child nodes along the vertical path of the tree structure until the end and then backtracks to sibling nodes;
[0051] A field mapping module, which is used to match the field identifier set with the target database field encoding through the Hungarian algorithm to generate a field mapping table, and transfer it to the logic driving module synchronously with the nested tag linked list;
[0052] The Hungarian algorithm is a polynomial-time algorithm for solving the task assignment problem in combinatorial optimization, which realizes the optimal matching by constructing a matrix and adjusting the row and column labels;
[0053] A logic driving module, which is used to construct a decision matrix through the field mapping table, perform logical operations on the field values in the nested tag linked list to generate a status bit identifier, generate a write strategy instruction set after triggering a threshold, and transfer it to the script generation module synchronously with the nested tag linked list;
[0054] The decision matrix defines the logical dependencies between fields through a truth table or a rule engine;
[0055] A script generation module is used to call the XSD schema validation algorithm to detect the closure of the nested tag linked list, recursively verify the node levels based on the write policy instruction set, output a standard script file after correcting errors, and write the standard script file to the target system transaction queue;
[0056] The XSD schema validation algorithm refers to the process of using the XML schema definition language specification to check the legality of the document structure, ensuring that the tag nesting levels and data types conform to predefined rules.
[0057] The carrier type identifier is specifically PDF, XML, JSON. The extraction policy parameter set includes offset, delimiter, and block size. The nested tag linked list includes parent-child relationship, hierarchical depth, and closure status. The field identifier set includes tag name, attribute hash, and path index. The field mapping table specifically refers to the encoding matching pairs, type conversion rules, and constraint conditions. The write policy instruction set includes batch operation flag, transaction isolation level, and rollback condition. The standard script file contains a syntax tree structure, validation log, and transaction handle.
[0058] The offset refers to the starting position value of the characteristic bytes of the file header. The delimiter is the ASCII code set of the field separator. The block size is the unit length for slicing the binary stream;
[0059] The parent-child relationship is realized through the association of DOM node pointers. The hierarchical depth is generated by counting the number of hops from the root node to the end node. The closure status is the identifier of the tag syntax integrity verification result;
[0060] The tag name is the XML / JSON node namespace string. The attribute hash is the 128-bit digest generated by the MD5 algorithm from the tag attribute key-value pairs. The path index is the XPath location expression of the node in the document;
[0061] The encoding matching pairs are the mapping relationship table between the source fields and the target database fields. The type conversion rules include integer-floating point conversion strategies and character encoding conversion tables. The constraint conditions refer to the field value range and non-null verification marks;
[0062] The batch operation flag indicates whether to enable the database batch write mode. The transaction isolation level includes predefined parameters such as read committed and repeatable read. The rollback condition is the error code threshold when the field value exceeds the limit or there is a logical conflict;
[0063] The syntax tree structure is the serialized storage format of the abstract syntax tree. The validation log records the error types and line numbers in the XSD schema verification. The transaction handle is the queue operation identifier returned by the target system.
[0064] See also Figure 2 , the data carrier adaptation module includes:
[0065] The type identifier comparison submodule obtains the MIME type identifier of the input file, compares the identifier character sequence with the extension character sequence in the preset type corresponding table item by item, determines the mapping relationship priority between the identifier and the extension, and completes the type matching according to the weight value in the preset table to generate a carrier type identifier;
[0066] First, process a specific file input, which is set as the spreadsheet file named Invoice_Q1_2025.xlsx uploaded by the user. The file type identification sub-module of the type identification comparison starts the file type identification process. It obtains the MIME type identifier of the input file as application / vnd.openxmlformats-officedocument.spreadsheetml.sheet by querying the operating system registration information or directly analyzing the starting bytes of the file. After obtaining this identifier, the sub-module makes an exact and comprehensive character sequence comparison between its character sequence application / vnd.openxmlformats-officedocument.spreadsheetml.sheet and a pre-set type correspondence table maintained internally (shown in Table 1 below). This table stores various file extensions that the system can process, the corresponding MIME type patterns, the internally unified carrier type identifiers, and the matching weight values for decision-making. During the specific comparison, the input MIME identifier exactly matches the MIME pattern application / vnd.openxmlformats-officedocument.*ml.sheet in the first row of Table 1 (the asterisk * acts as a wildcard to match the spreadsheet part). At the same time, the system also extracts the file extension.xlsx from the file name and finds a matching item in the table. Subsequently, the system determines the priority of the mapping relationship according to the pre-set rules. The rules stipulate that the priority of the exact match result of the MIME type is set to 1, and the priority of the file extension match is set to 2. When there are multiple matches, the match result with the smaller priority value is preferred. This rule ensures that even if the file name is modified incorrectly, as long as the file content itself meets the standards, it can still be accurately identified by the MIME type. Therefore, the MIME type match result is selected, and the corresponding weight value obtained from Table 1 is 0.9. This weight value is pre-set according to the uniqueness of the MIME type identifier and the recognition accuracy rate in actual applications (historical data analysis shows that its accurate recognition rate reaches 99%). The higher the value, the higher the credibility of the match. Next, compare the obtained weight value of 0.9 with the pre-set type confirmation threshold, which is set to 0.7. The setting basis is the statistical analysis of 10,000 files processed historically, and the lowest weight value that can cover 95% of the correct recognition samples is selected, that is, the weight at the 5th percentile is exactly 0.7. Since 0.9 is greater than 0.7, it is confirmed that this type match is valid, and the system determines that the type of the input file is credible. A weight value greater than 0.85 is classified as a high-confidence match, between 0.7 and 0.85 as a medium-confidence match, and below 0.7 as a low-confidence match. Currently, 0.9 belongs to the high-confidence interval. Finally, the type identification comparison sub-module generates an internally unified carrier type identifier as Excel_OOXML.
[0067] Table 1 Preset Type Correspondence Table
[0068]
[0069] As shown in Table 1, this table details the file extensions, MIME type matching patterns, carrier type identifiers used internally by the system, and corresponding matching weights for different file types, for the automated identification and verification of file types.
[0070] The feature location sub-module constructs a partial match table based on the KMP pattern matching algorithm, slides a window over the input file binary stream in byte order to compare it with the preset file header feature sequence, calculates the index position where the feature sequence first appears in the binary stream, and generates a feature offset.
[0071] The feature localization sub-module receives the carrier type identifier Excel_OOXML determined in the previous step and the input file Invoice_Q1_2025.xlsx. For the Excel_OOXML type, the system queries and retrieves its pre-set standard file header feature sequence, namely 504B0304 (represented in hexadecimal bytes, corresponding to the local file header marker of the ZIP file format). After obtaining this 4-byte feature sequence, the sub-module uses this sequence [50, 4B, 03, 04] to construct the internal data structure required by the KMP pattern matching algorithm - the partial match table (or next array). The role of this table is that in subsequent binary stream comparisons, when a mismatch occurs, it can indicate the most effective number of positions to shift the pattern string (feature sequence) to the right, avoiding unnecessary backtracking comparisons. For the sequence 504B0304, since there are no repeated prefix and suffix substrings inside it, the generated partial match table is [0, 0, 0, 0]. Then, the sub-module starts to read the binary data stream of the input file Invoice_Q1_2025.xlsx. Starting from the first byte (index 0) of the file, in byte order, it applies a sliding window of size 4 bytes and compares the byte sequence within the window with the pre-set feature sequence [50, 4B, 03, 04] byte by byte. For a standard, unmodified XLSX file, its first 4 bytes are exactly 504B0304. Therefore, in the first comparison, that is, when the window covers bytes 0 to 3, a complete match is achieved. The sub-module immediately calculates and records the starting index position where this feature sequence first appears in the file binary stream. In the scenario of this standard file, this index position is 0. If there is non-standard padding data at the file header, for example, the first two bytes are 0000, then the binary stream is 0000504B0304..., the sliding window first compares [00, 00, 50, 4B] with [50, 4B, 03, 04], and there is a mismatch. According to the partial match table [0, 0, 0, 0], the window slides one position to the right and compares [00, 50, 4B, 03], still with a mismatch. Then it slides one more position and compares [50, 4B, 03, 04], and at this time a complete match is achieved, and the recorded index position is 2. In this embodiment, the file is standard, so the finally generated feature offset is determined to be 0.
[0072] The parameter generation sub-module determines the starting position of the data block according to the segment length constraint condition corresponding to the carrier type identifier, divides the binary stream by byte step size and filters out non-continuous regions to generate a set of extraction strategy parameters.
[0073] The parameter generation sub-module receives the carrier type identifier Excel_OOXML and the feature offset 0. First, according to the Excel_OOXML identifier, the sub-module queries the segmentation length constraints and structure information defined in the system associated with this type. These information indicate that the Excel_OOXML file is essentially a ZIP compressed package, and its core data, such as worksheet content, is stored in specific XML file entries after decompression (for example, xl / worksheets / sheet1.xml), and the recommended data reading block size is set to 4096 bytes to optimize I / O performance. Combining the feature offset 0 (indicating that the file starts at the standard position), the sub-module confirms that the starting point of data extraction needs to first parse the ZIP file structure. It reads the central directory record of the ZIP file, locates the xl / worksheets / sheet1.xml file entry, and obtains the starting offset position of the data of this entry in the decompression stream. Through this process, the starting byte offset of the XML content of this worksheet is calculated to be 8192, which is the starting reading position of the data block. Then, according to the set byte step of 4096 bytes, starting from the starting position of 8192, the binary stream (the decompressed XML text stream) of the worksheet content is segmented. The first data block contains bytes 8192 to 12287, the second data block contains bytes 12288 to 16383, and so on. While segmenting and reading, the sub-module applies filtering rules to identify and filter out non-continuous or non-business data regions in the XML stream. The specific filtering operations include: removing all XML comment nodes (in the form of <!--comment-->), removing XML processing instruction nodes (in the form of <?targetinstruction?>), removing non-structural white spaces defined based on XML Schema (such as line breaks and indentation spaces purely for formatting between tags), and only retaining the elements and their contents directly related to table data, mainly <row>(row), <c>(cell) and containing the actual value of <v>Elements such as these, through the above processing, finally generate a structured set of extraction strategy parameters, the content of which is: {"carrier type": "Excel_OOXML", "data start offset": 8192, "reading step": 4096, "target data area rules": [" / worksheet / sheetData / row / c / v", " / worksheet / sheetData / row / c[@t='s'] / v"], "filtering rules": ["XML comments", "processing instructions", "unstructured whitespace"]}. This set precisely guides how the subsequent module extracts valid data.
[0074] Please refer to Figure 3 , the structure parsing module includes:
[0075] The metadata parsing sub-module matches the label separation rules in the DOM tree parsing algorithm based on the carrier type identifier, maps the metadata parameters in the extraction strategy parameter set according to the label name length and the attribute key value hash value, constructs a linked list of parent-child relationships between nodes, and generates a node tree structure;
[0076] The metadata parsing sub-module starts working. It receives the carrier type identifier Excel_OOXML and the previously generated set of extraction policy parameters {"carrier type": "Excel_OOXML", "data start offset": 8192, "reading step": 4096, "target data area rules": [" / worksheet / sheetData / row / c / v", " / worksheet / sheetData / row / c[@t='s'] / v"], "filtering rules": ["XML comments", "processing instructions", "non-structural whitespace"]}. Based on the Excel_OOXML type, it determines to adopt the XML tag separation rule, that is, regarding < as the tag start symbol and > as the tag end symbol, and strictly applies the filtering rules defined in the parameter set. The sub-module starts parsing the XML content concatenated by data blocks from the 8192-byte offset of the data stream. When the parser encounters an XML element, such as <row r="1" spans="1:3">, it extracts the tag name row and all the attribute key-value pairs of this tag, that is, r="1" and spans="1:3". The information extracted (tag name, attribute key, attribute value) constitutes the metadata parameters. Next, a mapping operation is performed to build the relationships between nodes. The specific mapping rule is: calculate the character length of the tag name and calculate the hash value of the key-value pair of its first attribute. A simple additive hash algorithm is used for demonstration (a more robust hash algorithm such as MurmurHash3 will be adopted in actual applications): for the tag row, its length is 3, and for its first attribute r="1", calculate its hash value as ASCII('r') + ASCII('=') + ASCII('1') = 114 + 61 + 49 = 224. The identifier of this node is defined as (3, 224). When parsing nested sub-tags, such as <c r="A1" t="s">, it is processed in the same way: the tag name c has a length of 1, and the hash value of the first attribute r="A1" is ASCII('r') + ASCII('=') + ASCII('A') + ASCII('1') = 114 + 61 + 65 + 49 = 289. The identifier of this sub-node is (1, 289). Since <c>The tag is in the XML structure <row>The direct children of the label, the parser processes when <c>When recording, its parent node is <row>(Identifier (3,224)), thus establishing a link from the parent node (3,224) to the child node (1,289). By continuously traversing all XML tags within the target data area and constantly establishing such parent-child links based on the parsing context, while ignoring the nodes specified by the filtering rules (comments, processing instructions, non-structural whitespace), a node tree reflecting the data hierarchy of sheet1.xml is finally constructed. The root node of the tree is the <sheetdata>Element, which contains multiple below <row>child node, each <row>The node further contains several <c>(Cell) child node, partial <c>The node also contains <v>Child node.
[0077] The traversal control sub-module visits the child nodes layer by layer according to the number of child nodes of the root node of the node tree structure, records the difference between the node access order and the level number, and generates a traversal index sequence according to the depth-first traversal rule;
[0078] The traversal control sub-module receives the node tree structure constructed in the previous step. This tree represents the core content of sheet1.xml, and its root node is <sheetdata>, assuming it directly contains 3 below <row>Child node, the sub-module first checks the root node <sheetdata>The number of child nodes is confirmed to be 3, and then the depth-first traversal (DFT) algorithm is started to visit the tree. The access information of each node is strictly recorded during the traversal process: First, visit the root node <sheetdata>, assign the access sequence number as 1, record its level number as 0 (the root node is defined as the 0th level), the level number difference is recorded as a non-applicable value (NaN) because it is the first node, and then access <sheetdata>The first child node <rowr="1"> is assigned an access sequence number of 2, the hierarchical number is recorded as 1, and the calculation is performed with the previous access node ( <sheetdata>The hierarchical number difference of ) is 1 - 0 = 1. Then, going deeper, access the first child node <cr="A1"> of <rowr="1">, assign the access sequence number 3, record the hierarchical number as 2, calculate the hierarchical number difference with the previous accessed node (<rowr="1">) as 2 - 1 = 1. Then, going deeper again, access the child node of <cr="A1"> <v>(Assume it exists and is a leaf node), assign the access sequence number 4, record the level number as 3, calculate the difference in level numbers from the previous accessed node (<cr="A1">) as 3 - 2 = 1. At this time, reaching the leaf node, start backtracking, return to <cr="A1">, it has no other unaccessed child nodes, then backtrack to <rowr="1">, access its second child node <cr="B1">, assign the access sequence number 5, record the level number as 2, calculate the difference from the previous accessed node ( <v>The hierarchical number difference of ( ) is 2 - 3 = -1 (indicating going back one level upward), and access the child node of <cr="B1"> <v>, allocate access sequence number 6, record the level number as 3, calculate the difference in level numbers with the previous access node (<cr="B1">) as 3 - 2 = 1, backtrack, access the third child node <cr="C1"> of <rowr="1">, allocate access sequence number 7, record the level number as 2, calculate with the previous access node ( <v>The hierarchical number difference of () is 2 - 3 = -1, and access the child node of <cr="C1"> <v>, assign access sequence number 8, record the level number as 3, calculate the difference in level numbers with the previous access node (<cr="C1">) as 3 - 2 = 1, complete the access of all subtrees, and backtrack to <sheetdata>, access its second child node <rowr="2">, assign the access sequence number 9, record the level number as 1, and calculate the distance from the previous accessed node (the last <v>) The hierarchical number difference of ( ) is 1 - 3 = -2 (indicating backtracking up two levels). The sub-module continuously traverses according to this depth-first rule until all nodes in the tree are visited. Record the identifier, access sequence number, hierarchical number of each node, and the difference in hierarchical numbers from its previous visited node. Finally, generate a complete traversal index sequence in the form of [( <sheetdata>,1,0,NaN),(<rowr="1">,2,1,1),(<cr="A1">,3,2,1),( <v>,4,3,1),(<cr="B1">,5,2,-1),( <v>,6,3,1),(<cr="C1">,7,2,-1),( <v>,8,3,1),(<rowr="2">,9,1,-2),(<cr="A2">,10,2,1),( <v>,11,3,1),(<cr="B2">,12,2,-1),( <v>,13,3,1),(<cr="C2">,14,2,-1),( <v>,15,3,1),(<rowr="3">,16,1,-2),(<cr="A3">,17,2,1),( <v>,18,3,1),(<cr="B3">,19,2,-1),( <v>,20,3,1),(<cr="C3">,21,2,-1),( <v>,22,3,1)]。
[0079] The nested relationship sub-module extracts the hierarchical number differences between adjacent nodes in the traversal index sequence, counts the difference between the number of parent node branches and the nested depth, calculates the nested level weight of the label and the hash value of the field identifier, and generates the nested level coefficient and the field identifier set.
[0080] The nested relationship sub-module uses the traversal index sequence generated in the previous step ([( <sheetdata>,1,0,NaN),(<rowr="1">,2,1,1),...,( <v>,22,3,1)] performs a deep analysis. First, it extracts the sequence of hierarchical number differences between adjacent nodes recorded in the sequence that reflects the traversal path changes: [1,1,1,-1,1,-1,1,-2,1,1,-1,1,-1,1,-2,1,1,-1,1,-1,1]. Next, the sub-module counts the actual number of branches of each parent node (non-leaf node). This information has been calculated and stored in the node attributes during the tree construction phase, or can be dynamically counted during traversal. For the node <rowr="1">, after it appears in the traversal sequence, its three child nodes <cr="A1">, <cr="B1">, <cr="C1"> (and their respective child nodes) are immediately visited. Therefore, it is determined that the number of its branches is 3. At the same time, the sub-module directly uses the hierarchical number differences recorded in the traversal index sequence as the nested depth differences for subsequent calculations. The next step is to calculate the label nested level weight of each node. This weight is used to quantify the relative importance of the node in the document structure. Its calculation rule is defined as: W = D + 0.5 * B, where D represents the absolute hierarchical depth of the node (counting from the root node 0), and B represents the number of branches of the parent node of this node. This rule aims to balance the depth position of the node and the local complexity of its environment (the breadth of the parent node). Calculation example: For the node <cr="A1">, its hierarchical depth D = 2, and the number of branches of its parent node <rowr="1"> is B = 3. Then the nested level weight of this node W = 2 + 0.5 * 3 = 3.5. At the same time, the sub-module calculates the field identifier hash value, mainly for leaf nodes containing actual data (such as <v>a label) or an attribute node with key identification information, such as <c>The r attribute value "A1" of the label can be regarded as the field coordinates), for <v>For a node, extract its internal text content, set it as "InvoiceNumber", and then apply a specified hash function (such as the FNV-1a algorithm) to calculate the 64-bit hash value of this string, obtaining a numerical identifier. For example, hash("InvoiceNumber") = 14695981039346656037 (example value). Perform this hash calculation on all identified field contents (such as "Invoice Number", "Invoice Date", "Total Amount", etc.). Finally, collect the calculated weight values of all nodes to generate a nested hierarchy coefficient set {W_node1, W_node2,...}, and collect all the calculated field content hash values to generate a field identifier set {14695981039346656037, hash("2025-04-15"), hash("1250.75"),...}. For the convenience of subsequent processing, map the hash values back to approximate field names that are easy to understand, obtaining a set {"Invoice Number", "Invoice Date", "Total Amount"}.
[0081] Please refer to Figure 4 , the field mapping module includes:
[0082] The weight calculation sub-module extracts the character hash values of multi-field names in the field identifier set and the character lengths of the target database field encodings, compares the character sequence coincidence degrees of the field names and the encodings, and calculates the weight values according to the semantic similarity algorithm to generate a field weight matrix;
[0083] The weight calculation sub-module receives the set of field identifiers generated in the previous step (whose corresponding original field names are approximately {"Invoice Number", "Invoice Date", "Total Amount"}) and the table structure information of the target database. The target table is set as Invoices, which contains fields InvoiceID (type VARCHAR, length 50), IssueDate (type DATE), and TotalAmount (type DECIMAL, precision 10, scale 2). The sub-module starts to calculate the matching weights between the source fields and the target fields. Taking the source field "Invoice Number" and the target field InvoiceID as an example: First, extract the source field name "Invoice Number" and the target field name InvoiceID, and record the defined length of the target field InvoiceID as 50 (here, we are more concerned about the name character length, which is 9). Then, conduct a comprehensive evaluation of the character sequence similarity and semantic similarity. The evaluation process is as follows: The first step, text preprocessing: Convert the source field name and the target field name to lowercase and perform word segmentation, obtaining source: ["invoice", "number"], target: ["invoice", "id"]. The second step, calculate semantic similarity: Utilize the pre-trained word vector model loaded by the system (such as the Word2Vec model trained based on a large-scale text library) to query the similarity between words, obtaining the cosine similarity between "invoice" and "invoice" as 0.82, and the similarity between "number" and "id" (possibly based on a thesaurus or model) as 0.91. Take the average semantic similarity SemanticSim = (0.82 + 0.91) / 2 = 0.865. The third step, calculate the length similarity factor: Use the formula LengthFactor = 1 - |len(src) - len(tgt)| / max(len(src), len(tgt)), where len(src) is the length of the source field name "Invoice Number", which is 3, and len(tgt) is the length of the target field name "InvoiceID", which is 9. Calculate LengthFactor = 1 - |3 - 9| / max(3, 9) = 1 - 6 / 9 = 1 - 0.667 = 0.333. The fourth step, weighted fusion to calculate the final weight: Set the fusion weight coefficients α and 1 - α for semantic similarity and length similarity. According to experience, semantic matching is usually more important than length matching. Set α = 0.7. This coefficient indicates that semantic similarity accounts for 70% of the total weight, and the length factor accounts for 30%. The calculation formula is Weight = α * SemanticSim + (1 - α) * LengthFactor. Substitute the values: Weight("Invoice Number", "InvoiceID") = 0.7 * 0.865 + 0.3 * 0.333 = 0.6055 + 0.0999 = 0.7054, the sub-module repeats this calculation process for all combinations of source field identifiers and target database field codes. For example, when calculating the weights of "Invoice Date" and IssueDate, since their semantics are highly related and their name lengths are close, the weight obtained is 0.958; when calculating the weights of "Total Amount" and TotalAmount, a relatively high weight of 0.932 is also obtained. For pairs that are obviously unrelated, such as "Invoice Number" and IssueDate, their semantic similarity is extremely low, and the calculated weight is close to 0, such as 0.115. Finally, a complete field weight matrix is generated, as shown in Table 2 below.
[0084] Table 2 Field Weight Matrix
[0085]
[0086] As shown in Table 2, this matrix accurately shows the quantitative matching weight values calculated through comprehensive semantic and length similarity between the fields extracted from the source file and the target database fields.
[0087] The matching execution sub-module splits the field weight matrix into row indices and column indices based on the Hungarian algorithm, traverses all nodes to construct the edge weight relationship of the bipartite graph, iteratively calculates the augmenting path and updates the matching status, and generates the initial matching pairs;
[0088] The matching execution sub-module receives the field weight matrix generated in the previous step (see Table 2). This matrix is regarded as the adjacency matrix representation of a weighted bipartite graph. Among them, the source fields {"Invoice Number", "Invoice Date", "Total Amount"} form the set of nodes U on one side of the graph, and the target database fields {"InvoiceID", "IssueDate", "TotalAmount"} form the set of nodes V on the other side. The value w(u, v) in the matrix represents the weight of the edge connecting nodes u ∈ U and v ∈ V. The goal is to find a perfect matching (or near-perfect matching if the number of sources and targets is inconsistent) with the maximum weight, that is, to find a set of edges connecting U and V such that each node is connected by at most one edge, and the sum of the weights of all selected edges reaches the maximum. The sub-module executes an iterative optimization algorithm (logically equivalent to the application of the Hungarian algorithm or the minimum-cost maximum-flow algorithm) to determine the best matching: Initialize an empty matching set M. At the beginning of the iteration, find the edges in the current weight matrix (or the matrix after internal transformation by the algorithm) that can be added to M to increase the total weight. The algorithm realizes this by maintaining the potential values (top labels) of the nodes and finding augmenting paths. In the first iteration, it is identified that the edge with the highest weight is ("Invoice Date", "IssueDate") with a weight of 0.9580, and it is added to the matching M. In the second iteration, considering the remaining unmatched nodes, the edge with the second-highest available weight is ("Total Amount", "TotalAmount") with a weight of 0.9320, and it is added to the matching M. In the third iteration, only the source field "Invoice Number" remains unmatched. Examine its connections with all unmatched target fields (at this time, all target fields are already matched, but the algorithm will handle this situation internally, or in this simple scenario, InvoiceID is the only target). The highest weight is ("Invoice Number", "InvoiceID") with a weight of 0.7054, and it is added to the matching M. The iteration end condition: all source fields are already matched (achieved in this example), or no more augmenting paths can be found that can increase the total weight of the matching. At this time, the obtained matching is the optimal solution under the current weight. Through this process, the system traverses all possible pairing relationships and makes decisions based on the principle of maximizing the weight, and finally generates the initial matching pair set: {("Invoice Number", "InvoiceID", 0.7054), ("Invoice Date", "IssueDate", 0.9580), ("Total Amount", "TotalAmount", 0.9320)}.
[0089] The table optimization sub-module filters out the mapping items with weight values lower than the field conflict threshold in the initial matching pairs according to the uniqueness constraint conditions of the target database field codes, reassigns the matching priorities of the conflicting fields, and generates a field mapping table.
[0090] The table optimization sub-module receives the initial matching pair set {("Invoice Number", "InvoiceID", 0.7054), ("Issue Date", "IssueDate", 0.9580), ("Total Amount", "TotalAmount", 0.9320)}, and at the same time obtains the constraint conditions of the Invoices table in the target database. In particular, the InvoiceID field has uniqueness (UNIQUE) and non-null (NOT NULL) constraints. The sub-module first applies a preset field conflict threshold to evaluate the matching quality. This threshold is set to 0.70, and the setting process of the threshold is as follows: Collect 10,000 field mapping records automatically completed by the system within the recent 6 months and their subsequent manual verification results, and count the mapping error rates corresponding to different weight intervals. It is found through analysis that when the matching weight is lower than 0.70, the mapping error rate significantly rises to 12%, while when the weight is 0.70 or above, the error rate is lower than 2%. To balance the automation efficiency and data accuracy, the inflection point 0.70 where the error rate starts to increase significantly is selected as the conflict threshold, that is, the conflict threshold = 0.70. The sub-module compares each weight value in the initial matching pair with this threshold: ("Invoice Number", "InvoiceID", 0.7054): The weight 0.7054 ≥ 0.70, initially judged to be acceptable. ("Issue Date", "IssueDate", 0.9580): The weight 0.9580 ≥ 0.70, judged to be a high-confidence match. ("Total Amount", "TotalAmount", 0.9320): The weight 0.9320 ≥ 0.70, judged to be a high-confidence match.In this example, the weights of all matches are not lower than the threshold. However, the sub-module also optimizes by combining database constraints. InvoiceID is the primary key or unique key, and its accuracy is crucial. Although the weight of 0.7054 just exceeds the threshold, the system checks whether there are other source fields that can be mapped to InvoiceID with a higher weight (none in this example), or whether "Invoice Number" can be mapped to other target fields with a significantly higher weight (its weights to IssueDate and TotalAmount are 0.115 and 0.130 respectively, far lower than 0.7054). Given the criticality of InvoiceID and that the current match is the best choice for "Invoice Number" and the weight reaches the threshold, the system decides to retain this match. However, considering that its weight is not very high (defining 0.70 - 0.80 as the medium confidence interval), a status flag will be attached to this match indicating that it requires attention or review. For matches with weights far exceeding the threshold (defining >0.80 as the high confidence interval), they are marked as confirmed. After this round of optimization screening and priority adjustment based on the threshold and database constraints, the final field mapping table is generated: {("Invoice Number", "InvoiceID", 0.7054, "Medium Confidence"), ("Invoice Date", "IssueDate", 0.9580, "High Confidence"), ("Total Amount", "TotalAmount", 0.9320, "High Confidence")}.
[0091] Please refer to Figure 5 , the logic-driven module includes:
[0092] The matrix construction sub-module calls the field codes in the field mapping table and the type codes of the field values in the nested label linked list, extracts the length parameters and priority weights of the field values, and arranges the mapping relationship between the field codes and type codes according to the row and column index rules to generate a decision matrix;
[0093] The matrix construction sub-module starts to operate. It calls the field mapping table generated in the previous step {("Invoice Number", "InvoiceID", 0.7054, "Medium Confidence"),..., ("Total Amount", "TotalAmount", 0.9320, "High Confidence")} and the specific field values and type information extracted from the nested tag linked list (node tree) parsed previously. Assume the first data record is being processed currently, and the extracted values are: Invoice Number = "INV1001", Issue Date = "2025-04-15", Total Amount = 1250.75, and the corresponding field type codes are: Invoice Number = String, Issue Date = Date, Total Amount = Decimal. The sub-module associates these values with the target database field codes InvoiceID, IssueDate, TotalAmount according to the mapping table. Then, it extracts the applicable length parameters for each field value: for the string "INV1001", the length is 7; for the date "2025-04-15", the standard format 'YYYY-MM-DD' is adopted; for the decimal number 1250.75, its precision and decimal digit information are recorded, which is consistent with the target field DECIMAL(10,2). Then, it calculates or obtains the priority weight for each field. This weight combines the mapping confidence (from the mapping table) and the business importance. The priority weight calculation rule is set as: PriorityWeight = MappingWeight + CriticalityBonus, where CriticalityBonus is set according to the business importance of the field: Total Amount (critical financial data) = 0.1, InvoiceID (primary key identifier) = 0.05, Issue Date (ordinary date) = 0.0. Calculate the priority weights for each field: Prio(InvoiceID) = 0.7054 + 0.05 = 0.7554 Prio(IssueDate) = 0.9580 + 0.0 = 0.9580 Prio(TotalAmount) = 0.9320 + 0.1 = 1.0320 Finally, the sub-module organizes and arranges all the above information (target field codes, field values, type codes, length / format, priority weights) according to the predetermined row and column index rules to generate a decision matrix row for a single record, as shown in Table 3 below. This matrix will be used for subsequent logical judgments.
[0094] Table 3 Example of Decision Matrix Row (Single Record)
[0095]
[0096] As shown in Table 3, this decision matrix row clearly integrates the values, types, formats, and processing priority information mapped to each field in the target database for a single data record.
[0097] The logic operation submodule compares the field value with the truth table of the preset logic condition based on the Boolean operation rules of the field value in the decision matrix, counts the number of fields that meet the conditions and the number of triggers, and generates a status bit identifier;
[0098] The logic operation submodule receives the decision matrix (including multiple rows of records shown in Table 3, assuming that a total of 100 records are processed) and a series of logic conditions and Boolean operation rules preset by the system. These rules are used to trigger specific business logic based on the data content. The following two logic conditions are set: Condition 1 (marking large invoices of the current month): TotalAmount>1000ANDIssueDate>='2025-04-01'ANDIssueDate<='2025-04-30' Condition 2 (marking test invoices): CONTAINS(InvoiceID,'TEST') The submodule processes the records in the decision matrix row by row and applies the following to each record: For all logical conditions, calculate their truth table matching degree (i.e. whether the condition is true): For the first record shown in Table 3 (InvoiceID="INV1001", IssueDate="2025-04-15", TotalAmount=1250.75): Evaluation condition 1: Check whether TotalAmount(1250.75) is greater than 1000, the result is true; check whether IssueDate("2025-04-15") is between April 1 and April 30, 2025, the result is true. Since it is a logical AND operation, both sub-conditions are true, so the matching result of condition 1 for this record is true. Evaluation condition 2: Check whether the string InvoiceID("INV1001") contains the substring "TEST", the result is false, so the matching result of condition 2 for this record is false. The submodule repeats this evaluation process for all 100 records. After processing all the records, it performs statistics: summarizes the total number of times each logical condition is triggered (that is, the matching result is true). Assume that the statistical results are: condition 1 (large amount of invoices for the current month) is triggered a total of 15 times, and condition 2 (test invoice) is triggered a total of 0 times. At the same time, the number of records that meet the conditions is recorded (in this case, equal to the number of triggers). These statistical data are used to generate a status bit flag, which is a collection or mapping that reflects whether the current batch of data meets specific business rules. Its content is: {"large amount of invoices for the current month": 15,"test invoice": 0}.
[0099] The instruction generation submodule filters the fields whose trigger times exceed the threshold in the status bit identification according to the write strategy trigger threshold, arranges the operation instructions according to the field coding order and the trigger time priority, and generates a write strategy instruction set.
[0100] The instruction generation sub-module receives the status bit identifiers {"Large amount of invoices in the current month": 15, "Test invoice": 0} and the write policy trigger thresholds configured in the system. These thresholds determine which status and its trigger frequency require the initiation of subsequent automated processing instructions. The thresholds are set as follows: For the status "Large amount of invoices in the current month", the trigger threshold is set to 10 times. The setting of this threshold refers to the company's financial risk control policy, which stipulates that when there are more than 10 invoices with an amount greater than 1000 yuan within the current month, an internal audit review process needs to be initiated. Calculation basis: Analyzing the data of the past 12 months, the average monthly number of large-amount invoices is 5, and the standard deviation is 2. To cover approximately 99% of the normal fluctuation range (using the mean plus 2.5 times the standard deviation), the calculated threshold is 5 + 2.5 * 2 = 10. For the status "Test invoice", the trigger threshold is set to 1 time. The setting basis is that any mixing of test data needs to be immediately identified and isolated. Therefore, as long as it appears 1 time, it must trigger processing. Calculation basis: The business requirement is zero tolerance, and the threshold is set to the minimum value of 1. The sub-module compares the trigger times in the status bit identifiers with the corresponding thresholds to screen the statuses that need to trigger instructions: "Large amount of invoices in the current month": Trigger times 15 times > threshold 10 times, this status is selected. "Test invoice": Trigger times 0 times < threshold 1 time, this status is not selected. Next, specific operation instructions are generated for the selected status "Large amount of invoices in the current month". The content of the instructions is determined according to the predefined business process, including: Instruction 1: Insert a record into the AuditLog audit log table, including the InvoiceID, TotalAmount of the triggering record, and the triggering reason "Large amount of invoices in the current month". Instruction 2: Update the NeedsReview field of the corresponding record in the Invoices invoice table and set its value to TRUE.Then, the sub-module arranges these operation instructions according to the logical order of the target database fields (the order of field processing within the instruction, such as recording the ID first and then the amount) and the priority of the trigger status (if there are multiple status triggers, the instruction execution order is determined according to the preset priority or the number of trigger times. Here, there is only one status trigger), and finally generates a structured write policy instruction set, the content of which is as follows: [{"trigger status": "Large-amount monthly invoice", "condition SQL": "TotalAmount>1000 AND IssueDate>='2025-04-01' AND IssueDate<='2025-04-30'", "instruction list": [{"action": "INSERT", "target_table": "AuditLog", "column_data": {"InvoiceID": "$SourceRecord.InvoiceID", "Amount": "$SourceRecord.TotalAmount", "Reason": "Large-amount monthly invoice"}}, {"action": "UPDATE", "target_table": "Invoices", "update_set": {"NeedsReview": true}, "where_clause_field": "InvoiceID", "where_clause_value": "$SourceRecord.InvoiceID"}]}, where $SourceRecord... represents dynamically obtaining the corresponding field values from the original data records that meet the conditions.
[0101] Please refer to Figure 6 , the script generation module includes:
[0102] The closure detection sub-module calls the XSD schema validation algorithm to parse the label closure rules of the nested label linked list, extracts the sequence matching degree of the label start symbol and end symbol, counts the number and position offset of the unclosed labels, and generates a closure check value;
[0103] The closure detection sub-module starts to work. It receives the nested label linked list (i.e., the node tree) representing the content of sheet1.xml constructed in the previous step and validates it according to the built-in XML basic specifications and the XSD (XML Schema Definition) schema rules for the specific structure of Excel_OOXML. The sub-module executes an algorithm that simulates the core logic of the XSD validator to check the closure of the labels: it maintains an internal stack structure. When traversing the node tree (or the original XML stream), when encountering a non-self-closing start label (such as <row> , <c>) it is pushed onto the stack; upon encountering an end tag (such as < / c> < / row> ,< / v> < / c> ), check whether its name matches the name of the start tag at the top of the stack. If it matches, pop the top element of the stack. If it doesn't match or the corresponding end tag is not found at the expected position (such as before the end of the parent node or before the start of the next sibling node), record it as a closing error. Through this process, the sub-module accurately calculates the sequence matching degree of the tag start and end delimiters. Assume that there are a total of 500 tag pairs to be closed in the processed sheet1.xml content, and it is detected that there are 2 places <c>If the tag is not closed correctly, the number of successfully matched pairs is 498, and the calculated matching degree is 498 / 500=0.996. The system sets the matching degree threshold: less than 0.99 is considered a serious structural problem, 0.99 to 0.999 means there are a few problems, and 1.0 is completely closed. The current 0.996 belongs to the interval with a few problems. At the same time, the submodule counts the specific number of unclosed tags as 2, and records their paths in the node tree and the estimated byte offset relative to the start content of the parent node: Error 1 occurs in the path / sheetData / row[3] / c [4], with an estimated offset of 150 bytes; Error 2 occurs at path / sheetData / row[5] / c[2], with an estimated offset of 280 bytes. Finally, a closure check value is generated, whose structure is: {"matching degree":0.996,"number of unclosed bytes":2,"error details":[{"path":" / sheetData / row[3] / c[4]","estimated offset":150},{"path":" / sheetData / row[5] / c[2]","estimated offset":280}]}.
[0104] The hierarchical verification submodule traverses the node hierarchy of the nested tag linked list based on the field encoding order of the write policy instruction set, recursively compares the hierarchical depth difference between the parent node and the child node, calculates the legitimacy weight of the node nesting path, and generates the node hierarchy depth;
[0105] The hierarchical verification submodule receives a set of write strategy instructions (which implies the order of the fields to be processed, such as InvoiceID, IssueDate, TotalAmount) and a nested tag list (node tree). The submodule's working goal is to verify whether the hierarchical structure of the data nodes in the tree conforms to the expected XSD architecture specification. It first focuses on the paths corresponding to the data fields involved in the instruction set in the node tree. For example, the InvoiceID value is usually located under the path pattern / sheetData / row / c[@r='A*'] / v, IssueDate is located under / sheetData / row / c[@r='B*'] / v, and TotalAmount is located under / sheetData / row / c[@r='C*'] / v. For each specific data node instance, such as the value node path of the first row of invoice number is / sheetData / row[1] / c[@r='A1'] / v, the submodule performs a recursive check, starting from the root node, comparing the hierarchical depth difference between the parent node and the child node level by level, and setting the root node <xml-document>(Virtual) level is 0, <sheetdata>is the first layer, <row>It is the second layer, <c>It is the third layer, <v>It is the 4th layer. For the path / sheetData / row[1] / c[@r='A1'] / v, check: (root -> sheetData) difference is 1, (sheetData -> row[1]) difference is 1, (row[1] -> c[@r='A1']) difference is 1, (c[@r='A1'] -> v) difference is 1. All level differences are 1, which conforms to the standard parent - child direct nesting relationship. Then, calculate the legality weight of the node nesting path. The calculation rule is defined as: the base weight W_base = 1.0. For each occurrence of a situation where the difference in the hierarchical depth of the parent - child nodes on the path is not equal to 1 (indicating skipping levels or abnormal nesting), multiply the current weight by a penalty factor PenaltyFactor = 0.5. The final weight W_legality = W_base * (PenaltyFactor ^ NumberOfViolations). For the above path, since all level differences are 1 and the number of violations NumberOfViolations = 0, the legality weight W_legality = 1.0 * (0.5 ^ 0) = 1.0. The system sets the legality weight threshold to 0.8. Values lower than this are determined to be illegal paths and need to be corrected because 0.8 means there is at least one serious level - skipping nesting (for example, the weight becomes 0.5). In this example, the weight 1.0 > 0.8, so the path is legal. The sub - module records the final hierarchical depth and its legality weight of all key data nodes, forming a set of node hierarchical depth information: {" / sheetData / row[1] / c[@r='A1'] / v": {"depth": 4, "legality weight": 1.0}, " / sheetData / row[1] / c[@r='B1'] / v": {"depth": 4, "legality weight": 1.0}, " / sheetData / row[1] / c[@r='C1'] / v": {"depth": 4, "legality weight": 1.0},...}, ensuring that subsequent scripts generate paths based on a correct structure.
[0106] The script correction sub - module corrects the termination position of unclosed tags according to the closure check value, adjusts the nesting order of illegal paths in the node hierarchical depth, reorganizes the label linked - list structure according to the XSD schema rules, generates a standard script file and writes it into the target system transaction queue, and outputs the written index number.
[0107] The script correction sub-module undertakes the final collation and output tasks. It receives the closure check value {"matching degree": 0.996, "number of unclosed": 2, "error details": [...]} and the node hierarchy depth information {...,"legitimacy weight": 1.0,...}, as well as the relevant XSD schema rules (as built-in logic). First, the sub-module corrects the node tree structure in memory according to the two unclosed tag errors reported in the closure check value and their location information / sheetData / row[3] / c[4] and / sheetData / row[5] / c[2]: it locates these two unclosed <c>nodes and according to the XSD rules (stipulating <c>The element must be within the parent <row>The element is internally closed and is typically after all its child elements and before the next sibling element <c>or the end tag of the parent element< / c> < / row> ), and add the corresponding end tag for it in the correct position< / c> , thus fixing the tag closure problem. Then, the sub-module checks whether there are paths in the node hierarchy depth information with a legal weight lower than 0.8. If there are (none in this example), it adjusts the nesting order of these illegal paths according to the element nesting rules defined by XSD. For example, it moves the incorrectly nested nodes under their correct parent nodes, or creates missing intermediate-level nodes as needed to meet the architecture requirements. After the above corrections, the sub-module conducts a final comprehensive check to ensure that the entire node tree structure strictly adheres to the XSD architecture rules, including the correct order of elements, the integrity of required sub-elements, the format compliance of attribute values, etc. Reorganize the tag linked list (node tree) structure according to this rule to obtain a memory data representation with a completely compliant structure. Then, based on this verified and corrected standard node tree data, combined with the previously generated write policy instruction set [{"trigger status": "large monthly invoice", "instruction list": [...]}], the sub-module generates a final standard script file, the format of which is determined according to the requirements of the target system, usually a set of SQL statements. For example, it generates 15 SQL statements for triggering the condition of "large monthly invoice", including 15 INSERT INTO AuditLog... and 15 UPDATE Invoices SET NeedsReview = TRUE WHERE InvoiceID =... statements. Take these generated SQL statements as a transaction unit and write them into the configured target database system transaction queue, waiting for the database to execute. As an operation voucher, the system records and outputs the unique identifier obtained by this write task in the transaction queue, that is, the write index number TXN_ID_987654321.
[0108] The above is only a preferred embodiment of the present invention and does not limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.< / c> < / v> < / c> < / row> < / sheetdata> < / c> < / v> < / v> < / sheetdata> < / v> < / v> < / v> < / v> < / v> < / v> < / v> < / v> < / v> < / sheetdata> < / v> < / sheetdata> < / v> < / v> < / v> < / v> < / v> < / sheetdata> < / sheetdata> < / sheetdata> < / sheetdata> < / row> < / sheetdata> < / v> < / c> < / c> < / row> < / row> < / sheetdata> < / row> < / c> < / row> < / c> < / v> < / c> < / row>
Claims
1. A data exchange system with a nested tag structure, characterized in that, The system comprises: The data carrier adaptation module is used to compare the MIME type identifier of the input file with the preset type correspondence table to generate a carrier type identifier, call the KMP pattern matching algorithm to locate the binary sequence features of the file header, generate an extraction strategy parameter set, and pass the carrier type identifier and the extraction strategy parameter set to the structure analysis module; A structure parsing module, used to call the DOM tree parsing algorithm through the carrier type identifier to process the metadata parameters in the extraction strategy parameter set, perform depth-first traversal on the file hierarchy path to generate a node sequence set, calculate the hierarchical nesting relationship and output a nested tag list and a field identifier set, and pass them to the field mapping module; A field mapping module, used to match the field identification set with the target database field code through the Hungarian algorithm, generate a field mapping table, and transmit it to the logic driving module synchronously with the nested tag chain table; The logic driving module is used to construct a decision matrix through the field mapping table, perform logical operations on the field values in the nested tag list to generate a status bit identifier, generate a write strategy instruction set after triggering a threshold, and transmit it to the script generation module synchronously with the nested tag list.
2. The data exchange system with a nested tag structure according to claim 1, characterized in that, The carrier type identifier is specifically PDF, XML, and JSON. The extraction strategy parameter set includes offset, delimiter, and block size. The nested tag list includes parent-child relationship, hierarchical depth, and closed state. The field identifier set includes tag name, attribute hash, and path index. The field mapping table specifically refers to encoding matching pairs, type conversion rules, and constraints. The write strategy instruction set includes batch operation flags, transaction isolation levels, and rollback conditions.
3. The data exchange system with a nested tag structure according to claim 2, characterized in that, The data carrier adaptation module comprises: The type identification comparison submodule obtains the MIME type identifier of the input file, compares the character sequence of the identifier with the character sequence of the extension in the preset type correspondence table item by item, determines the mapping relationship priority between the identifier and the extension, and completes the type matching according to the weight value in the preset table to generate a carrier type identifier; The feature location submodule builds a partial matching table based on the KMP pattern matching algorithm, compares the input file binary stream with the preset file header feature sequence in byte order through a sliding window, calculates the index position of the first appearance of the feature sequence in the binary stream, and generates a feature offset; The parameter generation submodule determines the starting position of the data block according to the segment length constraint corresponding to the carrier type identifier and the characteristic offset, divides the binary stream according to the byte step and filters the non-continuous area to generate an extraction strategy parameter set.
4. The data exchange system with a nested tag structure according to claim 3, characterized in that, The structure analysis module comprises: The metadata parsing submodule matches the tag separation rule in the DOM tree parsing algorithm based on the carrier type identifier, maps the metadata parameters in the extraction strategy parameter set according to the tag name length and the attribute key value hash value, constructs a parent-child relationship list between nodes, and generates a node tree structure; The traversal control submodule accesses the child nodes layer by layer according to the number of child nodes of the root node of the node tree structure according to the depth-first traversal rule, records the node access order and the difference between the level numbers, and generates a traversal index sequence; The nested relationship sub-module extracts the hierarchical number difference between adjacent nodes in the traversal index sequence, counts the number of parent node branches and the difference in nested depth, calculates the label nested hierarchy weight and the field identifier hash value, and generates the nested hierarchy coefficient and the field identifier set.
5. The data exchange system with a nested tag structure according to claim 4, characterized in that, The field mapping module includes: The weight calculation sub-module extracts the character hash value of multiple field names in the field identifier set and the character length of the target database field encoding, compares the coincidence degree of the character sequences of the field name and the encoding, and calculates the weight value according to the semantic similarity algorithm to generate the field weight matrix. The matching execution sub-module splits the field weight matrix into row indexes and column indexes based on the Hungarian algorithm, traverses all nodes to construct the edge weight relationship of the bipartite graph, iteratively calculates the augmenting path and updates the matching status, and generates the initial matching pairs. The table optimization sub-module filters out the mapping items with weight values lower than the field conflict threshold in the initial matching pairs according to the uniqueness constraint condition of the target database field encoding, reassigns the matching priorities of the conflicting fields, and generates the field mapping table.
6. The data exchange system with a nested tag structure according to claim 5, characterized in that, The logic driving module includes: The matrix construction sub-module calls the field encoding in the field mapping table and the type encoding of the field value in the nested label linked list, extracts the length parameter and the priority weight of the field value, and arranges the mapping relationship between the field encoding and the type encoding according to the row and column index rules to generate the decision matrix. The logic operation sub-module compares the matching degree of the field value with the truth table of the preset logic condition based on the Boolean operation rule of the field value in the decision matrix, counts the number of fields and the trigger times that meet the conditions, and generates the status bit identifier. The instruction generation sub-module filters out the fields with trigger times exceeding the threshold in the status bit identifier according to the write policy trigger threshold, arranges the operation instructions according to the field encoding order and the trigger time priority, and generates the write policy instruction set.
7. The data exchange system with a nested tag structure according to claim 6, characterized in that, The system further includes: The script generation module is used to call the XSD schema verification algorithm to detect the closure of the nested label linked list, recursively verify the node hierarchy based on the write policy instruction set, output the standard script file after correcting the errors, and write the standard script file into the target system transaction queue. The standard script file includes the syntax tree structure, the verification log, and the transaction handle.
8. The data exchange system with a nested tag structure according to claim 7, characterized in that The script generation module includes: The closure detection sub-module calls the XSD schema verification algorithm to parse the label closure rule of the nested label linked list, extracts the sequence matching degree of the label start symbol and the end symbol, counts the number and position offset of the unclosed labels, and generates the closure check value. The hierarchy verification sub-module traverses the node hierarchy of the nested label linked list based on the field encoding order of the write policy instruction set, recursively compares the difference in the hierarchical depth between the parent node and the child node, calculates the legality weight of the node nested path, and generates the node hierarchical depth. The script correction sub-module corrects the position of the end symbol of the unclosed label according to the closure check value, adjusts the nested order of the illegal paths in the node hierarchical depth, reorganizes the label linked list structure according to the XSD schema rules, generates the standard script file and writes it into the target system transaction queue, and outputs the write index number.
Citation Information
Patent Citations
Model file processing method and device and computer equipment
CN119047424A
System and method for applying development patterns for component based applications
CN1834908A