Off-line detection report automatic generation system and method based on intelligent template matching

CN122389843BActive Publication Date: 2026-09-25BAOTOU WATER QUALITY TESTING TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610757884.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-29
Publication Date
2026-09-25
Estimated Expiration
2046-05-29

AI Technical Summary

Technical Problem

[0007]针对现有技术的不足,本发明提供了基于智能模板匹配的离线检测报告自动生成系统及方法,解决了现有报告自动生成系统因依赖硬编码规则导致匹配准确率与容错能力低,在算力受限环境下运行不稳定,以及对复杂异构模板结构解析能力弱的问题

Benefits of technology

[0042]本发明通过设置锚点确立模块与仲裁推演模块,将模板文件转换为模板有向图,将结构化检测数据文件转换为数据实体图,利用量化词嵌入模型构建基准锚点集合,并以基准锚点集合作为原点进行拓扑路径推演,在推演路径发生冲突时,系统计算拓扑路径边数的倒数作为拓扑引力权重执行匹配结果仲裁,替代了传统依赖绝对位置绑定的匹配规则,在模板排版结构发生变动或检测数据嵌套层级出现差异的条件下,依靠图论拓扑关系自动确立最优匹配节点,提升了数据自动化注入的准确率与容错能力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122389843B_ABST
    Figure CN122389843B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, and discloses an offline detection report automatic generation system and method based on intelligent template matching, which comprises a template reconstruction module, an entity graph construction module, an anchor point establishment module, an arbitration deduction module and a report generation module. The system projects data containers in a template file to a two-dimensional atomic grid matrix to generate a template directed graph containing text nodes; key nodes and value nodes in a structured detection data file are extracted to generate a data entity graph; text nodes and key nodes are input into a quantitative word embedding model to map and construct a benchmark anchor point set; the benchmark anchor point set is taken as an origin to deduce in the template directed graph and the data entity graph, and when a deduction path conflicts, a topological path edge number reciprocal is calculated as a topological gravity weight. The application improves the accuracy and fault tolerance of data automatic injection through topological path deduction, avoids resource overload in an offline environment by combining an adaptive degradation mechanism, and considers matching reliability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, specifically to an automatic offline detection report generation system and method based on intelligent template matching. Background Technology

[0002] The offline test report automatic generation system is a basic software support tool in the industrial testing and quality control process. After completing the workpiece testing task, the testing equipment deployed in the industrial production environment will output structured or semi-structured test data files. The report automatic generation system is responsible for receiving the aforementioned test data files and automatically filling the test results into the pre-designed report template, and finally outputting a standardized test report document for archiving or distribution.

[0003] In the current data processing flow, existing systems generally adopt mapping rules based on absolute physical coordinates or hard-coded field names. When the software program starts up, it first reads the target template file and grabs the blank cells to be filled according to fixed pixel position identifiers or preset regular expressions. Then the system reads the parameter names in the test data file and compares the strings one by one with the set corresponding dictionary. After a successful match, the program directly writes the test value to a fixed memory offset address or cell object node to complete the one-way injection of data. However, this method has poor robustness.

[0004] For example, Chinese patent application CN202511717600.4 discloses a method and system for fast data filling based on hash indexes. This method achieves some optimization in matching speed by pre-establishing a hash mapping relationship between template placeholders and data fields. However, the essence of this scheme still falls into the category of hard-coded rules. The matching process heavily relies on the precise string consistency between the template placeholder text and the data field names. Once the template layout is slightly adjusted (such as changing cell positions, adding or deleting rows and columns), or the field naming of the data source undergoes slight changes (such as adding prefixes or using synonyms), the pre-set hash index will completely fail, leading to widespread data injection failures. The system lacks the necessary semantic understanding and structural fault tolerance capabilities.

[0005] Current technologies rely on absolute position binding or hard-coded rules when performing actual tasks. When encountering situations where the template layout structure is adjusted or the nesting level of the detection data changes, the system cannot perform dynamic node pairing, resulting in low data injection accuracy and insufficient fault tolerance. At the same time, some systems that integrate model inference calculations have high requirements for computing resources. When running on offline industrial control equipment with limited computing power, they are prone to excessive memory usage and computation timeouts. Due to the lack of adaptive degradation processing mechanisms, the operation stability is poor. Furthermore, existing file parsing schemes directly read the underlying physical pixel coordinates. When faced with complex heterogeneous templates containing cross-row and cross-list cells and multi-level nested text boxes, they cannot convert the physical layout into a unified discrete graph data structure. The weak template structure parsing capability leads to incomplete layout feature extraction, hindering the subsequent automated data entry process.

[0006] Therefore, the purpose of this invention is to provide an automatic offline detection report generation system and method based on intelligent template matching to address the shortcomings of the prior art. Summary of the Invention

[0007] To address the shortcomings of existing technologies, this invention provides an offline detection report automatic generation system and method based on intelligent template matching. This solves the problems of low matching accuracy and fault tolerance caused by the reliance on hard-coded rules in existing automatic report generation systems, unstable operation under computing power-limited environments, and weak ability to parse complex heterogeneous template structures.

[0008] To achieve the above objectives, the present invention provides the following technical solution:

[0009] The first aspect of this invention provides an offline detection report automatic generation system based on intelligent template matching, comprising:

[0010] The template reconstruction module is used to project the data containers in the template file onto a two-dimensional atomic mesh matrix to generate a directed template graph containing text nodes;

[0011] The entity graph construction module is used to extract key nodes and value nodes from structured detection data files to generate data entity graphs.

[0012] The anchor point establishment module is used to input the text nodes of the template directed graph and the key nodes of the data entity graph into the quantized word embedding model mapping to construct a set of benchmark anchor points;

[0013] The arbitration deduction module is used to generate deduction paths in the template directed graph and data entity graph by taking the set of reference anchor points as the origin. When the deduction paths conflict, the number of topological path edges is calculated. Based on the number of topological path edges, the topological gravity weight is calculated by an exponential decay function. The data carried in the value node pointed to by the deduction path with the largest topological gravity weight is extracted as the target matching result.

[0014] The report generation module is used to output a detection report file based on the target matching results.

[0015] Preferably, the template reconstruction module is specifically used to obtain the width and height values ​​of the data container, extract the smallest non-zero width value as the basic column width, and extract the smallest non-zero height value as the basic row height; and obtain the total page width, total page height, basic column width, and basic row height of the template file to construct a two-dimensional atomic mesh matrix;

[0016] The starting row index and starting column index of the data container in the two-dimensional atomic grid matrix are calculated based on the vertex x-coordinate and y-coordinate values ​​of the data container. The starting row index is calculated by dividing the y-coordinate value by the base row height and rounding down. The starting column index is calculated by dividing the x-coordinate value by the base column width and rounding down. The ending row index and ending column index are calculated using the same operation logic based on the width and height values.

[0017] Data containers that have a difference between the starting row index and the ending row index, or a difference between the starting column index and the ending column index, are grouped into bounding box structures, and data containers bound to bounding box structures are marked as macro nodes.

[0018] Preferably, the template reconstruction module is also used to traverse the set of data containers in the two-dimensional atomic mesh matrix to perform pairwise position comparisons, and to establish spatial orientation edges and hierarchical logical edges based on the boundary index data of the data containers. Spatial orientation edges include horizontal spatial orientation edges and vertical spatial orientation edges. Taking the first data container and the second data container as an example, under the condition that the starting column index of the second data container is equal to the ending column index of the first data container plus a constant one, and the maximum value between the starting row index of the first data container and the starting row index of the second data container is less than or equal to the minimum value between the ending row index of the first data container and the ending row index of the second data container, the horizontal adjacency state of the first data container and the second data container is established and a horizontal spatial orientation edge is generated.

[0019] Hierarchical logical edges include nested dependent edges established under the condition that the boundary index range of one data container completely contains another data container, as well as merged associated edges pointing from macro nodes to other data containers within the basic grid cell covered by macro nodes.

[0020] Preferably, the entity graph construction module is specifically used to establish key-value binding edges between key nodes and value nodes according to identifier matching rules; and to establish parent-child dependent edges between key nodes according to the hierarchy depth parameter and parent node pointing parameter of the key node. Taking the first key node and the second key node as an example, under the condition that the hierarchy depth parameter of the second key node is equal to the hierarchy depth parameter of the first key node plus a constant one, and the parent node pointing parameter of the second key node is equal to the node identifier of the first key node, a parent-child dependent edge from the first key node to the second key node is generated.

[0021] When multiple sibling key nodes or multiple sibling value nodes have the same parent node pointing parameter, assign an incrementing sequence index value and attach the sequence index value as an edge weight attribute to the parent-child node dependent edge or key-value bound edge.

[0022] Preferably, the anchor point establishment module is specifically used to perform word segmentation and stop word filtering operations on the layout string data in the template file and the field name string data in the structured detection data file to generate a lexical sequence; input the lexical sequence into the quantized word embedding model to perform forward propagation calculation to output a high-dimensional feature vector; and perform L2 normalization processing on the high-dimensional feature vector to generate template feature vector and key node feature vector.

[0023] The normalization process calculates the sum of the squares of the components of the original feature vector in all dimensions and takes the square root. The original feature vector is then divided by the square root result to obtain the normalized feature vector projected onto the unit hypersphere.

[0024] Preferably, when constructing the benchmark anchor point set, the anchor point establishment module is specifically used to calculate the cosine similarity between the template feature vector and the key node feature vector; filter text nodes and key node pairs with a cosine similarity greater than a preset similarity threshold to construct the benchmark anchor point set, wherein the preset similarity threshold is set to a value range between 0.75 and 0.90.

[0025] Preferably, the anchor point establishment module is also used to execute an adaptive degradation mechanism to monitor the model inference time and memory usage of the quantified word embedding model in physical memory; and to trigger the adaptive degradation mechanism to stop the feature mapping process when the model inference time exceeds a preset time threshold or the memory usage exceeds a safety threshold.

[0026] The Levenstein edit distance algorithm is used to calculate the minimum number of single-character edit operations between text nodes and key nodes. The minimum number of single-character edit operations is inversely mapped to the downgraded matching confidence, which replaces the feature mapping process as the basis for node matching and completes the adaptive downgrade mechanism.

[0027] Preferably, the anchor point establishment module is also used to trigger a minimal intervention mechanism to extract isolated node pairs whose downgraded matching confidence is lower than a preset judgment threshold; extract the context topology boundary data of isolated node pairs in the template directed graph and the data entity graph, and package the isolated node pairs and the context topology boundary data into an independent dataset.

[0028] Based on an independent dataset, an intervention task package is generated, and the received node pairing correction instructions are rendered to establish a mapping connection for isolated node pairs. The confirmed mapping state is then synchronously persisted to the local cache to complete the minimal intervention mechanism.

[0029] Preferably, the arbitration deduction module is specifically used to initialize the template search queue and the data search queue, push text nodes into the template search queue, and push key nodes into the data search queue; pop a current text node from the template search queue and a current key node from the data search queue, and locate the current value node bound to the current key node; probe adjacent data containers with empty internal data states along the current text node as target data containers to be filled; and write the measured detection data carried in the current value node into the memory address of the target data container, provided that the data structure type of the target data container is consistent with that of the current value node.

[0030] Preferably, the arbitration deduction module is also used to perform matching conflict detection of the deduction path. It counts the number of hierarchical logical edges in the directed graph of the template pointing to the next level data container as the template local out-degree parameter; counts the number of nested dependent edges in the key node of the statistical entity graph pointing to the next level child key node as the data local out-degree parameter; uses the absolute value of the difference between the template local out-degree parameter and the data local out-degree parameter as the structural error value; determines that there is a topological structure conflict in the deduction path when the structural error value is greater than zero; extracts the memory address of the template node, the data node identifier, and the conflict type code at the conflict location to generate a conflict tracing log; writes the conflict tracing log to the local storage device; and skips the deduction branch where the conflict occurs to complete the matching conflict detection.

[0031] Preferably, when performing matching conflict detection, the arbitration deduction module is also used to perform data type conflict detection, extract the preset field constraint rules of the target data container in the template directed graph and the physical storage type of the measured detection data in the value node; if the physical storage type of the measured detection data exceeds the enumeration range allowed by the field constraint rules, it is determined that there is a data type conflict in the deduction path, and the conflict type code corresponding to the data type conflict is written into the conflict tracking log.

[0032] Preferably, when handling path conflict, the arbitration deduction module also executes an arbitration strategy based on exponentially decaying gravity weights. When multiple candidate value nodes compete for a single target data container, or a single value node competes for multiple target data containers, the node pair that most recently successfully completed data injection is used as the reference anchor point; the shortest path algorithm is called to calculate the shortest topological hop count from each candidate node to the reference anchor point; an exponential decay function is used to map the shortest topological hop count to a gravity weight value, where the gravity weight value is equal to the product of the negative preset distance decay constant of the natural constant and the shortest topological hop count raised to the power of the power of the value of the gravity weight; the candidate node with the largest gravity weight value is selected as the target matching result; and when gravity weight values ​​are equal, the candidate node with the earliest parsing order in the original file is selected.

[0033] Preferably, the report generation module is specifically used to encode the target matching results into rich text data blocks with layout tags according to the layout specifications of the detection report file; allocate a contiguous byte buffer in physical memory, and calculate the starting offset address of the current rich text data block based on the starting offset address of the previously written rich text data block and the length of the data in bytes. The starting offset address of the current rich text data block is obtained by adding the starting offset address of the previously written rich text data block, the actual length of the data in bytes, and the length of the interval tag in bytes. The rich text data blocks are serialized and written to the contiguous byte buffer to construct a report document object stream.

[0034] Preferably, before constructing the report document object stream, the report generation module also performs pagination truncation calculations. It estimates the rendering height based on the total number of characters, font size, and physical length of the image in the rich text data block. The rendering height is equal to the sum of the product of the total number of characters, font size, and screen resolution conversion constant, and the physical length of the image. Based on the preset physical paper size, margin parameters, and the cumulative height value already occupied by the current page, it determines whether the remaining space of the current page is sufficient to accommodate the rich text data block. If the rendering height is greater than the total height parameter minus the sum of the upper and lower margin heights minus the cumulative height value, it determines that the remaining space of the current page is insufficient, and inserts a forced page break before the current writing position of the report document object stream.

[0035] A second aspect of the present invention provides a method for automatically generating offline detection reports based on intelligent template matching, comprising:

[0036] Project the data container in the template file onto a two-dimensional atomic mesh matrix to generate a template directed graph containing text nodes;

[0037] Extract key nodes and value nodes from structured detection data files to generate a data entity graph;

[0038] Input the text nodes of the template directed graph and the key nodes of the data entity graph into the quantized word embedding model to construct a set of benchmark anchor points;

[0039] Using the set of reference anchor points as the origin, a deduction path is generated in the template directed graph and the data entity graph. When the deduction path conflicts, the number of edges of the topological path is calculated. Based on the number of edges of the topological path, the topological gravity weight is calculated through an exponential decay function. The value node data pointed to by the deduction path with the largest topological gravity weight is extracted as the target matching result.

[0040] Output a detection report file based on the target matching results.

[0041] This invention provides an automatic offline detection report generation system and method based on intelligent template matching, which has the following beneficial effects:

[0042] This invention converts template files into directed template graphs and structured detection data files into data entity graphs by setting up an anchor point establishment module and an arbitration deduction module. It constructs a set of benchmark anchor points using a quantized word embedding model and uses the set of benchmark anchor points as the origin to perform topological path deduction. When a conflict occurs in the deduction path, the system calculates the reciprocal of the number of edges of the topological path as the topological gravity weight to arbitrate the matching result. This replaces the traditional matching rules that rely on absolute position binding. Under the condition that the template layout structure changes or the nesting level of the detection data differs, the optimal matching node is automatically established based on graph theory topological relationships, which improves the accuracy and fault tolerance of automated data injection.

[0043] This invention, by setting an adaptive degradation mechanism and a minimal intervention mechanism, monitors in real time the inference time and occupancy rate of the quantified word embedding model in physical memory. When the hardware resource load exceeds the safety threshold, the system actively stops feature vector mapping and switches to a degradation algorithm that calculates the minimum number of single-character editing operations between text nodes and key nodes. For isolated node pairs whose degradation matching confidence is lower than the judgment threshold, the system packages context topology boundary data to trigger local manual review. This avoids system resource overload caused by model calculations in offline industrial control environments with limited computing power. At the same time, it controls manual intervention to the scope of local nodes with mismatch risk, thus balancing the stability of system operation and the reliability of data matching.

[0044] This invention, by setting up a template reconstruction module, extracts the minimum non-zero width and height values ​​of the template file data container as basic units, constructs a two-dimensional atomic grid matrix, and establishes spatial orientation edges and hierarchical logical edges based on the index boundaries of the data container in the two-dimensional atomic grid matrix. This standardizes the absolute physical pixel coordinates of the document's underlying layer into a discrete graph data structure, eliminating data parsing obstacles caused by differences in formatting rules of different document types. This enables the system to uniformly process complex tables spanning rows and columns and multi-layered nested text boxes, improving the structural parsing capability of heterogeneous detection report templates. Attached Figure Description

[0045] Figure 1 This is a flowchart of the offline detection report automatic generation method of the present invention;

[0046] Figure 2 This is a diagram illustrating the template file parsing and directed graph generation process of the present invention.

[0047] Figure 3 This is a diagram illustrating the structured data parsing and data entity graph generation process of the present invention;

[0048] Figure 4 This is a diagram illustrating the feature vector mapping and reference anchor point establishment process of the present invention.

[0049] Figure 5 This is a diagram illustrating the bidirectional topology synchronization deduction and conflict detection process of the present invention;

[0050] Figure 6 This is a diagram illustrating the memory mapping conversion and report file output process of the present invention;

[0051] Figure 7 This is a system architecture diagram of the present invention;

[0052] Figure 8 This is a graph showing the decay of the gravity weight of the present invention with the number of hops along the shortest topological edge;

[0053] Figure 9 This is a line graph showing the load change of the offline environment adaptive degradation mechanism of the present invention.

[0054] Figure 10 This is a bar chart comparing the offline detection report automatic generation method of the present invention with traditional methods in terms of generation time and error rate. Detailed Implementation

[0055] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0056] See Figure 1 and Figure 7 This invention provides an automatic offline detection report generation system based on intelligent template matching, including: a template reconstruction module, an entity graph construction module, an anchor point establishment module, an arbitration deduction module, and a report generation module.

[0057] The template reconstruction module is used to read template files, parse document object models to extract data containers; generate a two-dimensional atomic mesh matrix based on the physical coordinates of the data containers, project the data containers into the two-dimensional atomic mesh matrix; establish spatial orientation edges and hierarchical logical edges based on the position boundaries of the data containers in the two-dimensional atomic mesh matrix, and generate a template directed graph.

[0058] The entity graph construction module is used to read structured detection data files, extract key nodes and value nodes, establish dependent edge relationships between key nodes and value nodes based on the nested hierarchical features within the structured detection data files, and generate a data entity graph.

[0059] The anchor point establishment module is used to load the quantized word embedding model. It inputs the text nodes in the template directed graph and the key nodes in the data entity graph into the quantized word embedding model for feature vector mapping. It calculates the cosine similarity and filters the corresponding node pairs whose cosine similarity meets the preset threshold to build a benchmark anchor point set. The preset threshold is set according to the experience data of historical matching tests, and is set between 0.75 and 0.90 to balance the matching accuracy and recall.

[0060] The arbitration deduction module is used to perform node expansion deduction synchronously along the edges using the anchor points in the set of benchmark anchor points as the origin and a breadth-first search algorithm. When multiple deduction paths point to the same blank node to be filled and correspond to different value nodes, the module calculates the number of topological path edges from the blank node to be filled to the benchmark anchor points of each conflicting path, and takes the reciprocal of the number of topological path edges as the topological gravity weight. The deduction path of the benchmark anchor point with the largest topological gravity weight is set as the valid path, and the value node data pointed to by the valid path is extracted as the target matching result.

[0061] The report generation module is used to write value node data into a template memory object based on the target matching results, render the template memory object, and output a detection report file.

[0062] See Figure 1This invention provides a method for automatically generating offline detection reports based on intelligent template matching, comprising the following steps:

[0063] S10: Read the template file, parse the document object model to extract the data container; generate a two-dimensional atomic mesh matrix based on the physical coordinates of the data container, project the data container into the two-dimensional atomic mesh matrix, establish the spatial orientation edge and hierarchical logical edge based on the position boundary of the data container in the two-dimensional atomic mesh matrix, and generate a template directed graph.

[0064] S20: Read the structured detection data file and extract key nodes and value nodes; establish the subordinate edge relationship between key nodes and value nodes based on the nested hierarchical features in the structured detection data file, and generate a data entity graph;

[0065] S30, Load the quantized word embedding model, input the text nodes in the template directed graph and the key nodes in the data entity graph into the quantized word embedding model to perform feature vector mapping and calculate the cosine similarity; filter the corresponding node pairs whose cosine similarity meets the preset threshold and construct the benchmark anchor point set.

[0066] S40: Using the anchor points in the set of benchmark anchor points as the origin, a breadth-first search algorithm is used to simultaneously expand and extrapolate nodes. When multiple extrapolation paths point to the same blank node to be filled and correspond to different value nodes, the number of topological path edges from the blank node to be filled to the benchmark anchor points of each conflicting path is calculated, and the reciprocal of the number of topological path edges is taken as the topological gravity weight. The extrapolation path of the benchmark anchor point with the largest topological gravity weight is set as the valid path, and the value node data pointed to by the valid path is extracted as the target matching result.

[0067] S50: Write the value node data into the template memory object based on the target matching result, render the template memory object and output the detection report file.

[0068] The technical content of reading the template file and parsing the document object model to extract the data container in step S10 is explained in detail. Step S10 is specifically divided into the following sub-steps:

[0069] S101 identifies the file format type of the template file and loads the data stream of the template file into the memory of the industrial control computer. The template file specifically includes Office Open XML standard format files and portable document format files. The Office Open XML standard format files include word processing document format files and spreadsheet format files.

[0070] S102, parse the underlying document object model of the template file, generate a document object model tree and store it in memory. For the unpacking of the underlying data stream of the document and the construction of the document object model tree, those skilled in the art can use existing open source parsing components.

[0071] S103: Traverse the nodes of the Document Object Model tree and extract data containers that can carry text data content through preset node label filtering rules. At the underlying data structure level, the data containers include table cell nodes and text box nodes in the Document Object Model tree. The system stores the filtered table cell nodes and text box nodes into the data container collection.

[0072] S104: Read the rendering attribute label of each data container in the data container set, obtain the coordinate and size attributes of each data container relative to the origin of the document page, and for each data container in the data container set, extract the x-coordinate, y-coordinate, width, and height values ​​of the top left corner vertex of the data container. The system combines the x-coordinate, y-coordinate, width, and height values ​​into a coordinate attribute structure, binds the coordinate attribute structure with the corresponding data container with a unique identifier, and saves it to the local cache as the reference positioning data for generating the atomic mesh matrix.

[0073] See Figure 2 The technical content of step S10, which involves generating a two-dimensional atomic mesh matrix based on the physical coordinates of the data container and projecting the data container onto the two-dimensional atomic mesh matrix, is explained in detail. The process of generating a two-dimensional atomic mesh matrix based on the physical coordinates of the data container and projecting the data container onto the two-dimensional atomic mesh matrix specifically includes the following sub-steps:

[0074] S105, traverse all coordinate attribute structures bound to data containers in the local cache, obtain the set of width values ​​and the set of height values ​​contained in the coordinate attribute structure, the system extracts the smallest non-zero width value from the set of width values ​​as the basic column width, and extracts the smallest non-zero height value from the set of height values ​​as the basic row height.

[0075] S106. Based on the total page width, total page height, basic column width, and basic row height of the template file, a two-dimensional atomic grid matrix is ​​constructed in memory. The underlying data structure of the two-dimensional atomic grid matrix is ​​a two-dimensional array. For memory allocation operations of the two-dimensional array data structure, those skilled in the art can use the built-in memory management functions of standard programming languages.

[0076] S107, Calculate the grid index coordinates of the data container in the two-dimensional atomic grid matrix. Convert the physical coordinates of the data container into grid indices. Based on the vertex x and y coordinate values ​​stored in the coordinate attribute structure, the system calculates the starting row index and starting column index of the data container in the two-dimensional atomic grid matrix. The formulas for calculating the starting row index and starting column index are as follows:

[0077] ;

[0078] ;

[0079] In the formula, Indicates the starting row index of the calculated output; Indicates the starting column index of the calculated output; Represents the x-coordinate value of the data container; Represents the vertical axis value of the data container; Indicates the height of the base row to be extracted; Indicates the base column width to be extracted; This indicates the floor function.

[0080] The system uses the same calculation logic: it adds the x-coordinate value of the data container to the width value to obtain the boundary x-coordinate, and adds the y-coordinate value of the data container to the height value to obtain the boundary y-coordinate. The system divides the boundary x-coordinate by the base column width and rounds down to obtain the terminating column index, and divides the boundary y-coordinate by the base row height and rounds down to obtain the terminating row index.

[0081] S108, for data containers where there is a difference between the starting row index and the ending row index, or a difference between the starting column index and the ending column index, the system combines the starting row index, the starting column index, the ending row index, and the ending column index into a bounding box structure. The bounding box structure is the bounding box, which is specifically represented by merging cells across rows and merging cells across columns. The system marks the data containers bound to the bounding boxes as macro nodes. Macro nodes occupy multiple consecutive basic grid cells in the two-dimensional atomic grid matrix. By establishing the two-dimensional atomic grid matrix and generating macro nodes, the physical coordinates are converted into grid block positions.

[0082] See Figure 2 This document provides a detailed explanation of the technical content of step S10, which involves establishing spatial orientation edges and hierarchical logical edges based on the position boundaries of the data container in the two-dimensional atomic mesh matrix, and generating a template directed graph. The process of establishing spatial orientation edges and hierarchical logical edges based on the position boundaries of the data container in the two-dimensional atomic mesh matrix and generating a template directed graph specifically includes the following sub-steps:

[0083] S109: Extract the boundary index data of all data containers from the two-dimensional atomic grid matrix. The boundary index data includes the starting row index, the ending row index, the starting column index, and the ending column index. The system traverses the set of data containers in the two-dimensional atomic grid matrix and performs pairwise position comparisons on the data containers to establish spatial orientation edges. The lower-level features of the spatial orientation edges include the horizontal spatial orientation edges that indicate the horizontal adjacency relationship of the data containers and the vertical spatial orientation edges that indicate the vertical adjacency relationship of the data containers.

[0084] S110, determine the horizontal and vertical adjacency status between data containers based on the boundary index data. Taking any two different data containers as an example, denoted as the first data container and the second data container, the system determines that the first data container and the second data container meet the horizontal adjacency condition using the following formula:

[0085] ;

[0086] ;

[0087] In the formula, Indicates the starting column index of the second data container; Indicates the index of the terminating column of the first data container; Indicates the starting row index of the first data container; Indicates the starting row index of the second data container; Indicates the index of the terminating row of the first data container; Indicates the index of the terminating row of the second data container; This represents the function that takes the maximum value. This represents the function that takes the minimum value.

[0088] Under the condition that the above two formulas are true at the same time, the system determines that the second data container is located to the right of the first data container and generates a horizontal spatial orientation edge from the first data container to the second data container. The system uses alternating row and column variables to determine the vertical adjacency condition. When it is determined that the second data container is located directly below the first data container, a vertical spatial orientation edge from the first data container to the second data container is generated.

[0089] S111, establish hierarchical logical edges based on the inclusion status of boundary index data. The lower-level features of the hierarchical logical edges include nested dependent edges indicating the nested dependent relationship of containers, and merged associated edges indicating the disassembly correspondence of macro nodes. Under the condition that the second data container is completely within the boundary index range of the first data container, the system generates nested dependent edges from the first data container to the second data container. For data containers marked as macro nodes, the system obtains all basic grid cells covered by the macro node and generates merged associated edges from the macro node to other data containers within the range of basic grid cells.

[0090] S112, the extracted data container is used as the vertex set of the graph data structure, and the established spatial orientation edges and hierarchical logical edges are used as the directed edge set of the graph data structure. A template directed graph is instantiated in memory. The template directed graph reflects the two-dimensional layout structure and data nesting logic of the detection report template. For the memory instantiation of the graph data structure and the creation of graph vertex objects, those skilled in the art can call the graph data structure implementation of existing programming languages.

[0091] See Figure 3 The technical content of reading the structured detection data file and extracting key nodes and value nodes in step S20 is explained in detail. The process of reading the structured detection data file and extracting key nodes and value nodes specifically includes the following sub-steps:

[0092] S201: Obtain the structured test data file generated by the offline test equipment, identify the text encoding format and data serialization type of the structured test data file. The lower-level features of the structured test data file include JavaScript object simplified spectrum format files and Extensible Markup Language format files. The system loads the data stream of the structured test data file into the memory of the industrial control computer and establishes an initial character buffer.

[0093] S202, the corresponding text parsing component is called to deserialize the data stream in the initial character buffer and generate an abstract syntax tree residing in memory. For the data stream deserialization and abstract syntax tree generation operations of JavaScript object musical notation files and Extensible Markup Language (EXPLAIN) format files, those skilled in the art can use existing data parsing libraries.

[0094] S203, traverse each data field node in the abstract syntax tree, extract the name attribute and data content attribute contained in the data field node, the system performs a depth-first traversal algorithm on the abstract syntax tree, identify the character content carrying the name of the detection item and the specific value carrying the detection result in the abstract syntax tree, the system instantiates the extracted name attribute as a key node in the graph data structure, and the extracted data content attribute as a value node in the graph data structure. The key node maps to the name of the detection item output by the offline detection device, and the value node maps to the actual measured data of the corresponding detection item.

[0095] S204. Assign a globally unique node identifier to each generated key node and each value node. The system extracts the hierarchical depth parameters and parent node pointing parameters of the key nodes and value nodes in the original structured detection data file. The system combines the node identifier, hierarchical depth parameters, and parent node pointing parameters into a node attribute dictionary. The system binds the node attribute dictionary with the corresponding key nodes and value nodes and saves it to the local cache as the base logical parameters for generating the topology of the data entity graph.

[0096] Reference Figure 3 This document provides a detailed explanation of the technical content of step S20, which involves establishing dependent edge relationships between key nodes and value nodes based on the nested hierarchical features within the structured inspection data file to generate a data entity graph. The process of establishing dependent edge relationships between key nodes and value nodes based on the nested hierarchical features within the structured inspection data file to generate a data entity graph is specifically divided into the following sub-steps:

[0097] S205, traverse the node attribute dictionary in the local cache, establish the mapping relationship between key nodes and value nodes according to the identifier matching rules in the node attribute dictionary, the subordinate features of the dependent edge relationship include key-value binding edges indicating the corresponding key-value pairs at the same level, and parent-child node dependent edges indicating the nested data containment level. The system pairs key nodes and value nodes belonging to the same data field and generates key-value binding edges from key nodes to corresponding value nodes.

[0098] S206, Based on the level depth parameter and parent node pointing parameter recorded in the node attribute dictionary, establish the nested topological relationship between key nodes. Define any two different key nodes as the first key node and the second key node. The system determines the generation formula for parent-child dependent edges pointing from the first key node to the second key node as follows:

[0099] ;

[0100] ;

[0101] In the formula, This represents the hierarchy depth parameter of the second key node; The hierarchy depth parameter represents the first key node; This indicates the parameter that the parent node of the second key node points to; This represents the node identifier of the first-key node. Under the condition that both of the above formulas are true, the system generates parent-child dependent edges pointing from the first-key node to the second-key node.

[0102] S207. For multidimensional array data structures and list data structures in structured detection data files, the system performs a node index appending operation. When multiple peer key nodes or multiple value nodes have the same parent node pointing parameter, the system assigns an incrementing sequence index value to the key node or value node according to the data flow parsing order. The system appends the sequence index value as an edge weight attribute to the parent-child node dependent edge or key-value binding edge to distinguish duplicate data items with the same name in the structured detection data file.

[0103] S208: The extracted key nodes and value nodes are used as the node set of the graph data structure, and the generated key-value binding edges and parent-child node dependent edges are used as the edge set of the graph data structure. The data entity graph is instantiated in memory and has tree-like directed acyclic topology. For the instantiation of the graph data structure and the memory allocation operation of the tree-like topology rules, those skilled in the art can use the graph theory algorithm library built into the standard programming language.

[0104] See Figure 4 This paper details the technical aspects of loading a quantized word embedding model and mapping its feature vectors by inputting text nodes from the template directed graph and key nodes from the data entity graph. The process of loading the quantized word embedding model and performing feature vector mapping is specifically divided into the following sub-steps:

[0105] S301, loads the quantized word embedding model stored on the local hard disk of the industrial control computer into memory. In order to adapt to the limited hardware environment without a graphics processor acceleration unit, the quantized word embedding model specifically includes at least one of the transformer-based bidirectional encoder representation model (BERT) that has been compressed in 8-bit integer (INT8) data format, or a compressed word vector mapping model (Word2Vec or FastText).

[0106] In its implementation, the quantized word embedding model is constructed using post-training quantization technology. It utilizes a linear quantization algorithm to map the original 32-bit floating-point weight parameters to an 8-bit integer range, reducing memory throughput and accelerating inference by calling the CPU's vector extended instruction set (such as AVX-512). The model is pre-trained on a general text corpus and fine-tuned using a vocabulary of specialized terms in the detection report domain, enabling it to possess specific semantic representation capabilities for "detection item names" and "formatting tags." The system calls the quantized word embedding model through the CPU's standard instruction set, establishing a text inference runtime environment offline.

[0107] S302, traverse the directed graph of the template to extract text nodes containing text content, traverse the data entity graph to extract key nodes, the system reads the format string data carried by the text nodes and the field name string data carried by the key nodes, the system performs word segmentation and stop word filtering on the format string data and field name string data, and generates a standardized lexical sequence. For the word segmentation and stop word filtering operations of string data, those skilled in the art can use existing natural language processing toolkits.

[0108] S303: The generated lexical sequence is input into the quantized word embedding model, forward propagation is performed, and a high-dimensional feature vector is output. The system converts the discrete words in the lexical sequence into continuous floating-point vectors, performs average pooling operation on the floating-point vectors at the word level, and generates a fixed-length feature vector representing the overall semantics of the node. The system converts the text nodes in the template directed graph into template feature vectors and the key nodes in the data entity graph into key node feature vectors.

[0109] S304, the generated template feature vector and key node feature vector are subjected to L2 normalization. This normalization process eliminates the interference of the absolute value of the feature vectors on subsequent similarity calculations by projecting the length of the feature vectors onto a unit hypersphere. For any feature vector, the system performs the normalization calculation using the following formula:

[0110] ;

[0111] In the formula, This represents the feature vector output after normalization. This represents the original feature vector generated by the quantized word embedding model; This represents the total number of dimensions of the original feature vector; This indicates that the original feature vector is at the th... The component values ​​in each dimension.

[0112] S305: The normalized template feature vector is data-bound to the corresponding text node, and the normalized key node feature vector is data-bound to the corresponding key node. The system persists the node object containing the feature vector mapping results to the local cache as the benchmark data source for calculating the similarity of topological anchor points.

[0113] Reference Figure 4 This paper provides a detailed explanation of the technical aspects of implementing adaptive degradation and minimal intervention mechanisms under operating conditions. The process of implementing adaptive degradation and minimal intervention mechanisms is specifically divided into the following sub-steps:

[0114] S306 monitors and quantifies the operational status data of the word embedding model in the industrial control computer. This operational status data includes model inference time and memory usage. If the model inference time exceeds a preset time threshold or the memory usage exceeds a safety threshold, the system triggers an adaptive degradation mechanism. The preset time threshold is set at 2000 milliseconds, and the safety threshold is set at 85% of the industrial control computer's physical memory. Exceeding these limits indicates a risk of resource overload. The underlying characteristic of the adaptive degradation mechanism is that the system stops relying on the semantic feature mapping process of the deep neural network and switches to executing a deterministic text matching algorithm based on discrete character comparison.

[0115] S307, a deterministic text matching algorithm is invoked to calculate the character similarity between text nodes and key nodes. The deterministic text matching algorithm includes the Levenstein edit distance algorithm. The system reads the layout string carried by the text node and the field name string carried by the key node, and calculates the minimum number of single-character editing operations required to convert the layout string into the field name string. Single-character editing operations include insertion, deletion, and replacement. For the specific calculation of the string edit distance and the solution of the matrix dynamic programming, those skilled in the art can use existing string processing libraries.

[0116] The system inversely maps the edit distance value to a downgraded matching confidence level with a value between 0 and 1, based on the calculated edit distance value and the maximum length of the two strings.

[0117] S308, based on the similarity of feature vectors or the downgraded matching confidence, triggers a minimum intervention mechanism. The lower-level feature of the minimum intervention mechanism is that the system only extracts isolated node pairs with matching confidence below a preset judgment threshold, generates a local manual review request, and blocks the fault-tolerant control logic of abnormal mapping. The preset judgment threshold is set to 0.6. If it is lower than the judgment threshold, it is considered to have a high risk of mismatch. The system directly allows and retains node pairs with matching confidence higher than or equal to the preset judgment threshold, and does not interrupt the mapping process of normal node pairs.

[0118] S309, for isolated node pairs that trigger manual review requests, the system extracts the context topology boundary data of the isolated node pairs in the template directed graph and data entity graph. The context topology boundary data includes adjacent text nodes connected by spatial orientation edges and adjacent key nodes connected by nested subordinate edges. The system packages the isolated node pairs and the context topology boundary data into an independent dataset, generates an intervention task package based on the dataset, and transmits the intervention task package to the graphical interface of the industrial control computer for highlight rendering.

[0119] S310 receives node pairing correction instructions issued by external input devices for intervention task packages. The system establishes correct mapping connections for isolated node pairs in memory based on the node pairing correction instructions, and synchronously persists the confirmed mapping status to the local cache. The system uses an adaptive degradation mechanism to avoid hardware computing power bottlenecks and controls computing resource consumption by limiting the range of nodes that can be manually reviewed.

[0120] Reference Figure 5 This paper details the technical aspects of performing a breadth-first search bidirectional topology synchronous deduction in the template directed graph and the data entity graph based on the matched node pairs to complete the data injection. The process of performing a breadth-first search bidirectional topology synchronous deduction based on the matched node pairs is specifically divided into the following sub-steps:

[0121] S401: Extract the set of text node and key node pairs that have been confirmed to match by feature vector mapping or minimization intervention mechanism from the local cache. The system uses the extracted set of text node and key node pairs as the initial anchor point. The lower-level feature of the bidirectional topology synchronous inference is the calculation process of calling graph theory traversal algorithm in both the template directed graph and the data entity graph, starting from the initial anchor point, and advancing layer by layer according to edge relationships to establish node mapping connections.

[0122] S402, initialize the template search queue and the data search queue in memory. The system pushes the text nodes in the initial anchor into the template search queue and pushes the key nodes in the initial anchor into the data search queue. For the memory initialization of the queue data structure and the enqueue and dequeue operations, those skilled in the art can use the linear list data structure built into the standard programming language.

[0123] S403, perform loop traversal operation. Under the condition that neither the template search queue nor the data search queue is empty, the system pops a current text node from the template search queue and a current key node from the data search queue. The system traverses the adjacent edges of the current key node in the data entity graph, locates the current value node associated with the current key node based on the key-value binding edge, and extracts the measured detection data carried in the current value node.

[0124] S404, in the directed graph of the template, the system searches for the target data container to be filled based on the spatial orientation edge. The system performs topological node detection along the horizontal or vertical spatial orientation edge of the current text node, extracts adjacent data containers with empty internal data states as target data containers, and the system determines whether the target data container meets the data injection conditions using the following formula:

[0125] ;

[0126] In the formula, Indicates the data injection flag bit of the calculation output; This represents the container for the target data to be extracted; This indicates the current value node in the location; This refers to the function that retrieves the data structure type inside a node.

[0127] S405, under the condition that the data injection flag is equal to one, the system writes the measured detection data carried in the current value node into the memory address of the target data container. After the system completes the measured detection data writing operation, the system marks the state of the target data container as filled. The system extracts the next-level child key node connected by nested subordinate edges in the data entity graph of the current key node and pushes the next-level child key node into the data search queue. The system synchronously extracts the next-level data container connected by hierarchical logical edges in the template directed graph of the target data container and pushes the next-level data container into the template search queue. The system uses the update and cyclic advancement of the queue state to complete the structured injection of detection data into the template grid space.

[0128] Reference Figure 5 This paper details the technical aspects of performing path matching conflict detection during bidirectional topology synchronous simulation. The process of performing path matching conflict detection is specifically divided into the following sub-steps:

[0129] S406 monitors the node features popped from the template search queue and data search queue during the breadth-first search algorithm's advancement. The lower-level features for path matching conflict detection are shown by comparing the number of out-degrees of the current data container in the template directed graph with the number of out-degrees of the current key node in the data entity graph, and verifying the validity of the memory format of the actual test data to be written.

[0130] S407: Extract the edge connection features of the current popped node and perform out-degree count verification. The system counts the number of hierarchical logical edges pointing from the current data container to the next-level data container in the directed graph of the template, and uses the number of hierarchical logical edges as the local out-degree parameter of the template. The system synchronously counts the number of nested dependent edges pointing from the current key node to the next-level child key node in the entity graph of the data, and uses the number of nested dependent edges as the local out-degree parameter of the data. The system's calculation formula for determining structural mismatch is as follows:

[0131] ;

[0132] In the formula, This represents the calculated structural error value. This represents the calculated local out-degree parameters of the template; This represents the local out-degree parameter of the calculated data; This represents an absolute value function. When the structural error value is greater than 0, the system determines that there is a topological conflict in the current deduction path.

[0133] S408, under the condition that the structural error value is equal to 0, the system performs a data format validity check. The system extracts the pre-set field constraint rules of the template directed graph for the target data container, extracts the physical storage type of the measured detection data in the current value node, and the field constraint rules limit the range of data types that the target data container can accept. If the physical storage type of the measured detection data exceeds the range allowed by the field constraint rules, the system determines that there is a data type conflict in the current deduction path.

[0134] S409, for the deduction path that triggers topology conflict or data type conflict, the system suspends the data injection process of the current branch to prevent illegal data from being written to the memory address of the target data container. The system extracts the template node memory address, data node identifier and conflict type code of the conflict location, and packages the template node memory address, data node identifier and conflict type code to generate a conflict tracing log.

[0135] S410 writes the conflict tracking log to the local storage device of the industrial control computer and marks the set of nodes that triggered the conflict as risk pending investigation. The system skips the inference branch where the conflict is currently occurring and continues to process other normal nodes in the template search queue and data search queue until the queues are empty. For the formatting and encapsulation of log files and hard disk read / write operations, those skilled in the art can call the standard input / output interface of the operating system to implement them. The formatting and encapsulation of log files and hard disk read / write operations are well known technologies in this field.

[0136] like Figure 5 As shown, the technical content of implementing the gravity weight arbitration strategy based on topological distance when node matching conflicts occur is described in detail. The process of implementing the gravity weight arbitration strategy based on topological distance is specifically divided into the following sub-steps:

[0137] S411, during the breadth-first search bidirectional topology synchronous deduction process, if the system detects that a target data container simultaneously satisfies the data injection condition with multiple candidate value nodes, or a current value node simultaneously satisfies the data injection condition with multiple candidate data containers, the system extracts the set of candidate nodes that trigger matching conflicts. The lower-level feature of the gravity weight arbitration strategy based on topological distance is as follows:

[0138] Using the most recently successfully matched and injected node pair as the reference anchor point, the control logic calculates the edge hop distance between the candidate nodes in the candidate node set and the reference anchor point in the topology, and converts the edge hop distance into numerical weights for priority sorting.

[0139] S412, the system extracts the text node and key node that have recently completed data injection from memory, and combines the text node and key node that have recently completed data injection into a reference anchor point. For each candidate node in the candidate node set, the system calls the shortest path algorithm to calculate the shortest topological edge hop count from the candidate node to the corresponding reference anchor point. For the specific calculation of the shortest path between two nodes in the graph data structure and the traversal search operation, those skilled in the art can use the existing Dijkstra algorithm library.

[0140] S413, the system obtains the calculated shortest topological edge hop count and substitutes it into the gravity weight calculation formula. The gravity weight is used to assign higher calculation priority to candidate nodes that are closer to the reference anchor point. The formula for calculating the gravity weight is as follows:

[0141] ;

[0142] In the formula, This represents the calculated gravitational weight value. Represents the natural constant; This represents the preset distance attenuation constant. It is a fixed value between 0.5 and 1.0; This represents the number of hops in the shortest topological edge calculated.

[0143] S414, the system traverses the gravity weight values ​​corresponding to all candidate nodes in the candidate node set, compares and selects the optimal candidate node with the largest gravity weight value. If there are multiple candidate nodes with the same maximum gravity weight value, the system extracts the parsing order of the candidate nodes in the structured detection data file or report template file, and selects the candidate node with the earliest parsing order as the optimal candidate node.

[0144] S415, the system confirms the optimal candidate node as the final matching node and performs the data injection operation. Under the condition that multiple value nodes compete for a single container, the system writes the measured detection data carried in the optimal candidate value node into the memory address of the target data container;

[0145] Under the condition of a single value node competing for multiple containers, the system writes the measured detection data carried in the current value node into the memory address of the optimal candidate data container. The system then removes the remaining unselected candidate nodes from the current mapping relationship and resumes the bidirectional topology synchronous deduction process.

[0146] Reference Figure 6 This paper details the technical aspects of converting the template directed graph after data injection into the target report document format and establishing memory mapping. The process of performing node data format conversion and memory mapping is specifically divided into the following sub-steps:

[0147] S501 involves traversing the template directed graph that completes the bidirectional topology synchronous deduction, extracting the target data container set whose internal state is marked as filled, and the lower-level feature of node data format conversion is the process by which the system identifies the physical storage structure of the measured detection data in the target data container and encodes the measured detection data into rich text data blocks with layout tags according to the underlying layout specifications of the target report document.

[0148] S502: Read the measured detection data from the target data container, perform encoding conversion based on the data storage format. If the measured detection data is numerical, the system truncates the numerical data according to preset significant digit retention rules, converting the truncated numerical data into a character array. If the measured detection data is an image stream, the system converts the image stream into a standard base encoded string. For the numerical truncation and standard base encoding conversion operations, those skilled in the art can use the built-in data conversion toolkit of a standard programming language.

[0149] S503, extract the spatial style attributes associated with the target data container. Based on the topological edge connections of the target data container in the directed graph of the template, the system obtains the font identifier, font size value, and color code at the corresponding position. The system concatenates and merges the spatial style attributes, layout tags, and the converted character array or standard base encoded string to generate a rich text data block. In order to perform pagination truncation calculations for underlying rendering, the system treats the characters as vertically arranged to obtain a conservative estimate of the maximum height. The system's estimation formula for calculating the rendering height of the rich text data block is as follows:

[0150] ;

[0151] In the formula, This represents the estimated rendering height of the rich text data block in the calculated output; This represents the total number of characters in the character array; Indicates the font size value; This represents the preset screen resolution conversion constant. The value is a conversion factor that approximates dots per inch; This indicates the physical length of the image mapped by the standard base encoded string at a specified scaling ratio.

[0152] S504 triggers the memory mapping operation of the document data stream. The lower-level characteristics of memory mapping are that the system allocates a continuous byte buffer in physical memory, calculates the node byte offset according to the topological traversal order of the template directed graph, and writes the generated rich text data blocks into the specified offset address of the byte buffer in sequence.

[0153] S505, calculate the physical memory starting offset address corresponding to the current rich text data block. The system calculates the starting offset address of the current rich text data block based on the starting offset address of the previously written rich text data block and the length of the data in bytes. The formula for calculating the offset address is: ;

[0154] In the formula, This indicates the starting offset address for writing the current rich text data block to the byte buffer; This indicates the starting offset address of the previously written rich text data block; This indicates the actual data length in bytes occupied by the previous rich text data block that has already been written; This indicates the byte length occupied by the separator label between two rich text data blocks. Based on the calculated starting offset address, the system serializes the generated rich text data blocks and writes them into a contiguous byte buffer, constructing a complete report document object stream in memory.

[0155] like Figure 6 As shown, the technical content of converting the report document object stream constructed in memory into an entity file and saving it to the local storage medium is explained in detail. The process of performing low-level rendering and file persistence is specifically divided into the following sub-steps:

[0156] S506 reads the completed report document object stream from a contiguous byte buffer in physical memory. The underlying rendering features are manifested as the system calling the graphics rendering engine of the industrial control computer to parse the rich text data blocks and spatial style attributes in the report document object stream, and converting the rich text data blocks into a two-dimensional typesetting instruction set containing font outlines and pixel coordinates.

[0157] S507 performs page break calculations based on preset physical paper size and margin parameters. The system extracts rendering height data from the 2D typesetting instruction set one by one to determine if there is sufficient remaining space on the current page to accommodate the current rich text data block. The formula for calculating the page break flag is as follows:

[0158] ;

[0159] In the formula, This indicates the page feed flag in the calculation output; This represents the estimated rendering height of the current rich text data block; This indicates the total height parameter of the preset physical paper. The total height parameter of the preset physical paper is set according to the standard specifications of the target printing paper, such as the height value of A4 paper. This represents the sum of the preset top and bottom page margins. This represents the cumulative height occupied by the current page's layout. When the page break flag is set to one, the system inserts a forced page break in the two-dimensional layout instruction set, resetting the current page's cumulative height.

[0160] S508 encapsulates the two-dimensional typesetting instruction set after adding page breaks into a binary data packet conforming to the target document specification. According to the preset report output format requirements, the system writes a file identifier header and document attribute metadata at the beginning of the binary data packet and a document end identifier at the end of the binary data packet. The binary data packet is encapsulated according to the document standard specification. Those skilled in the art can implement this using existing document processing framework code libraries.

[0161] S509 calls the operating system's input / output read / write interface to write the encapsulated binary data packet to the hardware storage device. The underlying characteristic of file persistence is that the system allocates physical disk sectors in non-volatile storage media through the file system, writes the binary data packet sequentially into the physical disk sectors in the form of a byte stream, and generates an entity report file that is detached from the memory environment.

[0162] S510: After the entity report file is generated, the system reads the entity report file, calculates the hash value, and compares it with the binary data packet in memory for data integrity. If the data comparison is consistent, the system updates the absolute storage path of the entity report file to the display log of the industrial control computer. The system clears the template search queue, data search queue, and continuous byte buffer allocated in memory, performs system memory reclamation operation, and ends the format conversion of the measured detection data to the report document.

[0163] Specific application examples are as follows:

[0164] To verify the effectiveness of the offline inspection report automatic generation method proposed in this invention in solving problems such as data node mismatch, report rendering and layout overflow, and high manual review costs when merging complex layout templates and massive unordered inspection data, this specific application embodiment is based on the application scenario of a large third-party testing institution batch processing offline structured inspection data of industrial equipment and outputting customized inspection report files, and combines... Figure 8 , Figure 9 as well as Figure 10 The data shown will be explained in detail.

[0165] Figure 8 , Figure 9 as well as Figure 10 The data in this document are comparisons between data features generated by real-time system computation and traditional data based on hard-coded mapping rules and manual review.

[0166] In the application scenario of this embodiment, the industrial control computer receives a structured detection data file containing 1500 discrete detection fields. The goal is to accurately inject all the measured detection data into a detection report template file containing 120 table cell nodes and text box nodes. To address the risks of complex data nesting levels and inconsistencies between device output names and report template requirements, the system has developed a joint automatic typesetting generation scheme that includes template two-dimensional grid mapping, gravity weight arbitration mechanism, adaptive degradation scheduling, and memory mapping conversion.

[0167] Template parsing and 2D atomic mesh mapping implementation:

[0168] Before performing the data injection task, the template reconstruction module first parses the underlying document object model of the template file. In this embodiment, the system is set to traverse and extract a target table cell node. The coordinate attribute of the top left corner vertex of the data container is: the x-coordinate value. y-axis value The system extracts the basic row height from the coordinate attribute structure set. Basic column width The system performs physical coordinate transformation of the two-dimensional atomic mesh matrix.

[0169] Calculation formula based on grid index coordinates: (Round down); Calculate the starting row index Calculate the starting column index. The system accurately projects the target table cell nodes into a two-dimensional atomic grid matrix in memory based on the starting row index and the starting column index, thus establishing the reference position parameters for the subsequent establishment of spatial orientation edges.

[0170] Implementation of Path Matching Conflict Detection and Gravity Weight Arbitration:

[0171] During the breadth-first search bidirectional topology synchronization simulation, the system monitored multiple measured detection data competing for the same target data container. To verify the effectiveness of the system's arbitration mechanism, the current simulation path parameters were extracted.

[0172] The system discovered during the deduction path that the template has the number of hierarchical logical edges pointing from the target data container to the next level data container, i.e., the template's local out-degree parameter. Meanwhile, the number of nested dependent edges pointing from the current key node to the next-level child key node in the target data container competes for this data, which is the local out-degree parameter of the data. .

[0173] System call structure mismatch determination formula: Calculate the structural error value If the structural error value is greater than 0, the system determines that there is a topological structure conflict in the current inference path and immediately suspends the data injection process of the current branch to avoid illegal writing of erroneous data.

[0174] The system then triggers the gravity weight arbitration strategy, by Figure 8 The data details are available. Figure 8 The horizontal axis represents the number of shortest topological edge hops, and the vertical axis represents the gravity weight value. Figure 8 It shows the distance decay constant. Gravitational weight decay characteristics for values ​​of 0.5 and 0.8 respectively (set in this embodiment). ).

[0175] The system extracts the set of candidate nodes that trigger matching conflicts and calculates the shortest topological edge hop count from the first candidate node to the most recent successfully matched reference anchor point. The number of topological edge hops from the second candidate node to the most recently successfully matched reference anchor point. The system calculates the gravity weights using the following formula: Calculate the gravitational weight value corresponding to the first candidate node. ≈0.3679; Calculate the gravitational weight value corresponding to the second candidate node. .

[0176] Combination Figure 8 As can be seen from the characteristic curve, the distance attenuation constant set in this embodiment... Under these conditions, the gravity weight value decreases exponentially with the increase of the shortest topological edge hop count. When the edge hop count increases from 2 hops for the first candidate node to 4 hops for the second candidate node, the corresponding gravity weight decreases from 0.36 to 0.13, a decrease of more than 60%. This trend shows that the exponential decay formula widens the weight difference between nodes with different topological distances. The system calculates based on this weight difference, weakening the competitive priority of nodes with greater topological distances, ensuring that nodes with close local structural connections and small edge hop counts obtain data injection permissions, thereby preventing data misalignment caused by distant irrelevant nodes.

[0177] Adaptive degradation and resource load monitoring implementation:

[0178] This embodiment runs in an offline industrial control computer environment without a dedicated graphics processor. During the anchor point establishment phase, the system continuously embeds the string input quantization words from the template directed graph and the data entity graph into the model and calls the formula. The high-dimensional feature vector output is normalized using the L2 norm. Due to the increasing number of concurrent nodes, this floating-point calculation matrix causes an accumulation of system resources.

[0179] To verify the stability of the adaptive degradation mechanism, the system output runtime resource monitoring logs, as shown in the attached figure. Figure 9 As shown, Figure 9 This is a dual Y-axis line chart. The horizontal axis represents the cumulative number of processed nodes, the left main vertical axis represents the physical memory usage rate, and the right secondary vertical axis represents the model inference time.

[0180] Depend on Figure 9 Data shows that system resources were within a manageable range when processing the first 700 nodes. However, as the cumulative number of nodes increased from 700 to 850, the memory usage rose from 74.8% to 86.8% due to the large amount of unreleased context cache generated by the depth-first traversal of the graph structure. Model inference time also increased from 1650 milliseconds to 2150 milliseconds. When the number of nodes reached 850, the system monitoring module determined that the current memory usage (86.8%) had exceeded the set 85% safety threshold, and the inference time (2150ms) had exceeded the 2000ms time threshold.

[0181] Upon meeting the threshold condition, the system, following instructions, terminates the forward propagation calculation of the quantized word embedding model and halts the high-dimensional feature mapping process. Subsequently, the system triggers a degradation switch, employing the Levenstein edit distance algorithm to perform discrete character comparison, such as... Figure 9As shown by the curve trend after node 850, after the degradation occurred, the system reclaimed the memory space occupied by the model parameters, and the memory usage rate dropped sharply and stabilized at around 50%. At the same time, the deterministic string comparison eliminated floating-point matrix operations, rapidly reducing the node inference time to less than 300 milliseconds. The load change characteristics confirm that the degradation mechanism can proactively cut off high-energy-consuming processes when facing hardware resource bottlenecks, achieving seamless hot switching between high-dimensional semantic feature mapping of the model and lightweight character editing distance algorithm. This solves the technical bias and technical difficulties of offline industrial computers crashing due to memory overflow, and balances the reliability of matching with the system availability in the offline environment.

[0182] Memory mapping conversion and rendering height estimation implementation:

[0183] After the data injection is completed, the system performs the underlying rendering pagination truncation calculation for the rich text data blocks.

[0184] The system extracts the total number of characters contained in the currently populated target data container. Font size value Screen resolution conversion constant It does not contain an image stream, therefore the physical length of the image is... .

[0185] The system estimates the rendering height of rich text data blocks using the following formula: Calculate the estimated rendering height of the rich text data block. The system compares the estimated height of the rich text data block rendering with the remaining height parameters of the preset physical paper. If the cumulative height occupied by the layout already completed on the current page results in insufficient page space to accommodate a height of 630, the system inserts a forced page break in the two-dimensional layout instruction set, completely eliminating the technical problem of text overflowing the page boundary.

[0186] To address the issues of mismatched matching and extensive manual review and rework caused by traditional hard-coded mapping rules, a comparison of the overall report generation effect was introduced during the complete cycle of processing 1500 discrete detection fields.

[0187] This method was tested in the same batch of structured detection data files and compared with real data using traditional empirical methods that rely on manually defined regular expressions and hard-coded scripts. Figure 10 The bar chart in the middle shows that Figure 10The left side compares the data injection matching error rate, while the right side compares the overall time taken to generate a single report containing 1500 data items. When using the traditional hard-coding scheme (light gray bars in the figure), the matching error rate is as high as 12.5% ​​due to the inability to adapt to minor changes and hierarchical nesting disturbances in the output fields of the detection device, and with a large amount of manual intervention for correction, the overall time takes 28 minutes. However, by using the method of this system (dark gray bars in the figure), relying on the quantized word embedding model and the gravity weight arbitration mechanism, the matching error rate is strongly suppressed to below 0.4%, and the overall time is reduced to 1.2 minutes.

[0188] This embodiment demonstrates the industrial-grade practicality of the offline detection report automatic generation method and system. On the one hand, it generates a two-dimensional atomic grid matrix based on the coordinates of the data container, calculates the grid index to construct a directed graph, and realizes accurate digital reconstruction of the template structure. On the other hand, it uses structural error values ​​to determine matching conflicts and combines the shortest topological edge jump number to calculate the gravity weight, overcoming the defect that relying solely on text similarity can easily lead to data misalignment.

[0189] When the system detects a conflict in the deduction path, it automatically suspends and performs gravity weight sorting to select the optimal matching node, thus avoiding the generation of incorrectly formatted data. At the same time, it performs adaptive degradation when hardware resources are on the verge of overload, maintaining the continuity of offline operation. Finally, by accurately estimating the rendering height and performing direct conversion of the memory buffer, it saves manual proofreading time and management costs.

Claims

1. An offline detection report automatic generation system based on intelligent template matching, characterized in that, include: The template reconstruction module is used to project the data containers in the template file onto a two-dimensional atomic mesh matrix to generate a directed template graph containing text nodes; The entity graph construction module is used to extract key nodes and value nodes from structured detection data files to generate data entity graphs. An anchor point establishment module is used to input the text nodes of the template directed graph and the key nodes of the data entity graph into the quantized word embedding model mapping to construct a set of benchmark anchor points; The arbitration deduction module is used to use the set of reference anchor points as the origin to generate deduction paths in the template directed graph and the data entity graph. When the deduction paths conflict, the number of topological path edges is calculated. Based on the number of topological path edges, the topological gravity weight is calculated using an exponential decay function. The data carried in the value node pointed to by the deduction path with the largest topological gravity weight is extracted as the target matching result. The report generation module is used to output a detection report file based on the target matching results; The arbitration deduction module is specifically used for: Initialize the template search queue and the data search queue, push the text node into the template search queue, and push the key node into the data search queue; Pop a current text node from the template search queue, pop a current key node from the data search queue, and locate the current value node bound to the current key node; The adjacent data containers with empty internal data states along the current text node are used as target data containers to be filled. Under the condition that the data structure type of the target data container is consistent with that of the current value node, the measured detection data carried in the current value node is written into the memory address of the target data container.

2. The offline detection report automatic generation system based on intelligent template matching according to claim 1, characterized in that, The template reconstruction module is specifically used for: Obtain the width and height values ​​of the data container, extract the minimum non-zero width value as the basic column width, and extract the minimum non-zero height value as the basic row height; The total page width, total page height, basic column width, and basic row height of the template file are obtained to construct the two-dimensional atomic grid matrix; The starting row index and starting column index of the data container in the two-dimensional atomic mesh matrix are calculated based on the x-coordinate and y-coordinate values ​​of the vertex of the data container, and the ending row index and ending column index are calculated based on the width and height values. Data containers that have a difference between the starting row index and the ending row index or a difference between the starting column index and the ending column index are combined into bounding box structures. The bounding box structure is a bounding box, and the data containers bound to the bounding box structure are marked as macro nodes.

3. The offline detection report automatic generation system based on intelligent template matching according to claim 1, characterized in that, The entity graph construction module is specifically used for: Establish key-value binding edges between the key nodes and the value nodes according to the identifier matching rules; Based on the hierarchy depth parameter of the key node and the parent node pointing parameter, establish the parent-child dependent edges between the key nodes; When multiple peer key nodes or multiple peer value nodes have the same parent node pointing parameter, an incrementing sequence index value is assigned, and the sequence index value is attached as an edge weight attribute to the parent-child node dependent edge or the key-value binding edge.

4. The offline detection report automatic generation system based on intelligent template matching according to claim 1, characterized in that, The anchor point establishment module is specifically used for: Perform word segmentation and stop word filtering operations on the layout string data in the template file and the field name string data in the structured detection data file to generate a lexical sequence; The input quantized word embedding model of the lexical sequence is subjected to forward propagation calculation to output a high-dimensional feature vector; The high-dimensional feature vector is subjected to L2 norm normalization to generate template feature vector and key node feature vector.

5. The offline detection report automatic generation system based on intelligent template matching according to claim 1, characterized in that, The anchor point establishment module is also used to execute an adaptive degradation mechanism: Monitor the inference time and memory usage of the quantitative word embedding model in physical memory; If the model inference time exceeds a preset time threshold or the memory usage exceeds a safety threshold, an adaptive degradation mechanism is triggered to stop the feature mapping process. The Levenstein edit distance algorithm is called to calculate the minimum number of single-character edit operations between the text node and the key node. The minimum number of single-character edit operations is inversely mapped to the downgraded matching confidence, which replaces the feature mapping process as the basis for node matching to complete the adaptive downgrade mechanism.

6. The offline detection report automatic generation system based on intelligent template matching according to claim 5, characterized in that, The anchor point establishment module is also used to trigger a minimal intervention mechanism: Extract isolated node pairs whose downgraded matching confidence is lower than a preset judgment threshold; Extract the context topology boundary data of the isolated node pairs in the template directed graph and the data entity graph, and package the isolated node pairs and the context topology boundary data into an independent dataset; An intervention task package is generated based on the independent dataset, and the intervention task package is highlighted and rendered in a graphical interface. Receive node pairing correction instructions issued for the intervention task package; The mapping connection of the isolated node pair is established according to the node pairing correction instruction, and the confirmed mapping state is synchronously persisted to the local cache to complete the minimal intervention mechanism.

7. The offline detection report automatic generation system based on intelligent template matching according to claim 1, characterized in that, The arbitration deduction module is also used to perform matching conflict detection of the deduction path: The number of hierarchical logical edges pointing from the data container to the next level data container in the directed graph of the template is counted and used as the local out-degree parameter of the template. The number of nested dependent edges pointing from the key nodes to the next-level child key nodes in the data entity graph is counted and used as the local out-degree parameter of the data. The absolute value of the difference between the template local out-degree parameter and the data local out-degree parameter is calculated as the structural error value. If the structural error value is greater than 0, it is determined that there is a topological structure conflict in the inference path. The memory address of the template node, the data node identifier, and the conflict type code at the conflict location are extracted to generate a conflict tracking log. The conflict tracking log is written to the local storage device. The inference branch where the conflict occurs is skipped to complete the matching conflict detection.

8. The offline detection report automatic generation system based on intelligent template matching according to claim 1, characterized in that, The report generation module is specifically used for: The target matching result is encoded into a rich text data block with layout tags according to the layout specifications of the detection report file; A contiguous byte buffer is allocated in physical memory, and the starting offset address of the current rich text data block is calculated based on the starting offset address of the previously written rich text data block and the length of the data in bytes. The rich text data blocks are serialized and written to the consecutive byte buffers to construct a report document object stream.

9. A method for automatically generating offline detection reports based on intelligent template matching, characterized in that, The offline detection report automatic generation system based on intelligent template matching, as described in any one of claims 1-8, includes the following steps: Project the data container in the template file onto a two-dimensional atomic mesh matrix to generate a template directed graph containing text nodes; Extract key nodes and value nodes from structured detection data files to generate a data entity graph; The text nodes of the template directed graph and the key nodes of the data entity graph are input into the quantized word embedding model mapping to construct a set of benchmark anchor points; Using the set of reference anchor points as the origin, a deduction path is generated in the template directed graph and the data entity graph. When the deduction path conflicts, the number of topological path edges is calculated. Based on the number of topological path edges, the topological gravity weight is calculated using an exponential decay function. The value node data pointed to by the deduction path with the largest topological gravity weight is extracted as the target matching result. Output a detection report file based on the target matching results.

Citation Information

Patent Citations

  • Industrial detection report generation method and device, electronic equipment and storage medium

    CN121581003A

  • Automatic document generation system

    CN120540748A

  • Energy equipment data report generation method based on artificial intelligence

    CN121882008A