AI semantic enhanced unstructured manufacturing document data structuring method

By extracting semantic anchors and logical gravitational field models from manufacturing documents, the problem of aligning spatial topology with semantic logic distribution in manufacturing documents is solved, realizing the transformation of unstructured text flow into hierarchical nested structured trees, and improving the robustness and consistency of data processing.

CN121579618AActive Publication Date: 2026-02-27FUJIAN YOUHEKE NETWORK TECH CO LTD

Patent Information

Application Number
CN202610086187.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-22
Publication Date
2026-02-27
Estimated Expiration
2046-01-22

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve precise alignment between spatial topological distribution and semantic logical distribution in manufacturing documents, leading to ambiguity and mapping conflicts during the structuring process. This is especially problematic in multi-station parallel process scenarios, where it is difficult to ensure the physical authenticity and logical consistency of the data.

Method used

By acquiring the character recognition stream of the document, semantic anchors with physical attributes and their spatial topological coordinates are extracted. Combined with feature vectorization processing and logical gravitational field model, the logical correlation strength of the vector to be processed is calculated. Through polarization procedure and boundary impedance modulation mechanism, the unique ownership of the vector to be processed is established, and a hierarchical nested structured tree is constructed.

Benefits of technology

It enables automatic merging of discrete fields in complex manufacturing documents, resolves the ownership conflict of similar semantically labeled data entities in densely arranged scenarios, ensures the topological consistency and logical integrity of data recognition, and improves the robustness of the data processing system under extreme working conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121579618A_ABST
    Figure CN121579618A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electrical digital data processing, and discloses an AI semantic enhanced unstructured manufacturing document data structuring method, which comprises the following steps: acquiring a character recognition stream of a target document to extract a semantic anchor point and topological coordinates thereof, and vectorizing a to-be-structured field to generate a to-be-processed vector; querying an engineering logic mapping table to determine association intensity, and calculating a logic gravitational field intensity value of the to-be-processed vector relative to the semantic anchor point according to the association intensity; associating the to-be-processed vector to a target semantic anchor point according to the field intensity value, starting a logic polarization program when the to-be-processed vector is identified to be interfered by the homogeneous semantic anchor point in the association stage, extracting bias characteristics to generate a polarization vector, and modulating the native gravitational field intensity; according to the method, the attribution ambiguity of similar semantic entities in a dense distribution scene is solved by introducing the logic gravitation constraint, the structured conflict caused by spatial distribution and engineering logic decoupling is eliminated, and the topology consistency in the data conversion process is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to an AI semantic enhanced unstructured manufacturing document data structuring method and belongs to the technical field of electric digital data processing. BACKGROUND

[0002] Currently, the digital management process of the manufacturing industry extracts unstructured document data by using optical character recognition combined with natural language processing technology. This path is universally applicable to processing format regular text streams and constitutes the main method of current data preprocessing. Manufacturing documents contain physical directional spatial information and engineering logic characteristics. In a multi-station parallel process scenario, there are a large number of engineering parameters with consistent semantic labels and dense arrangement in the document. Conventional data recognition schemes follow linear extraction logic of text streams, ignore the coupling relationship between document page topology distribution and manufacturing logic skeleton, and are difficult to determine the physical attribution of homogeneous data entities. Linear recognition methods increase the cost of logic reconstruction, making the data structuring process face attribution ambiguity and mapping conflicts, which leads to the need for manual intervention to correct errors in the later stage.

[0003] To address the above challenges, the industry attempts to alleviate the contradiction by increasing the size of the corpus or improving the model calculation dimension. This linear improvement method increases the load of computing power and cannot establish a nonlinear relationship between spatial topology coordinates and logic levels at the bottom. When facing complex layouts or noise interference, the system lacks physical background knowledge, leading to output results deviating from objective laws, and cannot solve the structuring conflict caused by the decoupling of spatial layout and engineering logic. Although there have been some progress in the optimization of document scanning hardware parameters and the improvement of basic algorithm accuracy in the current industrial data acquisition scheme, it is difficult to deal with the logic mapping problem caused by complex topological relationships by relying solely on the improvement of physical analysis power. The software processing mechanism and control logic are still insufficient. For example, a document data structuring storage and retrieval method based on AI driving is disclosed in Chinese Patent Publication No. CN119829723A. A structured question and answer library is constructed using a large language model and vector retrieval technology, and historical dialogue context is used to assist intent recognition. This method is limited to keyword matching and semantic vector indexing in the field of general natural language processing and does not address the deep coupling between spatial physical sites and engineering logic skeletons in manufacturing documents. When facing multi-station parallel process tables and high-density dense engineering parameters, it lacks awareness and dynamic adjustment of page visual boundaries, topological distances, and physical dimension constraints. It faces attribution ambiguity and mapping conflicts in conditions such as the attribution determination of homogeneous engineering entities and the dislocation of cross-page key anchors, making it difficult to ensure that the structured output conforms to the physical reality of the manufacturing site.

[0004] Therefore, how to accurately align the spatial topology distribution and semantic logic distribution in unstructured manufacturing documents and establish a data organization mechanism with physical logic calibration capability has become a technical problem to be solved by the present application. SUMMARY

[0005] To solve the problems presented in the background art, the technical solutions of the present application are as follows: An AI semantic enhancement unstructured manufacturing document data structuring method, comprising the following steps:

[0006] Step 101, obtaining the character recognition stream of the target document, extracting the keywords with physical attribute orientation and defining them as semantic anchor points through industry dictionary matching, and obtaining the spatial topological coordinates of each semantic anchor point in the target document page coordinate system;

[0007] Step 102, performing feature vectorization processing on the to-be-structured field in the target document to obtain a to-be-processed vector containing text semantic features and position distribution features;

[0008] Step 103, determining the logical correlation strength between the semantic anchor points and the to-be-processed vector, and calculating the logical gravitational field strength value of the to-be-processed vector relative to each semantic anchor point according to the Euclidean distance between the spatial topological coordinates and the logical correlation strength;

[0009] Step 104, associating the to-be-processed vector to the target semantic anchor point according to the order of the logical gravitational field strength value from large to small; during the association process, if the difference between the logical gravitational field strength values of the to-be-processed vector relative to multiple homogeneous semantic anchor points is less than a preset polarization threshold, a logical polarization program is started;

[0010] Step 105, extracting the logical bias features in the preset neighborhood of the to-be-processed vector, generating a polarization vector according to the logical bias features and the endogenous engineering correlation degree of each homogeneous semantic anchor point, and performing nonlinear weight bias modulation on the original gravitational field strength value of each homogeneous semantic anchor point using the polarization vector to determine the unique attribution of the to-be-processed vector.

[0011] Preferably, in step 103, it further comprises: identifying the visual segmentation lines and background mutation lines existing in the page of the target document and establishing them as logical impedance boundaries; detecting whether the logical connection between the to-be-processed vector and the semantic anchor point passes through the logical impedance boundary; if the logical connection passes through the logical impedance boundary, adjusting the equivalent logical distance between the to-be-processed vector and the target semantic anchor point according to the number of the logical impedance boundaries passed through to perform attenuation correction on the logical gravitational field strength value.

[0012] Preferably, the step further comprises: extracting and caching the semantic anchor points of the processed pages in the target document and defining them as virtual state features; projecting the virtual state features to the initial logical site of the current to-be-processed page, assigning a persistent weight according to the engineering properties of the virtual state features; calculating the compensation gravitational field strength value of the virtual state features to the to-be-processed vector in the current to-be-processed page according to the persistent weight and the page offset; and fusing the compensation gravitational field strength value and the original gravitational field strength value in the current to-be-processed page to perform cross-page logical structuring judgment.

[0013] Preferably, in step 105, the step further comprises: if the difference between the gravitational field strength of the to-be-processed vector relative to the plurality of homogeneous semantic anchors is less than a preset threshold, virtually associating the to-be-processed vector with each homogeneous semantic anchor to construct a corresponding candidate logical state space; obtaining an engineering logic consistency entropy value corresponding to each candidate logical state space; and determining the homogeneous semantic anchor corresponding to the candidate logical state space with the smallest engineering logic consistency entropy value as the final attribution of the to-be-processed vector.

[0014] Preferably, the step further comprises: extracting a recognition confidence residual corresponding to each to-be-structured field in the character recognition stream; constructing a convergence damping operator for each to-be-processed vector according to the recognition confidence residual; and modulating the convergence speed of each to-be-processed vector in the logical gravitational field by using the convergence damping operator; wherein the greater the recognition confidence residual, the stronger the retarding effect of the convergence damping operator on the association of the to-be-processed vector with the semantic anchor.

[0015] Preferably, the industry dictionary and the engineering logic mapping table are stored as a hexadecimal matrix, and the hexadecimal matrix is internally preset with a correlation coefficient representing the physical order of magnitude variation law of different manufacturing parameters over time; in step 103, the processor calls the pre-stored hexadecimal matrix in the memory to determine the logical correlation strength through cyclic iteration operation.

[0016] Preferably, in the logical polarization procedure, the logic of using the polarization vector to perform nonlinear weight bias modulation on the original gravitational field strength value satisfies the following formula: wherein, is the modulated target gravitational field strength value, is the original gravitational field strength value of the to-be-processed vector relative to a specific homogeneous semantic anchor, is a preset polarization response coefficient, is the generated polarization vector, is a feature vector representing the engineering logic correlation degree between the to-be-processed vector and the homogeneous semantic anchor.

[0017] Preferably, after step 104, the step further comprises: capturing process chain residual information generated by the to-be-processed vector in the gravitational field calculation process; when an orphan field without a corresponding semantic anchor is identified in the target document, performing topological matching in the current logical space by using the process chain residual information; and according to the matching result, deducing the missing semantic anchor category to which the orphan field belongs, and creating a virtual placeholder anchor to perform logical association compensation on the orphan field.

[0018] Preferably, the step further comprises: performing logical conflict scanning on the preliminary structured tree generated in step 104 by using a preset physical dimension constraint; when a physical dimension conflict is identified, extracting context features in the neighborhood of the conflict point and performing secondary path redirection to re-determine the hierarchical attribution of the to-be-processed vector.

[0019] Preferably, in the process of performing secondary path redirection, by calculating the local logical structure entropy value of the conflict point under different redirection paths, the path that makes the local logical structure tend to an ordered state is selected as the final merging path, and an updated memory memory mapping index is generated.

[0020] Compared with the prior art, the beneficial effects of the present application are:

[0021] 1. In the AI semantic enhanced unstructured manufacturing document data, a gravity field model is established between the physical pointing keyword semantic anchor point and the to-be-processed field, the spatial topological coordinate Euclidean distance is combined with the preset engineering logic coupling coefficient to perform weight mapping, so that the discrete fields are automatically merged to the potential energy value maximum target anchor point under the constraint of gravitational potential energy, solving the attribution conflict of similar semantic label data entities caused by spatial distribution interference in the dense arrangement scene, and realizing the logical conversion of unstructured text stream to hierarchical nested structured tree.

[0022] 2. A boundary impedance modulation mechanism is constructed, the page visual segmentation line and the background mutation edge are extracted as the impedance boundary, and when the gravity field path passes through the impedance boundary, a nonlinear attenuation impedance factor is modified to the equivalent logical distance, so that the semantic gravity preferentially conducts along the low resistance area, eliminates the risk of false capture caused by physical distance proximity and logical partition isolation in complex form design, and guarantees the topological consistency of data recognition under visual container interference; A state tunnel mechanism is established, the high confidence semantic anchor point of the processed page is encapsulated as a virtual state particle and projected to the initial site of the subsequent page, the page code offset and the persistent weight are combined to calculate the compensation gravity field strength, a cross-page dimension semantic association conduction path is established, the problem of orphan data fragmentation caused by the absence of key anchor points in the current window in long document processing is solved, and the global logical integrity of large-scale data governance is maintained in the limited processing window.

[0023] 3. A logic topological consistency potential dynamic self-calibration process is performed, a plurality of candidate attribution paths are simulated and generated for the logical to-be-arbitrated field, entropy value determination is performed by using an engineering evolution matrix, a path that makes the local logical structure tend to an ordered state is selected for final merging, the gravity field strength distribution is updated synchronously, the data inherent closed loop constraint is used as a judgment criterion, adaptive correction in a semantic feature highly overlapped area is realized, and the discrimination robustness of the data processing system in extreme working conditions is improved; The logical conflict scanning is performed on the initially generated structured tree by using the physical dimension constraint, the secondary path redirection is performed on the conflict point based on the neighborhood context, the attribution relationship is dynamically arbitrated by using the logical perturbation and the system consistency entropy value correlation determination, the semantic collapse risk of the traditional static mapping model when facing non-standard documents is avoided, and it is ensured that the output structured data stream conforms to the logical evolution law. BRIEF DESCRIPTION OF DRAWINGS

[0024] Fig. 1The structural processing flow chart of the logical gravitational field and polarization mechanism of the application;

[0025] Fig. 2 The performance comparison curve of the attribution determination accuracy of the application under different signal-to-noise ratio environments;

[0026] Fig. 3 The system hardware architecture and data flow direction schematic diagram of the integrated logical arbitration core of the application. DETAILED DESCRIPTION

[0027] The detailed description of the application is intended to enable a person skilled in the art to understand and implement the application, and the following description is for explanatory purposes only and does not constitute a limitation on the scope of protection of the application.

[0028] The application provides an AI semantic enhancement unstructured manufacturing document data structuring method, which runs in an electronic digital data processing environment, converts discrete manufacturing document data into a structured tree with hierarchical nested relationship by constructing a semantic anchor point based gravitational field model, and includes the following stages: obtaining a character recognition stream of a target document and establishing semantic anchor points and their spatial topological coordinates; performing feature vectorization processing on the field to be structured to generate a to-be-processed vector; calculating the logical gravitational field strength value according to the correlation strength determined by the engineering logic mapping table; performing logical merging according to the field strength value, and starting the logical polarization program to lock the unique attribution when the field strength interference is recognized; the processor receives the character recognition stream of the target document, performs keyword matching with the industry dictionary pre-stored in the memory, extracts the keywords with physical attribute direction and defines them as semantic anchor points, the semantic anchor points select equipment number, key part name or process procedure title, the processor synchronously obtains the spatial topological coordinates of each semantic anchor point in the target document page coordinate system, that is, obtains the center point coordinates (x, y) of the character bounding rectangle to which the semantic anchor point belongs, for the field to be structured in the document, the system performs feature vectorization processing to obtain a to-be-processed vector containing text semantic features and position distribution features, the text semantic features represent the physical quantity category to which the field belongs, and the position distribution features are established based on the topological site of the field in the page, so that the standardization representation of the discrete parameter entity in the feature space is realized.

[0029] Because the engineering parameters in the manufacturing documentation are logically constrained by physical entities, the system constructs a logical gravitational field calculation model. The execution steps are as follows: the processor calls a pre-stored hexadecimal engineering logic mapping table in memory, and a correlation coefficient is set to characterize the physical magnitude changes of different manufacturing parameters in the process timeline. The processor determines the logical correlation strength between semantic anchors and the vector to be processed based on the correlation coefficient. Based on the Euclidean distance between each semantic anchor and the vector to be processed and the logical correlation strength, the processor calculates the logical gravitational field strength value of the vector to be processed relative to each semantic anchor. The processor associates the vector to be processed with the logical level of the target semantic anchor with the highest gravitational potential energy according to the logical gravitational field strength values ​​in descending order. To ensure the determinism of the parameter determination process, this logical correlation strength... The calculations meet the preset calibration procedures, namely, querying the cosine similarity between two feature vectors based on the preset industry association matrix and multiplying it by the corresponding engineering weight coefficient; the processor retrieves the association coefficient from the hexadecimal engineering logic mapping table in memory; the Euclidean distance is calculated based on the spatial coordinates of the vector to be processed and the semantic anchor point; field strength mapping is performed using the gravity calculation operator; the parameter gradient is determined by the physical property directional experiment; if the vector to be processed is located in the overlapping area of ​​multiple anchor points, the system constructs a logical gravitational potential energy map, merges discrete fields into the level of the target anchor point with the maximum gravitational field strength, maps the association coefficient to the hexadecimal encoding position, and updates it in real time by combining the cosine similarity with the preset engineering weight to ensure that the calculation accuracy meets the monotonic evolution law, and completes the logical transformation of unstructured text flow into a hierarchical nested structured tree.

[0030] When processing high-density layouts such as complex equipment parameter tables, if the difference between the logical gravitational field strength values ​​of the vector to be processed and multiple homogeneous semantic anchor points is less than a preset polarization threshold, the processor initiates a logical polarization program. The processor extracts logical bias features within a preset neighborhood of the vector to be processed. These logical bias features include associated workstation numbers or cooling circuit identifiers. The processor generates a polarization vector based on the correlation between these logical bias features and each homogeneous semantic anchor point, and uses this polarization vector to perform nonlinear weighted offset modulation on the original gravitational field strength value. The modulation logic satisfies the following formula: ,in, This represents the modulated target gravitational field strength. This represents the original gravitational field strength of the vector to be processed relative to a specific homogeneous semantic anchor point. To preset the polarization response coefficients, For the generated polarization vector, The feature vector is used to represent the engineering logical correlation degree between the to-be-processed vector and the specific homogeneous semantic anchor point. The system establishes a logical impedance boundary to prevent cross-partition data from being captured by mistake in view of the visual segmentation line and background mutation line existing in the document page. The processor detects whether the logical connection between the to-be-processed vector and the semantic anchor point passes through the logical impedance boundary. If the logical connection passes through the logical impedance boundary, the processor adjusts the equivalent logical distance between the to-be-processed vector and the target semantic anchor point according to the number of the passed logical impedance boundaries, so as to perform decay correction on the logical gravity field strength value, so that the semantic gravity preferentially flows along the same cell equal low-resistance area, and the topological consistency of data recognition under visual interference is ensured.

[0031] In the processing of an industrial long document, the global logical integrity is maintained through a state tunnel mechanism. The semantic anchor points of the processed page are extracted and cached, and are defined as virtual state features. The virtual state features are projected to the initial logical site of the current to-be-processed page, that is, the starting coordinates of the upper left corner of the page. The processor assigns a persistent weight according to the engineering properties of the virtual state features. The processor calculates the compensation gravity field strength value of the virtual state features to the to-be-processed vector in the current page according to the persistent weight and the page offset, and fuses it with the original gravity field strength value in the current page to perform logical attribution judgment across pages. This procedure solves the problem of data fragmentation caused by the absence of key anchor points in the current window. The processor encapsulates the semantic anchor points of the processed page through the state tunnel protocol, generates virtual state particles carrying engineering properties and physical dimensional features, projects them to the initial logical site coordinates of the current page, calculates the compensation gravity field strength, and calibrates the persistent weight and the decay factor through long-range correlation experiments. The page offset between the current page and the anchor point source page is calculated. The system performs gravity fusion operation to solve the anchor point missing failure caused by window switching. The compensation field strength value is written into the structured tree node register synchronously to realize cross-page dimensional semantic correlation conduction and maintain the global logical integrity of large-scale data governance. In order to cope with the pseudo gravity field generated by the pollution or interference characters, the system constructs a convergence damping operator using the OCR recognition confidence residual. The processor extracts the recognition confidence residual corresponding to each to-be-structured field in the character recognition stream, and modulates the convergence speed of each to-be-processed vector in the logical gravity field according to the recognition confidence residual. The greater the recognition confidence residual, the stronger the retardation effect of the convergence damping operator on the association of the to-be-processed vector to the semantic anchor point, so that the noise field cannot complete effective association within the calculation iteration period, thereby improving the processing efficiency of low-quality document scans.

[0032] The system performs a logical conflict scan using physical dimension constraints after generating the preliminary structured tree, and when a physical dimension conflict is identified, the processor extracts context features within the conflict point neighborhood and performs secondary path redirection. In the process of performing secondary path redirection, the processor selects the path that makes the local logical structure tend to an ordered state as the final merging path by calculating the local logical structure entropy value of the conflict point under different redirection paths, and generates an updated memory memory mapping index. When performing secondary path redirection, the processor calls the gradient boundary parameters in the hexadecimal engineering evolution matrix, constructs multiple virtual logical branches for the to-be-processed vectors in the conflict point neighborhood, and obtains the corresponding physical dimension sequence. The system calculates the sum of the absolute values of the differences between adjacent sampling points in the sequence as a dispersion criterion, and converts the dispersion criterion into a path merging probability distribution using a Gaussian kernel function. The processor selects the path with the lowest entropy value in the path merging probability distribution, i.e., the path with the highest logical chain evolution stability, as the final merging path, and writes the determination result to the structured tree node register address to trigger the bus control signal to update the memory mapping index. For the problem of missing key anchor points generated in the digitization process of old paper documents, the processor captures process chain residual information generated by structured entities. When an orphan field without a corresponding semantic anchor point is identified, the processor performs topological matching in the current logical space using the process chain residual information, reverses the missing semantic anchor point category to which the orphan field belongs according to the matching result, and creates a virtual placeholder anchor point to perform logical association compensation, thereby realizing information completion under the condition of physical signal loss. The polarization vector generated by the logical polarization program provides a high-confidence anchor reference for the capture of process chain residual information. The processor uses the high-confidence anchor reference to define a process path search window in the memory addressing space, reducing the topological search radius for reversing the semantic anchor point category to which the orphan field belongs, so that the logical polarization program and the semantic residual accompanying utilization mechanism produce a synergistic effect when processing high-noise and densely arranged unstructured documents. The processor creates a virtual placeholder anchor point to perform logical association compensation after completing the reversal.

[0033] Example 1: In the processing of a large number of multi-page process record documents on a numerical control machine tool, two groups of technical parameters, main shaft one and main shaft two, exist side by side in the page, and the structured field, the bearing speed value 1500, is located in the geometric center region of the two groups of semantic anchors, the main shaft one number and the main shaft two number, resulting in a Euclidean distance difference value of the value relative to different semantic anchors less than a preset polarization threshold. The processor obtains the target document character recognition stream and extracts the center point coordinates of the character bounding rectangle belonging to the semantic anchor and establish a logical coordinate system, perform feature vectorization processing on the to-be-structured field to obtain a to-be-processed vector, and detect whether a logical connection between the vector and a semantic anchor point passes through a visual segmentation line of a bearing partition; if the logical connection passes through the visual segmentation line of the bearing partition, the processor establishes the visual segmentation line as a logical impedance boundary and calculates an impedance factor according to the number of passes to adjust the equivalent logical distance, realize decay correction of the logical gravitational field strength value, and when the corrected logical gravitational field strength value is still in the balance interval, extract a cooling loop identifier in the neighborhood of the to-be-processed vector as a logical bias feature to generate a polarization vector , and combine the feature vector representing the engineering logical correlation degree with a preset polarization response coefficient According to the formula , the modulated target gravitational field strength value is calculated, wherein is the modulated target gravitational field strength value, is the original gravitational field strength value of the to-be-processed vector relative to a specific semantic anchor point, is a preset polarization response coefficient, is the generated polarization vector, is a feature vector representing the engineering logical correlation degree between the to-be-processed vector and the specific semantic anchor point.

[0034] In the document cross-page processing stage, the processor extracts the main axis number as a high-confidence semantic anchor point and defines it as the initial logical site of the virtual state feature projection to the subsequent page, and according to the page number offset , combines the persistent weight to calculate the compensation gravitational field strength value, so that the orphan field located in the subsequent page, i.e., the instantaneous working current, is constrained by the compensation gravitational field strength value and merged into the corresponding logical level of the main axis one; finally, the processor calls the pre-stored hexadecimal engineering evolution matrix in the memory to perform topological consistency score calculation on the generated preliminary structured tree, and according to the monotonic evolution law of rotational speed and current in the physical dimension sequence, selects a merging path that minimizes the local logical structure entropy value, and the topological consistency score satisfies the formula , wherein is the topological consistency score, is the current measurement value in the physical dimension sequence, is the expected theoretical value predicted based on the engineering evolution matrix.

[0035] Embodiment 2: In the test environment simulating the digital governance of discrete manufacturing site process record documents, the test platform is deployed on a computing node configured with a 2.8GHz processor, which simulates the electrical-digital data processing process by calling the hexadecimal industry dictionary matrix stored in the memory, the test data is derived from 1000 multi-station parallel processing process records collected by the physical experimental platform, the sequence of scanned copies is superimposed with a Gaussian white noise with a signal-to-noise ratio of 20dB in the generation process, and a 5% nonlinear geometric distortion is introduced to simulate scanning deflection, in the test design, the setting of the polarization threshold considers the balance between the recognition sensitivity of homogeneous anchor points and the noise suppression performance, when the polarization threshold value is 5% of the average value of the original gravitational field strength, the response speed of the system to the high symmetry layout is in the optimal interval, this test sets the polarization threshold value to 5%, and sets the initial value of the persistence weight to 0.85, as described in the foregoing specific embodiments, the processor extracts the spatial topology coordinates (x, y) of the semantic anchor points and performs feature vectorization processing on the structured fields to be processed, and then starts the logical gravitational field model to perform data mapping; to verify the effectiveness of the method of the present application under complex working conditions, the test sets up a multi-dimensional control system composed of a test group, a control group A, a control group B and a control group C, wherein the control group A adopts a linear extraction method combining character recognition and named entity recognition, the control group B removes the logical polarization program and the state tunnel mechanism, and the control group C sets the polarization threshold to 30%, when processing the principal axis parameter table with high symmetry, it is observed that the control group A produces attribution conflict, and the control group B produces 54.2% of the judgment deadlock in the gravitational balance zone due to the lack of modulation of the polarization vector . The test group obtains an attribution accuracy of 98.6% by calculating the target gravitational field strength value

[0036] Table 1: Performance comparison table of each group under different scenarios

[0037]

[0038] In the gradient verification process, by adjusting the core problem variable, i.e. the page offset , it is observed that when increases from 1 to 5, due to the time step attenuation factor contained in the state tunnel mechanism, the topological consistency score of the test group presents an exponential decline trend, but the attribution accuracy remains above 92.0% at , compared with the control group B without configuring the persistence weight, which shows anchor dislocation phenomenon at , with an accuracy of less than 15%, when the identification confidence residual increases due to document contamination, the convergence damping operator enhances the blocking effect on noise characters, which makes the convergence speed decrease.

[0039] Embodiment 3: This embodiment combinesFigs. 1 to 3 An AI semantic enhanced unstructured manufacturing document data structuring method is described as follows, Fig. 1 As shown, the process starts with inputting the target document, and the system then enters the parallel processing stage, on one hand, it calls the industry dictionary to obtain the character recognition stream to parse the original text data of the document, and then locates the semantic anchor points and their topological coordinates based on the physical attribute orientation keywords, on the other hand, it performs feature vectorization processing on the to-be-structured field to generate a to-be-processed vector containing text semantics and location distribution features, the system combines the correlation strength determined by the engineering logic mapping table and the Euclidean distance calculation logic gravitational field strength value, and performs homogeneous anchor point field strength difference value judgment according to the field strength value, if the field strength difference value is less than the polarization threshold, the logic polarization program is started and the logic bias features are extracted, and then the polarization vector is generated and used to perform nonlinear weight bias modulation on the original gravitational field strength to determine the unique attribution based on the maximum field strength, if the field strength difference value is not less than the threshold, the to-be-processed vector is directly associated with the target semantic anchor point, and finally the structured data is output to end the process.

[0040] As shown, Fig. 2 The horizontal coordinate represents the signal-to-noise ratio in dB, and the vertical coordinate represents the attribution determination accuracy in %. The chart contains three change curves, the test group represented by the solid line, the control group A represented by the dotted line, and the control group B represented by the long dashed line. As can be seen from the data trend, the attribution determination accuracy of the test group remains at a high level and shows a steady upward trend when the signal-to-noise ratio increases from 10 dB to 40 dB, while the accuracy of the control group A is very low in a low signal-to-noise ratio environment but increases with the increase of signal-to-noise ratio, and the control group B shows an intermediate growth trend between the two. Fig. 3 As shown, the left side of the overall architecture of the system is deployed with a field data acquisition terminal, which integrates an industrial PC or a scanning workstation, an industrial document scanner, and a data transmission interface, responsible for inputting the original document image and transmitting the character recognition stream to the right side of the system. The core area of the architecture is the data center computing node, which is internally integrated with high-performance processors arranged from top to bottom as character recognition and feature extraction engines, logic gravitational field modeling units, a logic arbitration module as the core processing unit, a logic polarization program, and an internal control bus. The computing node is connected downward to a non-volatile memory, which is internally constructed with a knowledge base and mapping table containing a hexadecimal industry dictionary, an engineering logic mapping table, and a memory mapping index. After the matrix calling and index updating are completed, the computing node outputs the structured tree data in JSON / XML format to the right side.

[0041] Example 4: In processing a 150-page heavy equipment maintenance manual, the system's computing nodes face an edge condition where the signal-to-noise ratio (SNR) of the scanned document drops from the normal 35dB to 15dB. At this point, the character recognition confidence residual within the document page increases from 0.02 to 0.28, leading to a risk of uncontrolled collapse of the logic gravitational field. The processor monitors the quality indicators of the character recognition stream through the logic arbitration module and initiates an adaptive calibration procedure for the polarization threshold based on the SNR variable. When the average recognition confidence residual of a local page exceeds 0.15, the processor increases the polarization threshold to 0.25 to improve the recognition of homogeneous semantic anchors. Regarding the tolerance for point conflicts, in practice, the processor receives a 16-bit integer character feature sequence, calls the hexadecimal engineering logic mapping table stored in non-volatile memory to perform addressing operations, calculates the logical gravitational field strength value of the vector to be processed relative to each semantic anchor point, and when processing cross-page associations, extracts the homepage sensor group number as a semantic anchor point and projects it to the starting point of the memory address. To address the situation where the homepage anchor point is missing due to page corruption, the processor captures the residual information of the process chain generated by the structured entity and calls the hexadecimal engineering evolution matrix to perform topology matching, calculating the topology consistency score of the virtual logical branches. At that time, the processor compares the physical dimension sequence through a single loop iterative algorithm. When the field to be structured is modulated by the damping coefficient of the semantic friction operator, its convergence path in the gravitational field shows a nonlinear step convergence trend. Experimental measurements show that when the signal-to-noise ratio drops to 15dB and more than 3 key anchor points are missing, the virtual occupant anchor points reconstructed by the system maintain a topological alignment accuracy of 93.4%, and the data orphanage rate is reduced by 78.5% compared with the sample group without the process chain residual utilization mechanism.

[0042] Table 2: Parameter Calibration and Accuracy Records under Different Signal-to-Noise Ratio Gradients

[0043]

[0044] With page number offset Added, the logical arbitration module is based on topology consistency score The gravitational potential energy distribution is corrected, in which, ,in The current measurement value in the sequence of physical quantities. The expected theoretical value is based on the engineering evolution matrix prediction, when the measured value Compared with the expected theoretical value When the deviation fluctuates within a 5% error band, the processor writes the judgment result to the register address corresponding to the structured tree node. When processing complex nested forms, the system identifies the logical impedance boundary generated by the visual segmentation line and introduces an impedance factor. The equivalent logical distance is corrected to avoid the data of adjacent cells being incorrectly merged. After the system completes the processing of the entire manual, the processor triggers the bus control signal to update the memory mapping index, and the generated hierarchical nested structured tree is solidified into the memory partition.

[0045] In embodiment 5, the processor calls the memory addressing space to write the industry dictionary and the engineering logic mapping table initialization data. The system extracts 100 standard process card samples to analyze the features, calculates the coupling coefficient of the association between each physical entity name and the process parameter category, converts it to a hexadecimal code and solidifies it to the memory addressing site. The processor operates clustering on the feature vectors with a semantic similarity greater than 0.82 in the samples to determine the topological site coordinates in the multidimensional space. According to the monotonicity characteristics of physical parameters in the process timing, the processor assigns a gravitational weight to each clustered site coordinate. The processor calibrates the correlation coefficient in the mapping table representing the physical magnitude change law of different manufacturing parameters in the process timing. After data testing, it is determined that the dispersion of the initial logical gravitational field strength value generated by the to-be-processed vector relative to the theoretical expected value is within a deviation interval of 4.2%. When the system faces low-resolution scanning conditions, the pre-calibration program is started. The system calls 10 document reference segments containing visual segmentation lines, calculates the recognition confidence residual under different contrast levels and determines the polarization threshold dynamic correction step. The processor searches for the polarization threshold in the range of 0.05 to 0.30 through a cyclic iteration algorithm to improve the tolerance to homogeneous semantic anchor point conflicts. According to the page background noise power spectral density, the processor adjusts the persistent weight reference point to the 0.85 coordinate site, making the initial calculation deviation of topological consistency score within the preset tolerance interval. The processor writes the calibrated parameters into the logical arbitration module configuration register and verifies the impedance factor correction through the test logic connection across the impedance boundary. Under this program constraint, the system's perception sensitivity to page offset changes is improved by 15.6%, and the final structured tree node register address write action matches the preset memory mapping index.

[0046] In the pre-calibration scenario for systematic deployment of cross-industry manufacturing documents, the system faces a mixed stream of original process record data in metric units and imperial units. The processor starts the reference alignment procedure for physical quantity dimension units to eliminate the dimensional deviation of the to-be-processed vector in the gravitational field calculation process. The specific process is as follows: the processor extracts the numerical value bits and the corresponding unit descriptor bits in the character recognition stream, maps the non-standard units to a unified electronic digital processing scale by retrieving the pre-stored hexadecimal dimension conversion matrix in the memory, and the normalized calculation of the numerical value satisfies the linear transformation relationship , where is the dimensionless feature value after normalization processing, is the original measured value. is a preset lower limit of the range of the physical quantity, is a preset upper limit of the range of the physical quantity, the processor writes the transformed dimensionless characteristic value into the corresponding feature vector dimension; after completing the dimension alignment action, the system starts the polarization threshold optimization procedure for the page layout complexity, the processor calls 5 groups of typical high-density sample tables under the current deployment environment and calculates the local logical structure entropy value, according to the average distribution interval of the homogeneous semantic anchor points in the sample table determines the initial working point of the polarization threshold, the processor performs a cyclic iterative search near the initial working point with a step size of 0.01 to obtain the target polarization threshold that makes the judgment deadlock rate lower than 0.5%, and solidifies it and the corresponding gravitational weight coefficient to the configuration register partition of the logical arbitration module, under the constraint of the pre-calibration procedure, when the system processes subsequent unknown layout manufacturing documents, it realizes the merging of different semantic levels by using the calibrated logical gravitational field model, the processor triggers the bus control signal after determining the ownership, updates the memory mapping index to complete the logical conversion of the unstructured document to the hierarchical nested structured tree.

[0047] It is apparent for those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and the present application can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application.

[0048] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not limiting, although the present application is described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. A method for structuring unstructured manufacturing document data with AI semantic enhancement, characterized in that, Includes the following steps: Step 101: Obtain the character recognition stream of the target document, extract keywords with physical attributes by matching with industry dictionary and define them as semantic anchors, and obtain the spatial topological coordinates of each semantic anchor in the coordinate system of the target document page; Step 102: Perform feature vectorization processing on the fields to be structured in the target document to obtain a vector to be processed containing text semantic features and positional distribution features; Step 103: Determine the logical association strength between the semantic anchor point and the vector to be processed, and calculate the logical gravitational field strength value of the vector to be processed relative to each semantic anchor point based on the Euclidean distance between spatial topological coordinates and the logical association strength. Step 104: Associate the vector to be processed with the target semantic anchor point according to the order of logical gravitational field strength values ​​from largest to smallest; During the association process, if the difference between the logical gravitational field strength values ​​of the vector to be processed and multiple homogeneous semantic anchors is less than the preset polarization threshold, the logical polarization procedure is initiated. Step 105: Extract the logical bias features within the preset neighborhood of the vector to be processed, generate a polarization vector based on the logical bias features and the endogenous engineering correlation of each homogeneous semantic anchor point, and use the polarization vector to perform nonlinear weight offset modulation on the original gravitational field strength value of each homogeneous semantic anchor point to determine the unique affiliation of the vector to be processed.

2. The method for structuring unstructured manufacturing document data with AI semantic enhancement according to claim 1, characterized in that, Step 103 further includes: identifying visual segmentation lines and background abrupt changes in the target document's page and establishing them as logical impedance boundaries; detecting whether the logical connection between the vector to be processed and the semantic anchor point crosses the logical impedance boundary; if the logical connection crosses the logical impedance boundary, increasing the equivalent logical distance between the vector to be processed and the target semantic anchor point according to the number of logical impedance boundaries crossed, in order to perform attenuation correction of the logical gravitational field strength value.

3. The method for structuring unstructured manufacturing document data with AI semantic enhancement according to claim 1, characterized in that, The steps also include: extracting and caching semantic anchors of processed pages in the target document and defining them as virtual state features; projecting the virtual state features onto the initial logical position of the current page to be processed and assigning persistence weights according to the engineering attributes of the virtual state features; calculating the compensation gravitational field strength value of the virtual state features for the vector to be processed in the current page to be processed based on the persistence weights and page number offsets; and fusing the compensation gravitational field strength value with the original gravitational field strength value in the current page to be processed and performing cross-page logical structured judgment.

4. The method for structuring unstructured manufacturing document data with AI semantic enhancement according to claim 1, characterized in that, Also includes: In step 105, if the difference in gravitational field strength between the vector to be processed and multiple homogeneous semantic anchors is less than a preset threshold, the vector to be processed is virtually associated with each homogeneous semantic anchor to construct the corresponding candidate logical state space. Obtain the engineering logic consistency entropy value corresponding to each candidate logic state space; determine the homogeneous semantic anchor point corresponding to the candidate logic state space with the smallest engineering logic consistency entropy value as the final destination of the vector to be processed.

5. The method for structuring unstructured manufacturing document data with AI semantic enhancement according to claim 1, characterized in that, The steps also include: extracting the recognition confidence residuals corresponding to each field to be structured in the character recognition stream; constructing a convergence damping operator for each vector to be processed based on the recognition confidence residuals; and using the convergence damping operator to modulate the convergence speed of each vector to be processed in the logical gravitational field. The larger the recognition confidence residual, the stronger the hindering effect of the convergence damping operator on the association of the vector to be processed with the semantic anchor point.

6. The method for structuring unstructured manufacturing document data with AI semantic enhancement according to claim 1, characterized in that, The industry dictionary and engineering logic mapping table are stored as hexadecimal matrices, with pre-set correlation coefficients representing the physical magnitude changes of different manufacturing parameters in the process timing. In step 103, the processor calls the pre-stored hexadecimal matrix in the memory and determines the logical correlation strength through iterative calculation.

7. The method for structuring unstructured manufacturing document data with AI semantic enhancement according to claim 1, characterized in that, In the logic polarization procedure, the logic for performing nonlinear weighted offset modulation on the native gravitational field strength using the polarization vector satisfies the following formula: ,in, This represents the modulated target gravitational field strength. This represents the original gravitational field strength of the vector to be processed relative to a specific homogeneous semantic anchor point. The preset polarization response coefficients, For the generated polarization vector, The feature vector is used to characterize the degree of engineering logical correlation between the vector to be processed and homogeneous semantic anchors.

8. The method for structuring unstructured manufacturing document data with AI semantic enhancement according to claim 1, characterized in that, Also includes: After step 104, residual information of the process chain generated during the gravitational field calculation of the vector to be processed is captured; When an orphan field without a corresponding semantic anchor is identified in the target document, topological matching is performed in the current logical space using residual information from the process chain; the missing semantic anchor category to which the orphan field belongs is inverted based on the matching result, and virtual placeholder anchors are created to perform logical association compensation for the orphan field.

9. The method for structuring unstructured manufacturing document data with AI semantic enhancement according to claim 1, characterized in that, The steps also include: performing a logical conflict scan on the preliminary structured tree generated in step 104 using preset physical dimension constraints; when a physical dimension conflict is identified, extracting the context features in the neighborhood of the conflict point and performing a secondary path redirection to re-determine the hierarchical affiliation of the vector to be processed.

10. A method for structuring unstructured manufacturing document data with AI semantic enhancement according to claim 9, characterized in that, During the secondary path redirection process, the entropy value of the local logical structure of the conflict point under different redirection paths is calculated. The path that makes the local logical structure tend to be ordered is selected as the final merging path, and an updated memory mapping index is generated.

Citation Information

Patent Citations

  • Document data structured storage and retrieval method based on AI drive

    CN119829723A

  • Computer document intelligent compliance detection system based on deep learning

    CN120596657A

  • Semantic anchor point definition method based on guide word and topological structure

    CN120833606A

  • Automatic label labeling and classifying method and system for unstructured system documents

    CN121301571A

  • System and method for adaptive semantic parsing and structured data transformation of digitized documents

    US12417214B1

Cited By

  • Data asset standardized packaging method and system and storage medium

    CN122113144A

  • Database logic topology construction method based on large language model

    CN122153034A

  • Intelligent classification and archiving system for social security electronic records integrating RPA and NLP

    CN122412674A