BOM analysis method, system and equipment based on multi-modal semantic fusion and medium

By adopting multimodal semantic fusion technology in BOM analysis, including BERT model, Table Transformer model, graph neural network and reinforcement learning algorithm, the shortcomings of multimodal and unstructured BOM data analysis in the existing technology are solved, and more efficient and accurate BOM data processing and automation are achieved.

CN120197608APending Publication Date: 2025-06-24深圳一道创新技术有限公司
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510269959.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing BOM analytical methods are difficult to effectively process multimodal and unstructured data, and they rely too much on manual intervention, resulting in poor adaptability of analysis, low accuracy, low efficiency and high labor costs.

Method used

The BOM analysis method based on multimodal semantic fusion is adopted to extract the semantic features of text data through the BERT model, the Table Transformer model recognizes the table structure in the image data, and the graph neural network constructs the knowledge graph of materials and processes, and dynamically optimizes mapping rules through reinforcement learning algorithms to generate optimized BOM structured data.

Benefits of technology

It improves the efficiency and accuracy of BOM data analysis, reduces manual intervention, improves the level of automation and intelligence of data processing, and can effectively process multimodal and unstructured data, adapt to BOM data of different formats and complex properties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197608A_ABST
    Figure CN120197608A_ABST
Patent Text Reader

Abstract

The invention relates to a BOM analysis method, system and device based on multi-modal semantic fusion and a medium, the method comprises the following steps: receiving BOM data, the BOM data comprising text input data and image input data; performing semantic feature extraction on the text input data by adopting a BERT model to generate a semantic embedding vector of the text; performing row and column structure identification on the image input data by adopting a Table Transform model so as to obtain a corresponding table image, and extracting table content information from the table image; based on the semantic embedding vector and the table content information, analyzing the relationship between the material and the process, and adopting a graph neural network to construct a knowledge graph of the material and the process; dynamically optimizing a mapping rule in the knowledge graph through a reinforcement learning algorithm, adjusting a corresponding field mapping rule, and generating an optimized knowledge graph; and based on the optimized knowledge graph, performing standardized conversion on the corresponding BOM data, and outputting BOM structured data. The BOM analysis method and device have the effect of improving the BOM analysis efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of electronic manufacturing, and in particular, to a BOM parsing method, system, device, and medium based on multi-modal semantic fusion. Background Art

[0002] Currently, during the BOM (Bill of Materials) parsing process in electronic manufacturing enterprises, BOM data in various formats are often faced, including text and image data. BOM data usually contains complex material and process information, and this information must be correctly parsed and structured to support subsequent production scheduling and material management.

[0003] Existing BOM parsing methods mainly rely on template matching or rule matching techniques. These methods require pre-defining templates or rules and adapting to different formats of BOM data through manual intervention. However, with the increasing complexity of material types and process flows in the production process, existing parsing methods often struggle to handle multi-modal, multi-format, and large-scale BOM data, and rely on manual intervention, resulting in low parsing accuracy, low efficiency, and high labor costs.

[0004] The above-mentioned existing technical solutions have the following defects: Existing BOM parsing methods cannot effectively handle multi-modal and unstructured data, and rely too much on manual intervention, resulting in poor adaptability of parsing. Therefore, there is room for improvement. Summary of the Invention

[0005] To improve the efficiency of BOM parsing, the present application provides a BOM parsing method, system, device, and medium based on multi-modal semantic fusion.

[0006] The first invention object of the present application is achieved through the following technical solutions: A BOM parsing method based on multi-modal semantic fusion, the BOM parsing method based on multi-modal semantic fusion includes: Receiving BOM data, the BOM data including text input data and image input data; Using a BERT model to extract semantic features from the text input data to generate a semantic embedding vector of the text; Using a Table Transformer model to perform row-column structure recognition on the image input data, thereby obtaining a corresponding table image, and extracting table content information from the table image; Based on the semantic embedding vector and the table content information, analyzing the relationship between materials and processes, and using a graph neural network to construct a knowledge graph of the materials and the processes; Dynamically optimize the mapping rules in the knowledge graph through a reinforcement learning algorithm, adjust the corresponding field mapping rules, and generate an optimized knowledge graph; Based on the optimized knowledge graph, perform standard conversion on the corresponding BOM data and output BOM structured data.

[0007] By adopting the above technical solutions, by receiving BOM data and preprocessing it, it is possible to ensure the uniformity and cleaning of the data format, reduce errors in subsequent processing, and thus improve the efficiency and accuracy of data parsing; by using the BERT model to extract semantic features from text input data, complex semantic information in the text data can be transformed into embedding vectors, providing an efficient data representation for subsequent analysis and reasoning, and thus improving the accuracy of text data parsing; by using the Table Transformer model to identify the row-column structure of image input data, the table structure in the image can be accurately extracted, avoiding misidentification problems in traditional image processing methods, and thus improving the accuracy of image data parsing; by analyzing the relationship between materials and processes based on semantic embedding vectors and table content information and constructing a knowledge graph of materials and processes, the complex dependency relationship between materials and processes can be effectively captured, providing accurate data support for the optimization and management of subsequent production processes; by dynamically optimizing the mapping rules in the knowledge graph through a reinforcement learning algorithm, the rules can be adaptively adjusted during the parsing process, improving the parsing accuracy and reducing manual intervention, and thus improving the automation level and accuracy of BOM data processing.

[0008] In one example, this application can be further configured as follows: when using the BERT model to extract semantic features from the text input data to generate semantic embedding vectors of the text, it specifically includes: Perform adaptive preprocessing on the text input data to obtain standard text data, and input the standard text data into the BERT model; Extract context information from the standard text data through the multi-layer Transformer structure of the BERT model; Identify lexical relationships and semantic associations in the text through the self-attention mechanism in the BERT model, extract the semantic features of each field, and then generate the semantic embedding vectors according to the context relationship and the semantic features.

[0009] By adopting the above technical solution, by using the BERT model to extract semantic features from the text input data, valuable semantic information can be effectively extracted from unstructured text data, and each field in the text is converted into a high-dimensional semantic vector. These vectors can reflect the context relationship and lexical relevance in the text, enabling the system to more accurately understand and parse key information such as material descriptions and process requirements in BOM data, reducing the limitations of relying on manually set rules in traditional methods, improving the adaptability and accuracy of the system when processing diverse texts, and ultimately promoting the automation degree and intelligent level of BOM data parsing.

[0010] In one example, this application can be further configured as follows: using the Table Transformer model to perform row-column structure recognition on the image input data, thereby obtaining the corresponding table image, and extracting table content information from the table image, specifically including: Preprocess the image input data and use a preset DETR model to detect cross-page tables. The loss function of the DETR model is: total loss = λ L1 × L1 loss + λ giou × GIOU loss + λ ce × CE loss, where λ L1 、λ giou 、λ ce are loss weights, and then obtain the standard image data; Input the standard image data into the Table Transformer model, identify the row-column structure in the image through the self-attention mechanism, and calibrate the positions of table cells; Based on the row-column structure, extract the content information of the positions of the table cells, and perform context semantic coherence analysis on the content information through bidirectional LSTM, thereby obtaining the table content information.

[0011] By adopting the above technical solution, by using the Table Transformer model to perform row-column structure recognition on the image input data, the row-column information of the table can be accurately extracted from the image, and the positions of table cells can be calibrated, thus ensuring that the content of each cell is accurately extracted. This step reduces the misrecognition problem of traditional image processing methods in complex table structures, improves the reliability and accuracy of image data parsing. Especially when processing complex formats such as scanned images and PDF files, it can effectively avoid the loss or misalignment of table information, making the finally generated BOM structure more accurate and complete, and supporting more efficient subsequent data processing and decision-making support.

[0012] In one example, the present application can be further configured as follows: based on the semantic embedding vector and the table content information, analyze the relationship between materials and processes, and use a graph neural network to construct a knowledge graph of materials and processes, specifically including: Represent the materials and processes in the text according to the semantic embedding vector to generate a preliminary semantic relationship between the materials and the processes; According to the table content information, perform information fusion and relationship reasoning on the materials and processes to generate structural information between the materials and the processes; Input the preliminary semantic relationship and the structural information into the graph neural network, represent the materials and processes as nodes, then model the relationships between the nodes, and generate the knowledge graph of materials and processes through feature propagation of adjacent nodes and graph update.

[0013] By adopting the above technical solution, by analyzing the relationship between materials and processes based on semantic embedding vectors and table content information, it is possible to deeply explore the association between materials and processes on the basis of multi-modal data, and generate a more refined knowledge graph of materials and processes. In this process, the information in the text and table data is effectively integrated, and the system can identify the mutual influence and dependence relationship between the attributes, quantities, specifications of materials and the process flow, so as to provide accurate data support for subsequent production scheduling and material management. This knowledge graph generated based on the graph neural network can help enterprises make optimized decisions in actual production, improve production efficiency, reduce resource waste, and further enhance the accuracy and flexibility of supply chain management.

[0014] In one example, the present application can be further configured as follows: the construction of the reinforcement learning algorithm specifically includes: Obtain and define the working state of reinforcement learning based on the current field mapping rule, historical parsing accuracy trend, and manual intervention frequency, and then define the actions of reinforcement learning based on the working state. The actions include adjusting the field mapping weight, adding or deleting field mapping rules, and modifying regular expression parameters; Define a reward function based on the historical parsing accuracy trend and the manual intervention frequency. The reward function is R = α × Accuracy + β × (1 / Human_Intervention_Count), where Accuracy is the historical parsing accuracy, Human_Intervention_Count is the number of manual interventions, and α, β are loss weights; Optimize the current field mapping rule based on the PPO algorithm, and use the entropy regularization mechanism to expand and update the current field mapping rule, and then generate the reinforcement learning algorithm.

[0015] By adopting the above technical solution, through the construction of a reinforcement learning algorithm, the field mapping rules can be dynamically adjusted according to real-time feedback, thereby maximizing the parsing accuracy and reducing manual intervention. The reinforcement learning model can automatically identify the deficiencies in the current rules and optimize the field mapping by analyzing the historical parsing accuracy and the frequency of manual intervention in real time, thus improving the intelligence level of the system. The optimization based on the PPO algorithm can not only accurately adjust the mapping rules, but also prevent the model from overfitting through the entropy regularization mechanism, ensuring that the rules can adaptively change under various BOM data formats and diverse data scenarios, and improving the system's parsing ability for different BOM data. This method can significantly reduce manual intervention, improve the automation degree and long-term stability of BOM parsing, and enhance the overall data processing efficiency.

[0016] In one example, the present application can be further configured as follows: Dynamically optimizing the mapping rules in the knowledge graph through the reinforcement learning algorithm, adjusting the corresponding field mapping rules, and generating an optimized knowledge graph, specifically including: Taking the mapping rules in the knowledge graph as input and executing the reinforcement learning algorithm; Based on the reward function, calculating and adjusting the corresponding mapping rules, and then updating the mapping rules in the knowledge graph according to the adjusted mapping rules to generate the optimized knowledge graph.

[0017] By adopting the above technical solution, dynamically optimizing the mapping rules in the knowledge graph through the reinforcement learning algorithm can flexibly handle the changes and updates of different BOM data, automatically adjust the parsing rules, and reduce the cumbersome process of manually setting rules. The optimization of the mapping rules based on reinforcement learning can evaluate the effect of each rule adjustment through the reward function, thereby ensuring that while improving the parsing accuracy, the need for manual intervention is reduced. The optimized knowledge graph not only improves the accuracy of BOM data parsing, but also effectively supports long-term data updates and adaptations, maintaining the high-efficiency performance of the system in different application scenarios. By continuously iterating and updating the rules, the finally generated optimized knowledge graph will provide continuous high-quality data support for the decision-making of subsequent production processes, enhancing the intelligence and adaptability of the entire BOM parsing process, and greatly improving the work efficiency and the application scope of the system.

[0018] The above second invention object of the present application is achieved through the following technical solutions: A BOM parsing system based on multi-modal semantic fusion, the BOM parsing system based on multi-modal semantic fusion includes: A data receiving module for receiving BOM data, where the BOM data includes text input data and image input data; A text data processing module for extracting semantic features from the text input data using a BERT model to generate a semantic embedding vector of the text; An image data processing module for performing row-column structure recognition on the image input data using a Table Transformer model, thereby obtaining a corresponding table image and extracting table content information from the table image; A semantic relationship analysis module for analyzing the relationship between materials and processes based on the semantic embedding vector and the table content information, and constructing a knowledge graph of the materials and the processes using a graph neural network; A reinforcement learning optimization module for dynamically optimizing the mapping rules in the knowledge graph through a reinforcement learning algorithm, adjusting the corresponding field mapping rules, and generating an optimized knowledge graph; A data standardization module for performing standardized conversion on the corresponding BOM data based on the optimized knowledge graph and outputting BOM structured data.

[0019] By adopting the above technical solutions, by receiving BOM data and preprocessing it, it is possible to ensure the uniform format and cleaning of the data, reduce errors in subsequent processing, and thus improve the efficiency and accuracy of data parsing; by using a BERT model to extract semantic features from text input data, the complex semantic information in the text data can be transformed into embedding vectors, providing an efficient data representation for subsequent analysis and reasoning, and thus improving the accuracy of text data parsing; by using a Table Transformer model to perform row-column structure recognition on image input data, the table structure in the image can be accurately extracted, avoiding misrecognition problems in traditional image processing methods, and thus improving the accuracy of image data parsing; by analyzing the relationship between materials and processes based on the semantic embedding vector and the table content information and constructing a knowledge graph of materials and processes, the complex dependency relationship between materials and processes can be effectively captured, providing accurate data support for the optimization and management of subsequent production processes; by dynamically optimizing the mapping rules in the knowledge graph through a reinforcement learning algorithm, the rules can be adaptively adjusted during the parsing process, improving the parsing accuracy and reducing manual intervention, and thus improving the automation level and accuracy of BOM data processing.

[0020] The above object three of the present application is achieved through the following technical solutions: A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above BOM parsing method based on multi-modal semantic fusion are implemented.

[0021] The above object four of the present application is achieved through the following technical solutions: A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned BOM parsing method based on multimodal semantic fusion are implemented.

[0022] In summary, the present application includes the following beneficial technical effects: 1. By receiving BOM data and preprocessing it, the format uniformity and cleaning of the data can be ensured, reducing errors in subsequent processing, thereby improving the efficiency and accuracy of data parsing; by using the BERT model to extract semantic features from text input data, complex semantic information in the text data can be transformed into embedding vectors, providing an efficient data representation for subsequent analysis and reasoning, thereby improving the accuracy of text data parsing; by using the Table Transformer model to identify the row-column structure of image input data, the table structure in the image can be accurately extracted, avoiding misidentification problems in traditional image processing methods, thereby improving the accuracy of image data parsing; by analyzing the relationship between materials and processes based on semantic embedding vectors and table content information, and constructing a knowledge graph of materials and processes, the complex dependence relationship between materials and processes can be effectively captured, thereby providing accurate data support for the optimization and management of subsequent production processes; by dynamically optimizing the mapping rules in the knowledge graph through reinforcement learning algorithms, the rules can be adaptively adjusted during the parsing process, improving the parsing accuracy and reducing manual intervention, thereby improving the automation level and accuracy of BOM data processing; 2. By using the BERT model to extract semantic features from the text input data, valuable semantic information can be effectively extracted from unstructured text data, transforming each field in the text into a high-dimensional semantic vector. These vectors can reflect the context relationship and lexical relevance in the text, enabling the system to more accurately understand and parse key information such as material descriptions and process requirements in BOM data, reducing the limitations of relying on manually set rules in traditional methods, improving the adaptability and accuracy of the system when processing diverse texts, and ultimately promoting the automation and intelligence levels of BOM data parsing; 3. By using the Table Transformer model to identify the row-column structure of the image input data, the row-column information of the table can be accurately extracted from the image, and the positions of the table cells can be calibrated, thereby ensuring the accurate extraction of the content of each cell. This step reduces misidentification problems in traditional image processing methods for complex table structures, improving the reliability and accuracy of image data parsing. Especially when processing complex formats such as scanned images and PDF files, it can effectively avoid the loss or misalignment of table information, making the finally generated BOM structure more accurate and complete, supporting more efficient subsequent data processing and decision-making support. Description of the Drawings

[0023] Figure 1 is a flowchart of a BOM parsing method based on multimodal semantic fusion in an embodiment of the present application; Figure 2 is a flowchart for implementing step S20 in the BOM parsing method based on multimodal semantic fusion in an embodiment of the present application; Figure 3 is a flowchart for implementing step S30 in the BOM parsing method based on multimodal semantic fusion in an embodiment of the present application; Figure 4 is a flowchart for implementing step S40 in the BOM parsing method based on multimodal semantic fusion in an embodiment of the present application; Figure 5 is a flowchart for implementing the construction of a reinforcement learning algorithm in the BOM parsing method based on multimodal semantic fusion in an embodiment of the present application; Figure 6 is a flowchart for implementing step S50 in the BOM parsing method based on multimodal semantic fusion in an embodiment of the present application; Figure 7 is a schematic block diagram of a BOM parsing system based on multimodal semantic fusion in an embodiment of the present application; Figure 8 is a schematic diagram of a device in an embodiment of the present application. Detailed Description of the Embodiment

[0024] The present application will be further described in detail below with reference to the accompanying drawings.

[0025] In one embodiment, as Figure 1 shown, the present application discloses a BOM parsing method based on multimodal semantic fusion, which specifically includes the following steps: S10: Receive BOM data, where the BOM data includes text input data and image input data.

[0026] Specifically, the BOM data is composed of various types of inputs. The text input data may come from common data formats such as CSV, Excel, and TXT files, and contains fields such as material information, quantity, specifications, and process requirements. The image input data usually comes from scanned images or PDF files and contains information such as tables, graphics, or handwritten annotations. The received data will be processed and converted into standardized structured data according to its format and type to ensure the unity of the data and facilitate subsequent analysis and processing.

[0027] S20: Use the BERT model to extract semantic features from the text input data and generate semantic embedding vectors of the text.

[0028] Specifically, the text input data will first undergo adaptive preprocessing, including removing redundant punctuation marks, unifying the formats of numbers and dates, removing unnecessary spaces, etc. Then, the cleaned text data is input into the BERT model. The BERT model analyzes the relationships between the words and the context in the text through a multi-layer Transformer structure, thereby generating semantic embedding vectors for each word. These embedding vectors can represent the complex semantic information in the text, enabling the text data to be better understood and utilized by subsequent processing models.

[0029] S30: Use the Table Transformer model to perform row-column structure recognition on the image input data, thereby obtaining the corresponding table image, and extracting the table content information from the table image.

[0030] Specifically, preprocess the image input data, which includes operations such as image denoising, binarization, and region localization, to ensure that the table area in the image is accurately recognized. Then, input the processed image into the TableTransformer model. This model uses the self-attention mechanism to recognize the row-column structure in the image, extracts the content information of each cell by calibrating the position of each cell, and then restores the content of the table in the image, such as material numbers, names, quantities, etc., and ensures the correct parsing of multi-page tables.

[0031] S40: Based on the semantic embedding vectors and the table content information, analyze the relationship between materials and processes, and use a graph neural network to construct a knowledge graph of materials and processes.

[0032] Specifically, use the semantic embedding vectors generated by the BERT model to represent the material and process information in the text, and then analyze the preliminary relationship between materials and processes. At the same time, through the processing of the table content information, further extract the association between materials and processes, and generate the structured relationship between materials and processes through information fusion and relationship reasoning. Then, input these preliminary semantic relationships and structure information into the graph neural network. Materials and processes are represented as nodes in the graph. The graph neural network generates a complete knowledge graph of materials and processes through feature propagation between nodes and relationship modeling of adjacent nodes.

[0033] S50: Dynamically optimize the mapping rules in the knowledge graph through a reinforcement learning algorithm, adjust the corresponding field mapping rules, and generate an optimized knowledge graph.

[0034] Specifically, based on the historical parsing accuracy and the frequency of manual intervention, a reward function is defined. This reward function is used to evaluate whether the adjusted mapping rules improve the parsing accuracy and reduce manual intervention. Then, based on the Proximal Policy Optimization (PPO) algorithm, the mapping rules are optimized, and these rules are extended and updated through an entropy regularization mechanism to generate an optimized knowledge graph, so as to improve the accuracy and efficiency of BOM data parsing.

[0035] S60: Based on the optimized knowledge graph, the corresponding BOM data is subjected to a standardization transformation to output structured BOM data.

[0036] Specifically, the mapping rules in the optimized knowledge graph will be used to perform a standardization transformation on the BOM data, converting the unstructured or semi-structured BOM data into structured data, so that each piece of data conforms to a predetermined format standard, including a unified format for information such as material name, quantity, unit, and process steps. The transformed BOM data will be further processed and output to ensure that the output BOM data conforms to industry standards such as IPC-2581 and can be directly applied to subsequent manufacturing execution systems or other related systems.

[0037] In one embodiment, as Figure 2 shown, in step S20, that is, the BERT model is used to extract semantic features from the text input data to generate a semantic embedding vector of the text, specifically including: S21: The text input data is subjected to an adaptive preprocessing to obtain standard text data, and the standard text data is input into the BERT model.

[0038] Specifically, the text input data will undergo adaptive preprocessing, mainly including a cleaning process, such as removing irrelevant characters such as extra spaces, punctuation marks, etc., unifying the formats of numerical values and dates, removing stop words, etc.; after preprocessing, the standardized text data is sent into the BERT model for subsequent semantic feature extraction, ensuring that BERT can accurately extract valuable semantic information from the text and prepare for subsequent steps.

[0039] S22: The BERT model's multi-layer Transformer structure is used to extract context information from the standard text data.

[0040] Specifically, through the multi-layer Transformer architecture of the BERT model, the words in the standard text data are first vectorized, and the relationship between each word in the text and its context words is processed through the self-attention mechanism to generate the context information of each word. In this way, BERT can capture the long-range dependencies and complex semantic associations in the text, enabling the representation of each word to fully reflect its actual meaning in the context.

[0041] S23: Identify the word relationships and semantic associations in the text through the self-attention mechanism in the BERT model, extract the semantic features of each field, and then generate semantic embedding vectors based on the context relationship and semantic features.

[0042] Specifically, through the self-attention mechanism in the BERT model, the model will focus on the relative importance of each word in the text, and then identify the relationships and semantic associations between different words in the text. For example, the words "materials" and "processes" may have a specific connection in the context. The BERT model can identify this relationship and generate semantic feature vectors for each field, which will be used as the basis for further reasoning and analysis to help understand the specific meaning and structure of the text.

[0043] In one embodiment, as Figure 3 shown, in step S30, that is, the Table Transformer model is used to identify the row-column structure of the image input data, and then the corresponding table image is obtained, and the table content information is extracted from the table image, specifically including: S31: Preprocess the image input data and use a preset DETR model to detect cross-page tables. The loss function of the DETR model is: total loss = λ L1 × L1 loss + λ giou × GIOU loss + λ ce × CE loss, where λ L1 、λ giou 、λ ce are loss weights, and then the standard image data is obtained.

[0044] Specifically, preprocessing the image input data includes operations such as denoising, grayscale conversion, and binarization to enhance the contrast of the image and improve the clarity of the table area. Then, a preset DETR model is used to detect cross-page tables. This model is specifically designed to identify and locate table areas in complex images. The loss function of the DETR model includes L1 loss, GIOU loss, and CE loss. Among them, L1 loss measures the difference between the predicted bounding box and the true bounding box, GIOU loss is used to measure the overlap degree of the bounding boxes, CE loss is used to optimize the detection accuracy, and λ L1 、λgiou , λ ce is the loss weight, where λ L1 = 5, λ giou = 2, λ ce = 1. By combining these losses, the training of the model is optimized, and finally standardized image data is obtained.

[0045] S32: Input the standard image data into the Table Transformer model, identify the row and column structures in the image through the self-attention mechanism, and calibrate the positions of the table cells.

[0046] Specifically, the standard image data is input into the Table Transformer model. This model uses the self-attention mechanism to automatically identify the row and column structures in the image. First, through global attention to the image, the potential structural information in the image is extracted, and the rows and columns of the table are accurately calibrated. This model not only identifies the boundaries of the table but also can accurately locate the positions of the table cells. By adjusting the weights of each part in the image, the position information of each cell can be correctly identified and mapped to the actual table structure.

[0047] S33: Based on the row and column structures, extract the content information of the table cell positions, and perform context semantic coherence analysis on the content information through bidirectional LSTM, and then obtain the table content information.

[0048] Specifically, combined with the row and column structure information, the content of each cell position in the table is extracted to ensure that the data in all cells such as material numbers, names, quantities, etc. are correctly identified. Subsequently, a bidirectional LSTM model is used to perform context semantic analysis on the extracted content information. The bidirectional LSTM can consider the relationships of both the front and back contexts simultaneously, ensuring a more comprehensive semantic understanding of the table content. Especially when dealing with the possible complex associations between materials and processes, the LSTM can maintain the coherence and accuracy of information, thus extracting the complete table content information.

[0049] In one embodiment, as Figure 4 shown, in step S40, that is, based on the semantic embedding vector and the table content information, analyze the relationship between materials and processes, and use a graph neural network to construct a knowledge graph of materials and processes, specifically including: S41: According to the semantic embedding vector, represent the materials and processes in the text, and generate a preliminary semantic relationship between materials and processes.

[0050] Specifically, based on the semantic embedding vectors, the materials and processes in the text are represented in vector form. These vectors can capture the semantic connections between the materials and processes, helping to form the initial semantic relationships between them. For example, through the semantic embedding vectors, the similarity between the material names and related process operations can be quantified as a numerical value, thereby generating the semantic associations between the materials and processes, which is convenient for subsequent relationship reasoning and information fusion.

[0051] S42: According to the information in the table content, perform information fusion and relationship reasoning on the materials and processes to generate the structural information between the materials and processes.

[0052] Specifically, by analyzing the information in the table content, the relationship between the materials and processes will be further refined. By fusing the table content information with the semantic embedding vectors in the previous step, the system can infer more detailed relationships between the materials and processes, such as certain materials may require specific process steps for processing, or the types of materials used in specific process flows, thereby generating the structured information between the materials and processes, providing data support for the subsequent construction of the knowledge graph.

[0053] S43: Input the initial semantic relationship and structural information into the graph neural network, represent the materials and processes as nodes, and then model the relationships between the nodes. Through the feature propagation of adjacent nodes and the update of the graph spectrum, generate the knowledge graph of the materials and processes.

[0054] Specifically, input the initial semantic relationship and structural information into the graph neural network. In the graph neural network, the materials and processes are represented as nodes in the graph, and the edges between the nodes represent the relationships between them. The graph neural network models the relationships between the nodes to ensure that the dependence relationship between the materials and processes can be clearly reflected in the graph. At the same time, through the feature propagation of adjacent nodes and the update of the graph spectrum in the graph neural network, the knowledge graph can be gradually optimized in each iteration, generating an accurate and comprehensive knowledge graph of the materials and processes.

[0055] In one embodiment, as Figure 5 shown, in step S50, that is, the construction of the reinforcement learning algorithm, specifically includes: S501: Obtain and define the working state of the reinforcement learning based on the current field mapping rules, the historical parsing accuracy trend, and the frequency of manual intervention. Then, based on the working state, define the actions of the reinforcement learning. The actions include adjusting the field mapping weights, adding or deleting field mapping rules, and modifying the regular expression parameters.

[0056] Specifically, based on the current field mapping rules, the historical parsing accuracy trend, and the frequency of manual intervention, define the working state of the reinforcement learning. The working state includes the known field mapping rules and their performance in historical parsing, such as accuracy, error, etc., while taking into account the frequency of manual intervention. Then, based on this information, define the actions of the reinforcement learning. The actions can be adjusting the mapping weight of a certain field, adding or deleting the mapping rules of some fields, or modifying the parameters of the regular expression. These actions can help the model better adapt to the changes in BOM data during the learning process.

[0057] S502: Define the reward function based on the historical parsing accuracy trend and the frequency of manual intervention. The reward function is R = α × Accuracy + β × (1 / Human_Intervention_Count), where Accuracy is the historical parsing accuracy, Human_Intervention_Count is the number of times of manual intervention, and α and β are loss weights.

[0058] Specifically, the reward function is used to measure the optimization effect of the current field mapping rules. Its purpose is to drive the optimization process by weighing the parsing accuracy and the number of times of manual intervention. In this step, first, evaluate the accuracy of the current rules in the parsing task through the historical parsing accuracy Accuracy. Accuracy represents the ratio of the number of correctly parsed fields to the total number of fields, that is, the matching degree between the parsing result and the actual result. Then, the number of times of manual intervention Human_Intervention_Count is used to reflect the frequency of manual adjustment required by the system in actual application. Fewer manual interventions mean that the field mapping rules are more accurate, and vice versa, optimization is needed.

[0059] Furthermore, the reward function R combines these two factors. The α and β in the formula are weight coefficients, which are used to control the contribution degrees of the parsing accuracy and the number of times of manual intervention to the final reward respectively. Specifically, α = 0.7 means that the parsing accuracy has a greater impact on the reward. A higher accuracy will significantly increase the reward, thus driving the model to optimize in the direction of higher accuracy; β = 0.3 means that the impact of the number of times of manual intervention is relatively small. Reducing the frequency of manual intervention helps to increase the reward, but not as significantly as improving the parsing accuracy. This design enables the reinforcement learning algorithm to not only focus on improving the parsing accuracy during the optimization process but also moderately consider the need to reduce manual intervention.

[0060] S503: Optimize the current field mapping rules based on the PPO algorithm and use the entropy regularization mechanism to expand and update the current field mapping rules, and then generate the reinforcement learning algorithm.

[0061] Specifically, the current field mapping rules are optimized based on the PPO algorithm. The PPO algorithm updates the current policy through the policy gradient method, ensuring that each update is within a stable range, thereby improving the accuracy of BOM parsing. At the same time, an entropy regularization mechanism is adopted to encourage the system to explore new field mapping rules, avoiding overfitting by increasing the diversity of the policy, and finally generating an optimized reinforcement learning algorithm that can adaptively adjust the field mapping rules to accommodate different types of BOM data.

[0062] In one embodiment, as Figure 6 shown, in step S50, the mapping rules in the knowledge graph are dynamically optimized through the reinforcement learning algorithm, the corresponding field mapping rules are adjusted, and an optimized knowledge graph is generated, which specifically includes: S51: Use the mapping rules in the knowledge graph as input and execute the reinforcement learning algorithm.

[0063] Specifically, the mapping rules in the generated knowledge graph will be used as input and fed into the reinforcement learning algorithm for processing. The reinforcement learning algorithm will adjust these rules according to the current mapping rules, historical accuracy, and manual intervention frequency to make them more accurate and adaptable to different BOM data.

[0064] S52: Based on the reward function, calculate and adjust the corresponding mapping rules, and then update the mapping rules in the knowledge graph according to the adjusted mapping rules to generate an optimized knowledge graph.

[0065] Specifically, according to the reward function, the reinforcement learning algorithm calculates and adjusts the mapping rules, gradually optimizing the accuracy of the rules. When the rules are adjusted, the system will update the mapping rules in the knowledge graph according to the new rules to generate an optimized knowledge graph. This optimization process can improve the accuracy of BOM parsing, reduce manual intervention, enable the system to process more complex BOM data, and finally generate an optimized knowledge graph to support more efficient BOM parsing tasks.

[0066] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not indicate the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0067] In one embodiment, a BOM parsing system based on multi-modal semantic fusion is provided. The BOM parsing system based on multi-modal semantic fusion corresponds one-to-one with the BOM parsing method based on multi-modal semantic fusion in the above embodiments. As Figure 7 shown, the BOM parsing system based on multi-modal semantic fusion includes a data receiving module, a text data processing module, an image data processing module, a semantic relationship analysis module, a reinforcement learning optimization module, and a data standardization module. The detailed description of each functional module is as follows: A data receiving module for receiving BOM data, where the BOM data includes text input data and image input data; A text data processing module for extracting semantic features from the text input data using a BERT model to generate a semantic embedding vector of the text; An image data processing module for identifying row-column structures of the image input data using a Table Transformer model, thereby obtaining a corresponding table image, and extracting table content information from the table image; A semantic relationship analysis module for analyzing the relationship between materials and processes based on the semantic embedding vector and the table content information, and constructing a knowledge graph of the materials and the processes using a graph neural network; A reinforcement learning optimization module for dynamically optimizing the mapping rules in the knowledge graph through a reinforcement learning algorithm, adjusting the corresponding field mapping rules, and generating an optimized knowledge graph; A data standardization module for performing standardized conversion on the corresponding BOM data based on the optimized knowledge graph and outputting BOM structured data.

[0068] Optionally, the text data processing module specifically includes: A text preprocessing sub-module for adaptively preprocessing the text input data to obtain standard text data and inputting the standard text data into the BERT model; A context information extraction sub-module for extracting context information from the standard text data through a multi-layer Transformer structure of the BERT model; A semantic feature extraction sub-module for identifying lexical relationships and semantic associations in the text through the self-attention mechanism in the BERT model, extracting semantic features of each field, and generating the semantic embedding vector according to the context relationship and the semantic features.

[0069] Optionally, the image data processing module specifically includes: An image preprocessing sub-module for preprocessing the image input data and detecting cross-page tables using a preset DETR model. The loss function of the DETR model is: total loss = λ L1 ×L1 loss + λ giou ×GIOU loss + λ ce ×CE loss, where λ L1 、λ giou 、λ ce are loss weights, and then obtaining standard image data; The row-column structure recognition sub-module is used to input the standard image data into the Table Transformer model, recognize the row-column structure in the image through the self-attention mechanism, and calibrate the positions of table cells; The table content extraction sub-module is used to extract the content information of the table cell positions based on the row-column structure, and perform context semantic coherence analysis on the content information through bidirectional LSTM, so as to obtain the table content information.

[0070] Optionally, the semantic relationship analysis module specifically includes: The semantic relationship modeling sub-module is used to represent the materials and the processes in the text according to the semantic embedding vectors, and generate the preliminary semantic relationship between the materials and the processes; The information fusion sub-module is used to perform information fusion and relationship reasoning on the materials and the processes according to the table content information, and generate the structural information between the materials and the processes; The graph neural network construction sub-module is used to input the preliminary semantic relationship and the structural information into the graph neural network, perform node representation on the materials and the processes, then model the relationships between the nodes, and generate the knowledge graph of materials and processes through the feature propagation of adjacent nodes and the update of the graph spectrum.

[0071] Optionally, the construction of the reinforcement learning algorithm specifically includes: The working state definition module is used to obtain and define the working state of reinforcement learning based on the current field mapping rules, the historical parsing accuracy trend, and the frequency of manual intervention, and then define the actions of reinforcement learning based on the working state. The actions include adjusting the field mapping weights, adding or deleting field mapping rules, and modifying the regular expression parameters; The reward function design module is used to define the reward function based on the historical parsing accuracy trend and the frequency of manual intervention. The reward function is R = α × Accuracy + β × (1 / Human_Intervention_Count), where Accuracy is the historical parsing accuracy, Human_Intervention_Count is the number of times of manual intervention, and α, β are loss weights; The rule optimization module is used to optimize the current field mapping rules based on the PPO algorithm, and expand and update the current field mapping rules by using the entropy regularization mechanism, so as to generate the reinforcement learning algorithm.

[0072] Optionally, the reinforcement learning optimization module specifically includes: A mapping rule input module, configured to take the mapping rules in the knowledge graph as input and execute a reinforcement learning algorithm; A rule adjustment module, configured to calculate and adjust the corresponding mapping rules based on the reward function, and then update the mapping rules in the knowledge graph according to the adjusted mapping rules to generate the optimized knowledge graph.

[0073] For the specific limitations of the BOM parsing system based on multimodal semantic fusion, reference can be made to the limitations of the BOM parsing method based on multimodal semantic fusion in the foregoing text, which will not be elaborated here. Each module in the above-mentioned BOM parsing system based on multimodal semantic fusion can be implemented in whole or in part by software, hardware, and their combination. The above-mentioned modules can be embedded in the processor in the computer device in hardware form or independent of the processor, or stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.

[0074] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as Figure 8 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a BOM parsing method based on multimodal semantic fusion.

[0075] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: Receiving BOM data, where the BOM data includes text input data and image input data; Using the BERT model to extract semantic features from the text input data to generate a semantic embedding vector of the text; Using the Table Transformer model to identify the row-column structure of the image input data, and then obtaining the corresponding table image, and extracting table content information from the table image; Analyzing the relationship between materials and processes based on the semantic embedding vector and the table content information, and using a graph neural network to construct a knowledge graph of materials and processes; Dynamically optimize the mapping rules in the knowledge graph through a reinforcement learning algorithm, adjust the corresponding field mapping rules, and generate an optimized knowledge graph; Based on the optimized knowledge graph, perform standard conversion on the corresponding BOM data and output BOM structured data.

[0076] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: Receive BOM data, where the BOM data includes text input data and image input data; Use the BERT model to extract semantic features from the text input data and generate semantic embedding vectors of the text; Use the Table Transformer model to identify the row-column structure of the image input data, then obtain the corresponding table image, and extract table content information from the table image; Based on the semantic embedding vectors and the table content information, analyze the relationship between materials and processes, and use a graph neural network to construct a knowledge graph of materials and processes; Dynamically optimize the mapping rules in the knowledge graph through a reinforcement learning algorithm, adjust the corresponding field mapping rules, and generate an optimized knowledge graph; Based on the optimized knowledge graph, perform standard conversion on the corresponding BOM data and output BOM structured data.

[0077] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0078] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.

[0079] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A BOM parsing method based on multimodal semantic fusion, characterized in that: The BOM parsing method based on multimodal semantic fusion includes: Receiving BOM data, wherein the BOM data includes text input data and image input data; Using the BERT model to extract semantic features from the text input data to generate a semantic embedding vector of the text; Using a Table Transformer model to identify the row and column structure of the image input data, thereby obtaining a corresponding table image, and extracting table content information from the table image; Based on the semantic embedding vector and the table content information, the relationship between the material and the process is analyzed, and a knowledge graph of the material and the process is constructed using a graph neural network; Dynamically optimizing the mapping rules in the knowledge graph through a reinforcement learning algorithm, adjusting the corresponding field mapping rules, and generating an optimized knowledge graph; Based on the optimized knowledge graph, the corresponding BOM data is standardized and converted to output BOM structured data.

2. The BOM parsing method based on multimodal semantic fusion according to claim 1 is characterized in that: The BERT model is used to extract semantic features from the text input data to generate a semantic embedding vector of the text, specifically including: Adaptively preprocessing the text input data to obtain standard text data, and inputting the standard text data into the BERT model; Extracting context information from the standard text data through the multi-layer Transformer structure of the BERT model; The self-attention mechanism in the BERT model is used to identify lexical relationships and semantic associations in the text, extract semantic features of each field, and then generate the semantic embedding vector based on the contextual relationship and the semantic features.

3. The BOM parsing method based on multimodal semantic fusion according to claim 1 is characterized in that: The Table Transformer model is used to identify the row and column structure of the image input data, thereby obtaining a corresponding table image, and extracting table content information from the table image, specifically including: The image input data is preprocessed, and a preset DETR model is used to detect the cross-page table. The loss function of the DETR model is: total loss = λ L1 ×L1 loss+λ giou ×GIOU loss + λ ce ×CE loss, where λ L1 , giou , ce is the loss weight, and then the standard image data is obtained; Input the standard image data into the Table Transformer model, identify the row and column structure in the image through the self-attention mechanism, and mark the table cell positions; Based on the row and column structure, the content information of the table cell position is extracted, and the contextual semantic coherence analysis of the content information is performed through a bidirectional LSTM to obtain the table content information.

4. The BOM parsing method based on multimodal semantic fusion according to claim 1 is characterized in that: The analyzing the relationship between the material and the process based on the semantic embedding vector and the table content information, and constructing the knowledge graph of the material and the process using a graph neural network, specifically includes: Representing the material and the process in the text according to the semantic embedding vector, and generating a preliminary semantic relationship between the material and the process; According to the table content information, information fusion and relationship reasoning are performed on the material and the process to generate structural information between the material and the process; The preliminary semantic relationship and the structural information are input into the graph neural network, the materials and the processes are represented by nodes, and then the relationship between the nodes is modeled, and the knowledge graph of materials and processes is generated through feature propagation and graph updating of adjacent nodes.

5. The BOM parsing method based on multimodal semantic fusion according to claim 1 is characterized in that: The construction of the reinforcement learning algorithm specifically includes: Obtain and define the working state of reinforcement learning based on the current field mapping rules, historical parsing accuracy trends, and manual intervention frequency, and then define the actions of reinforcement learning based on the working state, the actions including adjusting the field mapping weight, adding or deleting the field mapping rules, and modifying the regular expression parameters; Based on the historical analysis accuracy trend and the human intervention frequency, a reward function is defined, the reward function is R=α×Accuracy+β×(1 / Human_Intervention_Count), wherein Accuracy is the historical analysis accuracy, Human_Intervention_Count is the number of human interventions, and α and β are loss weights; The current field mapping rule is optimized based on the PPO algorithm, and the current field mapping rule is expanded and updated using an entropy regularization mechanism, thereby generating the reinforcement learning algorithm.

6. The BOM parsing method based on multimodal semantic fusion according to claim 5 is characterized in that: The dynamically optimizing the mapping rules in the knowledge graph by using a reinforcement learning algorithm, adjusting the corresponding field mapping rules, and generating an optimized knowledge graph specifically includes: Taking the mapping rules in the knowledge graph as input, executing a reinforcement learning algorithm; Based on the reward function, the corresponding mapping rules are calculated and adjusted, and then according to the adjusted mapping rules, the mapping rules in the knowledge graph are updated to generate the optimized knowledge graph.

7. A BOM parsing system based on multimodal semantic fusion, characterized in that: The BOM parsing system based on multimodal semantic fusion includes: A data receiving module, used for receiving BOM data, wherein the BOM data includes text input data and image input data; A text data processing module, used to extract semantic features from the text input data using a BERT model to generate a semantic embedding vector for the text; An image data processing module, used for using a Table Transformer model to perform row and column structure recognition on the image input data, thereby obtaining a corresponding table image, and extracting table content information from the table image; A semantic relationship analysis module, used to analyze the relationship between materials and processes based on the semantic embedding vector and the table content information, and to construct a knowledge graph of the materials and processes using a graph neural network; A reinforcement learning optimization module, used to dynamically optimize the mapping rules in the knowledge graph through a reinforcement learning algorithm, adjust the corresponding field mapping rules, and generate an optimized knowledge graph; The data standardization module is used to standardize the corresponding BOM data based on the optimized knowledge graph and output BOM structured data.

8. The BOM parsing system based on multimodal semantic fusion according to claim 7 is characterized in that: The text data processing module specifically includes: A text preprocessing submodule, used for adaptively preprocessing the text input data to obtain standard text data, and inputting the standard text data into the BERT model; A context information extraction submodule, used to extract context information from the standard text data through the multi-layer Transformer structure of the BERT model; The semantic feature extraction submodule is used to identify the lexical relationship and semantic association in the text through the self-attention mechanism in the BERT model, extract the semantic features of each field, and then generate the semantic embedding vector according to the contextual relationship and the semantic features.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the BOM parsing method based on multimodal semantic fusion as described in any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the BOM parsing method based on multimodal semantic fusion as claimed in any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Data multi-protocol adaptive analysis method and system and storage medium

    CN120455570A

  • Production BOM file analysis method for satellite batch production line

    CN121328522A

  • XBOM data processing method, device and system based on ontology modeling driving

    CN121329290A

  • Equipment manufacturing configuration BOM management system and method based on dynamic constraint knowledge graph

    CN121458194A

  • Electric power material supply chain monitoring method and device based on WBS-BOM dynamic mapping

    CN121981694A