Multi-agent PCB design document automatic extraction and table completion method
Through the combination of the YOLOv11 model and multi-agent architecture, the problems of missing table lines and noise interference in PCB design documents are solved, and efficient and accurate table data extraction and completion are achieved, ensuring the integrity and consistency of circuit design information.
Patent Information
- Application Number
- CN202510396123.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-08
AI Technical Summary
It is difficult for the prior art to accurately extract table contents in PCB design documents, especially when the table lines are missing or affected by noise, resulting in inaccurate extraction of circuit parameter information.
The YOLOv11 model is used to combine the multi-agent architecture, and the missing table lines are completed through pixel statistics and interpolation algorithms, and the multi-agent collaboration is used to extract and synthesize text and tabular data to ensure efficient and accurate extraction of document information.
It realizes efficient and accurate extraction and completion of table data in PCB design documents in complex environments, improves the accuracy and efficiency of document processing, and ensures the integrity and consistency of circuit design information.
Smart Images

Figure CN120279575A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of document automation processing, and particularly to a method for automatically extracting tabular data from PCB design documents and completing missing table lines. Background Art
[0002] Currently, optical character recognition (OCR) technology and multi-modal models have been widely used in document information extraction. However, when the document contains tables, especially when the table lines are incomplete or the document is affected by noise, it is often difficult for the existing technologies to accurately extract the table content. In PCB design documents, tables usually contain key circuit parameter information, and the extraction accuracy is crucial for circuit design. Therefore, how to solve the problems of missing table lines and inaccurate data extraction has become a major technical challenge in PCB design document processing. Summary of the Invention
[0003] The present invention proposes a method for automatically extracting and table completing in PCB design documents based on the YOLOv11 model and multi-agent architecture, aiming to solve the problems of missing table lines, noise interference, and low recognition accuracy existing in the existing optical character recognition (OCR) technology during the table extraction process. This method innovatively combines the deep learning model YOLOv11 with image processing technology, completes the missing table lines through pixel statistical analysis, and at the same time uses the multi-agent architecture to effectively divide labor and cooperate to extract and synthesize text and tabular data, thus ensuring efficient and accurate document information extraction.
[0004] The core technical solutions of the present invention include the following steps:
[0005] 1. YOLOv11 Model Training and Table Area Extraction
[0006] The present invention uses the YOLOv11 (You Only Look Once) model to automatically identify and extract the table area in PCB design documents. YOLOv11 is an object detection model based on convolutional neural network (CNN), which can detect the target position in the image through a single forward propagation. The present invention customizes and trains the YOLOv11 model to enable it to accurately identify and mark the table area in PCB design documents, especially to maintain a high recognition accuracy even when the table lines are missing or incomplete. The training process uses a diverse dataset containing different types of tables and their missing parts, enhancing the generalization ability of the model in complex image environments.
[0007] In this step, the YOLOv11 model first performs image preprocessing on the PCB design document, including denoising, grayscale conversion, binarization, etc., to improve the detection accuracy of the table area. Subsequently, the model locates the table area through a regression framework. Especially in the case where table lines are missing or damaged, the model can accurately identify the boundaries of the table through the object detection box, providing precise area information for the subsequent table completion algorithm.
[0008] 2. Table Line Completion Algorithm
[0009] The present invention proposes a table line completion algorithm based on pixel statistics. This algorithm combines vertical pixel statistics of the image, convolution operations, and binarization processing techniques, and can efficiently complete the missing table lines. After extracting the table area, the image will first undergo binarization processing to separate the table and background areas, ensuring clear boundaries of the lines. Then, through convolution operations on the processed image, local image features are used to further enhance the recognizability of the table lines, especially for the missing parts of the table lines to be supplemented.
[0010] In order to accurately restore the missing table lines, the present invention analyzes the pixel density of each column in the image through the vertical pixel statistics method, and determines the most frequent pixel value as the reference line through statistical means. After determining the reference line, combined with the column width information of the context, an interpolation method is used to restore the missing table lines. This algorithm effectively identifies the positions of the table lines through statistical analysis and ensures the consistency of the completed table lines with the existing data, thereby improving the accuracy of table data extraction.
[0011] 3. Application of Multi-Agent Architecture
[0012] In order to improve the efficiency and accuracy of document processing, the present invention designs a multi-agent collaboration architecture, in which different agents are responsible for different tasks and cooperate to complete the automatic extraction and completion of the entire document. The core idea of this architecture is to improve the computing efficiency through parallel processing and ensure the comprehensiveness and accuracy of data processing.
[0013] Specifically, the multi-agent architecture of the present invention includes three main agents:
[0014] Text Extraction Agent: This agent is responsible for extracting complete text information from the PCB design document, including text, titles, annotations, etc. Through deep learning text detection of the document, this agent can effectively extract key information from complex documents.
[0015] Table Data Extraction Agent: This agent focuses on extracting detailed data from the extracted table areas. By splitting the table into rows and columns and identifying the content, this agent converts the extracted table data into Markdown format for easy subsequent display and processing. This agent can accurately identify each cell and its content in the table. Especially in the case of missing rows or columns in the table, the data integrity is ensured through a filling algorithm.
[0016] Data Merging Agent: The task of this agent is to integrate the information from the Text Extraction Agent and the Table Data Extraction Agent, and ensure the logical consistency of the two parts of information through cross-checking. Finally, this agent outputs a complete Markdown format document, which includes the processed text and table data, and ensures its clear structure and correct format.
[0017] 4. Document Output and Synthesis
[0018] After completing the table data extraction and text extraction, the system generates the final document output through a multi-agent architecture. This output uses the standard Markdown format, which is not only convenient for subsequent automated processing but also ensures the clear presentation of the document content. The finally generated document includes complete text information and accurate table data, providing efficient and accurate design document data support for PCB design engineers.
[0019] In addition, the multi-agent architecture of the present invention can ensure the integrity and accuracy of the document through cross-checking and feedback mechanisms. Each agent will feedback its current working status to other agents after completing its own task for necessary adjustment and optimization. In this way, the system can efficiently complete complex document extraction and filling tasks.
[0020] The present invention has the following beneficial effects:
[0021] The multi-agent architecture of the present invention can ensure the integrity and accuracy of the document through cross-checking and feedback mechanisms. Each agent will feedback its current working status to other agents after completing its own task for necessary adjustment and optimization. In this way, the system can efficiently complete complex document extraction and filling tasks.
[0022] 1. Based on the combination of the YOLOv11 model and the multi-agent architecture, the present invention can efficiently and accurately extract the table areas in PCB design documents, automatically complete the missing table lines, and is not affected by the complexity of the table areas, effectively improving the accuracy and efficiency of document processing;
[0023] 2. The present invention adopts pixel statistics and interpolation algorithms during the process of table line completion. By analyzing the pixel density of each column in the image, it reduces complex operations with high computational complexity, improves the computational speed of the completion algorithm, and optimizes the performance of the model when processing large-scale documents.
[0024] 3. The present invention uses a multi-agent architecture to process document content in parallel, separately extracts text information and table data, and performs intelligent merging to ensure the consistency and integrity of document content, providing an efficient solution for the automated processing of PCB design documents and a theoretical basis for further automated detection and analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a schematic diagram of table area extraction of YOLOv11 of the present invention.
[0026] Figure 2 It is a flowchart of the table line completion algorithm of the present invention.
[0027] Figure 3 It is the process of table line completion of the present invention
[0028] Figure 4 It is the PR curve and F1 score curve of YOLOv11 of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0029] The method for automatic extraction and table completion of PCB design documents based on the YOLOv11 model and multi-agent architecture proposed by the present invention aims to improve the automated processing ability of PCB design documents through efficient table area extraction, missing table line completion, and multi-agent collaboration. The following details the specific implementation methods of each step, including the adopted algorithms, data processing flow, and optimization strategies.
[0030] Refer to Figure 1 , Figure 2 As shown, first, binarization and convolution operations are performed to make the boundary between the table lines and the table content clear. Then, pixel point statistics in the vertical direction are carried out to obtain a clear dividing line, and the table lines are completed accordingly.
[0031] Figure 3 It is the actual operation process of the above algorithm. First, the picture is converted into a distribution map of vertical pixels, then binarized, followed by convolution, then pixel statistics are carried out to obtain the dividing line, and finally the table lines are completed.
[0032] Figure 4 It is the experiment conducted after training the YOLOv11 model, namely the PR curve and F1 score, used to evaluate the performance of the model.
[0033] 1. Training of YOLOv11 Model and Extraction of Table Areas
[0034] As an efficient object detection algorithm, the YOLOv11 model has been widely applied to various image analysis tasks. To enable YOLOv11 to adapt to the special requirements of PCB design documents, this invention conducts customized training on YOLOv11.
[0035] Step 1.1: Dataset Construction and Preprocessing
[0036] To train YOLOv11, an image dataset containing PCB design documents needs to be prepared first. This dataset includes different types of PCB design documents, with a focus on those containing incomplete or missing table lines. The annotation of the dataset adopts the frame annotation method, and each table area needs to be accurately framed.
[0037] The preprocessing of document images includes the following steps:
[0038] Denoising processing: Gaussian blur and median filtering methods are used to remove the noise in the document, ensuring the clarity of the table area in the image.
[0039] Grayscale conversion and binarization: The image is converted to a single channel through grayscale conversion to reduce the computational complexity. Subsequently, the Otsu binarization algorithm is used to process the image to separate the table and background areas in the image.
[0040] Step 1.2: Training of YOLOv11 Model and Extraction of Table Areas
[0041] Using the above preprocessed dataset, the YOLOv11 model is trained to recognize and extract the table areas in the document. The goal of the model is to locate the table areas in the document, especially the parts with missing table lines. During the training process, YOLOv11 outputs the position of each table area through a regression framework, marking the start and end positions of the table.
[0042] In this step, the output of the YOLOv11 model includes the coordinate box (x_min, y_min, x_max, y_max) of each table area and its confidence level. During the processing, YOLOv11 can quickly and accurately recognize the table areas. Even in the case of missing table lines, it can still effectively recognize the overall framework of the table.
[0043] 2. Table Line Completion Algorithm
[0044] After the extraction of the table areas, the next task of this invention is to complete the missing table lines. To achieve this goal, this invention designs a table line completion algorithm based on pixel statistics and convolutional operations.
[0045] Step 2.1: Extraction of Table Lines from Binary Images
[0046] After binarizing the image of the table area, a black-and-white image of the table is obtained, where the black part represents the table lines. By performing a convolution operation on the binary image, the saliency of the table lines can be enhanced and the noise in the image can be removed.
[0047] Step 2.2: Vertical Pixel Statistics and Determination of Reference Lines
[0048] To complete the missing table lines, the present invention adopts a vertical pixel statistics method. Specifically, for the pixel density of each column, the following formula is used to calculate the pixel statistic of each column:
[0049]
[0050] where P(i) is the pixel density of the i-th column, h is the height of the image, and I(i,j) is the pixel value of the image at the j-th row of the i-th column. If the pixel value is black (table line), the value is 1; otherwise, it is 0.
[0051] By statistically analyzing the density of pixels in each column, the algorithm identifies the areas with higher density and uses the pixel values of these areas to determine the reference lines. These reference lines are used to determine the final positions of the missing table lines.
[0052] Step 2.3: Completion of Table Lines
[0053] Based on the results determined by the reference lines, an interpolation method is used to complete the missing table lines. Specifically, based on the pixel distribution of the front and rear columns, the table lines of the missing columns are interpolated and completed to ensure that each column of the table is consistent in position.
[0054] For each missing table line, the following formula is used for completion:
[0055] L 补全 =L 前列
[0056] where L 补全 is the completed table line, and L 前列 is the initial position of the non-text area determined according to the reference line.
[0057] 3. Implementation of the Multi-Agent Architecture
[0058] To improve the efficiency and accuracy of document information extraction, the present invention designs a multi-agent architecture that can effectively process different parts of the document in parallel. Specifically, three agents are responsible for text extraction, table data extraction, and data merging respectively.
[0059] Step 3.1: Text Extraction Agent
[0060] The text extraction agent extracts the text content in the document through OCR technology. A deep learning text detection model is used, which can identify the title, body text, annotations, etc. in the document and accurately extract all the text in the document.
[0061] Step 3.2: Table data extraction agent
[0062] The table data extraction agent receives the table area information from the YOLOv11 model and further processes it through a table segmentation algorithm to extract the content of each table cell. For each cell, the agent extracts its content and converts it into a structured Markdown format for subsequent processing and display.
[0063] Step 3.3: Data merging agent
[0064] The data merging agent is responsible for cross-checking and merging the information extracted by the text extraction agent and the table data extraction agent. Finally, the merged document is converted into Markdown format, which contains the complete text content and table data. This agent is also responsible for ensuring the logical consistency between the text and the table, avoiding data duplication or omission.
[0065] 4. Document output and synthesis
[0066] After being processed by the multi-agent architecture, the final document will be output in Markdown format. Markdown format has good structural characteristics, which is convenient for subsequent display and processing. The final document includes the complete text content and the extracted table data, and ensures that its format and structure are clear.
[0067] The process of this document output is as follows:
[0068] (1) The text extraction agent extracts the text content;
[0069] (2) The table data extraction agent extracts the table data and converts it into Markdown format;
[0070] (3) The data merging agent integrates the text and table data into a complete Markdown format document.
[0071] Each table in the document will be presented in its original format, and it is ensured that the content of each row and each column can be accurately displayed, providing the user with a high-quality automated document.
Claims
1. A method for automatically extracting and table complementing of multi-agent PCB design documents, characterized in that: Automatically identify and extract the table areas in PCB design documents by training the YOLOv11 model. Based on the table area coordinate frames output by the YOLOv11 model, combined with the pixel statistical algorithm, complete the missing table lines; First, the YOLOv11 model accurately extracts the table areas in the document through image preprocessing and deep learning models, and calibrates the boundaries of the tables; Then, use the pixel statistical method to analyze the pixel density of each column, determine the reference position of the table lines by convolution operation, and complete the missing table lines through the interpolation algorithm. Adopt a multi-agent architecture, and complete the extraction and integration of document content through a text extraction agent, a table data extraction agent, and a data merging agent respectively; Finally, the text and table data are converted into a structured Markdown format for output.
2. The multi-agent PCB design document automatic extraction and table completion method according to claim 1, characterized in that: Training of the YOLOv11 model and automatic identification and extraction of table areas: The YOLOv11 (You Only Look Once) model is used for automatic identification and extraction of table areas in PCB design documents. The YOLOv11 is an object detection model based on the convolutional neural network CNN, which can detect the target positions in images through a single forward propagation. By customizing the training of the YOLOv11 model, it can accurately identify and mark the table areas in PCB design documents. Especially in the case of missing or incomplete table lines, it can still maintain a high recognition accuracy. The training process uses a diverse dataset containing different types of tables and their missing parts, enhancing the generalization ability of the model in complex image environments.
3. The multi-agent PCB design document automatic extraction and table completion method according to claim 2, characterized in that: Customized training refers to training the open-source YOLOv11x model on the basis of building one's own dataset to achieve the effect of accurate recognition.
4. The method for automatically extracting and table complementing of the multi-agent PCB design document according to claim 2, characterized in that: The YOLOv11 model first performs image preprocessing on the PCB design document, including denoising, grayscale conversion, and binarization, to improve the detection accuracy of the table areas. Subsequently, the model locates the table areas through a regression framework. Especially in the case of missing or damaged table lines, it accurately identifies the boundaries of the tables through the target detection boxes, providing accurate area information for the subsequent table completion algorithm. The document image preprocessing includes the following steps: Denoising processing: Use Gaussian blur and median filtering methods to remove the noise in the document, ensuring the clarity of the table areas in the image; Grayscale conversion and binarization: Convert the image into a single channel through grayscale conversion to reduce the computational complexity. Subsequently, use the Otsu binarization algorithm to process the image and separate the table and background areas in the image.
5. The method for automatically extracting and table complementing of multi-agent PCB design documents according to claim 1, wherein: The table line completion algorithm, combined with vertical pixel statistics, convolution operation, and binarization processing techniques of images, can efficiently complete the missing table lines. After extracting the table area, the image will first undergo binarization processing to separate the table and background areas. Then, a 3*3 convolution operation is performed on the processed image to further enhance the recognizability of the table lines using local image features, especially for supplementing the missing table line parts. After that, vertical pixel statistics are carried out to distinguish between the table lines and the data within the table.
6. The method for automatically extracting and table completing multi-agent PCB design documents according to claim 5, characterized in that: To accurately restore the missing table lines, the pixel density of each column in the image is analyzed through the vertical pixel statistics method, and the most frequent pixel value is determined as the reference line through statistics. After the reference line is determined, combined with the column width information of the context, the interpolation method is used to restore the missing table lines, effectively identifying the table line positions through statistical analysis and ensuring the consistency between the completed table lines and the existing data.
7. The method for automatically extracting and table complementing of the multi-agent PCB design document according to claim 1, characterized in that: The multi-agent architecture, where different agents undertake different tasks and cooperate to complete the automatic extraction and completion of the entire document.
8. The method for automatically extracting and table completing a multi-agent PCB design document according to claim 7, characterized in that: The multi-agent architecture includes three agents: Text extraction agent: This agent is responsible for extracting complete text information from the PCB design document, including title, electrical information, and package information content. By performing deep learning text detection on the document, it extracts key information from the complex document. Table data extraction agent: This agent focuses on extracting detailed data from the extracted table area. By performing row and column segmentation and content recognition on the table, this agent converts the extracted table data into Markdown format for subsequent display and processing. Data merging agent: The task of this agent is to integrate the information from the text extraction agent and the table data extraction agent, and ensure the logical consistency of the two parts of information through cross-checking. Finally, this agent outputs a complete Markdown format document, which includes processed text and table data, and ensures its clear structure and correct format.
9. The method for automatically extracting and table completing a multi-agent PCB design document according to claim 1, characterized in that: Document output and synthesis After completing the table data extraction and text extraction, the system will generate the final document output through the multi-agent architecture. The output adopts the standard Markdown format, which is not only convenient for subsequent automated processing but also ensures the clear presentation of the document content. The finally generated document includes complete text information and accurate table data, providing efficient and accurate design document data support for PCB design engineers.
Citation Information
Cited By
Incomplete data trend question and answer method based on multi-agent collaboration
CN121958360A
A Question Answering Method for Incomplete Data Based on Multi-Agent Collaboration
CN121958360B