BOM recognition method and device for IA based on RPA and AI

By combining RPA and AI, the BOM area in the design drawings is automatically identified, solving the problems of low efficiency and high error rate of manual identification, and achieving efficient and accurate BOM recognition.

CN114937279BActive Publication Date: 2025-10-10BEIJING BENYING NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210636992.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-07
Publication Date
2025-10-10
Estimated Expiration
2042-06-07

AI Technical Summary

Technical Problem

In the existing technology, when identifying bills of materials (BOMs) from design drawings, manual identification is time-consuming and error-prone, and general OCR systems have difficulty distinguishing table areas, resulting in low automated processing efficiency.

Method used

An intelligent automation method based on RPA and AI is adopted to obtain the target image, identify and determine the BOM area, use OCR and image segmentation network for table recognition, and combine natural language processing technology to automatically identify and extract BOM content.

Benefits of technology

It realizes intelligent and automatic positioning and identification of BOM from design drawings, improves identification efficiency and accuracy, and replaces manual operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114937279B_ABST
    Figure CN114937279B_ABST
Patent Text Reader

Abstract

The present disclosure provides a BOM recognition method and device based on RPA and AI to realize IA, wherein the method comprises: acquiring a target image; determining at least one target BOM region from the target image; for any target BOM region in the at least one target BOM region, performing table recognition on the target BOM region to obtain at least one candidate BOM; and determining a target BOM from the at least one candidate BOM. Thus, the BOM can be positioned and recognized from the target image intelligently and automatically, replacing manual operation and improving the BOM content recognition efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of artificial intelligence (AI), robotic process automation (RPA), and intelligent automation (IA), and in particular to a BOM recognition method and device for implementing IA based on RPA and AI. Background Art

[0002] Robotic Process Automation (RPA) uses specific "robot software" to simulate human operations on computers and automatically execute process tasks according to rules.

[0003] Artificial Intelligence (AI) is a technical science that studies and develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence.

[0004] Intelligent Automation (IA) is a general term for a range of technologies from robotic process automation to artificial intelligence. It combines RPA with multiple AI technologies, including optical character recognition (OCR), intelligent character recognition (ICR), process mining, deep learning (DL), machine learning (ML), natural language processing (NLP), automatic speech recognition (ASR), text to speech (TTS), and computer vision (CV), to create end-to-end business processes that can think, learn, and adapt. This covers the entire process from process discovery and automation to automatic and continuous data collection, understanding the meaning of data, and using data to manage and optimize business processes.

[0005] Currently, downstream manufacturers need to identify Bill of Materials (BOMs) from drawings to meet their needs for extracting and entering them into production and procurement systems, as well as for comparing and verifying multiple design versions. Related technologies typically identify BOMs manually and enter them into spreadsheets. However, manual BOM identification is time-consuming and error-prone. Summary of the Invention

[0006] The present disclosure aims to solve one of the technical problems in the above-mentioned technologies at least to some extent.

[0007] To this end, the present disclosure proposes a BOM recognition method and device based on RPA and AI to implement IA, so as to intelligently and automatically locate and identify the BOM from the target image, replacing manual operations and improving the efficiency and accuracy of BOM recognition.

[0008] An embodiment of the first aspect of the present disclosure proposes a bill of materials (BOM) recognition method for implementing intelligent automation (IA) based on robotic process automation (RPA) and artificial intelligence (AI), including: acquiring a target image; determining at least one target BOM area from the target image; performing table recognition on any target BOM area in the at least one target BOM area to obtain at least one candidate BOM; and determining a target BOM from the at least one candidate BOM.

[0009] An embodiment of the second aspect of the present disclosure proposes a bill of materials (BOM) recognition device for realizing intelligent automation (IA) based on robotic process automation (RPA) and artificial intelligence (AI), including: an acquisition module for acquiring a target image; a first determination module for determining at least one target BOM area from the target image; an identification module for performing table recognition on any target BOM area in the at least one target BOM area to obtain at least one candidate BOM; and a second determination module for determining a target BOM from at least one candidate BOM.

[0010] The third aspect embodiment of the present disclosure proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in the first aspect embodiment of the present disclosure is implemented.

[0011] The fourth embodiment of the present disclosure provides a non-temporary computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method described in the first embodiment of the present disclosure is implemented.

[0012] A fifth aspect embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the method described in the first aspect embodiment of the present disclosure.

[0013] The technical solution provided by the embodiments of the present disclosure has the following beneficial effects:

[0014] The technical solution disclosed herein acquires a target image; determines at least one target BOM area from the target image; performs table recognition on any target BOM area in the at least one target BOM area to obtain at least one candidate BOM; and determines the target BOM from the at least one candidate BOM. Thus, table recognition is performed on the target BOM area in the target image, and the target BOM is determined from the recognized candidate BOMs. This can realize intelligent and automatic positioning and identification of the BOM from the target image, replace manual operations, and improve the efficiency and accuracy of BOM content recognition.

[0015] The above summary is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments and features described above, further aspects, embodiments and features of the present application will be readily apparent by reference to the accompanying drawings and the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments disclosed in this application and should not be construed as limiting the scope of this application.

[0017] Figure 1 A schematic diagram of the distribution of the BOM area in a drawing of an embodiment of the present disclosure;

[0018] Figure 2 This is a flowchart of a BOM identification method for implementing IA based on RPA and AI according to an embodiment of the present disclosure;

[0019] Figure 3 This is a flowchart of a BOM identification method for implementing IA based on RPA and AI according to another embodiment of the present disclosure;

[0020] Figure 4 is a schematic diagram of pixels of each horizontal line segment in a target image according to an embodiment of the present disclosure;

[0021] Figure 5 is a schematic diagram of pixels of each vertical line segment in a target image according to an embodiment of the present disclosure;

[0022] Figure 6is a schematic diagram of a set of horizontal line segments in a region according to an embodiment of the present disclosure;

[0023] Figure 7 is a schematic diagram of a set of vertical line segments in a region according to an embodiment of the present disclosure;

[0024] Figure 8 is a reference table diagram of a region according to an embodiment of the present disclosure;

[0025] Figure 9 is a schematic diagram of a connected region with the largest area corresponding to a reference table according to an embodiment of the present disclosure;

[0026] Figure 10 This is a schematic diagram of a process for determining a target BOM area according to an embodiment of the present disclosure;

[0027] Figure 11 This is a flowchart of a BOM identification method for implementing IA based on RPA and AI according to another embodiment of the present disclosure;

[0028] Figure 12 This is a flowchart of a BOM identification method for implementing IA based on RPA and AI according to another embodiment of the present disclosure;

[0029] Figure 13 This is a flowchart of a BOM identification method for implementing IA based on RPA and AI according to another embodiment of the present disclosure;

[0030] Figure 14 This is a schematic diagram of the structure of a BOM recognition device for implementing IA based on RPA and AI according to an embodiment of the present disclosure;

[0031] Figure 15 The present invention is a block diagram of an electronic device for implementing a bill of materials (BOM) identification method for IA based on RPA and AI, according to an exemplary embodiment. DETAILED DESCRIPTION

[0032] Embodiments of the present application / disclosure are described in detail below. Examples of the embodiments are illustrated in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are illustrative and are intended only to explain the present application / disclosure and are not to be construed as limiting the present application / disclosure.

[0033] Computer-aided design (CAD) technology is now widely used in the manufacturing industry. During the design process, a bill of materials (BOM) is typically generated by a CAD system, automatically traversing and capturing product structure information from design drawings. This information is then presented in a tabular format on the drawing. The design BOM generated by design companies using CAD systems serves as a fundamental data source for manufacturers in their manufacturing, management, and procurement activities.

[0034] In order to protect their intellectual property rights from infringement, design companies convert design CAD files into image formats and provide them to the outside world. In addition, according to existing industry standards in many industries, drawings with formal legal effect must be stamped or signed. After printing, stamping, signing, and re-scanning, the drawing files can only be in image format rather than the original CAD file format. Therefore, in many cases, when the production and procurement departments of manufacturing companies receive design CAD drawings in electronic file format, they are no longer drawings in the original design format (such as dwg), but only files in image format (such as jpg, png, etc., or pure image format pdf files). In order to meet the needs of downstream manufacturing companies to extract the contents of the Bill of Materials (BOM) and enter them into the production and procurement systems, as well as to compare and verify between multiple design versions of BOMs, it is necessary to identify the contents of the BOM from the drawings.

[0035] In the related art, there are two main methods for identifying BOM from design drawing images: the first method is to obtain BOM entirely by manual recognition and enter the recognized BOM into a table; the second method is to try to use the table recognition capabilities provided by various general OCR systems; however, the first method of obtaining BOM entirely by manual recognition and entering the recognized BOM into a table is not only time-consuming but also prone to errors; since the content on a drawing is very rich, the table recognition of the general OCR system is difficult to distinguish between the table and other content in the design drawing, and the BOM area and other table areas. In addition, Figure 1 As shown, as the number of parts in a design increases, the BOM area is split into several sections, appearing in a specific order above and to the left of the legend in the lower right corner of the drawing, often completely integrated with the legend table. Manually capturing each BOM section and feeding it into the OCR system for table recognition in multiple steps is one possible approach. However, this step still involves significant manual effort, making it inconvenient for scenarios requiring automated processing of large numbers of drawings, such as when building RPA processes.

[0036] To address the above issues, the present disclosure proposes a BOM recognition method and device for implementing IA based on RPA and AI.

[0037] The following describes the BOM recognition method and device for implementing IA based on RPA and AI in accordance with an embodiment of the present disclosure with reference to the accompanying drawings. Before describing the embodiment of the present disclosure in detail, for ease of understanding, the following common technical terms are introduced:

[0038] "Target image": refers to the image obtained by scanning the CAD drawing, which can be in the following formats: JPG, PNG or PDF;

[0039] "Candidate horizontal line segment set": refers to the set of horizontal line segments identified from the target image;

[0040] "Candidate vertical line segment set": refers to the set of vertical line segments identified from the target image;

[0041] "Candidate horizontal line segment subset": refers to the set of horizontal line segments in different regions of the target image;

[0042] "Candidate vertical line segment subset": refers to the set of vertical line segments in different regions of the target image;

[0043] "Reference table": refers to the table obtained by merging the subset of candidate horizontal line segments and the corresponding subset of candidate vertical line segments in the same area;

[0044] "Candidate BOM area": ​​refers to the area corresponding to the circumscribed rectangle of the largest connected area in the reference table;

[0045] "Target BOM area" refers to an area in the candidate BOM area where the number of horizontal lines is greater than or equal to a first set number threshold, and the number of vertical line segments is greater than or equal to a second set number threshold;

[0046] "First element window" refers to setting a horizontal rectangular window. For example, if the width of the target image is w, set the width of the horizontal rectangular window to w / 40 and the height to 1;

[0047] "Second element window" refers to setting a vertical rectangular window. For example, if the height of the target image is h, set the height of the vertical rectangular window to h / 80 and the width to 1;

[0048] "Candidate BOM" refers to the table identified from the target BOM area;

[0049] "Target BOM" refers to the BOM identified from the candidate BOM. For example, if the candidate BOM contains the legend content and the BOM, the target BOM is the candidate BOM with the legend content removed, which only contains the BOM.

[0050] Optical Character Recognition (OCR) refers to the process by which an electronic device examines printed characters on paper, determines their shape by detecting dark and light patterns, and then uses character recognition methods to translate these shapes into computer text. In other words, for printed characters, the text in a paper document is optically converted into a black and white dot matrix image file, and recognition software is used to convert the text in the image into text format for further editing and processing by word processing software.

[0051] These and other aspects of the embodiments of the present application / disclosure will become apparent with reference to the following description and accompanying drawings. In these descriptions and accompanying drawings, some specific implementations of the embodiments of the present application / disclosure are disclosed in detail to illustrate some ways of implementing the principles of the embodiments of the present application / disclosure, but it should be understood that the scope of the embodiments of the present application / disclosure is not limited thereby. On the contrary, the embodiments of the present application / disclosure include all changes, modifications and equivalents that fall within the spirit and connotation of the appended claims.

[0052] Figure 2 This is a flowchart of a BOM recognition method for implementing IA based on RPA and AI according to an embodiment of the present disclosure.

[0053] In a possible implementation method provided by an embodiment of the present disclosure, the present disclosure uses the example of the BOM recognition method for implementing IA based on RPA and AI being configured in a BOM recognition device for implementing IA based on RPA and AI. The BOM recognition device for implementing IA based on RPA and AI can be applied to any electronic device with computing capabilities.

[0054] The electronic device may be a personal computer, a mobile terminal, etc. The mobile terminal may be, for example, a mobile phone, a tablet computer, a personal digital assistant, or other hardware device with various operating systems.

[0055] In another possible implementation of the embodiment of the present disclosure, the BOM recognition device for implementing IA based on RPA and AI can be applied to an RPA robot, wherein the RPA robot can run in any electronic device with computing capabilities.

[0056] like Figure 2 As shown, the method may include the following steps:

[0057] Step 201: Acquire a target image.

[0058] In the embodiment of the present disclosure, the target image may be an image of a design drawing, and the design drawing may be scanned to obtain the target image, wherein the format of the target image may be JPG, PNG, or PDF, etc.

[0059] Step 202: Determine at least one target BOM area from the target image.

[0060] As a possible implementation method of an embodiment of the present disclosure, at least one candidate BOM area can be determined from the target image, and at least one target BOM area can be determined from the at least one candidate BOM area, wherein the candidate BOM area can be the area corresponding to the circumscribed rectangle of the connected area with the largest area in a reference table in the area, and the reference table refers to a table obtained by merging the set of horizontal line segments and the set of vertical line segments in the area.

[0061] Step 203 : performing table recognition on any target BOM area in the at least one target BOM area to obtain at least one candidate BOM.

[0062] As a possible implementation of the embodiment of the present disclosure, table recognition can be performed on any target BOM area in the at least one target BOM area based on OCR to obtain at least one candidate BOM. The table recognition may include recognition of each cell in the table and recognition of the content of each cell.

[0063] Step 204: Determine a target BOM from at least one candidate BOM.

[0064] Since the candidate BOM may include the BOM and its contents other than the BOM, wherein the contents other than the BOM may be, for example, legend contents.

[0065] In the embodiment of the present disclosure, a BOM is identified from at least one candidate BOM and the identified BOM is used as the target BOM content. For example, the legend content in each candidate BOM can be discarded to obtain the corresponding target BOM.

[0066] In summary, by acquiring a target image; determining at least one target BOM area from the target image; performing table recognition on any target BOM area in the at least one target BOM area to obtain at least one candidate BOM; and determining the target BOM from the at least one candidate BOM, thereby automatically performing table recognition on the target BOM area in the target image and determining the target BOM from the recognized candidate BOMs, it is possible to intelligently and automatically locate and identify the BOM from the target image, replacing manual operations and improving the efficiency and accuracy of BOM recognition.

[0067] In order to accurately determine at least one target BOM area from the target image, such as Figure 3 As shown, Figure 3This is a flowchart of a BOM recognition method for implementing IA based on RPA and AI according to another embodiment of the present disclosure. In this embodiment of the present disclosure, candidate BOM areas corresponding to each reference table can be determined based on the circumscribed rectangle corresponding to the connected area with the largest area in the reference table. At least one target BOM area can be determined from the candidate BOM areas. Figure 3 The illustrated embodiment may include the following steps:

[0068] Step 301: Acquire a target image.

[0069] Step 302 : Identify, from the target image, a set of candidate horizontal line segments and a set of candidate vertical line segments corresponding to the target image.

[0070] Optionally, a horizontal rectangular window is set as the first element window, and a first filtering process is performed on the target image to obtain a first image, wherein the first image includes the pixel points of each candidate horizontal line segment in the target image; a vertical rectangular window is set as the second element window, and a second filtering process is performed on the target image to obtain a second image, wherein the second image includes the pixel points of each candidate vertical line segment in the target image; according to the coordinates of each pixel point in the first image, the first circumscribed rectangular coordinates of each candidate horizontal line segment are determined, wherein the first circumscribed rectangular coordinates are used to indicate the endpoint coordinates of each candidate horizontal line segment; according to the coordinates of each pixel in the second image, the second circumscribed rectangular coordinates of each candidate vertical line segment are determined, wherein the second circumscribed rectangular coordinates are used to indicate the endpoint coordinates of each candidate vertical line segment; according to the first circumscribed rectangular coordinates of each candidate horizontal line segment, a set of candidate horizontal line segments are determined, and according to the second circumscribed rectangular coordinates of each candidate vertical line segment, a set of candidate vertical line segments are determined.

[0071] That is to say, in order to facilitate the display of the candidate horizontal line segment set and the candidate vertical line segment set in the target image, the target image can be converted into a single-channel grayscale image and adaptively binarized. A horizontal rectangular window is set as the first element window on the binarized image, and the binarized target image is subjected to morphological filtering (e.g., first erosion and then dilation) to obtain the first image, that is, Figure 4 As shown in , the text pixels in the target image can be filtered out, and only the pixels of the horizontal line segments in the target image are retained in the first image; similarly, a vertical rectangular window is set as the second element window on the binarized image, and the binarized target image is subjected to morphological filtering (first erosion and then expansion) to obtain the second image, that is, as shown in Figure 5 As shown, the text pixels in the target image can be filtered out, and only the pixels of the vertical line segments in the target image can be retained in the second image. Furthermore, the coordinates of the first circumscribed rectangle of each candidate horizontal line segment can be determined based on the coordinates of each pixel in the first image; and the coordinates of the second circumscribed rectangle of each candidate vertical line segment can be determined based on the coordinates of each pixel in the second image.

[0072] Next, as an example, based on the first circumscribed rectangle coordinates of each candidate horizontal line segment, the longest candidate horizontal line segment that matches the boundary horizontal frame segment of the target image can be deleted from each candidate horizontal line segment to determine a set of candidate horizontal line segments, and the set of candidate horizontal line segments can be recorded as allRow; at the same time, based on the second circumscribed rectangle coordinates of each candidate vertical line segment, the longest candidate vertical line segment that matches the boundary vertical frame segment of the target image can be deleted from each candidate vertical line segment to determine a set of candidate vertical line segments, and the set of candidate vertical line segments can be recorded as allCow.

[0073] Among them, taking the width and height of the target image as w and h respectively as an example, the size of the horizontal rectangular window can be set to [w′, 1] (e.g., w′=w / 80), and the size of the vertical rectangular window can be set to [1, h′] (e.g., h′=h / 80). It should be noted that the pixel coordinate system of the target image can be established in advance, with the upper left corner of the target image as the origin of the pixel coordinate system. The coordinates of the first circumscribed rectangle can be {(x1, y1), (x2, y1), (x1, y2), (x2, y2)}, and the coordinates of the second circumscribed rectangle can be {(x3, y3), (x4, y3), (x3, y4), (x4, y4)}.

[0074] Step 303 : From the candidate horizontal line segment set and the candidate vertical line segment set, determine each candidate horizontal line segment subset and each candidate vertical line segment subset in sequence according to a set direction.

[0075] The candidate horizontal line segment subset includes multiple horizontal line segments located in the same area, and the target vertical line segment set includes multiple vertical line segments located in the same area.

[0076] It should be understood that since the area containing the BOM is usually divided into several vertical columns and arranged in the drawing image according to a set direction (for example, from right to left), each candidate horizontal line segment subset and each candidate vertical line segment subset can be determined in sequence from the candidate horizontal line segment set and the candidate vertical line segment set according to the set direction.

[0077] As an example, according to the set direction, the starting point of each horizontal coordinate and the ending point of each vertical coordinate corresponding to the starting point of each horizontal coordinate can be determined in sequence. From the set of candidate horizontal line segments, according to the starting point of each horizontal coordinate, a subset of each candidate horizontal line segment can be determined; from the set of candidate vertical line segments, according to the ending point of each vertical coordinate, a subset of each candidate vertical line segment can be determined.

[0078] For example, the horizontal coordinate starting point of the first region can be determined from the candidate horizontal line segment set according to the set direction. For example, the width and height of the target image are w and h respectively, and the direction is set from right to left. The horizontal line segments whose starting point horizontal coordinate x1 satisfies x1>wth (for example, wth=w / 3) are selected as the horizontal line segment set in the first region (for example, Figure 6Similarly, the vertical coordinate end point of the first region can be determined from the candidate vertical line segment set. For example, the vertical line segments whose end point vertical coordinate y1 satisfies y1>hth (for example, hth=h / 2) are selected as the vertical line segment set in the first region (for example, Figure 7 As shown), the candidate horizontal line segment subset and the candidate vertical line segment subset in the first (lower right corner) area of ​​the target image can be screened out from the candidate horizontal line segment set and the candidate vertical line segment set. Figure 6 and Figure 7 The dotted box in the figure is only used to illustrate the image area and does not exist in actual applications.

[0079] It should be understood that since there are multiple areas containing BOM in the drawing image, each area has a roughly the same width because it is arranged with consistent serial numbers, numbers, materials, quantities and remarks information. Therefore, the starting point of the horizontal coordinate of the next area can be the difference between the starting point of the horizontal coordinate of the first area and the width of the first area (such as the width dw of the first area, the starting point of the horizontal coordinate of the first area is x1, and the starting point of the horizontal coordinate of the next area is x1-dw). The end point of the vertical coordinate of the next area can be the same as the end point of the vertical coordinate of the first area. Then, according to the starting point of the horizontal coordinate of each area and the corresponding end point of the vertical coordinate, each candidate horizontal line segment subset and each candidate vertical line segment subset can be determined. Among them, the width of the area can be determined by the difference between the maximum horizontal coordinate and the minimum horizontal coordinate of each candidate horizontal line segment in the candidate horizontal line segment subset in the first area. As an example, in order to make the width of the area more accurate, the horizontal coordinate difference corresponding to each horizontal line segment in the multiple candidate horizontal line segments in the first area can be added, the addition result can be compared with the number of the multiple candidate horizontal line segments, and the comparison result can be used as the width of the area.

[0080] Step 304 : Merge each candidate horizontal line segment subset with each candidate vertical line segment subset located in the same region of each candidate horizontal line segment subset to obtain at least one reference table.

[0081] Furthermore, if Figure 8 As shown, the candidate horizontal line segment subsets in each region are merged with the candidate vertical line segment subsets in the corresponding region to obtain a reference table in each region. Figure 8 The dotted box in the figure is only used to illustrate the image area and does not exist in actual applications.

[0082] Step 305 : For each reference table in the at least one reference table, determine a candidate BOM area corresponding to each reference table according to a circumscribed rectangle corresponding to a connected region with the largest area in each reference table.

[0083] Furthermore, if Figure 9As shown, for the reference table corresponding to each region, the connected regions in each reference table are determined, and the circumscribed rectangle of the connected region with the largest area among the connected regions is determined (e.g., Figure 9 The black area in the middle), and the bounding rectangle of the largest connected area in each reference table is used as the candidate BOM area corresponding to each reference table. Figure 9 The dotted box in the figure is only used to illustrate the image area and does not exist in actual applications.

[0084] Step 306 : Determine at least one target BOM area according to the number of horizontal line segments and the number of vertical line segments corresponding to each candidate BOM area.

[0085] It should be understood that if the number of horizontal and vertical line segments in the candidate BOM area exceeds a certain threshold, the candidate BOM area is considered to constitute a reasonable BOM area; if the number of horizontal or vertical line segments in the candidate BOM area is too small, it can be considered that a valid BOM cannot be formed, and the candidate BOM area cannot become a reasonable BOM area.

[0086] As a possible implementation method of an embodiment of the present disclosure, the number of horizontal line segments and the number of vertical line segments in any candidate BOM area are obtained; when the number of horizontal line segments in any candidate BOM area is greater than or equal to a first set number threshold, and the number of vertical line segments in any candidate BOM area is greater than or equal to a second set number threshold, any candidate BOM area is taken as the target BOM area.

[0087] It should be noted that, as a possible implementation of the embodiment of the present disclosure, Figure 10 As shown, the candidate horizontal line segment subset and the candidate vertical line segment subset corresponding to the first area can be determined from the candidate horizontal line segment set according to the set direction, and the candidate horizontal line segment subset and the candidate vertical line segment subset of the first area are merged to obtain the corresponding reference table, so that the circumscribed rectangle of the connected area with the largest area in the reference table is used as the candidate BOM area, and it is determined whether the number of horizontal line segments in the candidate BOM area is greater than or equal to the first set number threshold, and whether the number of vertical line segments is greater than or equal to the second set number threshold. When the number of horizontal line segments in the candidate BOM area is greater than or equal to the first set number threshold, and the number of vertical line segments in the candidate BOM area is greater than or equal to the second set number threshold, the candidate BOM area is used as the first target BOM area.

[0088] Next, the position coordinates of the next area are determined based on the position coordinates of the first target BOM area. For example, the position coordinates of the first target BOM area are {(Bleft, Btop), (Bright, Btop), (Bright, Bbottom), (Bleft, Bbottom)}, and the width dw of the first target BOM area is Bright-Bleft; the position coordinates of the next area are {(Bleft-dw, Btop), (Bright-dw, Btop), (Bright-dw, Bbottom), (Bleft-dw, Bbottom)}.

[0089] Next, based on the horizontal coordinate starting point and the vertical coordinate ending point in the next region, a subset of candidate horizontal line segments and a subset of candidate vertical line segments in the next region are determined, and the subset of candidate horizontal line segments and the subset of candidate vertical line segments in the next region are merged to obtain a reference table corresponding to the next region. Thus, the circumscribed rectangle of the connected region with the largest area in the reference table is used as the next candidate BOM region, and it is determined whether the number of horizontal line segments in the next candidate BOM region is greater than or equal to a first set number threshold, and whether the number of vertical line segments is greater than or equal to a second set number threshold. If the number of horizontal line segments in the next candidate BOM region is greater than or equal to the first set number threshold, and the number of vertical line segments in the next candidate BOM region is greater than or equal to the second set number threshold, the next candidate BOM region is used as the next target BOM region. This cycle continues until the number of horizontal line segments in the candidate region is less than the first set number threshold, or the number of vertical line segments in the candidate region is less than the second set number threshold. At the end of this process, all target BOM regions in the target image can be obtained.

[0090] Step 307 : performing table recognition on any target BOM area in the at least one target BOM area to obtain at least one candidate BOM.

[0091] Step 308: Determine a target BOM from at least one candidate BOM.

[0092] It should be noted that the execution process of step 301 and steps 307 to 308 can be implemented in any way in the embodiments of the present disclosure, and the embodiments of the present disclosure do not limit this and will not be described in detail.

[0093] In summary, according to each candidate horizontal line segment sub-set and each candidate vertical line segment sub-set in the candidate horizontal line segment set in the target image, the reference table in each region can be determined, according to the circumscribed rectangle corresponding to the largest connected region in the reference table, the corresponding candidate BOM region can be determined, and the number of horizontal line segments and the number of vertical line segments in the candidate BOM region are verified, whether the candidate BOM region is the target BOM can be determined.

[0094] In order to accurately determine the candidate BOM in the target BOM region, as shown in Figure 11 Figure 11 is a flowchart of a BOM identification method based on RPA and AI to realize IA according to another embodiment of the present disclosure, in the embodiment of the present disclosure, the image segmentation network can be used to determine a plurality of cells in any target BOM region, the text detection network can be used to determine a plurality of text segments in any target BOM region, and according to the plurality of cells and the corresponding plurality of text segments in any target BOM region, the candidate BOM in any target BOM region can be determined, Figure 11 As shown in the embodiment, the method can include the following steps:

[0095] Step 1101, acquiring a target image.

[0096] Step 1102, determining at least one target BOM region from the target image.

[0097] Step 1103, inputting any target BOM region into the image segmentation network and the text detection network respectively, to obtain a plurality of cells in any target BOM region output by the image segmentation network, and a plurality of text segments in any target BOM region output by the text detection network.

[0098] In the embodiment of the present disclosure, any target BOM region can be input into the image segmentation network and the text detection network respectively, the image segmentation network can segment the table line in the image of any target BOM region to obtain a plurality of cells in any target BOM region, and the text detection network can output a plurality of text segments in any target BOM region. It should be noted that the image segmentation network can be Unet, U2Net, etc., and the text detection network can be a differentiable binarization network (Differentiable Binarization Networking, DBNet for short).

[0099] Step 1104, merging each cell in the plurality of cells in each target BOM region and the text segment corresponding to the position of each cell to obtain the text segment corresponding to each cell.

[0100] ​Furthermore, according to the position of each cell in the multiple cells in each target BOM area and the position of each text fragment, each cell in the multiple cells in each target BOM area and the text fragment corresponding to each cell are merged to obtain the text fragment corresponding to each cell, and each cell is indexed and recorded.

[0101] In addition, when the position of any text segment among multiple text segments coincides with the positions of multiple adjacent cells, that is, any text segment among multiple text segments spans at least two adjacent cells, the dividing line in at least two adjacent cells can be determined, and according to the position of the dividing line, any text segment can be segmented to obtain multiple text sub-segments, and according to the positions of the multiple text sub-segments and the positions of the multiple adjacent cells, the multiple text sub-segments and the multiple adjacent cells are merged.

[0102] Step 1105 : Use a text recognition network to perform content recognition on the text segments corresponding to each cell in each target BOM area to obtain the text content corresponding to each cell.

[0103] As a possible implementation of the present disclosure, a text recognition network can be used to perform content recognition on the text fragments corresponding to each cell in each target BOM area to obtain the text content of each cell. The text recognition network can be a convolutional recurrent neural network (CRNN).

[0104] Step 1106 : Each cell in each target BOM area and the text content corresponding to each cell may be used to determine a candidate BOM corresponding to each target BOM area.

[0105] Then, by combining the cells in each target BOM area and the text content corresponding to each cell, the candidate BOM corresponding to each target BOM area can be obtained.

[0106] For example, if Figure 12 As shown, Figure 12 This is a flowchart of a method for implementing BOM content recognition for IA based on RPA and AI according to another embodiment of the present disclosure. In the first step, the input is an image containing a wired table area (an image of a candidate BOM table);

[0107] The second step is to use a trained image segmentation network (such as Unet, U2Net, etc.) to segment the table lines in the image and obtain the cell structure;

[0108] The third step is to use a text detection network (such as DBNet) to detect text fragments (text lines) in the image and obtain several text fragments;

[0109] Step 4: Combine the results of steps 2 and 3 to obtain the text segments contained in each table cell and record the index. If a text segment is found to cross a table line, the table line will split the text segment into two sub-segments, each belonging to two different adjacent cells.

[0110] In the fifth step, a text recognition network (e.g., CRNN) is used to identify the content of each text line segment processed in the fourth step. Multiple text segments can be grouped into a batch and recognized in one inference process to improve recognition speed.

[0111] In the sixth step, the table cell structure and the text fragment recognition results in the table cells are combined to give the table recognition result of the input image.

[0112] Step 1107: Determine a target BOM from at least one candidate BOM.

[0113] It should be noted that the execution process of steps 1101 to 1102 and step 1107 can be implemented in any way in the embodiments of the present disclosure, and the embodiments of the present disclosure do not limit this and will not be described in detail.

[0114] In summary, by inputting any target BOM area into the image segmentation network and the text detection network respectively, a plurality of cells in any target BOM area output by the image segmentation network and a plurality of text fragments in any target BOM area output by the text detection network are obtained; each cell in the multiple cells in each target BOM area and the text fragments corresponding to the position of each cell are merged to obtain the text fragment corresponding to each cell; the text recognition network is used to perform content recognition on the text fragment corresponding to each cell in each target BOM area to obtain the text content corresponding to each cell; each cell in each target BOM area and the text content corresponding to each cell can be used to determine the candidate BOM corresponding to each target BOM area. Thus, through the image segmentation network and the text detection network, a plurality of cells in each target BOM area and the text content corresponding to each cell are obtained, so that the candidate BOM in each target BOM area can be determined.

[0115] Since the candidate BOM may contain other contents besides the BOM, the BOM in the candidate BOM can be identified to determine the target BOM. Figure 13This is a flowchart of a BOM recognition method for implementing IA based on RPA and AI according to another embodiment of the present disclosure. In the embodiment of the present disclosure, based on OCR, the BOM header of the candidate BOM is recognized to determine the position of the header of at least one target sub-BOM. Then, according to the set content direction and the position of the header of each target sub-BOM in the header of at least one target sub-BOM, each target sub-BOM is determined. Based on each target sub-BOM, the target BOM can be determined. Figure 13 The illustrated embodiment may include the following steps:

[0116] Step 1301: Acquire a target image.

[0117] Step 1302: Determine at least one target BOM area from the target image.

[0118] Step 1303 : For any target BOM area in the at least one target BOM area, perform table recognition on the target BOM area to obtain at least one candidate BOM.

[0119] Step 1304 : Based on the natural language processing technology NLP, perform BOM header recognition on the candidate BOM to determine the location of the header of at least one target sub-BOM.

[0120] It should be understood that since the header usually contains fixed text fields, for example, the header usually contains fields such as "number, serial number, material, quantity", based on natural language processing technology (NLP), the candidate BOM is recognized according to the set header text to obtain the location of the header of at least one target sub-BOM.

[0121] Step 1305 : determining each target sub-BOM according to the set content direction and the position of the header of each target sub-BOM in the header of at least one target sub-BOM.

[0122] As a possible implementation method of an embodiment of the present disclosure, after determining the position of the header of at least one target sub-BOM, each target sub-BOM can be determined according to the set content direction (from bottom to top), wherein the area below the position of the header of the first target sub-BOM can be the legend area, and the table recognition result corresponding to the legend area can be discarded.

[0123] Step 1306: Merge the target sub-BOMs to determine the target BOM.

[0124] Then, each target sub-BOM is merged to determine the target BOM. For example, multiple target sub-BOMs are arranged and merged from right to left and the items in the same target sub-BOM are arranged and merged from bottom to top to obtain the target BOM.

[0125] It should be noted that the execution process of steps 1301 to 1303 can be implemented in any of the embodiments of the present disclosure, and the embodiments of the present disclosure do not limit this and will not be described in detail.

[0126] In summary, each target sub-BOM can be determined from the candidate BOM according to the set content direction and the position of the header of each target sub-BOM, and then the target sub-BOMs can be merged to determine the target BOM.

[0127] The disclosed embodiments of the BOM recognition method for implementing IA based on RPA and AI acquire a target image; determine at least one target BOM area from the target image; perform table recognition on any of the at least one target BOM area to obtain at least one candidate BOM; and determine the target BOM from the candidate BOMs. Thus, by automatically performing table recognition on the target BOM area in the target image and determining the target BOM from the identified candidate BOMs, the BOM can be intelligently and automatically located and identified from the target image, replacing manual operations and improving BOM recognition efficiency and accuracy.

[0128] With the above Figures 2 to 13 The BOM recognition method based on RPA and AI for implementing IA provided in the embodiment corresponds to this. The present disclosure also provides a BOM recognition device based on RPA and AI for implementing IA. Since the BOM recognition device based on RPA and AI for implementing IA provided in the embodiment of the present disclosure is similar to the above-mentioned BOM recognition method based on RPA and AI for implementing IA, the present disclosure also provides a BOM recognition device based on RPA and AI for implementing IA. Figures 2 to 13 The BOM identification method for implementing IA based on RPA and AI provided in the embodiment corresponds to the BOM identification method, so the implementation method of the BOM identification method for implementing IA based on RPA and AI is also applicable to the BOM identification device for implementing IA based on RPA and AI provided in the embodiment of the present disclosure, and will not be described in detail in the embodiment of the present disclosure.

[0129] Figure 14 This is a structural diagram of a BOM recognition device for implementing IA based on RPA and AI according to an embodiment of the present disclosure.

[0130] like Figure 14 As shown, the BOM identification device 1400 for implementing IA based on RPA and AI includes: an acquisition module 1410, a first determination module 1420, an identification module 1430 and a second determination module 1440.

[0131] Among them, the acquisition module 1410 is used to acquire the target image; the first determination module 1420 is used to determine at least one target BOM area from the target image; the identification module 1430 is used to perform table recognition on any target BOM area in at least one target BOM area to obtain at least one candidate BOM; the second determination module 1440 is used to determine the target BOM from at least one candidate BOM.

[0132] As a possible implementation of the embodiment of the present disclosure, a BOM recognition device 1400 for implementing IA based on RPA and AI is applied to an RPA robot.

[0133] As a possible implementation method of the embodiment of the present disclosure, the first determination module 1420 is also used to: identify, from the target image, a set of candidate horizontal line segments and a set of candidate vertical line segments corresponding to the target image; determine, from the set of candidate horizontal line segments and the set of candidate vertical line segments, each candidate horizontal line segment subset and each candidate vertical line segment subset in sequence according to a set direction, wherein the candidate horizontal line segment subset includes multiple horizontal line segments located in the same area, and the target vertical line segment set includes multiple vertical line segments located in the same area; merge each candidate horizontal line segment subset with each candidate vertical line segment subset located in the same area to obtain at least one reference table; for each reference table in the at least one reference table, determine the candidate BOM area corresponding to each reference table according to the circumscribed rectangle corresponding to the connected area with the largest area in each reference table; and determine at least one target BOM area according to the number of horizontal line segments and the number of vertical line segments corresponding to each candidate BOM area.

[0134] As a possible implementation method of the embodiment of the present disclosure, the first determination module 1420 is further used to: set a horizontal rectangular window as a first element window, perform a first filtering process on the target image to obtain a first image, wherein the first image includes the pixel points of each candidate horizontal line segment in the target image; set a vertical rectangular window as a second element window, perform a second filtering process on the target image to obtain a second image, wherein the second image includes the pixel points of each candidate vertical line segment in the target image; determine the first circumscribed rectangle coordinates of each candidate horizontal line segment according to the coordinates of each pixel point in the first image, wherein the first circumscribed rectangle coordinates are used to indicate the endpoint coordinates of each candidate horizontal line segment; determine the second circumscribed rectangle coordinates of each candidate vertical line segment according to the coordinates of each pixel in the second image, wherein the second circumscribed rectangle coordinates are used to indicate the endpoint coordinates of each candidate vertical line segment; determine a set of candidate horizontal line segments according to the first circumscribed rectangle coordinates of each candidate horizontal line segment, and determine a set of candidate vertical line segments according to the second circumscribed rectangle coordinates of each candidate vertical line segment.

[0135] As a possible implementation method of the embodiment of the present disclosure, the first determination module 1420 is also used to: delete the longest candidate horizontal line segment that matches the boundary horizontal frame segment of the target image from each candidate horizontal line segment according to the first circumscribed rectangle coordinates of each candidate horizontal line segment, and delete the longest candidate vertical line segment that matches the boundary vertical frame segment of the target image from each candidate vertical line segment according to the second circumscribed rectangle coordinates of each candidate vertical line segment; determine a set of candidate horizontal line segments based on each candidate horizontal line segment after deleting the longest candidate horizontal line segment, and determine a set of candidate vertical line segments based on each candidate vertical line segment after deleting the longest candidate vertical line segment.

[0136] As a possible implementation method of an embodiment of the present disclosure, the first determination module 1420 is also used to: obtain the number of horizontal line segments and the number of vertical line segments in any candidate BOM area among the candidate BOM areas; when the number of horizontal line segments in any candidate BOM area is greater than or equal to a first set number threshold, and the number of vertical line segments in any candidate BOM area is greater than or equal to a second set number threshold, take any candidate BOM area as the target BOM area.

[0137] As a possible implementation method of the embodiment of the present disclosure, the first determination module 1420 is further used to: determine, in sequence, the starting point of each horizontal coordinate and the ending point of each vertical coordinate corresponding to the starting point of each horizontal coordinate according to the set direction; determine, from the set of candidate horizontal line segments, a subset of each candidate horizontal line segment according to the starting point of each horizontal coordinate; and determine, from the set of candidate vertical line segments, a subset of each candidate vertical line segment according to the ending point of each vertical coordinate.

[0138] As a possible implementation method of the embodiment of the present disclosure, the recognition module 1430 is also used to: input any target BOM area into the image segmentation network and the text detection network respectively to obtain multiple cells in any target BOM area output by the image segmentation network, and multiple text fragments in any target BOM area output by the text detection network; merge each cell in the multiple cells in each target BOM area and the text fragments corresponding to each cell position to obtain the text fragment corresponding to each cell; use the text recognition network to perform content recognition on the text fragments corresponding to each cell in each target BOM area to obtain the text content corresponding to each cell; determine the candidate BOM corresponding to each target BOM area based on each cell in each target BOM area and the text content corresponding to each cell.

[0139] As a possible implementation of the embodiment of the present disclosure, the BOM identification device 1400 for implementing IA based on RPA and AI further includes: a third determination module, a segmentation module, and a processing module.

[0140] Among them, the third determination module is used to determine at least one separator line in multiple adjacent cells when the position of any text segment in multiple text segments coincides with the position of multiple adjacent cells, wherein the separator line is used to separate multiple adjacent cells; the segmentation module is used to segment any text segment according to the position of at least one separator line to obtain multiple text sub-segments; the processing module is used to merge multiple text sub-segments and multiple adjacent cells according to the positions of multiple text sub-segments and the positions of the multiple adjacent cells.

[0141] As a possible implementation method of the embodiment of the present disclosure, the second determination module is also used to perform BOM header recognition on the candidate BOM based on natural language processing technology NLP to determine the position of the header of at least one target sub-BOM; determine each target sub-BOM according to the set content direction and the position of the header of each target sub-BOM in the header of at least one target sub-BOM; merge each target sub-BOM to determine the target BOM.

[0142] The disclosed embodiment of the bill of materials (BOM) recognition device, which uses RPA and AI to implement IA, acquires a target image; identifies at least one target BOM area within the target image; performs table recognition on any of the at least one target BOM area to obtain at least one candidate BOM; and determines the target BOM from the candidate BOMs. Thus, by automatically performing table recognition on the target BOM area within the target image and determining the target BOM from the identified candidate BOMs, the BOM can be intelligently and automatically located and identified from the target image, replacing manual operations and improving BOM recognition efficiency and accuracy.

[0143] In order to implement the above embodiments, the embodiments of the present disclosure also propose an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it implements the BOM identification method for implementing IA based on RPA and AI as described in any of the above method embodiments.

[0144] In order to implement the above embodiments, the embodiments of the present disclosure also propose a non-temporary computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the BOM identification method for implementing IA based on RPA and AI as described in any of the above method embodiments is implemented.

[0145] In order to implement the above embodiments, the embodiments of the present disclosure also propose a computer program product. When the instruction processor in the computer program product is executed, the BOM recognition of IA based on RPA and AI as described in any of the above method embodiments is implemented.

[0146] In order to implement the above embodiment, the present application also proposes an electronic device, such as Figure 15 As shown, Figure 15 The present invention is a block diagram of an electronic device for implementing a BOM identification method for IA based on RPA and AI, according to an exemplary embodiment.

[0147] like Figure 15 As shown, the electronic device 1500 includes:

[0148] The memory 1510 and the processor 1520, a bus 1530 connecting different components (including the memory 1510 and the processor 1520), the memory 1510 stores a computer program, and when the processor 1520 executes the program, the BOM identification method for implementing IA based on RPA and AI as described in the embodiment of the present disclosure is implemented.

[0149] Bus 1530 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0150] The electronic device 1500 typically includes a variety of electronic device-readable media. These media can be any available media that can be accessed by the electronic device 1500, including volatile and non-volatile media, removable and non-removable media.

[0151] The memory 1510 may also include computer system readable media in the form of volatile memory, such as random access memory (RAM) 1540 and / or cache memory 1550. The electronic device 1500 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 1560 may be used to read and write non-removable, non-volatile magnetic media ( Figure 15 Not shown, often called a "hard drive"). Although Figure 15Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk"), and an optical drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 1530 via one or more data medium interfaces. Memory 1510 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of various embodiments of the present disclosure.

[0152] A program / utility 1580 having a set (at least one) of program modules 1570 may be stored, for example, in memory 1510. Such program modules 1570 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 1570 generally implement the functions and / or methods of the embodiments described herein.

[0153] The electronic device 1500 may also communicate with one or more external devices 1590 (e.g., a keyboard, a pointing device, a display, etc.), one or more devices that enable a user to interact with the electronic device 1500, and / or any device that enables the electronic device 1500 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed through an input / output (I / O) interface 1592. Furthermore, the electronic device 1500 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 1593. Figure 15 As shown, the network adapter 1593 communicates with other modules of the electronic device 1500 via the bus 1530. Figure 15 Not shown, other hardware and / or software modules may be used in conjunction with electronic device 1500, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0154] The processor 1520 executes various functional applications and data processing by running programs stored in the memory 1510 .

[0155] It should be noted that the implementation process and technical principles of the electronic device of this embodiment can be found in Figures 2 to 13 The explanation of the BOM identification method for implementing IA based on RPA and AI in the embodiment of the present disclosure will not be repeated here.

[0156] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and features of different embodiments or examples, unless they are mutually inconsistent.

[0157] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the present disclosure, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0158] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present disclosure includes additional implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present disclosure belong.

[0159] The logic and / or steps represented in flow diagrams or otherwise described herein, for example, can be considered as a sequence of instructions to implement logic functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus, or device, such as a computer-based system, processor- containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions. In the context of this specification, a "computer-readable medium" can be any means that can contain, store, communicate, propagate or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The computer-readable medium can be a machine-readable storage device (e.g., magnetic, optical or other) a machine-readable storage diskette (e.g., floppy disk, optical disk, CD- ROM, etc.), a machine- readable storage card (e.g., PCMCIA card, etc.), a machine-readable storage tape (e.g., magnetic tape, optical tape, etc.), a machine-readable storage medium (e.g., RAM, ROM, etc.), a machine-readable signal (e.g., electrical, optical, etc.), a machine-readable medium (e.g., carrier wave, etc.) or any other suitable medium or means of embodying the program. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a RAM, a ROM, an EPROM, a FLASH memory card, an optical fiber, and a portable compact disc read-only memory (CD-ROM). Additionally, the computer-readable medium can be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example, via optical scanning of the paper or other medium, then compiled, interpreted or otherwise processed in a suitable manner if necessary, and stored in a computer memory.

[0160] It should be understood that portions of the present disclosure can be implemented in hardware, software, firmware or combinations thereof. In the above embodiments, the various steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. As such, if implemented in hardware, and in another embodiment, any of the following technologies, known in the art, or their combinations can be used: discrete logic circuitry having logic gates for implementing logic functions on data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and the like.

[0161] Those skilled in the art can understand that all or part of the steps carried out by the above-mentioned embodiment methods can be completed by programs instructing related hardware, and the programs can be stored in a computer-readable storage medium. When the programs are executed, one or a combination of the steps of the method embodiments is included.

[0162] In addition, the functional units in the various embodiments of the present disclosure may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.

[0163] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present disclosure have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. A person of ordinary skill in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.

Claims

1. A bill of materials (BOM) recognition method based on robotic process automation (RPA) and artificial intelligence (AI) to achieve intelligent automation (IA), characterized in that: include: Acquire the target image; Determining at least one target BOM area from the target image; For any target BOM area in the at least one target BOM area, performing table recognition on the any target BOM area to obtain at least one candidate BOM; Determine a target BOM from the at least one candidate BOM; Determining at least one target BOM area from the target image includes: identifying, from the target image, a set of candidate horizontal line segments and a set of candidate vertical line segments corresponding to the target image; Determining, from the candidate horizontal line segment set and the candidate vertical line segment set, each candidate horizontal line segment subset and each candidate vertical line segment subset in sequence according to a set direction of the BOM area, wherein the candidate horizontal line segment subset includes multiple horizontal line segments located in the same area, and the target vertical line segment set includes multiple vertical line segments located in the same area, including: sequentially determining, according to the set direction, each horizontal coordinate starting point and each vertical coordinate ending point corresponding to each horizontal coordinate starting point, determining, from the candidate horizontal line segment set, each candidate horizontal line segment subset based on each horizontal coordinate starting point; and determining, from the candidate vertical line segment set, each candidate vertical line segment subset based on each vertical coordinate ending point; Merging each of the candidate horizontal line segment subsets with each of the candidate vertical line segment subsets located in the same area as the candidate horizontal line segment subsets to obtain at least one reference table; For each reference table in the at least one reference table, determining a candidate BOM area corresponding to each reference table according to a circumscribed rectangle corresponding to a connected region with the largest area in each reference table; At least one target BOM area is determined according to the number of horizontal line segments and the number of vertical line segments corresponding to each candidate BOM area.

2. The method according to claim 1, characterized in that The method is executed by an RPA robot.

3. The method according to claim 1, characterized in that The step of identifying, from the target image, a set of candidate horizontal line segments and a set of candidate vertical line segments corresponding to the target image includes: Setting a horizontal rectangular window as a first element window, performing a first filtering process on the target image to obtain a first image, wherein the first image includes pixel points of each candidate horizontal line segment in the target image; Setting a vertical rectangular window as a second element window, performing a second filtering process on the target image to obtain a second image, wherein the second image includes pixel points of each candidate vertical line segment in the target image; Determining first circumscribed rectangle coordinates of each candidate horizontal line segment according to the coordinates of each pixel point in the first image, wherein the first circumscribed rectangle coordinates are used to indicate the endpoint coordinates of each candidate horizontal line segment; determining, according to the coordinates of each pixel in the second image, the coordinates of a second circumscribed rectangle of each candidate vertical line segment, wherein the coordinates of the second circumscribed rectangle are used to indicate the coordinates of the endpoints of each candidate vertical line segment; A candidate horizontal line segment set is determined based on the first circumscribed rectangle coordinates of each candidate horizontal line segment, and a candidate vertical line segment set is determined based on the second circumscribed rectangle coordinates of each candidate vertical line segment.

4. The method according to claim 3, characterized in that The determining of a set of candidate horizontal line segments according to the first circumscribed rectangle coordinates of each candidate horizontal line segment, and the determining of a set of candidate vertical line segments according to the second circumscribed rectangle coordinates of each candidate vertical line segment, includes: Deleting the longest candidate horizontal line segment that matches the horizontal boundary frame segment of the target image from each of the candidate horizontal line segments according to the first circumscribed rectangle coordinates of each of the candidate horizontal line segments, and deleting the longest candidate vertical line segment that matches the vertical boundary frame segment of the target image from each of the candidate vertical line segments according to the second circumscribed rectangle coordinates of each of the candidate vertical line segments; A candidate horizontal line segment set is determined based on each candidate horizontal line segment after deleting the longest candidate horizontal line segment, and a candidate vertical line segment set is determined based on each candidate vertical line segment after deleting the longest candidate vertical line segment.

5. The method according to claim 1, wherein The determining of at least one target BOM area according to the number of horizontal line segments and the number of vertical line segments corresponding to each candidate BOM area includes: Obtaining the number of horizontal line segments and the number of vertical line segments in any candidate BOM area among the candidate BOM areas; When the number of horizontal line segments in any candidate BOM area is greater than or equal to a first set number threshold, and the number of vertical line segments in any candidate BOM area is greater than or equal to a second set number threshold, any candidate BOM area is taken as a target BOM area.

6. The method according to claim 1, characterized in that The step of sequentially determining target horizontal line segment sets and target vertical line segment sets from the candidate horizontal line segment sets and the candidate vertical line segment sets according to a set direction includes: According to the set direction, sequentially determining each horizontal coordinate starting point and each vertical coordinate ending point corresponding to each horizontal coordinate starting point; Determining a subset of candidate horizontal line segments from the candidate horizontal line segment set according to the starting points of the horizontal coordinates; From the candidate vertical line segment set, a subset of candidate vertical line segments is determined according to the end points of the vertical coordinates.

7. The method according to any one of claims 1 to 6, characterized in that The step of performing table recognition on any target BOM area in the at least one target BOM area to obtain at least one candidate BOM includes: Inputting the any target BOM area into an image segmentation network and a text detection network respectively to obtain a plurality of cells in the any target BOM area output by the image segmentation network and a plurality of text fragments in the any target BOM area output by the text detection network; Merging each of the plurality of cells in each target BOM area and the text segments corresponding to the positions of the cells to obtain text segments corresponding to the cells; Using a text recognition network to perform content recognition on the text fragments corresponding to each cell in each target BOM area to obtain the text content corresponding to each cell; The candidate BOM corresponding to each target BOM area is determined according to each cell in each target BOM area and text content corresponding to each cell.

8. The method according to claim 7, characterized in that The method further comprises: When a position of any text segment among the plurality of text segments coincides with positions of a plurality of adjacent cells, determining at least one separation line among the plurality of adjacent cells, wherein the separation line is used to separate the plurality of adjacent cells; Segmenting any one of the text segments according to the position of the at least one dividing line to obtain a plurality of text sub-segments; The multiple text sub-segments and the multiple adjacent cells are merged according to the positions of the multiple text sub-segments and the positions of the multiple adjacent cells.

9. The method according to any one of claims 1 to 6, characterized in that The determining of the target BOM from the at least one candidate BOM includes: Based on natural language processing technology (NLP), the BOM header of the at least one candidate BOM is recognized to determine the location of the header of the at least one target sub-BOM. Determine each target sub-BOM according to the set content direction and the position of the header of each target sub-BOM in the header of the at least one target sub-BOM; The target sub-BOMs are merged to determine the target BOM.

10. A device for realizing intelligent automation IA bill of materials (BOM) recognition based on robotic process automation (RPA) and artificial intelligence (AI), characterized in that: include: An acquisition module, used to acquire a target image; A first determining module is configured to determine at least one target BOM area from the target image; an identification module, configured to perform table identification on any target BOM area in the at least one target BOM area to obtain at least one candidate BOM; A second determining module, configured to determine a target BOM from the at least one candidate BOM; The first determining module is further configured to: identifying, from the target image, a set of candidate horizontal line segments and a set of candidate vertical line segments corresponding to the target image; Determining, from the candidate horizontal line segment set and the candidate vertical line segment set, each candidate horizontal line segment subset and each candidate vertical line segment subset in sequence according to a set direction of the BOM area, wherein the candidate horizontal line segment subset includes multiple horizontal line segments located in the same area, and the target vertical line segment set includes multiple vertical line segments located in the same area, including: sequentially determining, according to the set direction, each horizontal coordinate starting point and each vertical coordinate ending point corresponding to each horizontal coordinate starting point, determining, from the candidate horizontal line segment set, each candidate horizontal line segment subset based on each horizontal coordinate starting point; and determining, from the candidate vertical line segment set, each candidate vertical line segment subset based on each vertical coordinate ending point; Merging each of the candidate horizontal line segment subsets with each of the candidate vertical line segment subsets located in the same area as the candidate horizontal line segment subsets to obtain at least one reference table; For each reference table in the at least one reference table, determining a candidate BOM area corresponding to each reference table according to a circumscribed rectangle corresponding to a connected region with the largest area in each reference table; At least one target BOM area is determined according to the number of horizontal line segments and the number of vertical line segments corresponding to each candidate BOM area.

11. The device according to claim 10, characterized in that The first determining module is further configured to: Setting a horizontal rectangular window as a first element window, performing a first filtering process on the target image to obtain a first image, wherein the first image includes pixel points of each candidate horizontal line segment in the target image; Setting a vertical rectangular window as a second element window, performing a second filtering process on the target image to obtain a second image, wherein the second image includes pixel points of each candidate vertical line segment in the target image; Determining first circumscribed rectangle coordinates of each candidate horizontal line segment according to the coordinates of each pixel point in the first image, wherein the first circumscribed rectangle coordinates are used to indicate the endpoint coordinates of each candidate horizontal line segment; determining, according to the coordinates of each pixel in the second image, the coordinates of a second circumscribed rectangle of each candidate vertical line segment, wherein the coordinates of the second circumscribed rectangle are used to indicate the coordinates of the endpoints of each candidate vertical line segment; A candidate horizontal line segment set is determined based on the first circumscribed rectangle coordinates of each candidate horizontal line segment, and a candidate vertical line segment set is determined based on the second circumscribed rectangle coordinates of each candidate vertical line segment.

12. The device according to claim 11, characterized in that The first determining module is further configured to: Deleting the longest candidate horizontal line segment that matches the horizontal boundary frame segment of the target image from each of the candidate horizontal line segments according to the first circumscribed rectangle coordinates of each of the candidate horizontal line segments, and deleting the longest candidate vertical line segment that matches the vertical boundary frame segment of the target image from each of the candidate vertical line segments according to the second circumscribed rectangle coordinates of each of the candidate vertical line segments; A candidate horizontal line segment set is determined based on each candidate horizontal line segment after deleting the longest candidate horizontal line segment, and a candidate vertical line segment set is determined based on each candidate vertical line segment after deleting the longest candidate vertical line segment.

13. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method according to any one of claims 1 to 9 is implemented.

14. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Form identifying and positioning method based on image form analysis technology

    CN106203397A

  • Method for extracting lab test result from medical lab sheet image

    CN106446881A

  • Character recognition method and device combining RPA and AI, electronic equipment and storage medium

    CN112115774A