Three-Coordinate Report Extraction and Analysis Method, Equipment and Medium for Component Manufacturing

Through the automated three-coordinate report extraction and analysis method, the traditional analysis method is solved, which is time-consuming and error-prone, and efficient and accurate data extraction and analysis is achieved, which is suitable for different enterprises and equipment brands.

CN119578394BActive Publication Date: 2025-06-13HIMILE PRECISION MASCH (SHANDONG) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510142892.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-13
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

The traditional three-coordinate report analysis method relies on manual data export, which is time-consuming and labor-intensive, and is prone to errors, especially when the data volume is large.

Method used

An automated three-coordinate report extraction and analysis method is proposed. By obtaining report files and determining data extraction templates, element data in the report is extracted and stored, and summarized and displayed.

Benefits of technology

It realizes the automated extraction and analysis of three-coordinate report data, reduces manual labor, improves data processing efficiency and accuracy, and is suitable for different enterprises and equipment brands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119578394B_ABST
    Figure CN119578394B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device and medium for extracting and analyzing a coordinate measuring machine (CMM) report for parts manufacturing, which relates to the field of data analysis. The method includes: obtaining a report file corresponding to the CMM report and determining a data extraction template corresponding to the report file; determining a text box in the report file and extracting the content in the abscissa interval in the text box based on the data extraction template; storing the extracted content as corresponding element data according to the data extraction template; summarizing and displaying according to the element data. Through the data extraction template, there is no need to manually view the PDF report page by page and manually extract data, and it is possible to automatically extract the element data in the CMM report, process a large number of reports in a short time, reduce the degree of manual labor, reduce the labor cost, speed up the data processing process, and improve the overall work efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data analysis, and particularly to a method, device, and medium for extracting and analyzing CMM reports for component manufacturing. Background Art

[0002] A CMM report refers to a report that measures and records the shape, size, and position of a workpiece through a high-precision CMM during the machining process. Analyzing CMM reports is of great significance for machining, which can help enterprises improve machining accuracy, quality, and efficiency, reduce costs, and improve product reliability.

[0003] In traditional solutions, CMM reports are usually in PDF format. Traditional analysis methods generally use the method of manually exporting data, which requires a lot of time and manpower. And when the report data volume is large, manual analysis is prone to errors. Summary of the Invention

[0004] To solve the above problems, this application proposes a method for extracting and analyzing CMM reports for component manufacturing, including:

[0005] Obtain the report file corresponding to the CMM report, and determine the data extraction template corresponding to the report file; at least the abscissa intervals corresponding to each element data in the report file are included in the data extraction template;

[0006] Determine the text boxes in the report file, and extract the content in the abscissa intervals in the text boxes based on the data extraction template;

[0007] According to the data extraction template, store the extracted content as the corresponding element data;

[0008] Summarize and display according to the element data.

[0009] In one example, before determining the data extraction template corresponding to the report file, the method further includes:

[0010] For the report file corresponding to the CMM report, determine each element data included in the CMM report, and determine the alignment method of the element data;

[0011] For each element data, determine the corresponding abscissa undetermined interval according to its position order in the text box of the report file;

[0012] According to the alignment method of the element data, select the corresponding interval from the abscissa undetermined intervals as the abscissa interval corresponding to the element data.

[0013] In one example, the method further includes:

[0014] For each element data, determine the corresponding upper limit coordinate and lower limit coordinate according to its abscissa interval;

[0015] According to the position order of the element data, determine the interval range between each element data through the upper limit coordinate of the previous element data and the lower limit coordinate of the next element data;

[0016] For the interval range, determine the first preset reasonable range corresponding to the interval range according to the alignment methods respectively corresponding to the element data on both sides thereof;

[0017] If the interval range exceeds the first preset reasonable range, modify the abscissa interval of the previous element data and / or the abscissa interval of the next element data; wherein, the modification priorities from high to low are as follows: translate the abscissa interval of the element data whose interval range on the other side conforms to the corresponding first preset reasonable range to the other side; shrink the abscissa interval of the element data whose abscissa interval conforms to the corresponding second preset reasonable range; report and modify manually.

[0018] In an example, determine the text box in the report file, and extract the content in the abscissa interval in the text box based on the data extraction template, which specifically includes:

[0019] Parse the report file to determine the position coordinates of the text box in the report file;

[0020] In the position coordinates corresponding to the text box, use each preset ordinate interval as a component of a single dimension information, and within the ordinate interval, extract the content in the abscissa interval through file parsing according to the data extraction template;

[0021] If the extraction is successful, use the extracted content as the corresponding element data;

[0022] If the extraction fails, perform extraction recognition through OCR and use the extracted content as the corresponding element data.

[0023] In an example, using each preset ordinate interval as a component of a single dimension information specifically includes:

[0024] For each preset ordinate interval, obtain the unification mark corresponding to the current first ordinate interval;

[0025] If the unification mark of the first vertical coordinate interval is the same as that of the previous second vertical coordinate interval, and the vertical coordinate difference between the first vertical coordinate interval and the second vertical coordinate interval is lower than the preset maximum error value, then perform unification processing on the first vertical coordinate interval and the second vertical coordinate interval to update the value of the first vertical coordinate interval to the value of the second vertical coordinate interval.

[0026] In one example, for each preset vertical coordinate interval, obtaining the unification mark corresponding to the current first vertical coordinate interval specifically includes:

[0027] For each preset vertical coordinate interval, perform real-time parsing on the content extracted from the current first vertical coordinate interval to obtain a first parsing result; and obtain the second parsing result of the previous second vertical coordinate interval; wherein, the first parsing result and the second parsing result include the parsing results of each element data in each horizontal coordinate interval;

[0028] Compare the formats of the parsing results of the element data in the same position order in the first parsing result and the second parsing result;

[0029] If there are multiple different formats in the comparison result, assign the same unification mark as the second vertical coordinate interval to the first vertical coordinate interval.

[0030] In one example, based on the data extraction template, extract the content in the horizontal coordinate interval in the text box, specifically including:

[0031] Start content extraction from the lower limit coordinate of the horizontal coordinate interval corresponding to the data extraction template and determine the first character extracted;

[0032] Determine the first interval distance between the first character and the lower limit coordinate of the horizontal coordinate interval;

[0033] If the first interval distance is lower than the preset distance, lower the lower limit coordinate of the horizontal coordinate interval and re-perform content extraction until the first interval distance is higher than the preset distance;

[0034] Continue content extraction and determine the height of each character;

[0035] If the height of a specified character is higher than the preset height, determine whether the specified character is a text character or a graphic character;

[0036] If the specified character is a graphic character, convert the graphic character to the corresponding text character and then continue content extraction;

[0037] If there are multiple consecutive specified characters that are all text characters, the lower limit coordinate of the abscissa interval is increased according to the number of the specified characters, and content extraction is performed;

[0038] Through the content extraction, until the last character is recognized, and the second interval distance between the last character and the upper limit coordinate of the abscissa interval is determined;

[0039] If the first interval distance is lower than the preset distance, the lower limit coordinate of the abscissa interval is increased, and content extraction continues until the second interval distance is higher than the preset distance.

[0040] In one example, the method further includes:

[0041] If there are multiple consecutive dimension information and the same single element data exists in each dimension information, and the abscissa interval corresponding to the single element data has been adjusted, then in the remaining dimension information, the abscissa interval corresponding to the same single element data is adjusted;

[0042] Wherein, based on the adjusted upper limit coordinate and lower limit coordinate in the abscissa intervals of the multiple consecutive dimension information that have been adjusted, the upper limit coordinate and lower limit coordinate of the abscissa intervals of the remaining dimension information are respectively adjusted.

[0043] On the other hand, the present application also proposes a three - coordinate report extraction and analysis device for parts manufacturing, including:

[0044] At least one processor; and,

[0045] A memory communicatively connected to the at least one processor; wherein,

[0046] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the three - coordinate report extraction and analysis method for parts manufacturing as described in any of the above examples.

[0047] On the other hand, the present application also proposes a non - volatile computer storage medium storing computer - executable instructions, and the computer - executable instructions are set as: the three - coordinate report extraction and analysis method for parts manufacturing as described in any of the above examples.

[0048] The three - coordinate report extraction and analysis method for parts manufacturing proposed by the present application can bring the following beneficial effects:

[0049] Through the data extraction template, it is possible to automatically extract the element data in the CMM report without manually viewing the PDF report page by page and manually extracting the data. This can process a large number of reports in a short time, reduce the degree of manual labor, lower the labor cost, speed up the data processing flow, and improve the overall work efficiency.

[0050] It avoids situations such as negligence, misreading, and transcription errors that may occur during manual data extraction, ensures the consistency of the extracted data with the original report, improves the accuracy and reliability of the data, and provides a more accurate basis for subsequent analysis and decision-making.

[0051] By setting corresponding data extraction templates for different enterprises and equipment brands, the corresponding extraction of element data can be achieved, thus adapting to different enterprises and equipment brands, and having high versatility. Brief Description of the Drawings

[0052] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0053] Figure 1 It is a schematic flowchart of the method for extracting and analyzing the CMM report for parts manufacturing in the embodiment of the present application;

[0054] Figure 2 It is a schematic diagram of the CMM report in one case in the embodiment of the present application;

[0055] Figure 3 It is a line chart shown in summary in one case in the embodiment of the present application;

[0056] Figure 4 It is a bar chart shown in summary in one case in the embodiment of the present application;

[0057] Figure 5 It is a schematic diagram of some graphic characters in one case in the embodiment of the present application;

[0058] Figure 6 It is a schematic diagram of some training sets in one case in the embodiment of the present application;

[0059] Figure 7 It is a schematic diagram of the device for extracting and analyzing the CMM report for parts manufacturing in the embodiment of the present application. Detailed Embodiments

[0060] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the scope of protection of this application.

[0061] The following will detail the technical solutions provided by each embodiment of this application in conjunction with the drawings.

[0062] As Figure 1 shown, an embodiment of this application provides a method for extracting and analyzing a three-coordinate report for parts manufacturing, including:

[0063] S101: Obtain the report file corresponding to the three-coordinate report, and determine the data extraction template corresponding to the report file; at least the horizontal coordinate intervals corresponding to each element data in the report file are included in the data extraction template.

[0064] Generally speaking, for different enterprises and different equipment brands, the formats of their corresponding three-coordinate reports are also different. Therefore, data extraction templates can be set in advance for different enterprises, equipment brands, etc. When obtaining the report file, the corresponding pre-set data extraction template can be determined based on the source of the report file.

[0065] There are multiple dimension information in the three-coordinate report, and each dimension information may contain multiple element data. The horizontal data of each row except the first row can be regarded as the same dimension information (also called dimension type), which has the same vertical coordinate interval (i.e., Y coordinate) in the three-coordinate report. In this vertical coordinate area, different horizontal coordinate intervals correspond to the element data of different products under this dimension information.

[0066] The vertical data of each column except the first column corresponds to the same product or a component (hereinafter simply referred to as a product). Different vertical coordinate intervals correspond to different products. The same product has the same horizontal coordinate interval in the three-coordinate report. In this horizontal coordinate interval, different vertical coordinate intervals correspond to the element data of the product or component under different dimension information.

[0067] Based on this, for the convenience of representation, within the same vertical coordinate interval, the data recorded in different horizontal coordinate intervals is regarded as different element data. Each element data represents the data of different products under the corresponding dimension information.

[0068] As Figure 2As shown, it is a schematic diagram of a three - coordinate report under an example. The dimension information can include: form tolerance (FORM), distance measurement (DM), X - coordinate, Y - coordinate (where the X - coordinate and Y - coordinate refer to the records in the measurement data, rather than their position information in the three - coordinate report), etc. Of course, the dimension information can also include data such as dimension name, measured dimension value, theoretical dimension value, upper dimension tolerance, lower dimension tolerance, etc. Each dimension information contains multiple element data (for example, the form tolerance FORM includes multiple element data such as 0.0002, 0.0000, 0.0500, 0.0000, etc.), and the abscissa intervals of each element data are different.

[0069] Furthermore, the data extraction template can be set manually in advance. For example, manually set the abscissa intervals corresponding to each element data by summarizing the rules and combining expert experience based on the historical data of the three - coordinate reports of the corresponding enterprise or equipment brand. However, this setting method is not only time - consuming and laborious but also prone to errors.

[0070] Based on this, when setting the data extraction template, for the report file corresponding to the three - coordinate report, determine the element data included in the three - coordinate report and determine the alignment method of the element data. The alignment methods mainly include left alignment, right alignment, center alignment, etc. In this application, left alignment and right alignment are mainly taken as examples. Center alignment is not essentially different from left alignment and right alignment, but it is more centered in the abscissa interval.

[0071] Since the data content in the three - coordinate report is of different lengths, especially when the text box is too long, it is difficult to give a reasonable range for the abscissa interval, and there are situations of missing data extraction or data stringing. To ensure the accuracy of data extraction, select the corresponding abscissa interval according to the alignment method of the specific data text box. Here, the abscissa intervals corresponding to left alignment and right alignment can be called Left_X_Pos coordinate and Right_X_Pos coordinate respectively.

[0072] For each element data, determine its corresponding undetermined abscissa interval according to its position sequence in the text box of the report file. The undetermined abscissa interval means that the abscissa of the entire text box is evenly divided (or weighted evenly divided) according to the conventional table size (which can be statistically obtained from the historical report files of this enterprise or brand), and the abscissa undetermined intervals are closely connected. All the abscissa undetermined intervals together form the entire abscissa of the text box.

[0073] At this time, according to the alignment of the element data, select the corresponding interval in the horizontal coordinate interval selection as the horizontal coordinate interval corresponding to the element data. For example, assuming that there are three horizontal coordinate intervals, corresponding to the ranges: a~b, c~d, e~f, for the current element data, its position order is third, and select e~f as its corresponding horizontal coordinate interval. Then, according to its own alignment mode of right alignment, select the right area g~h in the horizontal coordinate interval (where h is close to f, and in some cases, g~f can also be directly selected) as its corresponding horizontal coordinate interval.

[0074] For example, as shown in Table 1 below, in one case, when the dimension name is left-aligned, the data text box coordinates use the Left_X_Pos coordinates, and the final horizontal axis range is 70~90. The measured value of the dimension is right-aligned, the data text box coordinates use the Right_X_Pos coordinates, and the final horizontal axis range is 180~200. The theoretical value of the dimension is right-aligned, the data text box coordinates use the Right_X_Pos coordinates, and the final horizontal axis range is 240~260. The upper tolerance value of the dimension is right-aligned, the data text box coordinates use the Right_X_Pos coordinates, and the final horizontal axis range is 295~315. The lower tolerance value of the dimension is right-aligned, the data text box coordinates use the Right_X_Pos coordinates, and the final horizontal axis range is 350~370.

[0075] Table 1

[0076]

[0077] Furthermore, after obtaining the data extraction template, in order to prevent unreasonable setting of the horizontal axis interval, the data extraction template can be checked accordingly.

[0078] For each element data, the corresponding upper limit coordinate and lower limit coordinate are determined according to its horizontal coordinate interval, wherein the upper limit coordinate and the lower limit coordinate refer to the left endpoint coordinate and the right endpoint coordinate in the horizontal coordinate interval, respectively.

[0079] According to the position sequence of element data, the interval range between each element data is determined by the upper limit coordinate of the previous element data and the lower limit coordinate of the next element data. For example, if the previous element data is the dimension name, its corresponding horizontal axis range is 70~90, and the lower limit coordinate is 90, and the next element data is the measured dimension value, its corresponding horizontal axis range is 180~200, and the upper limit coordinate is 180. At this time, the interval range is 180-90=90.

[0080] For an interval range, determine the first preset reasonable range corresponding to the interval range according to the alignment methods respectively corresponding to the element data on both sides of it. When the previous element data is left-aligned and the next element data is right-aligned, the first preset reasonable range between the two is the highest. When the alignment methods of the element data on both sides are the same, the first preset reasonable range between the two is medium. When the previous element data is right-aligned and the next element data is left-aligned, the first preset reasonable range between the two is the lowest. Of course, if alignment methods such as centered alignment and average alignment are considered, more refined settings can also be made.

[0081] At this time, if the interval range exceeds the first preset reasonable range, it indicates that the currently set abscissa interval may not be reasonable enough, and the abscissa interval of the previous element data and / or the abscissa interval of the next element data should be modified.

[0082] Among them, the modification priorities from high to low are as follows: translate the abscissa interval of the element data whose interval range on the other side conforms to the corresponding first preset reasonable range to the other side. At this time, let the previous element data be A, the next element data be B, and the next element data of element data B be C. Assume that the interval range on the other side of the next element data B (that is, the interval range between element data B and element data C) conforms to the corresponding first preset reasonable range

[0083] Assume that the previous element data, the interval range on the other side (that is, between the previous element data of the previous element data) conforms to the first preset reasonable range, then translate the abscissa interval of this element data B to the other side (that is, the side closer to element data C). Of course, if the interval ranges on the other side of the element data on both sides of the current interval range both conform to the first preset reasonable range, the interval range that exceeds the first preset reasonable range higher on the other side can be selected for translation. In this way, corresponding adjustments can be made without reducing the abscissa interval, making the range setting more reasonable.

[0084] If the element data on both sides do not meet this condition, or after maximizing the adjustment, the current interval range still cannot meet the condition of the preset interval range, the abscissa interval of the element data whose abscissa interval conforms to the corresponding second preset reasonable range (of course, the preset reasonable range of the abscissa interval and the preset reasonable range of the interval range are different and are set in advance respectively, so they are distinguished and described by the first and the second here) can be further reduced. At this time, an upper limit for reduction can be set so that the reduced abscissa interval still conforms to the corresponding second preset reasonable range.

[0085] If it is still difficult to solve by this method, it can be reported and modified manually.

[0086] S102: Determine the text boxes in the report file, and extract the content in the abscissa interval in the text boxes based on the data extraction template.

[0087] Specifically, when actually extracting content, first parse the report file to determine the position coordinates of the text boxes in the report file.

[0088] Among the position coordinates corresponding to the text boxes, each preset ordinate interval (this preset ordinate interval is usually divided manually, or divided according to the data in the historical records) is used as a component of a single dimension information.

[0089] Take the data in the same ordinate interval as a set of data for the same dimension information. Taking the example above, the text in the interval of 70 - 90 for the Left_X_Pos coordinate is the dimension name, the text in the interval of 180 - 200 for the Right_X_Pos coordinate is the measured value of the dimension, the text in the interval of 240 - 260 for the Right_X_Pos coordinate is the measured value of the dimension, the text in the interval of 295 - 315 for the Right_X_Pos coordinate is the measured value of the dimension, and the text in the interval of 350 - 370 for the Right_X_Pos coordinate is the measured value of the dimension.

[0090] Within the ordinate interval, according to the data extraction template, extract the content in the abscissa interval through file parsing.

[0091] At this time, taking the report file as a PDF file as an example, the extracted content includes the content of each text box in the PDF file. The data includes the label content of the text box, the Y coordinate of the text box (determined by the preset ordinate interval), and the X coordinate of the text box (including Left_X_Pos and Right_X_Pos, determined by the abscissa interval). Here, PDF parsing technologies in multiple programming languages can be used, such as the python language and the C# language. When parsing fails for some PDF files using one programming language, use another language for parsing. The two language parses complement each other, which can ensure the normal parsing of the vast majority of report files. The final extraction result graph can be as shown in Table 2 below, which gives an extraction result under an example.

[0092] Table 2

[0093]

[0094] If the extraction through file parsing is successful, take the extracted content as the corresponding element data, and each abscissa interval corresponds to one element data.

[0095] If the extraction fails, perform extraction recognition through OCR and take the extracted content as the corresponding element data.

[0096] Furthermore, due to the large variety of CMM report data, there are some special cases that need to be processed. For example, for some very long data, it needs to be displayed in multiple lines. At this time, the vertical coordinate intervals (i.e., Y coordinates) of the data in the same line may be different. At this time, a corresponding uniform marking is set for each vertical coordinate interval in advance. Different vertical coordinate intervals with the same uniform marking actually belong to the same line of data. The setting of the uniform marking can be automatically set based on the repeatability of the data.

[0097] For example, during the data extraction process, real-time parsing is performed on the content. For each preset vertical coordinate interval, the content extracted from the current first vertical coordinate interval is parsed in real time to obtain a first parsing result; and the second parsing result of the previous second vertical coordinate interval is obtained; wherein, the first parsing result and the second parsing result include the parsing results of each element data in each horizontal coordinate interval.

[0098] Compare the formats of the parsing results of the element data in the same position order in the first parsing result and the second parsing result. For example, compare the element data in the first position order in the first parsing result with the element data in the first position order in the second parsing result.

[0099] If there are multiple different formats (the format can include the number of digits of the data, the position of the decimal point, the data description method, etc.) in the comparison result (the specific value of multiple here can be set according to requirements, such as set to 2 or more), it is considered that the element data in the two vertical coordinate intervals this time belongs to the same data, and it is very likely due to the data continuing to be displayed on the next line during the process. Therefore, at this time, the first vertical coordinate interval is given the same uniform marking as the second vertical coordinate interval.

[0100] After setting the uniform marking, for each preset vertical coordinate interval, obtain the uniform marking corresponding to the current first vertical coordinate interval.

[0101] If the uniform marking of the first vertical coordinate interval is the same as the uniform marking of the previous second vertical coordinate interval, and the vertical coordinate difference between the first vertical coordinate interval and the second vertical coordinate interval is lower than the preset maximum error value (for example, set to 3 to prevent mis-setting of the uniform marking), then the first vertical coordinate interval and the second vertical coordinate interval are uniformly processed to update the value of the first vertical coordinate interval to the value of the second vertical coordinate interval, that is, the first vertical coordinate interval and the second vertical coordinate interval are merged into the same vertical coordinate interval.

[0102] Among them, if there are multiple second vertical coordinate intervals, as long as the vertical coordinate difference between the first second vertical coordinate interval and the first vertical coordinate interval among these multiple second vertical coordinate intervals is lower than the preset maximum error value, the first vertical coordinate interval can be merged with these multiple second vertical coordinate intervals. If it is higher than the preset maximum error value, it can be reported for manual processing.

[0103] In addition, there may also be a situation where a text box is extracted as multiple text boxes. At this time, the Left_X_Pos coordinates can be judged and spliced to merge into one text box. The allowable coordinate difference for splicing should be less than 1.5 times the character width, and then the spliced text boxes are deleted to avoid duplicate extraction of data.

[0104] S103: Store the extracted content as corresponding element data according to the data extraction template.

[0105] For the original data parsed and obtained from the PDF file, using the types of each element data set in the data extraction template, this original data is parsed and classified. The content such as the dimension name, measured dimension value, theoretical dimension value, upper dimension tolerance, lower dimension tolerance, etc. of each dimension is parsed out in a loop, and the data is summarized and exported to a table file and imported into the database for storage.

[0106] S104: Summarize and display according to the element data.

[0107] For the element data collected in the database, data analysis is carried out and summarized for display. The qualification status of each dimension of various workpieces can be displayed, such as Figure 3 and Figure 4 as shown, line charts and bar charts are made to conveniently and intuitively see the quality status of each dimension of the workpiece, which is used to monitor the machining quality of the workpiece. When it is monitored that the quality of a certain dimension is unstable (for example, there are out-of-tolerance data, or qualified data near the limit difference frequently appears), a warning is issued, and technical and quality personnel conduct problem analysis and troubleshooting to avoid a large number of dimensions being out of tolerance.

[0108] Through the data extraction template, there is no need to manually view the PDF report page by page and manually extract data. It can automatically extract the element data in the CMM report, process a large number of reports in a short time, reduce the degree of manual labor, reduce the labor cost, speed up the data processing process, and improve the overall work efficiency.

[0109] It avoids situations such as negligence, misreading, and copying errors that may occur during manual data extraction, ensures that the extracted data is consistent with the original report, improves the accuracy and reliability of the data, and provides a more accurate basis for subsequent analysis and decision-making.

[0110] By setting corresponding data extraction templates for different enterprises and equipment brands, the corresponding extraction of element data can be achieved, thus adapting to different enterprises and equipment brands and having high versatility.

[0111] In one embodiment, during content extraction, some special situations may also be encountered. For example, the abscissa interval is set unreasonably, resulting in incomplete content extraction. Or, due to system failures or software anomalies, the font sizes of some characters change. Or, some characters are presented in the form of graphic characters and are difficult to directly recognize.

[0112] Based on this, during content extraction, start content extraction from the lower limit coordinate of the abscissa interval corresponding to the data extraction template and determine the first character extracted. Among them, the determination of characters can be carried out by real-time parsing of the file during the content extraction process.

[0113] Determine the first interval distance between the first character and the lower limit coordinate of the abscissa interval, which can be determined by the horizontal distance interval between the position of the leftmost pixel point of the first character and the lower limit coordinate.

[0114] If the first interval distance is lower than the preset distance, it means that the currently recognized first character is too close to the edge, and there may be other characters before it. Therefore, lower the lower limit coordinate of the abscissa interval, that is, move the lower limit coordinate to the left, and re-perform content extraction until the first interval distance is higher than the preset distance.

[0115] Then continue with content extraction and determine the height of each character (of course, also determine the height of the first character). If the height of a specified character is higher than the preset height, determine whether the specified character is a text character or a graphic character. Among them, graphic characters are as Figure 5 shown. The graphic character on the left represents the meaning of position tolerance, and the graphic character on the right represents the meaning of flatness. Of course, coaxiality, parallelism, perpendicularity, straightness, cylindricity, and roundness can all be set with corresponding graphic characters. And the recognition of graphic characters can be achieved using corresponding image recognition algorithms or deep learning neural networks. For example, as Figure 6 shown, some pictures of the training set for position tolerance are provided, and various form and position tolerance pictures are extracted and transformed. A total of 100,000 pictures are constructed as the training set, and 10,000 pictures are used as the test set.

[0116] If the specified character is recognized as a graphic character, it is considered reasonable that its height is relatively high. After converting the graphic character into the corresponding text character through the neural network, continue with content extraction.

[0117] If there are multiple consecutive specified characters that are all text characters, it is considered that the font size of the current character is set to be larger than the regular font size. At this time, the space it occupies in the three - coordinate report is also larger, and it is necessary to widen the abscissa interval. According to the number of specified characters, increase the lower - limit coordinate of the abscissa interval (generally speaking, no matter how large the font size of the specified characters is set, it will not be higher than the corresponding ordinate interval, so only the quantity needs to be considered. The more the quantity, the larger the increased interval), and perform content extraction.

[0118] Through content extraction, until the last character is recognized, and similar to the first character, determine the second interval distance between the last character and the upper - limit coordinate of the abscissa interval. If the first interval distance is lower than the preset distance, it is considered that there may be other characters following the last character. Similarly, increase the lower - limit coordinate of the abscissa interval and continue content extraction until the second interval distance is higher than the preset distance.

[0119] Furthermore, if there are multiple consecutive dimension information and there is the same single - element data in each dimension information, and the abscissa interval corresponding to this single - element data has been adjusted, it is considered that the setting of the abscissa interval of this single - element data is unreasonable. At this time, in the remaining dimension information, adjust the abscissa interval corresponding to this same single - element data.

[0120] Among them, when adjusting the remaining dimension information, based on the upper - limit coordinate and lower - limit coordinate adjusted in the abscissa intervals of the multiple consecutive dimension information that have been adjusted, respectively adjust the upper - limit coordinate and lower - limit coordinate of the abscissa interval of the remaining dimension information. For example, calculate the average value or maximum value of the upper - limit coordinate and lower - limit coordinate adjusted in the abscissa intervals of the multiple dimension information that have been adjusted, and use them as the adjustment content of the upper - limit coordinate and lower - limit coordinate of the abscissa interval of the remaining dimension information.

[0121] As Figure 7 shown, an embodiment of the present application also proposes a three - coordinate report extraction and analysis device for parts manufacturing, including:

[0122] At least one processor; and,

[0123] A memory communicatively connected to the at least one processor; wherein,

[0124] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the three - coordinate report extraction and analysis method for parts manufacturing as described in any of the above embodiments.

[0125] An embodiment of the present application also provides a non-volatile computer storage medium storing computer-executable instructions, where the computer-executable instructions are configured to be the method for extracting and analyzing a three-coordinate report for component manufacturing according to any one of the above embodiments.

[0126] The various embodiments in the present application are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device and the medium, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the partial description of the method embodiments for the relevant parts.

[0127] The device and the medium provided in the embodiments of the present application correspond one-to-one to the method. Therefore, the device and the medium also have beneficial technical effects similar to those of the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the device and the medium will not be elaborated here.

[0128] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the present application.

Claims

1. A three-coordinate report extraction and analysis method for parts manufacturing, characterized in that: include: Get the report file corresponding to the three-coordinate report; For a report file corresponding to a three-coordinate report, determine each element data contained in the three-coordinate report, and determine the alignment of the element data; for each element data, determine its corresponding horizontal coordinate pending interval according to its position sequence in the text box of the report file; according to the alignment of the element data, select a corresponding interval in the horizontal coordinate pending interval selection as the horizontal coordinate interval corresponding to the element data; For each element data, the corresponding upper limit coordinate and lower limit coordinate are determined according to its horizontal coordinate interval; according to the position sequence of the element data, the interval range between the element data is determined through the upper limit coordinate of the previous element data and the lower limit coordinate of the next element data; for the interval range, according to the alignment methods corresponding to the element data on both sides thereof, the first preset reasonable range corresponding to the interval range is determined; if the interval range exceeds the first preset reasonable range, the horizontal coordinate interval of the previous element data and / or the horizontal coordinate interval of the next element data is modified; Determining a data extraction template corresponding to the report file; The data extraction template at least includes the horizontal axis interval corresponding to each element data in the report file; Determine a text box in the report file, and extract the content in the horizontal axis interval in the text box based on the data extraction template; According to the data extraction template, the extracted content is stored as corresponding element data; The element data are summarized and displayed.

2. The three-coordinate report extraction and analysis method for parts manufacturing according to claim 1 is characterized in that: The priority of modification is from high to low as follows: shift the horizontal axis interval of the element data on the other side so that the interval range matches the corresponding first preset reasonable range to the other side; reduce the horizontal axis interval so that the horizontal axis interval matches the corresponding second preset reasonable range of the element data; report and modify manually.

3. The three-coordinate report extraction and analysis method for parts manufacturing according to claim 1 is characterized in that: Determine a text box in the report file, and extract the content in the horizontal axis interval in the text box based on the data extraction template, specifically including: Parsing the report file to determine the position coordinates of the text box in the report file; In the position coordinates corresponding to the text box, each preset vertical coordinate interval is used as a component of a single size information, and within the vertical coordinate interval, the content in the horizontal coordinate interval is extracted through file parsing according to the data extraction template; If the extraction is successful, the extracted content is used as the corresponding element data; If the extraction fails, the OCR recognition is used to extract and recognize the content, and the extracted content is used as the corresponding element data.

4. The three-coordinate report extraction and analysis method for parts manufacturing according to claim 3 is characterized in that: Each preset vertical axis interval is used as a component of a single size information, specifically including: For each preset vertical coordinate interval, obtain the uniform mark corresponding to the current first vertical coordinate interval; If the consistency mark of the first vertical coordinate interval is the same as the consistency mark of the previous second vertical coordinate interval, and the vertical coordinate difference between the first vertical coordinate interval and the second vertical coordinate interval is lower than the preset maximum error value, the first vertical coordinate interval and the second vertical coordinate interval are consistent with each other to update the value of the first vertical coordinate interval to the value of the second vertical coordinate interval.

5. The three-coordinate report extraction and analysis method for parts manufacturing according to claim 4 is characterized in that: For each preset vertical coordinate interval, obtaining the uniform mark corresponding to the current first vertical coordinate interval specifically includes: For each preset vertical axis interval, the content extracted from the current first vertical axis interval is parsed in real time to obtain a first parsing result; and a second parsing result of the previous second vertical axis interval is obtained; wherein the first parsing result and the second parsing result include the parsing results of each element data in each horizontal axis interval; Comparing the formats of the parsing results of the element data in the same position sequence in the first parsing result and the second parsing result; If there are multiple different formats in the comparison result, the first vertical axis interval is assigned the same uniform mark as the second vertical axis interval.

6. The three-coordinate report extraction and analysis method for parts manufacturing according to claim 1 is characterized in that: Extracting the content in the horizontal axis interval in the text box based on the data extraction template specifically includes: Starting content extraction from the lower limit coordinate of the horizontal axis interval corresponding to the data extraction template, and determining the first character extracted; Determine a first spacing distance between the first character and the lower limit coordinate of the horizontal axis interval; If the first interval distance is lower than the preset distance, lowering the lower limit coordinate of the horizontal axis interval, and re-extracting the content until the first interval distance is higher than the preset distance; Continue with content extraction and determine the height of each character; If there is a designated character whose height is higher than the preset height, determining whether the designated character is a text character or a graphic character; If the designated character is a graphic character, convert the graphic character into a corresponding text character and then continue to extract the content; If there are multiple consecutive designated characters that are all text characters, then according to the number of the designated characters, the lower limit coordinate of the horizontal axis interval is increased, and content extraction is performed; Extracting the content until the last character is identified, and determining a second interval distance between the last character and the upper limit coordinate of the horizontal axis interval; If the first interval distance is lower than a preset distance, the lower limit coordinate of the horizontal axis interval is increased, and content extraction is continued until the second interval distance is higher than a preset distance.

7. The three-coordinate report extraction and analysis method for parts manufacturing according to claim 6 is characterized in that: The method further comprises: If there are multiple consecutive size information, and the same single element data exists in each size information, and the horizontal axis interval corresponding to the single element data is adjusted, then in the remaining size information, the horizontal axis interval corresponding to the same single element data is adjusted; Here, based on the adjusted upper limit coordinates and lower limit coordinates in the horizontal axis intervals of a plurality of continuous size information that have been adjusted, the upper limit coordinates and lower limit coordinates of the horizontal axis intervals of the remaining size information are adjusted respectively.

8. A three-coordinate report extraction and analysis device for parts manufacturing, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the three-coordinate report extraction and analysis method for component manufacturing as described in any one of claims 1 to 7.

9. A non-volatile computer storage medium storing computer executable instructions, characterized in that: The computer executable instructions are configured as: a three-coordinate report extraction and analysis method for component manufacturing as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Product quality data processing method and device, computer equipment and storage medium

    CN118260240A