Automatic radiographic inspection report checking and analyzing method and system

Through the automated ray detection report verification and analysis method, the PaddleOCR model and static database comparison technical standards are used to solve the problem of inefficient manual verification, and efficient and accurate ray detection report verification is achieved.

CN120579533APending Publication Date: 2025-09-02CHONGQING UNIV +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510717247.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The verification of existing ray detection reports depends on manual comparison, is inefficient and prone to errors, and is difficult to meet the verification requirements of complex and diverse parameters and various technical standards.

Method used

The automated method is adopted, and the PaddleOCR model is used to extract key information, and the comparison is combined with a static database is performed to generate an analysis report, supporting a variety of technical standards and specifications, and equipment adjustment is performed through the interactive interface.

Benefits of technology

It improves the work efficiency of radiological detection report verification, reduces human errors, ensures the accuracy and consistency of verification, and enhances the scope of application and flexibility of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579533A_ABST
    Figure CN120579533A_ABST
Patent Text Reader

Abstract

The invention provides an automatic radiographic inspection report checking and analyzing method and system, and the method comprises the steps: obtaining a report file, and carrying out the preprocessing of the report file, and obtaining a processed file; based on the processing file, key information is extracted through an identification model based on a PaddleOCR model; obtaining a static database, and comparing the key information with the static database to obtain comparison information; and generating an analysis report according to the comparison information, and performing equipment adjustment through the analysis report. According to the invention, the working efficiency of ray detection report checking is obviously improved, checking errors caused by human factors are reduced, and the accuracy and consistency of checking work are ensured. Meanwhile, various technical standards and specifications are supported, and the application range and flexibility of the system are improved. And through a visual user interaction interface, user operation is facilitated. And a splitting tool is provided, so that formatting of the input file is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of radiation report detection and analysis in the nuclear power field, and in particular to an automatic radiation detection report verification and analysis method and system. Background Art

[0002] In the nuclear power sector, radiographic testing is a key nondestructive testing technology for ensuring weld quality and safeguarding the safe operation of nuclear power equipment. Radiographic testing reports serve as a crucial basis for evaluating the quality and compliance of testing operations, and their accuracy and completeness are crucial to the overall safety and reliability of nuclear power projects. However, the verification of radiographic testing reports currently relies primarily on manual comparison and verification, a method that is inefficient and prone to errors when faced with complex and diverse parameters, cumbersome calculations, and numerous technical standards.

[0003] Specifically, radiographic inspection reports involve a wide variety of parameters, including but not limited to the type of X-ray source, exposure time, focal length, inspection location, defect type, and size. The accuracy and compliance of each parameter must be individually verified against specific technical specifications (such as 1516AT0012) and industry standards (such as RCCM, NB-20001, and GB-3323). This manual verification method is not only time-consuming and labor-intensive, but also susceptible to human factors, leading to uncertainty in the verification results and increased risk. Summary of the Invention

[0004] Based on this, it is necessary to provide an automatic X-ray detection report verification and analysis method and system to address the above technical issues.

[0005] An automatic radiographic inspection report verification and analysis method comprises the following steps:

[0006] Obtaining a report file, and preprocessing the report file to obtain a processed file;

[0007] Extracting key information based on the processed file through a recognition model based on the PaddleOCR model;

[0008] Obtaining a static database, and comparing the key information with the static database to obtain comparison information;

[0009] An analysis report is generated based on the comparison information, and equipment adjustments are performed based on the analysis report.

[0010] In one embodiment, generating an analysis report based on the comparison information, and adjusting the device based on the analysis report further includes:

[0011] A report splitting message is received, and report splitting is performed based on the analysis report to obtain a split report.

[0012] In one embodiment, obtaining a report file and preprocessing the report file to obtain a processed file includes:

[0013] Converting the report file into an image format to obtain an initial image;

[0014] Obtaining a text recognition result based on the initial image using a text recognition tool;

[0015] The position of the initial image is adjusted according to the text recognition result to obtain a processed file.

[0016] In one embodiment, before extracting key information based on the processed file by a recognition model based on a PaddleOCR model, the method further includes:

[0017] Acquire a prepared data set, and preprocess the prepared data set to obtain a preprocessed data set;

[0018] The preprocessed data set is labeled by PPOCRLabel to obtain a data set, and the data set is divided into a training set, a test set, and a validation set according to a preset ratio;

[0019] Training the PaddleOCR model using the training set, the test set, and the validation set to obtain a training model;

[0020] The training model is evaluated, and in response to the evaluation index of the training model reaching a preset index, a recognition model is obtained.

[0021] In one embodiment, extracting key information based on the processed file using a recognition model based on a PaddleOCR model includes:

[0022] Performing initialization preprocessing on the processing file to obtain a standardized image;

[0023] Based on the standardized image, feature extraction is performed using the DB algorithm to predict the text bounding box and obtain the text detection area;

[0024] Perform text extraction processing on the text detection area using the SVTR_LCNet algorithm to obtain a text list;

[0025] Using the TableAttn ​​algorithm and the attention mechanism, the table cells are located and aligned across modalities according to the text list to obtain text data;

[0026] The key information is obtained by fusing the visual layout with the text data through the LayoutXLM algorithm.

[0027] In one embodiment, performing initialization preprocessing on the processed file to obtain a standardized image includes:

[0028] Obtaining a preset optimal resolution, adjusting the processed file according to the preset optimal resolution, and obtaining a display image;

[0029] Performing perspective transformation on the displayed image using a tilt correction algorithm to eliminate image tilt and adjusting the image size according to a preset size requirement to obtain a corrected image;

[0030] The corrected image is subjected to grayscale processing, Gaussian blur processing, adaptive threshold adjustment and histogram equalization adjustment to obtain a standardized image.

[0031] In one embodiment, a static database is obtained, and the key information is compared with the static database to obtain comparison information including:

[0032] The static database includes: RCCM standard, NB-20001 standard, GB-3323 standard and 1516AT0012 technical specification.

[0033] In one embodiment, obtaining a static database, comparing the key information with the static database, and obtaining the comparison information further comprises:

[0034] Performing preliminary format and value constraint verification on the key information, and obtaining key parameters in response to the key information meeting preset standards;

[0035] receiving a static database selection message, and selecting a static database according to the static database selection message;

[0036] Comparing the key parameters with the parameters in the static database one by one, and in response to the key parameters meeting the standard requirements, recording the key parameters that do not meet the tolerance range as abnormal parameters, recording the error type, standard required value and actual extracted value of the abnormal parameters, and retaining non-abnormal parameters;

[0037] According to the parameter dependency graph, the abnormal parameters are replaced with verified correct preset associated parameters and then compared one by one with the parameters in the static database, and an abnormal comparison report is generated based on the error type of the abnormal parameter, the standard required value, the actual extracted value and the non-abnormal parameter;

[0038] In response to the key parameters meeting the standard requirements, retaining the key parameters and obtaining a normal comparison report;

[0039] The comparison information consists of the abnormal comparison report and the normal comparison report.

[0040] In one embodiment, generating an analysis report based on the comparison information, and adjusting the equipment based on the analysis report includes:

[0041] An analysis report is generated by a document generation tool according to the comparison information, and equipment adjustments are performed based on the analysis report.

[0042] An automatic radiographic inspection report verification and analysis system, for implementing the above-mentioned automatic radiographic inspection report verification and analysis method, comprises:

[0043] A report processing module is used to obtain a report file, pre-process the report file, and obtain a processed file;

[0044] An information extraction module, configured to extract key information based on the processed file through a recognition model based on a PaddleOCR model;

[0045] An information comparison module is used to obtain a static database, compare the key information with the static database, and obtain comparison information;

[0046] A report generation module is used to generate an analysis report based on the comparison information and perform equipment adjustments based on the analysis report.

[0047] Compared to existing technologies, the advantages and benefits of this invention include: significantly improving the efficiency of radiographic inspection report verification, reducing verification errors caused by human factors, and ensuring the accuracy and consistency of verification work. It also supports multiple technical standards and specifications, increasing the system's applicability and flexibility. Its intuitive user interface facilitates user operation, and it provides a splitting tool for easy formatting of input files. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 A schematic flow chart of an automatic radiographic inspection report verification and analysis method according to an embodiment;

[0049] Figure 2 A schematic diagram summarizing the principle of a comparison algorithm in one embodiment;

[0050] Figure 3 A schematic diagram of a key parameter comparison process in an embodiment;

[0051] Figure 4 A schematic diagram of a split visualization interface in one embodiment;

[0052] Figure 5 It is a structural schematic diagram of an automatic radiographic detection report verification and analysis system in one embodiment;

[0053] Figure 6 A schematic diagram of a visualization interface in one embodiment;

[0054] Figure 7 A schematic diagram of the detection results in one embodiment;

[0055] Figure 8 A pie-shaped schematic diagram of the test results in one embodiment;

[0056] Figure 9 Schematic diagram of the test results in one embodiment. DETAILED DESCRIPTION

[0057] Before describing the specific embodiments of the present invention, the overall concept of the present invention is described as follows:

[0058] The present invention is mainly developed for the X-ray report detection process. Currently, the verification of X-ray detection reports mainly relies on manual comparison and verification. This method is inefficient and prone to errors when faced with complex and diverse parameters, tedious calculation processes and numerous technical standards.

[0059] Therefore, the present invention proposes an automatic X-ray inspection report verification and analysis method and system, which automatically outputs the analysis results of all reports in the folder and detailed analysis information of each report by inputting a folder of PDF versions of reports, so as to verify whether the report content and inspection results meet the requirements of technical specifications and industry standards.

[0060] After introducing the overall concept of the present invention, in order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below through specific implementation methods in conjunction with the accompanying drawings.

[0061] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in one or more embodiments of this specification should have the usual meanings understood by people with ordinary skills in the field to which the invention belongs. The words "first", "second" and similar terms used in one or more implementations of this specification do not indicate any order, quantity or importance, but are only used to distinguish different components. Words such as "include" or "comprise" mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Words such as "connect" or "connected" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0062] In one embodiment, Figure 1 As shown, a method for automatically checking and analyzing radiographic inspection reports is provided, comprising the following steps:

[0063] Step S101: Acquire a report file, and pre-process the report file to obtain a processed file.

[0064] Specifically, the user inputs a folder of scanned PDF files through an interactive interface, obtains the PDF files in the folder as report files, pre-processes the report files, and converts them into processed files in image format.

[0065] On this basis, a report file is obtained and pre-processed to obtain a processed file including:

[0066] Converting the report file into an image format to obtain an initial image;

[0067] Obtaining a text recognition result based on the initial image using a text recognition tool;

[0068] The position of the initial image is adjusted according to the text recognition result to obtain a processed file.

[0069] Specifically, after converting the PDF report file into an image using a file format conversion tool, the text is recognized using a text recognition tool to obtain a text recognition result. Based on the text recognition result, the image is then rotated and corrected using the affine transformation function of OpenCV to improve the accuracy of subsequent OCR recognition. In this embodiment, the text recognition tool is PaddleHub, but the present invention is not limited to PaddleHub. All text recognition tools that can achieve the above functions are within the scope of the present invention.

[0070] Step S102: extract key information based on the processed file through a recognition model based on the PaddleOCR model.

[0071] Specifically, a self-trained recognition model based on the PaddleOCR model is used to automatically identify and extract key information from the processed files, such as report number, detection range, and type of radiation source.

[0072] On this basis, before extracting key information based on the processed file through the recognition model based on the PaddleOCR model, it also includes:

[0073] Acquire a prepared data set, and preprocess the prepared data set to obtain a preprocessed data set;

[0074] The preprocessed data set is labeled by PPOCRLabel to obtain a data set, and the data set is divided into a training set, a test set, and a validation set according to a preset ratio;

[0075] Training the PaddleOCR model using the training set, the test set, and the validation set to obtain a training model;

[0076] The training model is evaluated, and in response to the evaluation index of the training model reaching a preset index, a recognition model is obtained.

[0077] Specifically, a preparation data set is obtained, which is an existing report file. The preparation data set is preprocessed to convert the report file into an image, and then the tilted image in the data set is preprocessed using an opencv-based affine transformation to obtain a preprocessed data set.

[0078] Data is labeled using PPOCRLabel. After labeling, the data is divided into training set, test set, and validation set according to a preset ratio. In this embodiment, the ratio of training set, test set, and validation set is 6:2:2.

[0079] After downloading the official PaddleOCR pre-trained model, perform model training to obtain the trained model. The training evaluation metrics for the trained model are precision, recall, and F1 score. The specific preset metrics are as follows: Precision: 0.98851, Recall: 0.93478, F1 score: 0.96089. When the evaluation metrics of the trained model meet these preset metrics, the current trained model is used as the recognition model.

[0080] On this basis, the key information extracted from the processed file through the recognition model based on the PaddleOCR model includes:

[0081] Performing initialization preprocessing on the processing file to obtain a standardized image;

[0082] Based on the standardized image, feature extraction is performed using the DB algorithm to predict the text bounding box and obtain the text detection area;

[0083] Perform text extraction processing on the text detection area using the SVTR_LCNet algorithm to obtain a text list;

[0084] Using the TableAttn ​​algorithm and the attention mechanism, the table cells are located and aligned across modalities according to the text list to obtain text data;

[0085] The key information is obtained by fusing the visual layout with the text data through the LayoutXLM algorithm.

[0086] Specifically, the initialization preprocessing of the processing file is first performed, including image loading, image size adjustment and standardization, to ensure that the image is suitable for subsequent processing, and a standardized image is obtained after preprocessing.

[0087] Then, based on the standardized image, the DB algorithm is used for feature extraction, predicting text bounding boxes, and performing necessary post-processing operations to obtain the text detection area. The DB algorithm locates text by predicting a probability map and a threshold map for the text area. The probability map identifies the probability that each pixel belongs to text, while the threshold map provides a dynamic threshold to distinguish text from background. The DB algorithm adaptively determines the binarization threshold for each location to generate clear text boundaries. In addition, the DB algorithm also performs a series of post-processing operations, such as non-maximum suppression (NMS) and connected component analysis, to remove redundant bounding boxes and merge scattered text blocks, ultimately generating an accurate text region bounding box to obtain the text detection area. If rotated text is present, an additional text orientation classifier is required to identify the orientation of the text for correct subsequent processing.

[0088] Specifically, the DB algorithm first uses deep convolutional neural networks (such as ResNet, MobileNetV3, etc.), whose core function is to encode the input image into a multi-scale feature map to provide semantic and spatial information support for subsequent text detection.

[0089] Taking ResNet-50 as an example, the backbone network outputs four levels of feature maps:

[0090] C2 (1 / 4 resolution): Contains rich edge and texture information and is suitable for detecting small-sized text

[0091] C3 (1 / 8 resolution): captures medium-scale text features

[0092] C4 (1 / 16 resolution): Extract semantic features of large text regions

[0093] C5 (1 / 32 resolution): high-level semantic features, but with large loss of spatial information

[0094] Then, the feature pyramid is constructed and fused through upsampling and element-by-element addition operations:

[0095] P5=Conv(C5)

[0096] P4=UpSample(P5)+Conv(C4)

[0097] P3=UpSample(P4)+Conv(C3)

[0098] P2=UpSample(P3)+Conv(C2)

[0099] The fused feature maps (P2-P5) have the following characteristics: high-resolution features retain local details (P2), mid- and high-level features capture global semantics (P3-P5), and each layer of features shares the same number of channels (usually P2, P5, and P6).

[0100] Based on the above feature pyramid, the DB algorithm generates two key feature maps - probability map and threshold map through two parallel 1x1 convolutional layers.

[0101] Define the probability distribution of text regions:

[0102] P i,j =σ(Conv prob (F i,j ))

[0103] Among them, F i,j The feature vector representing the feature map coordinate (i, j) has an output value range of [0, 1], which indicates the probability that the pixel belongs to text.

[0104] Dynamically adjust the binarization threshold:

[0105] T i,j =Conv thresh (F i,j )+β

[0106] Among them, β represents a learnable bias term, the output value range is unrestricted, and is subsequently constrained by differentiable binarization.

[0107] Then, end-to-end training is achieved by fusing the probability map and the threshold map. The core formula is:

[0108]

[0109] Where k represents the amplification factor (the default is 50), which can make the curve approximate a step function. i,j Represents the final binarization result (continuously differentiable between 0 and 1).

[0110] Then comes the post-processing process, which mainly includes text region generation, non-maximum suppression (NMS), and bounding box optimization.

[0111] Text area generation, using OpenCV's findContours algorithm to extract connected areas, and using the minimum bounding rectangle to generate candidate boxes.

[0112] Non-maximum suppression (NMS) uses an improved local-sensitive NMS algorithm to determine whether two rotated text boxes overlap excessively. This is only used in rotated box NMS and takes the text box orientation into account when calculating IoU:

[0113]

[0114] Finally, the candidate box with the highest score is retained, and the threshold is usually set to 0.5.

[0115] Bounding box optimization is achieved by adjusting the polygon box through the Vatti clipping algorithm:

[0116] Δ=d×r

[0117] Where d is the original bounding box size and r is the expansion ratio (which can be controlled by the det_db_unclip_ratio parameter, default is 2.0).

[0118] The detected text detection area is then cropped using the SVTR_LCNet algorithm, and features are further extracted and a sequence model is constructed to obtain a text list. The SVTR_LCNet algorithm first fine-tunes the detected text area to ensure that each text line or word is processed separately. A lightweight LCNet network is then used to extract features from the image. Subsequently, the extracted features are processed using sequence modeling techniques (such as the Transformer architecture) to construct a model that can effectively recognize character sequences. The sequence modeling stage utilizes a sequence-to-sequence learning mechanism, enabling the model to understand the context of the text and convert the characters in the image into readable text strings.

[0119] Get the coordinates of the text detection area and use geometric correction (affine transformation) to correct the tilted text:

[0120]

[0121] Among them, (x, y) represents the original image pixel coordinates, (x', y') represents the transformed coordinates, a, b, d, e represent linear transformation parameters (control rotation, scaling, shearing), and c and f represent translation parameters.

[0122] The matrix parameters are optimized by the least square method to minimize the angle between the text region boundary and the horizontal axis.

[0123] Depthwise Separable Convolution is used for lightweight feature extraction (LCNet) to build a lightweight backbone network. The computational complexity is:

[0124] C depthwise =K 2 ·C in ·H·W

[0125] C pointwise =C in ·C out ·H·W

[0126] Among them, K represents the convolution kernel size, C in Indicates the number of input feature channels, C out Represents the number of output feature channels, H and W represent the height and width of the feature map.

[0127] Compared with the standard convolution calculation (K 2 ·C in ·C out ), the depth-wise separable convolution reduces about C out / K 2 times the computational effort.

[0128] Sequence modeling by introducing the axial attention mechanism (SVTR improved architecture):

[0129]

[0130] Among them, Q, K and V represent query, key and value matrices, and D k represents the scaling factor, Q row , Q col Represents row-wise query projection and column-wise query projection respectively, and ⊕ represents broadcast addition operation. This algorithm can reduce the complexity of the original attention mechanism to O((HW) 2 ) is reduced to O(HW(H+W)).

[0131] Use the TableAttn ​​algorithm to process data from text lists. TableAttn ​​uses an attention mechanism to precisely locate cells in a table and extract data from them. This algorithm effectively handles row and column alignment issues in tables. Even with complex or irregular table structures, it can accurately identify and extract information from the table, ensuring data integrity and accuracy.

[0132] The TableAttn ​​algorithm mainly includes: table structure modeling, cell positioning, and cross-modal alignment.

[0133] Table structure modeling uses a two-way attention mechanism to construct a row-column relationship matrix:

[0134]

[0135] Among them, Q row represents the row-wise query vector, K col represents the column direction key vector, P geo Represents a geometric position encoding item.

[0136] Cell localization,boundary regression based on deformable convolution:

[0137]

[0138] Among them, p k represents the preset anchor point offset, and Δpk represents the learned deformation offset.

[0139] Cross-modal alignment, jointly optimizing text content and table structure:

[0140]

[0141] Among them, E text represents the text embedding vector, E struct Represents the structure embedding vector.

[0142] Finally, the LayoutXLM algorithm was used to accurately capture the key information in the report. LayoutXLM combines visual layout information with textual content, using the Transformer architecture to understand and parse the document's overall structure. This algorithm identifies the location and relationships between different text blocks within a document, enabling more accurate extraction of key information and ensuring comprehensive and accurate information.

[0143] The core of the LayoutXLM algorithm includes: multimodal fusion, layout-aware attention, and information importance prediction.

[0144] Multimodal fusion, visual-text joint encoding:

[0145] H (0) =[E text ||E visual ||PE layout ]W proj

[0146] Among them, E text represents the text embedding vector, E visual Represents visual feature vector (CNN extraction), PE layout Represents the layout position encoding vector.

[0147] Layout-aware attention, improved attention weight calculation:

[0148]

[0149] Among them, L i =(x i ,y i ,w i ,h i ) represents the text block coordinates, and α=0.3 represents the layout weight factor.

[0150] Information importance prediction and key field identification:

[0151] P key (s)=σ(W kAvgPool(H (L) [s]))

[0152] Among them, W k represents the learnable weight matrix, H (L) Represents the L-th layer Transformer output.

[0153] On this basis, the processing file is initialized and preprocessed to obtain a standardized image including:

[0154] Obtaining a preset optimal resolution, adjusting the processed file according to the preset optimal resolution, and obtaining a display image;

[0155] Performing perspective transformation on the displayed image using a tilt correction algorithm to eliminate image tilt and adjusting the image size according to a preset size requirement to obtain a corrected image;

[0156] The corrected image is subjected to grayscale processing, Gaussian blur processing, adaptive threshold adjustment and histogram equalization adjustment to obtain a standardized image.

[0157] Specifically, we first processed the image loading process. The dpi parameter in the convert_from_path function of the pdf2image library controlled the output quality when converting PDF to image. This parameter directly impacted the subsequent OCR accuracy. After multiple tests, we ultimately selected the optimal preset resolution of 300 dpi. Furthermore, we introduced a memory optimization mechanism. By enabling the Intel math library through the mkldnn provided by PaddleHub, this optimization increased image loading speed by 2-3 times.

[0158] The image is then resized and normalized. The calculate_average_angle function in the tilt correction algorithm calculates the rotation angle using the text box coordinates. By using atan2 to calculate the slope angle between two points, the computational overhead of the traditional Hough transform is avoided. The average angle strategy outperforms extreme angles and is more robust to local distortions. Experimental data shows that this algorithm achieves 98% accuracy within a ±15° tilt range.

[0159] Finally, the processing is performed according to the following process: A[original image]-->B[grayscale]-->C[Gaussian blur]-->D[adaptive threshold]-->E[histogram equalization].

[0160] Compared to traditional Z-Score normalization, this solution can adaptively adjust the threshold to overcome uneven illumination (better than global thresholding) and retain more text details through CLAHE histogram equalization, making it more suitable for this OCR scan scenario. In addition, to improve the accuracy of subsequent OCR recognition, a post-processing hybrid enhancement strategy is introduced, as shown in the following table:

[0161] Table 1 Post-processing hybrid enhancement strategy

[0162]

[0163]

[0164] Step S103: Acquire a static database, compare the key information with the static database, and obtain comparison information.

[0165] Specifically, it is used to compare and analyze the extracted key information with the data in the locally built static database.

[0166] On this basis, a static database is obtained, and the key information is compared with the static database to obtain comparison information including:

[0167] The static database includes: RCCM standard, NB-20001 standard, GB-3323 standard and 1516AT0012 technical specification.

[0168] Specifically, the static database contains a rule engine library of multiple technical standards and specifications, including but not limited to RCCM, NB-20001, GB-3323 and other standards and 1516AT0012 technical specifications.

[0169] To import standards such as RCCM, NB-20001, and GB-3323 into a local static database, we used the lightweight SQLite database, eliminating the need for server configuration. We abstracted the relevant standard requirements into a two-dimensional table and imported them into the static database whenever possible. For standards that could not be abstracted into a two-dimensional table, we symbolized them using sympy and other tools. This eliminated the need to search for relevant standards in the static database for parameter comparison. The static database is automatically constructed during the initial implementation of this method. Even if the database is deleted for other reasons, it will be reinitialized during the parameter comparison process. If the static database already exists, it does not need to be rebuilt.

[0170] On this basis, a static database is obtained, and the key information is compared with the static database to obtain the comparison information, which also includes:

[0171] Performing preliminary format and value constraint verification on the key information, and obtaining key parameters in response to the key information meeting preset standards;

[0172] receiving a static database selection message, and selecting a static database according to the static database selection message;

[0173] Comparing the key parameters with the parameters in the static database one by one, and in response to the key parameters meeting the standard requirements, recording the key parameters that do not meet the tolerance range as abnormal parameters, recording the error type, standard required value and actual extracted value of the abnormal parameters, and retaining non-abnormal parameters;

[0174] According to the parameter dependency graph, the abnormal parameters are replaced with verified correct preset associated parameters and then compared one by one with the parameters in the static database, and an abnormal comparison report is generated based on the error type of the abnormal parameter, the standard required value, the actual extracted value and the non-abnormal parameter;

[0175] In response to the key parameters meeting the standard requirements, retaining the key parameters and obtaining a normal comparison report;

[0176] The comparison letter consists of the abnormal comparison report and the normal comparison report.

[0177] Specifically, such as Figure 2 As shown in the interactive interface, first obtain the key information. In order to prevent the inaccurate recognition accuracy from causing the extracted key parameter values ​​to be incorrect and affecting the correct judgment of the key parameters, all key information will be preliminarily checked. The main purpose is to check whether the parameter values ​​of the key information meet the standard format requirements or value constraints. If not, it will return a data parsing failure, carrying the key information with the recognition error; if it passes the preliminary check, it will enter the verification stage of the key parameter algorithm, and the key parameters will be compared one by one. When performing a standard check on a key parameter, use the sqlite3 module to extract relevant standards such as NB-20001 from the local static database, and then calculate or obtain the correct results in the standard requirements based on the relevant parameters of the comparison parameters (reference values ​​of corresponding key parameters in the static database), and then compare with the extracted key parameters. If the extracted value meets the standard requirements, the next key parameter will be compared. The process is as follows: Figure 3 If all key parameters meet the standard requirements, the report is correct; if one or more key parameters do not meet the standard requirements, the report is abnormal.

[0178] For key parameters that need to participate in calculations, the numbers module is mainly used to determine whether the extracted value is a numeric value. When the key parameter is a numeric value, it can participate in calculations and comparisons with comparison parameters normally. For specification-type key parameters, regular matching is mainly used to ensure the correctness of the data format and ensure that the data meets the segmentation standards. For related key parameters whose parameter values ​​can only be specific values, value constraints are added to them. When the extracted parameter value is not a specific value, data parsing will fail to ensure that it can be correctly compared with the standard in the static database.

[0179] Related key information can influence each other. To minimize the correlation between key information and the potential for a series of errors caused by a single key information error, when a key parameter does not meet standard requirements, the non-compliant parameter value will be replaced with a parameter value calculated using the correct key information. For example, if the geometry is not clear, it will not directly affect the pass / fail status, but will affect the selection criteria of another key parameter. For key parameters that do not meet standard requirements, the abnormal parameter and the reason for the abnormality will be returned.

[0180] Step S104: generating an analysis report based on the comparison information, and performing report analysis based on the analysis report.

[0181] Specifically, an analysis report is generated by a document generation tool according to the obtained comparison information, and equipment adjustments are performed based on the analysis report.

[0182] On this basis, an analysis report is generated according to the comparison information, and equipment adjustment is performed based on the analysis report, including:

[0183] An analysis report is generated by a document generation tool according to the comparison information, and equipment adjustments are performed based on the analysis report.

[0184] In this embodiment, the document generation tool is python-docx. Python-docx is used to generate a radiographic inspection analysis report based on the comparison information, allowing the user to make relevant adjustments to the equipment based on abnormal parameters and causes of the abnormality.

[0185] The key parameters meet the standard requirements, the key parameters are retained, and a normal comparison report is obtained. The comparison information is composed of the abnormal comparison report and the normal comparison report.

[0186] On this basis, an analysis report is generated according to the comparison information, and after the equipment is adjusted according to the analysis report, the following steps are further included:

[0187] A report splitting message is received, and report splitting is performed based on the analysis report to obtain a split report.

[0188] Specifically, such as Figure 4 In the interactive interface shown, when there are multiple reports in an analysis report, a custom-configured report splitting message is received, and the report is split to obtain a split report.

[0189] The present invention provides an automatic X-ray detection report verification and analysis method, which restores the file layout format to the greatest extent by calculating the coordinate relationship of the text detection box recognized by OCR, so as to facilitate subsequent data extraction. It can intelligently match the corresponding technical standards to achieve efficient comparison, display the OCR extraction results through a visual interface, automatically mark the extracted possible abnormal data in red, and provide a one-click opening function for the corresponding report to facilitate correction. The user interaction interface is simple and intuitive, and supports rapid positioning of a report and correction of anomalies. The present invention significantly improves the work efficiency of X-ray detection report verification, reduces verification errors caused by human factors, and ensures the accuracy and consistency of verification work. It also supports multiple technical standards and specifications, which improves the scope of application and flexibility of the system. Through an intuitive user interaction interface, user operation is convenient. A splitting tool is provided to facilitate formatting of input files.

[0190] It should be noted that the method of the embodiment of the present invention can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied in a distributed scenario, where multiple devices cooperate to perform the method. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiment of the present invention, and the multiple devices will interact with each other to complete the method.

[0191] It should be noted that the above description is limited to some embodiments of the present invention. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0192] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present invention also provides an automatic radiographic detection report verification and analysis system.

[0193] refer to Figure 5 The automatic radiographic detection report verification and analysis system comprises:

[0194] The report processing module 501 is used to obtain a report file and pre-process the report file to obtain a processed file;

[0195] An information extraction module 502 is configured to extract key information based on the processed file through a recognition model based on a PaddleOCR model;

[0196] An information comparison module 503 is used to obtain a static database, compare the key information with the static database, and obtain comparison information;

[0197] The report generation module 504 is configured to generate an analysis report based on the comparison information and perform equipment adjustments based on the analysis report.

[0198] For the convenience of description, the above system is described as being divided into various modules according to their functions. Of course, when implementing the present invention, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0199] The system of the above embodiment is used to implement a corresponding automatic radiographic detection report verification and analysis method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be repeated here.

[0200] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions of the present invention. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0201] Example 1

[0202] The specific interactive interface is as follows Figure 6 As shown, the user Figure 6 Enter the content in the edit boxes below I "Operator Name" and J "New Task Name", then click K to create the task output directory, then select L execution standard and M evaluation standard, then click N to open the report directory, select the PDF folder to obtain the report file, and click O to start processing the extracted data (P means you can stop processing). H is the system status, which will show whether the system is currently working. After the extraction process is completed, a pop-up window will appear. Figure 4 Interface. You can correct the wrong data by following this visual interface (shortcut keys Ctrl+E, Ctrl+S, Ctrl+Left, Ctrl+Right can respectively realize edit, save, previous copy, next copy operation). After the correction is completed, you can return to Figure 6 Interface, click Q data details to quickly locate the folder selected in step N, Q is abnormal view to quickly view the test results (such as Figure 7 ), Q data statistics can view the bar chart and pie chart information of this test result (such as Figure 8 、 Figure 9 ). R is a PDF splitting tool that allows users to customize the splitting of PDF files for input into the system.

[0203] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present invention (including the claims) is limited to these examples. Within the scope of the present invention, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present invention as described above, which are not provided in detail for the sake of simplicity.

[0204] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0205] In addition, while specific details are set forth to describe exemplary embodiments of the present invention, it will be apparent to those skilled in the art that embodiments of the present invention may be practiced without or with variations in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0206] While the invention has been described in conjunction with specific embodiments thereof, many alternatives, modifications and variations of these embodiments will be apparent to those skilled in the art in light of the foregoing description.

[0207] The embodiments of the present invention are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for automatically checking and analyzing radiographic inspection reports, characterized in that: include: Obtaining a report file, and preprocessing the report file to obtain a processed file; Extracting key information based on the processed file through a recognition model based on the PaddleOCR model; Obtaining a static database, and comparing the key information with the static database to obtain comparison information; An analysis report is generated based on the comparison information, and equipment adjustments are performed based on the analysis report.

2. The method for automatic radiographic inspection report verification and analysis according to claim 1, characterized in that: After generating an analysis report based on the comparison information and adjusting the equipment based on the analysis report, the following further comprises: A report splitting message is received, and report splitting is performed based on the analysis report to obtain a split report.

3. The method for automatic radiographic inspection report verification and analysis according to claim 1, characterized in that: The obtaining of the report file and preprocessing the report file to obtain the processed file include: Converting the report file into an image format to obtain an initial image; Obtaining a text recognition result based on the initial image using a text recognition tool; The position of the initial image is adjusted according to the text recognition result to obtain a processed file.

4. The method for automatic radiographic inspection report verification and analysis according to claim 1, characterized in that: Before extracting key information based on the processed file through a recognition model based on a PaddleOCR model, the method further includes: Acquire a prepared data set, and preprocess the prepared data set to obtain a preprocessed data set; The preprocessed data set is labeled by PPOCRLabel to obtain a data set, and the data set is divided into a training set, a test set, and a validation set according to a preset ratio; Training the PaddleOCR model using the training set, the test set, and the validation set to obtain a training model; The training model is evaluated, and in response to the evaluation index of the training model reaching a preset index, a recognition model is obtained.

5. The method for automatic radiographic inspection report verification and analysis according to claim 1, characterized in that: The extracting key information based on the processed file through a recognition model based on a PaddleOCR model includes: Performing initialization preprocessing on the processing file to obtain a standardized image; Based on the standardized image, feature extraction is performed using the DB algorithm to predict the text bounding box and obtain the text detection area; Perform text extraction processing on the text detection area using the SVTR_LCNet algorithm to obtain a text list; Using the TableAttn ​​algorithm and the attention mechanism, the table cells are located and aligned across modalities according to the text list to obtain text data; The key information is obtained by fusing the visual layout with the text data through the LayoutXLM algorithm.

6. The method for automatic radiographic inspection report verification and analysis according to claim 5, characterized in that: The initializing preprocessing of the processing file to obtain a standardized image comprises: Obtaining a preset optimal resolution, adjusting the processed file according to the preset optimal resolution, and obtaining a display image; Performing perspective transformation on the displayed image using a tilt correction algorithm to eliminate image tilt and adjusting the image size according to a preset size requirement to obtain a corrected image; The corrected image is subjected to grayscale processing, Gaussian blur processing, adaptive threshold adjustment and histogram equalization adjustment to obtain a standardized image.

7. The method for automatic radiographic inspection report verification and analysis according to claim 1, characterized in that: The obtaining of the static database and comparing the key information with the static database to obtain the comparison information includes: The static database includes: RCCM standard, NB-20001 standard, GB-3323 standard and 1516AT0012 technical specification.

8. The method for automatic radiographic inspection report verification and analysis according to claim 1, characterized in that: The obtaining of the static database and comparing the key information with the static database to obtain the comparison information further includes: Performing preliminary format and value constraint verification on the key information, and obtaining key parameters in response to the key information meeting preset standards; receiving a static database selection message, and selecting a static database according to the static database selection message; Comparing the key parameters with the parameters in the static database one by one, and in response to the key parameters meeting the standard requirements, recording the key parameters that do not meet the tolerance range as abnormal parameters, recording the error type, standard required value and actual extracted value of the abnormal parameters, and retaining non-abnormal parameters; According to the parameter dependency graph, the abnormal parameters are replaced with verified correct preset associated parameters and then compared one by one with the parameters in the static database, and an abnormal comparison report is generated based on the error type of the abnormal parameter, the standard required value, the actual extracted value and the non-abnormal parameter; In response to the key parameters meeting the standard requirements, retaining the key parameters and obtaining a normal comparison report; The comparison information consists of the abnormal comparison report and the normal comparison report.

9. The method for automatic radiographic inspection report verification and analysis according to claim 1, characterized in that: Generating an analysis report based on the comparison information, and adjusting the equipment based on the analysis report includes: An analysis report is generated by a document generation tool according to the comparison information, and equipment adjustments are performed based on the analysis report.

10. An automatic radiographic inspection report verification and analysis system, characterized in that: A method for automatically checking and analyzing radiographic inspection reports according to any one of claims 1 to 9, comprising: A report processing module is used to obtain a report file, pre-process the report file, and obtain a processed file; An information extraction module, configured to extract key information based on the processed file through a recognition model based on a PaddleOCR model; An information comparison module is used to obtain a static database, compare the key information with the static database, and obtain comparison information; A report generation module is used to generate an analysis report based on the comparison information and perform equipment adjustments based on the analysis report.

Citation Information

Patent Citations

  • Otology report target detection and analysis method and system based on deep learning

    CN118747895A

  • Character recognition and automatic auditing method and system

    CN119091460A

  • Method and system for automatically generating radiation report based on radiation priori knowledge

    CN119274736A

  • Radiographic detection compliance determination digital system based on digital negative film

    CN119643600A