A logging report structuring method based on depth reference reconstruction and topological fingerprint
Patent Information
- Application Number
- CN202610969111.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-09-25
AI Technical Summary
该类方法适用于规则线框表格,但对于弱线框、断线表、无线表以及图表混排的测井报告,常将曲线坐标框、图例框或注释框误判为表格边界,从而产生大量错映射
本发明提供了一种基于深度基准重建与拓扑指纹的测井报告结构化方法,不以通用二维表格网格恢复为中心,而是以海洋测井报告特有的深度基准轴重建为核心,将深度数字序列、轨道刻度、井段边界和续页信息统一纳入纵向主索引恢复过程,在弱线框、断线表和无线表场景下仍能稳定恢复深度主轴;通过参数族拓扑指纹对参数列进行统一建模,不仅考虑表头文字,还综合考虑单位、缩写、别名、邻接关系和曲线名称对应关系,能够有效应对多级表头、跨列合并、局部表头缺省和续页表头不完整等场景;采用“深度基准轴+参数族拓扑指纹”的双锚点归属机制,可在高密度数值块分布条件下更准确地恢复“深度-参数-取值”三元关系,降低相邻参数列间的错归属率;引入跨页重叠深度窗口和列漂移补偿机制,可有效补偿扫描偏移、页边裁切和续页表头缺省所导致的页间错位,提高跨页连续矩阵拼接的稳定性;输出对象不仅包括参数和值,还包括来源页码、坐标位置、标准化记录、修正原因和置信度,便于后续数据库入库、人工抽查、质量评估和知识资产治理。
Smart Images

Figure CN122821577A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the fields of artificial intelligence, industrial document understanding, optical character recognition, computer vision, oil and gas industry data governance and well logging data digitization technology, and specifically relates to a well logging report structuring method based on depth benchmark reconstruction and topological fingerprint. Background Technology
[0002] The exploration and development of offshore oil and gas generates a large amount of well logging interpretation and evaluation data. This data typically records well number, well type, stratigraphic position, well interval, well logging curve name, reservoir parameters, fluid properties, interpretation conclusions, and stratigraphic information. It is an important source of basic data for well history archiving, reservoir modeling, production dynamic analysis, and knowledge asset accumulation.
[0003] Compared to ordinary office documents, financial forms, or general reports, marine logging reports have significantly different layouts and business characteristics. First, logging reports typically use depth as the primary vertical business index, with multiple areas such as the curve trajectory area, parameter table area, well basic information area, and legend annotation area coexisting on the same page, and these areas are strongly interrelated in business. Second, the parameter table area is characterized by high density, multiple columns, small fonts, many decimal places, multi-level header merging, unit attachment, and complex parameter abbreviations. Third, a large amount of historical data comes from scanned documents or image-based PDFs, which are prone to problems such as background shadows, background noise, page tilt, page margin clipping, weak borders, character adhesion, and local blurring. Finally, cross-page continuation of tables is quite common; subsequent pages often do not completely repeat the table header, but only retain partial abbreviations, a small amount of unit information, or a "continued" label, making it difficult to recover the complete parameter semantics from a single page when parsed independently.
[0004] In existing technologies, one type of solution uses general OCR as its core, first detecting text and numbers on the page, and then outputting the results according to the reading order. While this method can extract characters, it struggles to reliably restore the "depth-parameter-value" structural relationship, typically only producing discrete text fragments. Another type of solution employs table detection and grid recovery methods, reconstructing a two-dimensional grid by recognizing table frames and cells, and then mapping the OCR results back to the cells. This method is suitable for tables with regular frame borders, but for well logging reports with weak frame borders, broken frame borders, no frame borders, or mixed charts, it often misinterprets curve coordinate frames, legend frames, or annotation frames as table boundaries, resulting in numerous mismappings. Yet another type of solution relies on template matching, pre-defining the well number position, depth column position, and parameter order. This is only suitable for documents with a fixed layout height and is difficult to adapt to the complex sources, long time spans, and variable layouts of marine well logging reports.
[0005] Therefore, the existing technology has at least the following shortcomings: First, it lacks a dedicated model for the main business index of depth, resulting in unstable depth axis recovery; second, it lacks a horizontally unified description mechanism for the semantics of well logging parameter columns, making it difficult to correctly identify parameter columns when the header is missing or abbreviations are mixed; third, it lacks a precise attribution mechanism suitable for high-density well logging numerical distribution, resulting in frequent misalignment between adjacent parameter columns; fourth, it lacks column drift compensation and depth window alignment mechanisms for continuation page scenarios, resulting in unstable cross-page splicing results; fifth, the output results mostly remain at the text or table level, lacking source location, correction records, and traceable information, making it difficult to meet the requirements of enterprise-level data governance and audit review.
[0006] Therefore, it is necessary to propose a new method for structuring marine logging reports that differs from general document parsing and ordinary table recovery. Summary of the Invention
[0007] Based on the aforementioned background technology, the technical problems to be solved by this invention include: difficulty in accurately distinguishing the parameter table area from the curve track area and the legend annotation area under conditions of weak wireframes, broken lines, non-wireless tables, and mixed chart layouts; difficulty in stably establishing the mapping relationship between the vertical position of the page and the actual depth value under conditions of missing depth numbers, skipped numbers, misidentification, omitted continuation pages, and scan clipping; difficulty in uniformly modeling the semantics of parameter columns when there are multi-level headers, cross-column merged titles, mixed parameter abbreviations and units, missing local headers, and inconsistent parameter naming across different pages; difficulty in restoring the unique correspondence between "depth-parameter-value" based solely on geometric position under conditions of high-density numerical block distribution; and difficulty in correctly splicing continuation pages caused by slight offsets in column positions between pages, incomplete header inheritance, discontinuous depth connections, and page edge clipping in cross-page continuation scenarios. Furthermore, existing output results lack source page numbers, original coordinates, correction reasons, standardized records, and confidence information, making it difficult to meet the requirements of subsequent warehousing, sampling, auditing, and backtracking analysis.
[0008] This invention is proposed to solve the aforementioned technical problems. Its purpose is to provide a dedicated structured method with depth reference axis reconstruction as the vertical core, parameter family topological fingerprint as the horizontal semantic skeleton, and dual anchor point attribution and cross-page overlapping depth window calibration as enhancement mechanisms.
[0009] This invention is achieved through the following technical solution: A well logging report structuring method based on depth benchmark reconstruction and topological fingerprinting includes the following steps: S1. Page Preprocessing and Layout Partitioning: Obtain the well logging report page to be processed, preprocess the well logging report page to obtain a standardized page, and divide the standardized page into regions. Step S1 specifically includes the following steps: S11. Obtain the page of marine logging reports to be processed; The marine logging report page is an image or a rendered PDF of the marine logging report page. S12. Preprocess the well logging report page obtained in step S11 to obtain a standardized page; The preprocessing includes noise reduction, grayscale equalization, background shadow suppression, tilt correction, weak line enhancement, character connectivity repair, and page margin clipping correction. The preprocessing also includes normalizing the coordinates of the well logging report page; normalizing the coordinates of the well logging report page can eliminate the influence of different resolutions and scan scales. The coordinate normalization formula for the well logging report page is:
[0010] In the formula: , These are the original coordinates; , Normalized coordinates; This is the width of the well logging report page; This refers to the height of the well logging report page; S13. Perform area division on the well logging report page, dividing the standardized page into the well basic information area, curve track area, parameter table area, legend annotation area, header and footer area, and continuation table identifier area; The method for dividing the well logging report page into regions is as follows: S131. Calculate the number density, text density, curve connectivity, scale periodicity, vertical elongated structure distribution, and row-column alignment for each candidate region; wherein, the number density, text density, curve connectivity, scale periodicity, and row-column alignment are all normalized and have a value range of 0 to 1. S132. The candidate region is divided based on quantitative feature thresholds and prior location information, specifically including: S1321. Identify header and footer areas: Based on the prior knowledge of page position, extract the top and bottom edge areas of the page. If the text density in the area is greater than 0.5 and contains consecutive page number features, then the area is identified as the header and footer area. S1322. Identify the well basic information area: In the area above the header and footer, if the text density is greater than 0.4 and contains preset keywords such as well number, well type, and layer, it is identified as the well basic information area. S1323. Identify curved track regions: If the curve connectivity of a candidate region is greater than 0.6 and the scale periodicity is greater than 0.7, it is determined to have high curve connectivity and scale periodicity, and is identified as a curved track region. S1324. Identify the parameter table area: If the number density of the candidate area is greater than 0.4, the row and column alignment is greater than 0.8, and it has column cluster distribution characteristics, then it is determined to have high number density and column cluster distribution and row and column alignment characteristics, and is identified as the parameter table area. S1325. Identify the legend annotation area: For the remaining area, if the text density is greater than 0.3, the scale periodicity is less than 0.7, the row and column alignment is less than 0.8, and it contains discrete text blocks and symbol graphics, then it is identified as the legend annotation area. S1326. Identify the continuation table identifier area: If the text in the area contains preset continuation table keywords such as "continued table" or "continued from previous table", then the area is identified as the continuation table identifier area.
[0011] S2. Deep Multi-Source Evidence Extraction and Page-Level Depth Reference Axis Reconstruction: Extract depth-related multi-source evidence from the parameter table area and curve track area divided in step S1; establish the mapping relationship between the page's vertical coordinates and the actual depth value based on the multi-source evidence to form a page-level depth reference axis; The multi-source evidence includes at least the vertical depth digital sequence, track scale lines, scale numbers, well section boundary markers, layer boundary locations, and continuous table depth prompts; Step S2 specifically includes the following steps: S21. Extract multiple candidate depth digital sequences from the parameter table area; S22. Based on the monotonicity of the candidate depth digital sequence, the consistency of adjacent depth differences, the correspondence with the position of the scale line in the curved track area, and the coincidence with the position of the well section boundary, select the target depth reference sequence. The specific method is as follows: calculate the depth benchmark confidence of the candidate depth digital sequence, and select the candidate sequence with the highest depth benchmark confidence as the target depth reference sequence; The formula for calculating the depth baseline confidence of the candidate depth digital sequence is as follows:
[0012] In the formula: For depth digital sequence The depth baseline confidence level; It is a monotonicity indicator; This is a deep incremental consistency indicator; This is a consistency indicator corresponding to the track scale; To ensure consistency with the well section boundary; For cross-page continuity consistency metrics; , , , , These are the weighting coefficients for the monotonicity index, depth increment consistency index, consistency index corresponding to the orbital scale, coincidence index with the well section boundary position, and cross-page continuity consistency index. ; S23. When the target depth reference sequence is missing, skipped, or misidentified, the target depth reference sequence is completed or corrected by using the statistical main step length of the difference between adjacent depths, the track scale spacing, and the continuation depth relationship in adjacent pages. S24. Based on the corrected target depth reference sequence, establish a segmented mapping relationship between the page's vertical coordinates and the actual depth values to form a page-level depth reference axis; Let the vertical division of the page be as follows: Then the first The depth mapping function within each interval is expressed as:
[0013] In the formula: and It is obtained by fitting a reference depth point with the corresponding scale point.
[0014] S3. Construction of parameter family topology fingerprint: Extract fingerprint information from the parameter table area divided in step S1, standardize and hierarchically expand the extracted fingerprint information, and construct the parameter family topology fingerprint of each parameter column. The fingerprint information includes multi-level header text, unit text, parameter abbreviations, full parameter names, curve aliases, and column adjacency relationships. The parameter family topological fingerprint includes header path features, unit features, parameter abbreviation or alias standard encoding, left and right adjacent column topological relationship features, and corresponding features with curve track names; For the The parameter family topological fingerprint is represented as follows:
[0015] In the formula: This refers to the header path characteristics; As a unit characteristic; Standard encoding for parameter abbreviations or aliases; Features representing the topological relationships with the left and right adjacent columns; Features corresponding to the names of curved tracks; For pages with incomplete headers or where subsequent pages do not repeat the complete header, the semantics of the parameter columns on the current page are inherited and restored using the topological fingerprint of the parameter family already determined on adjacent pages, the relative order of adjacent columns, and the curve track label information.
[0016] S4. Dual Anchor Point Attribution and In-Page Structured Restoration: Using the page-level depth reference axis constructed in step S2 as the vertical anchor point and the parameter family topology fingerprint constructed in step S3 as the horizontal anchor point, the content blocks in the parameter table area divided in step S1 are assigned, and the depth-parameter-value ternary structure record is restored. The content blocks include text blocks, numerical blocks, and symbol blocks; all text blocks, numerical blocks, and symbol blocks are extracted from the parameter table area, and their bounding rectangles, center point coordinates, and other information are recorded as the base objects for subsequent belonging operations. The method for assigning content blocks in the parameter table area in step S4 is as follows: S41. Determine the candidate depth row to which the content block belongs based on the central ordinate of the content block, the degree of overlap between its vertical projection and the page-level depth reference axis; S42. Determine the candidate parameter column to which the content block belongs based on the degree of overlap between the center horizontal coordinate of the content block and its horizontal projection and the parameter column interval; "Parameter column" refers to the data column that exists physically or is logically divided in the parameter table area of the well logging report (e.g., the vertical column representing density DEN and porosity PHI); the aforementioned "parameter family topological fingerprint" is a set of features constructed for each "parameter column" to describe its semantics and structure; the parameter family topological fingerprint specifically includes header path features, unit features, standard encoding of parameter abbreviations or aliases, topological relationship features of left and right adjacent columns, and corresponding features with curve trajectory names; these are the "attribute features" of the fingerprint, rather than the "feature columns" that actually exist on the page.
[0017] In step S42, the physical range of the "parameter column" to which the content block actually belongs is matched based on the spatial location of the content block (the center horizontal coordinate and its horizontal projection). S43. When a content block corresponds to multiple candidate parameter columns or multiple candidate depth rows, the candidate with the highest matching score is selected first. For content blocks with multiple candidate assignments, the candidate with the highest matching score is selected first. If multiple candidate matching scores are close, the unit compatibility, dimensional rationality and local context continuity are further considered to determine the unique assignment result. For any content block It belongs to deep lines and parameter list Matching score Represented as:
[0018] In the formula: The vertical projection matching degree between the content block and the candidate depth row; The horizontal positional matching degree between the content block and the candidate parameter column range; To ensure consistency between field types and parameter semantics; To maintain continuity with neighboring records; , , , Weighting coefficients for the vertical projection matching degree between the content block and the candidate depth row, the horizontal position matching degree between the content block and the candidate parameter column interval, the consistency between field type and parameter semantics, and the continuity consistency with neighboring records. S44. After completing the attribution, generate a three-element record of depth-parameter-value; The depth-parameter-value ternary structure record includes depth value, parameter identifier, original text, standardized text, unit, source page number, and source coordinates.
[0019] S5. Intra-page standardization, verification and ambiguity resolution: Perform unit normalization, value range verification, sampling interval consistency verification, adjacent depth continuity verification and parameter physical constraint verification on the ternary structure records generated in step S4 to correct misidentification results and improve the quality of intra-page structuring, and complete intra-page standardization and ambiguity resolution. Step S5 specifically includes the following steps: S51. Convert the different units of the same parameter appearing on different pages to the preset standard unit; S52. Identify outliers based on the legal value range of the parameters; S53. Identify misidentification results caused by missing digits, misaligned decimal points, or stuck characters based on the changing trends between adjacent depth records; S54. Based on the physical relationship between at least two parameters, perform consistency verification on the candidate values and complete the correction. The consistency verification of the candidate values is achieved through a candidate value correction cost function. The candidate value correction cost function is:
[0020] In the formula: Cost of range deviation; Cost of inconsistent units; The cost of continuous destruction; The cost of violating the physical relationship between parameters; , , , These are the weighting coefficients for range deviation cost, unit inconsistency cost, continuity disruption cost, and physical relationship violation cost between parameters, respectively. For multiple candidate interpretation results of the same original identification block, the one with the smallest candidate value correction cost function value is selected as the final standardized result.
[0021] S6. Cross-page continuity determination, overlap depth window matching, and column drift compensation: Specifically, the following steps are included: S61. For adjacent pages, construct an overlap depth window and calculate the continuity of adjacent pages; Construct an overlapping depth window using the depth ranges at the tail and head of adjacent pages; The expression for the continuity of adjacent pages is:
[0022] In the formula: and Adjacent pages; The continuity between adjacent pages; The depth matching degree within the overlap depth window; For the topological fingerprint similarity of the parameter family; For the connection between the beginning and end of the well section; To ensure consistency in the table of contents; , , , These are the weighting coefficients for depth matching degree within the overlapping depth window, parameter family topological fingerprint similarity, well segment start-end connectivity, and continuation table identifier consistency. S62. Determine whether adjacent pages belong to the same continuous parameter matrix based on the continuity of adjacent pages; When the continuity of adjacent pages When the value exceeds a preset threshold, adjacent pages are determined to belong to the same continuous parameter matrix. S63. When adjacent pages belong to the same continuous parameter matrix, the difference in the horizontal position of the parameter column corresponding to the common depth point of the two pages is statistically analyzed to obtain the inter-page column drift compensation amount. The formula for calculating the inter-page column drift compensation is as follows:
[0023] In the formula: This is the amount of compensation for inter-page column drift. For the first The center x-coordinate of the page target parameter column For the first The horizontal coordinate of the center of the parameter column corresponding to the page; This is the amount of compensation for inter-page column drift. This is the total number of corresponding parameter columns contained in the common depth points of the two pages (i.e., the number of parameter column pairs participating in the horizontal position difference statistics).
[0024] S64. Based on the inter-page column drift compensation, the parameter column range of the next page is shifted as a whole, and then the parameter records of adjacent pages are spliced together in depth order to complete the inter-page splicing; that is, the first... Page column range shifted as a whole Then, complete the page splicing.
[0025] S7. Object-oriented output and backtracking log generation: The results after the page verification in step S5 and the cross-page splicing in step S6 are object-oriented output to obtain structured results, thereby enabling the structured results to have the characteristics of being able to be stored in the database, being able to be sampled, being auditable, and being traceable.
[0026] The structured results include well base information objects, depth interval objects, parameter definition objects, parameter observation objects, source location objects, and correction log objects; The source location object includes the page number, the original coordinate frame, the page area to which it belongs, and the original recognized text; The correction log object includes the value before correction, the value after correction, the triggering rule, the reason for correction, and the confidence level.
[0027] The beneficial effects of this invention are: This invention provides a well logging report structuring method based on depth benchmark reconstruction and topological fingerprinting. Instead of focusing on general two-dimensional table grid recovery, it centers on the reconstruction of the depth benchmark axis unique to marine well logging reports. It integrates depth numerical sequences, track scales, well section boundaries, and continuation page information into the vertical master index recovery process, ensuring stable recovery of the depth master axis even in scenarios with weak wireframes, broken tables, and non-wireless tables. Furthermore, it uses parameter family topological fingerprinting to uniformly model parameter columns, considering not only header text but also units, abbreviations, aliases, adjacency relationships, and curve name correspondences. This effectively handles multi-level headers, cross-column merging, local header omissions, and continuation page issues. In scenarios such as incomplete headers, a dual-anchor point attribution mechanism of "depth baseline axis + parameter family topological fingerprint" is adopted, which can more accurately restore the "depth-parameter-value" ternary relationship under high-density numerical block distribution conditions, and reduce the misattribution rate between adjacent parameter columns. A cross-page overlap depth window and column drift compensation mechanism are introduced to effectively compensate for page misalignment caused by scan offset, page edge clipping and missing continuation page headers, and improve the stability of cross-page continuous matrix splicing. The output objects include not only parameters and values, but also source page number, coordinate position, standardized record, correction reason and confidence level, which facilitates subsequent database entry, manual sampling, quality assessment and knowledge asset governance. Attached Figure Description
[0028] Figure 1 This is a flowchart of the method of the present invention.
[0029] For those skilled in the art, other related figures can be obtained from the above figures without any creative effort. Detailed Implementation
[0030] To enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.
[0031] Example 1
[0032] This embodiment uses a marine well logging interpretation report as an example. The well logging interpretation report consists of 12 pages, with pages 3 to 6 being continuous parameter matrix pages. The page resolution is 2480×3508 pixels, the depth range is approximately 3180.0m to 3220.0m, and the main sampling interval is 0.5m. The left side of the page is the curve trajectory area, and the right-center part is the parameter table area. The parameters mainly include GR, RT, AC, DEN, CNL, and SW, etc. Some parameters are represented by English abbreviations, and some by their full Chinese names, and the units are distributed in different levels of table headers. like Figure 1 As shown, a well logging report structuring method based on depth benchmark reconstruction and topological fingerprinting includes the following steps: S1. Page Preprocessing and Page Layout: S11. Page reading and image standardization: First, scanned images from pages 3 to 6 are read. The original images are converted to grayscale, and a background smoothing map is generated using a local background estimation method. The background smoothing map is then subtracted from the original image to reduce the scan texture and large-area shadows. Finally, an adaptive contrast enhancement method is used to improve the distinguishability of decimal characters and thin lines. S12, Tilt Correction and Edge Repair: Based on the joint tilt angle estimation of the text line direction and the vertical long line segment direction, a clockwise deflection of about 0.8° was detected on page 3, so a counterclockwise rotation correction was performed; after correction, the black border of the page and the local cropping gap were repaired to ensure the measurement accuracy of the subsequent horizontal column interval and vertical depth row interval. S13. Regional Division: The numerical density, text density, curve connectivity, scale periodicity, and elongated structure distribution were calculated for candidate areas on the page. The results showed that the left area exhibited stable curve energy and scale rhythm in the vertical direction, thus it was classified as the curve track area. The middle-right area had a significant high-density numerical distribution and column alignment structure, therefore it was classified as the parameter table area. The top of the page was classified as the well basic information area, and the bottom page number area was classified as the header and footer area. Finally, the horizontally normalized range of the parameter table area on page 3 was [value missing]. The lateral range of the curved track area is .
[0033] S2. Deep Multi-Source Evidence Extraction and Page-Level Depth Reference Axis Reconstruction: S21, Deep Digital Sequence Extraction: Extracting the vertically arranged numerical string from the parameter table area yields a candidate depth sequence with approximately 81 candidate depth points; the candidate values are roughly 3180.0, 3180.5, 3181.0, 3181.5 up to 3220.0. Simultaneously, extracting the track scale lines and scale numbers from the curved track area; S22, Deep Candidate Sequence Scoring: Calculate the monotonicity index, depth increment consistency, trajectory scale correspondence consistency, well section boundary fit, and page continuity consistency for the candidate depth sequence; let the candidate sequence be... Then its comprehensive depth benchmark confidence level is:
[0034] In this embodiment, , , , , After comprehensive calculation It is higher than other candidate sequences, therefore it will Selected as the target depth reference sequence; S23. Depth outlier correction and depth mapping fitting: The location "3194.5m" was found to be misidentified as "31945". Based on the distribution of differences between adjacent depth points, the main step size of this page was determined to be 0.5m. Then, based on the equal spacing of the track scale, a piecewise mapping function from the longitudinal coordinate to the depth value was fitted.
[0035] According to the fitting results, the theoretical depth corresponding to the outlier should fall around 3194.48m, so it was corrected to 3194.5m. After the correction, the system established 81 stable depth row intervals, forming a page-level depth reference axis.
[0036] S3, Parameter Family Topological Fingerprint Construction: S31. Restore header path: Extracting the top header block of the parameter table area revealed that the header contained Chinese terms such as "resistivity," "acoustic transit time," "density," and "neutron porosity," as well as English abbreviations such as RT, AC, DEN, and CNL. The multi-level header path was then reconstructed based on the inter-block coverage relationship. S32, Parameter Normalization Mapping: By combining industry parameter dictionaries, "Resistivity / RT" is mapped to the unified parameter code RT, "Acoustic Transit Time / AC" to the unified parameter code AC, "Density / DEN" to the unified parameter code DEN, and "Neutron Porosity / CNL" to the unified parameter code CNL. Then, by combining the unit identification results, corresponding units are assigned to each column, such as RT corresponding to Ω·m, AC to μs / ft, and DEN to g / cm³. 3 CNL corresponds to % S33. Integration of curve aliases and adjacency relationships: Further reading of the curve labels in the curve track area revealed that GR, RT, AC, DEN, and CNL correspond to the parameter names in the table header. Taking the DEN column as an example, its parameter family topological fingerprint is represented as follows:
[0037] in, The header path is "Reservoir Parameters - Density". The unit is "g / cm". 3 " The standard abbreviation is "DEN". Left neighbor AC, right neighbor CNL This corresponds to the curve label DEN.
[0038] S4. Dual anchor point attribution and in-page structured restoration: S41. Content block candidate generation: Extract all text blocks, numerical blocks, and symbol blocks from the parameter table area, and record their bounding rectangles, center point coordinates, and local OCR results. S42, Depth Anchor Point Assignment: For any numerical block, its candidate depth row is determined based on the degree of overlap between its center ordinate and the longitudinal projection of its bounding box and the depth reference axis. Taking the numerical block "2.31" as an example, its central ordinate falls near row 3198.0m, therefore its relation to row 3198.0m is... maximum; S43. Parameter anchor point assignment: Based on the x-coordinate of the center of the numerical block and the overlap between the bounding box and the parameter column interval, calculate its relationship to each candidate parameter column. ; For "2.31", its horizontal position is close to both the DEN and CNL columns, but considering the consistency of field types... It can be seen that: if it belongs to the DEN column, its value range is reasonable; if it belongs to the CNL column, the porosity of 2.31% or 0.0231 is inconsistent with the context; therefore, the candidate with the largest score is finally obtained according to the following formula:
[0039] The system classifies this block as "3198.0m-DEN-2.31g / cm". 3 ”; S44. Generation of intra-page ternary records: After all content blocks are assigned, the system outputs a single-page ternary record. Each record includes at least: depth value, parameter encoding, original recognition text, standardized text, unit, page number, and source coordinates.
[0040] S5. In-page standardization, verification, and ambiguity resolution: S51. Value Range Verification: Set the range for GR to 0 to 250 API; Set the DEN range to 1.0 to 3.5 g / cm³. 3 ; Set the CNL range to 0 to 60%; When a DEN value of "231" is identified, the cost of calculating its range deviation is too high, triggering local re-identification of the original image; S52, Continuity Check: Compare the parameter variation trends between adjacent depth points; When the DEN value at a certain depth point suddenly changes to 23.1 or 0.231, it is considered to be a decimal point recognition error because it is seriously inconsistent with the continuity of the depth points before and after it, and it is corrected to 2.31. S53. Verification of physical relationships between parameters: Perform cross-consistency checks on parameters such as DEN, CNL, and SW within local well sections; If the combination of porosity and density does not conform to the typical reservoir physical relationship, the original image should be checked first and the candidate identification results should be reselected. In the end, a total of 8 decimal point errors and 3 column classification errors were corrected on page 3.
[0041] S6. Cross-page continuity determination, overlap depth window matching, and column drift compensation: S61. Construct the overlap depth window: For the depth window at the end of page 3 and the first deep window on page 4 Construct a common depth matching interval; it was found that the two pages share a common depth point at 3219.0m and 3219.5m; S62. Continuity calculation: Based on the common depth point matching degree, parameter family topological fingerprint similarity, well segment start-end connectivity degree, and continuation table identifier consistency degree, the following calculations are performed:
[0042] The calculation result is 0.944, which is higher than the continuous judgment threshold of 0.85. Therefore, it is determined that pages 3 and 4 belong to the same continuous parameter matrix. S63, Train Drift Compensation: The system calculates the lateral position difference of the parameter columns corresponding to points at common depths to determine the optimal drift compensation amount.
[0043] Calculated Pixels; therefore, the parameter column of page 4 was shifted 4.1 pixels to the left before stitching was performed. After stitching, pages 3 to 6 formed a continuous depth matrix.
[0044] S7. Object-oriented output and backtracking log generation: Ultimately, this embodiment outputs 1 well basic information object, 4 depth interval objects, 6 parameter definition objects, 486 parameter observation value objects, 486 source location objects, and 27 correction log objects; Each parameter observation includes the original page number, coordinate frame, original identification text, standardized value, and correction reason, which can be directly used for database entry and subsequent manual sampling.
[0045] Example 2
[0046] This embodiment describes a historical marine well logging interpretation report. Pages 7 and 8 of the report are consecutive continuation pages with a page resolution of approximately 2339×3308 pixels, a depth range of 4062.0m to 4090.0m, and a main sampling interval of 0.2m. The boundary between the curve track area and the parameter table area on the page is blurred, some parameter columns lack explicit separator lines, and the complete table header on page 8 is not repeated, only retaining local abbreviations such as "RT", "DEN", and "Φ".
[0047] A well logging report structuring method based on depth benchmark reconstruction and topological fingerprinting includes the following steps: S1. Page Preprocessing and Layout: Perform background modeling and frequency domain filtering on the page to remove scan patterns and stripe noise; Since the page lacks complete table boundaries, instead of relying on closed wireframes as the sole basis for table recognition, the layout is divided using a combination of features such as curve connectivity, track scale periodicity, row and column alignment of number blocks, and text density. This successfully identifies the left side as the curve track area and the middle right side as the parameter table area.
[0048] S2. Deep Multi-Source Evidence Extraction and Page-Level Depth Reference Axis Reconstruction: During the depth reference axis recovery phase, the parameter table area only identifies some depth figures, such as 4068.4, 4068.8, 4069.0, 4069.6, and 4069.8. Directly processing these figures using ordinary OCR results would result in a broken depth axis. Therefore, an initial vertical grid is established using the equidistant scale lines in the curved track area. Combined with the statistical results of a main step size of 0.2m, the missing depth points such as 4068.6, 4069.2, and 4069.4 are filled in, and "40722" is automatically corrected to 4072.2. This forms a continuous page-level depth reference axis.
[0049] S3, Parameter Family Topological Fingerprint Construction: During the parameter semantic recovery stage, with the header of page 7 complete, the system first constructs parameter family topological fingerprints for eight types of parameters, including GR, SP, Rxo, Rt, AC, DEN, PHI, and SW. Among them, the PHI column is simultaneously bound to three aliases: "porosity", "Φ", and "PHI". When the header of page 8 is missing, the system uses the fingerprint set of the previous page, the column adjacency order, and curve label information to inherit and recover the semantics of the parameter columns on the current page, thereby completing parameter column identification without requiring the complete header to be repeated.
[0050] S4. Dual anchor point attribution and in-page structured restoration: During the dual-anchor point attribution phase, some numerical blocks are located near the boundary of adjacent parameter columns; For example, the value block "0.18" is close to both the DEN and PHI columns. If judged only by the horizontal coordinate distance, it may be mistakenly assigned to the DEN column. However, the system considers the semantic consistency of parameters and the continuity of local context, and believes that 0.18 is more consistent with the porosity value distribution and is consistent with the PHI values of 0.17 and 0.19 of the adjacent depth points. Therefore, it is assigned to the PHI column. For example, if the value block "85" is adjacent to both columns GR and SW, the system will assign "85" to column SW based on the distribution of GR values of approximately 72 and 74 and SW values of approximately 83 and 86 in the adjacent depth windows.
[0051] S5. In-page standardization, verification, and ambiguity resolution: During the in-page verification stage, a physical relationship verification between parameters is introduced; When the combination "PHI=0.82, DEN=2.65" was identified, it was considered that the porosity was too high and inconsistent with the density value. Therefore, the original image was checked first and the decimal point position was corrected. Through this mechanism, a total of 9 decimal point errors and 4 column attribution errors were corrected on page 8.
[0052] S6. Cross-page continuity determination, overlap depth window matching, and column drift compensation: During the cross-page stitching stage, the system constructs an overlap depth window for the end of page 7 and the beginning of page 8, and detects that there is a common depth point between the two pages, but the parameter column of page 8 is shifted to the left as a whole; By fitting the horizontal positions of the same parameter column at the common depth point, the optimal drift compensation pixel is obtained; the continuity between pages before compensation is 0.781, and the continuity after compensation is improved to 0.931, which significantly improves the problem of misconnected pages.
[0053] S7. Object-oriented output and backtracking log generation: Ultimately, this embodiment outputs one well basic information object, two depth interval objects, eight parameter definition objects, 1120 parameter observation value objects, 1120 source location objects, and 41 correction log objects. All records include page numbers, original coordinate frames, standardized values, and correction descriptions, supporting subsequent auditing and review.
[0054] As can be seen from this embodiment, even in complex historical document scenarios with weak wireframes, default table headers, and mixed charts, the present invention can still stably achieve deep axis restoration, parameter semantic inheritance, page attribution, and cross-page splicing.
[0055] This invention provides a structured processing method for marine logging reports based on depth benchmark reconstruction and topological fingerprinting, applicable to marine logging interpretation reports, comprehensive logging reports, logging evaluation reports, completion test parameter reports, and their historical scan archives. It is based on page preprocessing and layout partitioning, with depth benchmark axis reconstruction as the vertical core, parameter family topological fingerprint construction as the horizontal semantic core, dual anchor point attribution as the structured recovery bridge, and in-page standardization verification, ambiguity resolution, and cross-page overlapping depth window splicing as enhancement mechanisms. This distinguishes it from traditional general OCR, ordinary table detection and recovery, and fixed template extraction methods. This invention is applicable to scenarios with weak wireframes, multi-level headers, and cross-page table continuations, improving the accuracy and traceability of structured logging report processing. The method of this invention can stably restore the "depth-parameter-value" structural relationship and output objectified data with source coordinates, standardized records, correction logs, and backtracking information, despite the existence of problems widely found in existing marine logging reports, such as mixed layout of curve track areas and parameter table areas, weak wireframes or no table, high-density parameter distribution, missing or misidentified depth numbers, complex multi-level table headers, mixed units and abbreviations, missing local table headers and cross-page continuation table headers, unstable scan quality, and inter-page column drift. It has high engineering application value, data governance value, and significance for promotion and application.
[0056] The applicant declares that the above description is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Those skilled in the art should understand that any changes or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention fall within the protection and disclosure scope of the present invention.
Claims
1. A well logging report structuring method based on depth benchmark reconstruction and topological fingerprinting, characterized in that: Includes the following steps: S1. Obtain the logging report page to be processed, preprocess the logging report page to obtain a standardized page, and divide the standardized page into regions. S2. Extract multi-source evidence related to depth from the parameter table area and curve track area divided in step S1; establish the mapping relationship between the vertical coordinate of the page and the actual depth value based on the multi-source evidence to form a page-level depth reference axis; S3. Extract fingerprint information from the parameter table area divided in step S1, standardize and hierarchically expand the extracted fingerprint information, and construct the parameter family topology fingerprint of each parameter column. S4. Using the page-level depth reference axis constructed in step S2 as the vertical anchor point and the parameter family topology fingerprint constructed in step S3 as the horizontal anchor point, perform the attribution of the content blocks in the parameter table area divided in step S1 and restore the depth-parameter-value ternary structure record. S5. Perform in-page standardization and ambiguity resolution on the ternary structure records generated in step S4; S6, cross-page continuity determination, overlap depth window matching and column drift compensation; S7. The results after the page-level verification in step S5 and the cross-page splicing in step S6 are objectified and output to obtain a structured result.
2. The well logging report structuring method based on depth benchmark reconstruction and topological fingerprinting according to claim 1, characterized in that: Step S1 specifically includes the following steps: S11. Obtain the page of logging reports to be processed; The marine logging report page is an image or PDF rendering of the logging report page; S12. Preprocess the well logging report page obtained in step S11 to obtain a standardized page; The preprocessing includes noise reduction, grayscale equalization, background shadow suppression, tilt correction, weak line enhancement, character connectivity repair, page margin clipping correction, and coordinate normalization. The coordinate normalization formula is: In the formula: , These are the original coordinates; , Normalized coordinates; This is the width of the well logging report page; This refers to the height of the well logging report page; S13. Perform area division on the well logging report page, dividing the standardized page into the well basic information area, curve track area, parameter table area, legend annotation area, header and footer area, and continuation table identifier area.
3. The well logging report structuring method based on depth benchmark reconstruction and topological fingerprinting according to claim 2, characterized in that: The method for dividing the well logging report page into regions is as follows: S131. Calculate the number density, text density, curve connectivity, scale periodicity, vertical long and thin structure distribution, and row and column alignment for each candidate region. S132. The candidate region is divided based on quantitative feature thresholds and prior location information, specifically including: S1321. Identify header and footer areas: Based on the prior knowledge of page position, extract the top and bottom edge areas of the page. If the text density in the area is greater than 0.5 and contains consecutive page number features, then the area is identified as the header and footer area. S1322. Identify the well basic information area: In the area above the header and footer, if the text density is greater than 0.4 and includes preset keywords for well basic information, it is identified as the well basic information area. The preset keywords for well foundation information include well number, well type, and stratigraphic level. S1323. Identify curved track regions: If the curve connectivity of a candidate region is greater than 0.6 and the scale periodicity is greater than 0.7, it is determined to have high curve connectivity and scale periodicity, and is identified as a curved track region. S1324. Identify the parameter table area: If the number density of the candidate area is greater than 0.4, the row and column alignment is greater than 0.8, and it has column cluster distribution characteristics, then it is determined to have high number density and column cluster distribution and row and column alignment characteristics, and is identified as the parameter table area. S1325. Identify the legend annotation area: For the remaining area, if the text density is greater than 0.3, the scale periodicity is less than 0.7, the row and column alignment is less than 0.8, and it contains discrete text blocks and symbol graphics, then it is identified as the legend annotation area. S1326. Identify the continuation table identifier area: If the text in the area contains the preset continuation table keyword, then the area is identified as the continuation table identifier area. The preset continuation table keyword is either "continuation table" or "continuation of the previous table".
4. The well logging report structuring method based on depth benchmark reconstruction and topological fingerprinting according to claim 1, characterized in that: The multi-source evidence includes at least the vertical depth digital sequence, track scale lines, scale numbers, well section boundary markers, layer boundary locations, and continuation table depth prompts.
5. The well logging report structuring method based on depth benchmark reconstruction and topological fingerprinting according to claim 1, characterized in that: Step S2 specifically includes the following steps: S21. Extract multiple candidate depth digital sequences from the parameter table area; S22. Calculate the depth reference confidence of the candidate depth digital sequence, and select the candidate sequence with the highest depth reference confidence as the target depth reference sequence; The formula for calculating the depth baseline confidence of the candidate depth digital sequence is as follows: In the formula: For depth digital sequence The depth baseline confidence level; It is a monotonicity indicator; This is a deep incremental consistency indicator; This is a consistency indicator corresponding to the track scale; To ensure consistency with the well section boundary; For cross-page continuity consistency metrics; , , , , These are the weighting coefficients for the monotonicity index, depth increment consistency index, consistency index corresponding to the orbital scale, coincidence index with the well section boundary position, and cross-page continuity consistency index. ; S23. When the target depth reference sequence is missing, skipped, or misidentified, the target depth reference sequence is completed or corrected by using the statistical main step length of the difference between adjacent depths, the track scale spacing, and the continuation depth relationship in adjacent pages. S24. Based on the corrected target depth reference sequence, establish a segmented mapping relationship between the page's vertical coordinates and the actual depth values to form a page-level depth reference axis; Let the vertical division of the page be as follows: Then the first The depth mapping function within each interval is expressed as: In the formula: and It is obtained by fitting a reference depth point with the corresponding scale point.
6. The well logging report structuring method based on depth benchmark reconstruction and topological fingerprinting according to claim 1, characterized in that: The fingerprint information includes multi-level header text, unit text, parameter abbreviation, full parameter name, curve alias, and column adjacency relationship; the parameter family topological fingerprint includes header path features, unit features, standard encoding of parameter abbreviation or alias, topological relationship features of left and right adjacent columns, and corresponding features with curve track name; For pages with incomplete headers or where subsequent pages do not repeat the complete header, the semantics of the parameter columns on the current page are inherited and restored using the topological fingerprint of the parameter family already determined on adjacent pages, the relative order of adjacent columns, and the curve track label information.
7. The well logging report structuring method based on depth benchmark reconstruction and topological fingerprinting according to claim 1, characterized in that: The method for assigning content blocks in the parameter table area in step S4 is as follows: S41. Determine the candidate depth row to which the content block belongs based on the central ordinate of the content block, the degree of overlap between its vertical projection and the page-level depth reference axis; S42. Determine the candidate parameter column to which the content block belongs based on the degree of overlap between the center horizontal coordinate of the content block and its horizontal projection and the parameter column interval; S43. When a content block corresponds to multiple candidate parameter columns or multiple candidate depth rows, the candidate with the highest matching score is selected first. For content blocks with multiple candidate assignments, the candidate with the highest matching score is selected first. If multiple candidate matching scores are close, the unit compatibility, dimensional rationality and local context continuity are further considered to determine the unique assignment result. For any content block It belongs to deep lines and parameter list Matching score Represented as: In the formula: The vertical projection matching degree between the content block and the candidate depth row; The horizontal positional matching degree between the content block and the candidate parameter column range; To ensure consistency between field types and parameter semantics; To maintain continuity with neighboring records; , , , These are the weighting coefficients for the vertical projection matching degree between the content block and the candidate depth row, the horizontal position matching degree between the content block and the candidate parameter column interval, the consistency between field type and parameter semantics, and the continuity consistency with neighboring records. S44. After completing the attribution, generate a three-element record of depth-parameter-value; The depth-parameter-value ternary structure record includes depth value, parameter identifier, original text, standardized text, unit, source page number, and source coordinates.
8. The well logging report structuring method based on depth benchmark reconstruction and topological fingerprinting according to claim 1, characterized in that: The in-page standardization and ambiguity resolution include unit normalization, value range verification, sampling interval consistency verification, adjacent depth continuity verification, and physical constraint verification between parameters.
9. The well logging report structuring method based on depth benchmark reconstruction and topological fingerprinting according to claim 1, characterized in that: Step S6 specifically includes the following steps: S61. For adjacent pages, construct an overlap depth window and calculate the continuity of adjacent pages; Construct an overlapping depth window using the depth ranges at the tail and head of adjacent pages; The expression for the continuity of adjacent pages is: In the formula: and Adjacent pages; The continuity between adjacent pages; The depth matching degree within the overlap depth window; For the topological fingerprint similarity of the parameter family; For the connection between the beginning and end of the well section; To ensure consistency in the table of contents; , , , These are the weighting coefficients for depth matching degree within the overlapping depth window, parameter family topological fingerprint similarity, well segment start-end connectivity, and continuation table identifier consistency. S62. Determine whether adjacent pages belong to the same continuous parameter matrix based on the continuity of adjacent pages; When the continuity of adjacent pages When the value exceeds a preset threshold, adjacent pages are determined to belong to the same continuous parameter matrix. S63. When adjacent pages belong to the same continuous parameter matrix, the difference in the horizontal position of the parameter column corresponding to the common depth point of the two pages is statistically analyzed to obtain the inter-page column drift compensation amount. The formula for calculating the inter-page column drift compensation is as follows: In the formula: This is the amount of compensation for inter-page column drift. For the first The center x-coordinate of the page target parameter column For the first The horizontal coordinate of the center of the parameter column corresponding to the page; This is the amount of compensation for inter-page column drift. For the first Page and The total number of corresponding parameter columns contained in the common depth point of the two pages; S64. After shifting the parameter column range of the next page as a whole according to the inter-page column drift compensation amount, the parameter records of adjacent pages are spliced together in depth order to complete the inter-page splicing.
10. The well logging report structuring method based on depth benchmark reconstruction and topological fingerprinting according to claim 1, characterized in that: The structured results include well base information objects, depth interval objects, parameter definition objects, parameter observation objects, source location objects, and correction log objects; the source location objects include page numbers, original coordinate frames, the page region to which they belong, and original identification text; The correction log object includes the value before correction, the value after correction, the triggering rule, the reason for correction, and the confidence level.