A data extraction error self-correction method based on external text semantic analysis

CN122819232APending Publication Date: 2026-09-25SHANGHAI DIGITAL IND DEVELOPMENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610699319.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]鉴于上述的分析,本发明公开一种基于外部文本语义解析的数据提取误差自校正方法;解决现有技术中图表数据提取结果缺乏有效校正手段、外部文本信息难以自动化利用的技术问题

Benefits of technology

本发明的方法通过图表数据提取结果与外部文本语义信息的协同利用,从外部描述文本中获取可作为参照的关键数据点,并据此对提取数据进行偏差估计和修正,从而提高图表数据提取结果的可校验性、修正效率和数据复用价值。在部分实施方式中,还可结合误差评估指标、置信度信息和人工确认机制,形成从文本解析到数值修正的闭环优化流程。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819232A_ABST
    Figure CN122819232A_ABST
Patent Text Reader

Abstract

The application discloses a data extraction error self-correction method based on external text semantic analysis, and belongs to the cross field of artificial intelligence and data science. The method comprises the following steps: acquiring an initial data point set containing chart metadata and external description text; constructing a prompt word or analysis rule based on chart information, extracting key data points and evaluating confidence, and outputting structured key data points; searching for optimal matching points in the initial data point set, establishing a mapping relationship and calculating horizontal and vertical coordinate deviation; statistically aggregating and solving correction parameters for the deviation, and outputting a corrected data set after performing coordinate translation transformation. Through the cooperation of chart data and external text semantics, the application improves the checkability, correction efficiency and data reuse value of the extraction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of artificial intelligence and data science, and in particular to a self-correction method for data extraction errors based on external text semantic parsing. Background Technology

[0002] In scientific literature, technical reports, and experimental records, a large amount of experimental data is presented in the form of charts and graphs. To facilitate subsequent data analysis, knowledge mining, and model training, it is usually necessary to convert the data curves in these charts and graphs into structured numerical sets. Existing chart and graph data extraction techniques are mainly based on computer vision methods, obtaining the coordinate values ​​of data points through steps such as coordinate calibration, color segmentation, and pixel extraction. However, due to the inevitable distortion introduced during the shooting, scanning, compression, and display of chart and graph images, coupled with interference from factors such as coordinate axis occlusion, curve intersections, and color degradation, data extracted solely based on visual features often contains systematic biases and random noise.

[0003] More importantly, the charts and graphs themselves do not contain true numerical values ​​that can be cross-validated, so the quantitative assessment and correction of extraction errors often rely on manual review, which is inefficient and highly subjective. Although some literature records the key data features in the charts and graphs in text form in the main text, this textual information is scattered throughout the various chapters of the paper and is expressed in a flexible and varied manner. Traditional rule-based or template-based methods are difficult to effectively utilize this information for automated error correction.

[0004] In recent years, large language models have demonstrated powerful semantic parsing capabilities in natural language understanding, extracting structured information from unstructured text; multimodal large models, on the other hand, possess the ability to simultaneously understand images and text. This provides new insights into overcoming the aforementioned bottlenecks. However, current technologies lack effective mechanisms to organically integrate these capabilities with the chart data extraction process. Therefore, how to construct an automated technical chain from "text semantic parsing" to "data error correction" remains a pressing technical problem to be solved. Summary of the Invention

[0005] Based on the above analysis, this invention discloses a data extraction error self-correction method based on external text semantic parsing; solving the technical problems of lack of effective correction methods for chart data extraction results and difficulty in automatically utilizing external text information in the prior art.

[0006] This invention discloses a data extraction error self-correction method based on external text semantic parsing, comprising: Obtain an initial set of data points, including chart metadata, obtained after processing the chart image through coordinate calibration, color analysis, and multi-feature extraction, as well as external descriptive text containing the key data features of the chart; Based on chart-related information, construct prompt words or parsing rules, extract key data points from external descriptive text, and perform format checks or confidence assessments on the extraction results to output structured key data points containing horizontal and vertical coordinate values. Search for the matching point with the best geometric proximity to each key data point in the initial data point set, establish a one-to-one mapping relationship between key data points and matching points, and calculate the deviation of each matching pair in the horizontal and vertical coordinate directions. Perform statistical aggregation on the deviation to calculate the correction parameters, perform coordinate translation transformation on the initial data point set based on the correction parameters, and output the corrected data set.

[0007] Furthermore, after outputting the corrected data set, a visualization result comparing the results before and after the correction is generated, the key data points are overlaid as external text reference benchmarks, and a statistical report on the correction effect is output.

[0008] Furthermore, The acquisition of the initial data point set includes: Obtain an initial set of data points generated by the chart data extraction process, the initial set of data points including at least one data series and its physical coordinate values; Obtain the chart metadata associated with the initial set of data points. The chart metadata includes one or more of the following: axis information, data series information, axis binding relationships, chart type information, or unit information.

[0009] Furthermore, the chart data extraction process includes: Perform region labeling preprocessing on the chart image to divide it into data region, legend region, and interference region; Construct a multi-coordinate system mapping model to establish the binding relationship between each data series and its corresponding coordinate axis in a multi-axis scenario; [Targeting...] The chart image is filtered to generate a binary mask, and the data series is determined by color clustering of the legend area. Complex topological scenes are constrained and filtered within the data area, and the pixel coordinate set is extracted by fusing line center path, shape detection and connected component features. The pixel coordinate set is converted into physical values ​​through a multi-coordinate system mapping model, and then the initial data point set is obtained through density control, noise removal, and duplicate point detection and merging.

[0010] Furthermore, the complex topological scene includes occlusion, breakage, and overlap, and the constraint filtering includes: For areas where data curves are obscured, extraction constraints are applied within user-specified lined areas, and a mask for extraction is generated using a color filtering strategy. For regions with broken data curves or discontinuous topology, continuous sorting and path construction are performed on mask pixels, and continuous line points are sampled and retained according to density patterns. To address the aliasing issue caused by overlapping multiple manually drawn areas, the overlapping parts are first detected, and discrete overlapping points are connected and clustered based on a distance threshold. Then, Laplacian local sharpness detection is performed on the unclear parts within the drawn areas to generate an enhanced extraction mask.

[0011] Furthermore, the step of constructing prompts or parsing rules based on chart-related information and extracting key data points from external descriptive text includes: Inject chart metadata, define task roles, constrain output format, and enhance semantic alignment between text and chart metadata; construct dedicated prompt words for semantic parsing of large models. Based on the prompt words, the large language model is invoked to extract structured data from the external descriptive text and obtain a set of candidate key data points. When the model returns an incomplete result or fails to parse, a step-by-step rollback strategy is triggered, which includes JSON repair, rule parsing rollback, and text value extraction. The candidate key data points are sequentially checked for numerical rationality, semantic consistency, and structural integrity. Based on the check results, the credibility of the candidate key data points is determined, and structured credible key data points are output.

[0012] Furthermore, by injecting chart metadata, defining task roles, constraining output formats, and enhancing the semantic alignment between text and chart metadata, dedicated prompt words for large-scale model semantic parsing are constructed, including: Inject chart metadata into the context of prompt words to establish the association constraints between numerical values, physical meaning, and coordinate location; Define the model task and expert roles, and clearly extract key data points that can be used as truth references; Constrain structured JSON output and provide guidance with few samples to standardize model output format and recognition patterns; Establish semantic alignment between external descriptive text and chart metadata, add cue constraints to descriptions that are obviously inconsistent or uncertain, and finally form special cue words for semantic parsing of large models.

[0013] Further, in the initial set of data points, a matching point with the optimal geometric proximity to each key data point is searched, establishing a one-to-one mapping relationship between key data points and matching points, and the deviation of each matching pair in the horizontal and vertical coordinate directions is calculated; including: Based on the series identifiers of key data points, filter the corresponding data series. If there are no identifiers, perform nearest neighbor matching in the entire series to determine the candidate data point set. In the candidate subset, the nearest matching point to each key data point is searched using distance calculation, and the corresponding matching relationship between key data points and matching points is established; Calculate the deviation between key data points and their corresponding matching points in the horizontal and vertical coordinate directions to quantify the systematic offset in data extraction.

[0014] Further, statistical aggregation is performed on the deviation to calculate correction parameters, coordinate translation transformation is performed on the initial data point set based on the correction parameters, and a corrected data set is output; including: Calculate the global correction parameter based on the set of deviations for each matching pair; The correction parameters are applied to the initial data point set to perform a coordinate translation transformation, resulting in corrected data; Calculate the offset statistics or error assessment indicators before and after the correction to evaluate the effect of the correction.

[0015] Furthermore, a visualization of the before-and-after correction results is generated, overlaying the key data points as an external text reference benchmark, and a statistical report on the correction effect is output; including: Generate multi-layered visual charts; present the original data and corrected data in a spatial comparison manner, showing the relative positional relationship between the data and external text reference points. Differentiate and label each data layer; use differentiated visual styles to distinguish original data points, corrected data points, key data points and their matching points, with key data points highlighted. Output structured correction results; generate structured correction results containing correction parameters, matching details and effect indicators, used to record the main correction process and results.

[0016] This invention can achieve one of the following beneficial effects: The method of this invention utilizes the synergistic effect of chart data extraction results and external text semantic information to obtain key data points from external descriptive text that can be used as references. Based on these points, the extracted data is used to estimate and correct deviations, thereby improving the verifiability, correction efficiency, and data reuse value of the chart data extraction results. In some embodiments, error assessment indicators, confidence information, and manual verification mechanisms can be combined to form a closed-loop optimization process from text parsing to numerical correction. Attached Figure Description

[0017] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. Figure 1 This is a flowchart of the data extraction error self-correction method based on external text semantic parsing in an embodiment of the present invention. Detailed Implementation

[0018] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and, together with the embodiments of the present invention, serve to illustrate the principles of the present invention.

[0019] One embodiment of the present invention discloses a data extraction error self-correction method based on external text semantic parsing, such as... Figure 1 As shown, it includes the following steps: S1. Obtain an initial set of data points, including chart metadata, obtained after processing the chart image through coordinate calibration, color analysis, and multi-feature extraction, as well as external descriptive text containing the key data features of the chart. The acquisition of the initial data point set includes: Obtain an initial set of data points generated by the chart data extraction process, the initial set of data points including at least one data series and its physical coordinate values; Obtain the chart metadata associated with the initial set of data points. The chart metadata includes one or more of the following: axis information, data series information, axis binding relationships, chart type information, or unit information.

[0020] S1 specifically includes: S101. Obtain an initial data point set obtained through coordinate calibration and pixel extraction. The initial data point set contains at least one data series, and each data point has physical coordinate values, providing data objects to be corrected for subsequent correction. S1011. Perform region labeling preprocessing on the chart image, dividing it into data region, legend region and interference region; Specifically, it includes: S10111 Image Format Standardization Conversion: Receives chart images uploaded by users in PNG, JPG, BMP, or TIFF formats, automatically performs format standardization conversion, and uniformly converts them to the internally processed RGB color space, providing a unified data foundation for subsequent image analysis.

[0021] S10112. Interactive Area Annotation: Employing a combination of interactive rectangular selection and polygon drawing, key areas are accurately annotated. Data area: The smallest bounding rectangle containing the curve / marker point; Legend area: Color coding and text description area; Interference areas include grid lines, coordinate axis ticks, and title text.

[0022] It supports direct color selection in the merged legend preview, and allows users to accurately select the target color without background noise by cropping the legend area and zooming in.

[0023] S10113. Save preprocessing state snapshot: The labeled coordinates, preprocessing parameters or extracted configuration can be saved in the form of structured data for subsequent configuration reuse, result tracking or batch processing; the structured data can be in the format of JSON or other formats.

[0024] An image quality assessment step can be added between image format standardization conversion and interactive region annotation: for user-specified lined or data regions, unclear areas are detected based on Laplacian local sharpness, and overlapping parts between multiple manually lined regions are detected and connected to prompt or generate parameters required for subsequent extraction strategies.

[0025] Optionally, an image enhancement step can be added between interactive region annotation and saving the preprocessed state snapshot: when the image contrast is insufficient or light-colored data points are difficult to identify, enhancement operations such as contrast enhancement, sharpening, and bilateral filtering noise reduction are performed to improve the data point detection rate.

[0026] S1012. Construct a multi-coordinate system mapping model and establish the binding relationship between each data series and the corresponding coordinate axis in a multi-coordinate axis scenario; Specifically, it includes: S10121. Pick at least two non-overlapping calibration points in the visible area of ​​each coordinate axis of the preprocessed chart image, and record the pixel coordinates and corresponding real data values ​​simultaneously to obtain a set of calibration points. Through a step-by-step guided interactive interface, the user is prompted sequentially within the pre-processed chart image to select at least two non-overlapping calibration points within the visible area of ​​each coordinate axis, simultaneously recording the pixel coordinates and their corresponding actual data values. For example, the user is guided to click the left and right endpoints of the X-axis, simultaneously recording the pixel coordinates. Compared with the actual data value input by the user .

[0027] S10122. Based on the set of calibration points, each coordinate axis is abstracted into an independent mathematical object containing type attributes, reference points and transformation functions. The master-slave relationship and geometric relationship of the multi-axis or multi-subgraph system are uniformly managed through the coordinate axis container, and linear axis transformation functions and logarithmic coordinate axis transformation functions are established. Each coordinate axis is abstracted as an independent mathematical object, containing type attributes (linear axis or logarithmic axis) and a pixel space reference point. Data space corresponding values ; The coordinate axis container is used to uniformly manage the master-slave relationship and geometric relationship in single X-axis, dual Y-axis, dual X-axis or multi-subgraph systems; Establish linear axis transformation functions and logarithmic coordinate axis transformation functions to achieve interpolation from pixel space to data space; The linear axis transformation function is: ; The logarithmic coordinate axis transformation function is: ; in, For the current pixel, These are the actual data values ​​of two known calibration points on the coordinate axis. ; For the current pixel The linear interpolation result of the corresponding linear axis; For the current pixel The corresponding logarithmic interpolation results for the logarithmic coordinate axes.

[0028] S10123. When the number of calibration points for any coordinate axis in the calibration point set exceeds two, the least squares fitting algorithm is used to optimize the linear transformation parameters of the coordinate axis, minimize the sum of squared calibration errors, and obtain the optimal linear transformation parameters. When the user selects more than two calibration points, the system uses a least-squares fitting algorithm to optimize the scaling parameters and minimize the sum of squared calibration errors, thereby improving the coordinate mapping accuracy within the observable region. The objective function of the least-squares estimation is: ; in, Let be the linear transformation parameters to be optimized, where This is the proportionality coefficient (slope). The optimal value is determined by minimizing the sum of squared errors, where the offset (intercept) is the value. The number of calibration sampling points, i.e., the total number of calibration points picked by the user on the coordinate axis; For sampling point index, For the first The pixel spatial coordinates of each sampling point, i.e., the pixel position when the user clicks on the calibration point in the image; For the first The actual data space value corresponding to each sampling point, that is, the actual physical quantity value corresponding to that pixel input by the user; By minimizing the true data values ​​of all sampling points Compared with linear predicted values By summing the squared errors between them, the optimal coordinate transformation parameters can be fitted. and This improves calibration accuracy.

[0029] S10124. Calculate the angular deviation between the actual tilt angle and the theoretical tilt angle of each coordinate axis based on the pixel coordinates in the calibration point set, classify the level to verify verticality and consistency, perform consistency check between the primary and secondary coordinate axes in the dual-axis system, and trigger a recalibration warning when the evaluation is poor. Based on the pixel coordinates of the coordinate axis calibration points, the difference between the actual tilt angle of the coordinate axis and the theoretical vertical or horizontal tilt angle is calculated using the arctangent function to obtain the angle deviation. The angle deviation is divided into four levels: excellent (deviation ≤ 1°), good (deviation ≤ 3°), average (deviation ≤ 5°), and poor (deviation > 5°) to verify verticality and consistency. Dual-axis systems require additional consistency checks between the primary and secondary axes; A recalibration warning is triggered when the evaluation is rated as "poor". The system summarizes the number of calibration points, verticality deviation, coordinate axis type, and related parameters, and generates a readable calibration quality assessment report.

[0030] S1013. Background filtering is performed on the chart image to generate a binary mask. Data series are determined by color clustering of the legend area. Constraints are applied to complex topological scenes within the data area. Pixel coordinate sets are extracted by fusing line center path, shape detection and connected component features. Specifically, it includes: S10131, Background Filtering and Foreground Mask Generation: Convert the chart image from RGB color space to HSV color space, and generate a foreground / target mask based on HSV threshold and region constraint to suppress background and interference areas. The RGB format chart image is converted to the device-independent HSV color space; white, black, and gray background areas are removed based on a combination of brightness and saturation thresholds to generate a foreground pixel set, achieving initial separation of the data series from the complex background. Specifically: Background mask ; The masks are white, black, and gray, respectively. " is the logical OR operator, representing the union operation between masks.

[0031] Foreground Mask ; " is the logical NOT operator; White mask ; black mask ; gray mask ; For pixel saturation, HSV color space components; For pixel brightness / luminance, HSV color space components; This is the minimum saturation threshold used to distinguish between color and gray. This is the maximum brightness threshold used to identify bright white areas; This is a conditional operator that returns 1 if the internal condition is met, and 0 otherwise. " is the logical AND operator.

[0032] S10132, Color Clustering and Data Series Determination: Determine the target color by combining auxiliary color sampling of the legend clipping area; generate HSV threshold and tolerance range for the target color, and generate a binary mask for single or multiple series extraction under constraints such as data area / interference area / exclusion area; Optionally, K-Means color clustering can be performed on the foreground pixels of the image to help discover candidate color series. A binary mask is generated for the candidate colors using the same method. The K-means clustering objective function is: ; This represents the total number of foreground pixels. , The number of cluster centers. ; For the first Feature vectors in each cluster For all A set consisting of cluster centers; It is the square of the Euclidean distance.

[0033] S10133, Series - Coordinate Axis Binding and Mapping Verification; Based on the already differentiated data series, establish the correspondence between each data series and the coordinate axis in the series configuration, including the horizontal and vertical axis identifiers used by each data series and the axis object corresponding to the range (for example, in a multi-vertical axis scenario, specify the series as the left Y-axis or the right Y-axis, and record the binding information that is consistent with the axis ID in the coordinate system object). Geometric relationship verification is performed on the coordinate axis mapping based on the calibration points; when the verification result does not meet the preset constraints, the mapping parameters are corrected or a recalibration prompt is output to the user so that subsequent pixel-to-data conversion adopts the updated mapping relationship; Based on the calibration points and the current mapping parameters, calibration reference information is generated. The calibration reference information includes at least accuracy prompts related to the range and the pixel spacing of the calibration points, which are used to help determine whether the calibration meets the extraction accuracy requirements.

[0034] S10134. Complex Topology Scene Constraint Screening: For complex topology conditions such as occlusion, fracture, and overlap, an interactive constraint and strategy-based processing approach is adopted to provide robust constraints for subsequent extraction; including: Occlusion handling: For areas where the data curve is occluded, extraction constraints are executed within the user-specified line area, and a mask for extraction is generated by combining a color filtering strategy, so that the target pixel set can still be stably obtained under occlusion conditions. Breakage handling: For broken or discontinuous regions of data curves, continuous sorting and path construction are performed on mask pixels, and continuous line points are sampled and preserved according to density patterns, so that an ordered sequence of curve pixels can still be obtained under the condition of breakage. Overlap processing: To address the aliasing problem caused by overlapping multiple manually drawn areas, the overlapping parts are first detected, and discrete overlapping points are connected and clustered based on a distance threshold; then, Laplacian local sharpness detection is performed on the unclear parts within the drawn area to generate an enhanced extraction mask; thereby improving the reliability of data extraction in overlapping and unclear scenarios. After the sharpness detection, strategies such as "prioritizing the extraction of overlapping and unclear areas, extracting only overlapping areas, extracting only unclear areas, or excluding problem areas" are used to generate an enhanced extraction mask.

[0035] S10135, Multi-feature pixel-level extraction; For line-type data series, the morphological skeletonization algorithm is used first to extract the center line and downsample according to the sampling step size. When skeletonization is not available, the graph is scanned column by column along the horizontal direction of the X-axis pixels, and the median of each column is taken as the center path. The center line or center path is then converted into the first pixel coordinate set. For discrete point data series, a circular detection and contour approximation / centroid localization algorithm is used to identify various types of marker points, and the marker points are converted into a second pixel coordinate set; Merge the first set of pixel coordinates with the second set of pixel coordinates to output a high-precision set of extracted data points.

[0036] S1014. The pixel coordinate set is converted into physical values ​​through a multi-coordinate system mapping model, and the initial data point set is obtained through density control, noise removal, and duplicate point detection and merging. Specifically, it includes: S10141, Coordinate Transformation: The obtained pixel coordinate set is batch transformed using the constructed unified transformation function, and the corresponding coordinate axis parameters are selected according to the determined data series and coordinate axis correspondence. If the coordinate axis mapping parameters have been corrected, the corrected parameters are used for transformation, thereby mapping the pixel coordinates to real data values ​​and forming a structured initial dataset. On the linear axis, reference point: pixel coordinates With corresponding data values ; The coordinate transformation relationship is as follows: ,in, , .

[0037] The converted physical quantity values ​​can be associated with the unit attributes of the corresponding coordinate axes or data series; in the structured output, the unit information can be saved as coordinate axis metadata, series metadata, or data point fields to supplement the semantic expression of the physical quantity values.

[0038] S10142. Density control and functional relationship constraints: Density control is performed based on the preset maximum number of points threshold and minimum point distance threshold, and dense point sets exceeding the threshold are sparsified; after clustering by x-coordinate, the median of y-value is taken to ensure that each x-coordinate corresponds to only a single y-value to achieve functional relationship constraints.

[0039] The constraints for sparsification are: ; This is the set of raw data points extracted from the chart. The subset of retained data points after density-controlled filtering satisfies all constraints; It is a subset relation; To preserve subsets Two different data points in the data; For point The Euclidean distance between them; This is the preset minimum point spacing threshold; To preserve subsets The cardinality; This is the preset maximum number of points threshold.

[0040] The functional relationship of a single-valued mapping is: ; The processed output y-value. It is a median function; For the first data point in the original data set The y-coordinate values ​​of each point. For the first data point in the original data set x-coordinates of each point; Let x be the x-coordinate of the target to be solved. The x-value is the preset clustering window radius; " is the condition separator.

[0041] S10143, Duplicate Point Detection and Merging: Redundant points with a distance less than a threshold are identified through a duplicate point detection mechanism, and the redundant points are merged to eliminate noise; Redundant points with a detection distance less than a set distance threshold (e.g., 3 pixels) can be merged using five merging strategies: average, median, centroid, and retaining the first / last point. For smooth curves, the average value is preferred. For data with high noise, the median should be used first. For markers that require precise positioning, the centroid is used; For the start and end points of time series data, the first and last points are retained respectively.

[0042] S10144. Save status snapshot; All cleaning parameters and structured data are saved as JSON snapshots via State objects, supporting subsequent batch processing and configuration replay.

[0043] S1015. Establish chart metadata and confirm the initial data point set; Summarize the structured information generated during S1011 to S1014 to establish chart metadata, and embed it in the state snapshot of the initial data point set in JSON format; The chart metadata may include the physical meaning and range of the coordinate axes, the data series name and the bound coordinate axis object, the chart type identifier, the spatial location information of the data area and the legend area, and the unit attributes of each data series; The initial data point set contains at least one data series, and each data point has physical coordinate values. When executing S2, the acquired chart metadata or user input information can be used as part of the prompt word context.

[0044] S102. Obtain the external description text associated with the chart, wherein the external description text contains a semantic description of the chart data features, serving as a source of benchmark truth values ​​required for correction; The external descriptive text includes, but is not limited to, the main body of the paper, figure and table annotations, experimental notes, figure and table titles and captions, and textual descriptions of the figure and table data in the main body.

[0045] S2. Construct prompt words or parsing rules based on chart-related information, extract key data points from external descriptive text, and perform format checks or confidence assessments on the extraction results to output structured key data points containing horizontal and vertical coordinate values. Specifically, it includes: S201. Inject chart metadata, define task roles, constrain output format and enhance semantic alignment between text and chart metadata, and build dedicated prompt words for semantic parsing of large models; Specifically, it includes: S2011. Inject the chart metadata into the context of the prompt words to establish the association constraints between the numerical values, physical meaning and coordinate position; The chart metadata obtained in S1 is injected into the context of the prompt words. The chart metadata includes at least the physical meaning and range of the coordinate axes, the data series name and the bound coordinate axis object, and the chart type identifier. Through metadata constraints, the large language model establishes the association between "numerical value - physical meaning - coordinate position" during the semantic parsing process, avoiding mismatch of numerical values ​​across ranges or series.

[0046] S2012, Define the model task and expert roles, and clearly extract key data points that can be used as truth references; The extraction target is limited to extracting key data points from scientific literature descriptions that can serve as numerical references. The model role is set as a "scientific data verification assistant," which prioritizes the identification of experimental results, peak records, critical conditions, or comparison benchmarks containing clearly labeled numerical values.

[0047] S2013, constrains structured JSON output and provides guidance with few samples, standardizes model output format and recognition patterns; The model can be required to return extraction results in structured JSON format, with each key data point containing at least horizontal and vertical coordinates. When sufficient textual information is available, it can also include fields for identifying the data series, physical units, and traceability or description of the original text fragment. A small number of sample examples can be provided based on the target chart domain to improve the model's stability in recognizing scientific literature representations.

[0048] S2014. Establish cross-modal alignment between text and charts, identify contradictions and add reverse constraints, and finally form special prompt words for semantic parsing of large models; When the external descriptive text contains a chart reference identifier, a connection can be established between the text paragraph and the chart image. When a multimodal model is available, the chart image or chart metadata, along with the corresponding text paragraph, can be used as input to the model to help it understand the chart context. For obviously contradictory or uncertain descriptions, hints or constraints can be added to the prompt words to reduce the risk of incorrect extraction.

[0049] S202. Based on the prompt words, call the large language model to perform structured extraction of the external description text to obtain a set of candidate key data points; when the model returns an incomplete result format or fails to parse, trigger the step-by-step rollback strategy of JSON repair, rule parsing rollback and text value extraction in sequence. In one enhanced implementation, two channels, direct numerical extraction and semantic reconstruction extraction, can be used to generate candidate results respectively, and the two sets of candidate results can be matched and fused or conflict-reduced. When extraction fails, the system automatically recovers by repairing JSON, rolling back rules, and extracting pure numerical values ​​step by step to ensure that the parsing process is not interrupted. When the model returns an incomplete result or fails to parse it, the following fallback strategies are triggered in sequence: JSON Repair: Perform format repair, field parsing, or type conversion on non-standard JSON text returned by the model; Rule parsing rollback: Re-extract candidate values ​​using regular expression templates or text rules; Text numerical extraction: When structured semantic results cannot be obtained, it degenerates into pure numerical extraction mode and marks the results as having low confidence.

[0050] S203. Perform three checks on the candidate key data points in sequence: numerical rationality, semantic consistency, and structural integrity. Adaptively determine the credibility of the candidate key data points and output structured and credible key data points that meet the correction requirements. Specifically, it includes: S2031. Perform coordinate range, numerical format, or outlier validation on candidate key data points to ensure numerical rationality. If the candidate key data points or chart metadata contain explicit unit information, check whether the units of the key data points are consistent with the coordinate axis units, and perform unit conversion according to preset rules or prompt the user for confirmation.

[0051] S2032. Text source tracing or trend consistency checks can be used to help confirm the semantic consistency between key data points and the original text description. Text source tracing can be used to confirm whether key data points originate from text fragments containing explicit numerical descriptions; trend consistency checks can be used to determine whether there are obvious contradictions between the extraction results and the qualitative trend descriptions in the text.

[0052] S2033. Check the series coverage and horizontal axis density, and complete the structural integrity verification of key data points; Series coverage checks are used to determine whether key data points cover the target data series; The x-axis distribution density check is used to determine whether the x-axis distribution is too sparse; When the inspection results do not meet the preset requirements, prompt the user to supplement external description text or perform manual confirmation.

[0053] S2034. Based on the combined results of the three verifications, perform confidence level classification and weight mapping, and output structured and reliable key data points; A confidence score can be calculated for each key data point based on the results of checks for numerical reasonableness, semantic consistency, or structural integrity. The confidence score can be represented as a continuous numerical value or mapped to discrete levels such as high, medium, low, or discard, depending on application needs. Different levels or confidence scores can be converted into corresponding weights during subsequent bias aggregation. The specific number of levels, weight range, and normalization method can be configured according to the application scenario.

[0054] S3. Search for the matching point with the best geometric proximity to each key data point in the initial data point set, establish a one-to-one mapping relationship between key data points and matching points, and calculate the deviation of each matching pair in the horizontal and vertical coordinate directions. Specifically, it includes: S301. Filter the corresponding data series based on the series identifier of the key data points. If there is no identifier, perform nearest neighbor matching in the entire series to determine the candidate data point set. Based on the series of identifiers in the key data points, locate the candidate matching range in the initial data point set: When a key data point contains a series identifier, the corresponding data series is filtered based on the series identifier; When a key data point does not contain a series identifier, nearest neighbor matching can be performed across multiple data series to determine the set of candidate data points corresponding to the key data point.

[0055] S302. In the candidate subset, use distance calculation to search for the nearest matching point to each key data point, and establish the corresponding matching relationship between key data points and matching points; The nearest neighbor search searches for the closest data points to each key data point within the candidate subset, establishing a matching relationship. The nearest neighbor search can use two-dimensional Euclidean distance, or, depending on the chart type or data distribution, can use horizontal axis distance, vertical axis distance, or weighted distance. In an optional implementation, a mismatch threshold can be set; when the minimum distance exceeds the threshold, the key data point is marked as low confidence or the user is prompted for confirmation.

[0056] S303. Calculate the deviation between key data points and corresponding matching points in the horizontal and vertical coordinate directions, and quantify the systematic offset of data extraction.

[0057] S4. Perform statistical aggregation on the deviation to calculate the correction parameters, perform coordinate translation transformation on the initial data point set based on the correction parameters, and output the corrected data set; Specifically, it includes: S401. Calculate the global correction parameters based on the set of deviations for each matching pair. The uncertainty of the correction parameters is assessed based on the degree of dispersion of the deviation. When the deviation shows obvious regional differences, the local correction parameters are calculated separately for each horizontal coordinate interval, and the correction results of adjacent intervals are smoothed by interpolation.

[0058] S402. Apply the correction parameters to the initial data point set to perform a coordinate translation transformation to obtain the corrected data; S403. Calculate the offset statistics or error assessment indicators before and after the correction, and evaluate the correction effect. When the evaluation result is lower than the preset requirements, the user is prompted to verify the external text, key data points or correction parameters; the output includes the correction parameters, key data point matching information, offset statistics, confidence information and a structured correction result comparing the before and after correction.

[0059] S5. Generate a visualization of the before-and-after correction results, overlay the key data points as an external text reference benchmark, and output a statistical report on the correction effect. Specifically, it includes: S501, Generate multi-layered overlay visualization charts; The relative positional relationship between the original data, the corrected data, and the external text reference points is presented intuitively using spatial comparison. S502. Differentiate and label each data layer; Differentiated visual styles are used to distinguish original data points, corrected data points, key data points and their matching points, with key data points highlighted to emphasize their reference role.

[0060] S503, Output structured correction results; Generate structured calibration results containing correction parameters, matching details, and performance metrics to record the main calibration process and results.

[0061] In summary, the method disclosed in this embodiment, through the synergistic utilization of chart data extraction results and external text semantic information, obtains key data points from external descriptive text that can serve as references, and uses these points to estimate and correct deviations in the extracted data, thereby improving the verifiability, correction efficiency, and data reuse value of the chart data extraction results. In some implementations, error assessment indicators, confidence information, and manual verification mechanisms can also be combined to assist in the verification of text parsing results and numerical correction results.

[0062] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A data extraction error self-correction method based on external text semantic parsing, characterized in that, include: Obtain an initial set of data points, including chart metadata, obtained after processing the chart image through coordinate calibration, color analysis, and multi-feature extraction, as well as external descriptive text containing the key data features of the chart; Based on chart-related information, construct prompt words or parsing rules, extract key data points from external descriptive text, and perform format checks or confidence assessments on the extraction results to output structured key data points containing horizontal and vertical coordinate values. Search for the matching point with the best geometric proximity to each key data point in the initial data point set, establish a one-to-one mapping relationship between key data points and matching points, and calculate the deviation of each matching pair in the horizontal and vertical coordinate directions. Perform statistical aggregation on the deviation to calculate the correction parameters, perform coordinate translation transformation on the initial data point set based on the correction parameters, and output the corrected data set.

2. The data extraction error self-correction method based on external text semantic parsing according to claim 1, characterized in that, After outputting the corrected data set, a visualization result comparing the results before and after the correction is generated, the key data points are overlaid as external text reference benchmarks, and a statistical report on the correction effect is output.

3. The data extraction error self-correction method based on external text semantic parsing according to claim 1 or 2, characterized in that, The acquisition of the initial data point set includes: Obtain an initial set of data points generated by the chart data extraction process, the initial set of data points including at least one data series and its physical coordinate values; Obtain the chart metadata associated with the initial set of data points. The chart metadata includes one or more of the following: axis information, data series information, axis binding relationships, chart type information, or unit information.

4. The data extraction error self-correction method based on external text semantic parsing according to claim 3, characterized in that, The chart data extraction process includes: Perform region labeling preprocessing on the chart image to divide it into data region, legend region, and interference region; Construct a multi-coordinate system mapping model to establish the binding relationship between each data series and the corresponding coordinate axis in a multi-coordinate axis scenario; perform background filtering on the chart image to generate a binary mask, combine the legend area color clustering to determine the data series, and perform constraint filtering on complex topological scenes within the data area, and integrate line center path, shape detection and connected component feature extraction to extract the pixel coordinate set; The pixel coordinate set is converted into physical values ​​through a multi-coordinate system mapping model, and then the initial data point set is obtained through density control, noise removal, and duplicate point detection and merging.

5. The data extraction error self-correction method based on external text semantic parsing according to claim 4, characterized in that, The complex topology scene includes occlusion, fracture, and overlap, and the constraint filtering includes: For areas where data curves are obscured, extraction constraints are applied within user-specified lined areas, and a mask for extraction is generated using a color filtering strategy. For regions with broken data curves or discontinuous topology, continuous sorting and path construction are performed on mask pixels, and continuous line points are sampled and retained according to density patterns. To address the aliasing issue caused by overlapping multiple manually drawn areas, the overlapping parts are first detected, and discrete overlapping points are connected and clustered based on a distance threshold. Then, Laplacian local sharpness detection is performed on the unclear parts within the drawn areas to generate an enhanced extraction mask.

6. The data extraction error self-correction method based on external text semantic parsing according to claim 1 or 2, characterized in that, The step of constructing prompt words or parsing rules based on chart-related information and extracting key data points from external descriptive text includes: Inject chart metadata, define task roles, constrain output format, and enhance semantic alignment between text and chart metadata; construct dedicated prompt words for semantic parsing of large models. Based on the prompt words, the large language model is invoked to extract structured data from the external descriptive text and obtain a set of candidate key data points. When the model returns an incomplete result or fails to parse, a step-by-step rollback strategy is triggered, which includes JSON repair, rule parsing rollback, and text value extraction. The candidate key data points are sequentially checked for numerical rationality, semantic consistency, and structural integrity. Based on the check results, the credibility of the candidate key data points is determined, and structured credible key data points are output.

7. The data extraction error self-correction method based on external text semantic parsing according to claim 6, characterized in that, Injecting chart metadata, defining task roles, constraining output format, and enhancing text-chart alignment, we construct dedicated prompt words for large-scale model semantic parsing, including: Inject chart metadata into the context of prompt words to establish the association constraints between numerical values, physical meaning, and coordinate location; Define the model task and expert roles, and clearly extract key data points that can be used as truth references; Constrain structured JSON output and provide guidance with few samples to standardize model output format and recognition patterns; Establish semantic alignment between external descriptive text and chart metadata, add cue constraints to descriptions that are obviously inconsistent or uncertain, and finally form special cue words for semantic parsing of large models.

8. The data extraction error self-correction method based on external text semantic parsing according to claim 1 or 2, characterized in that, Searching for matching points with the best geometric proximity to each key data point in the initial data point set, establishing a one-to-one mapping relationship between key data points and matching points, and calculating the deviation of each matching pair in the horizontal and vertical coordinate directions; including: Based on the series identifiers of key data points, filter the corresponding data series. If there are no identifiers, perform nearest neighbor matching in the entire series to determine the candidate data point set. In the candidate subset, the nearest matching point to each key data point is searched using distance calculation, and the corresponding matching relationship between key data points and matching points is established; Calculate the deviation between key data points and their corresponding matching points in the horizontal and vertical coordinate directions to quantify the systematic offset in data extraction.

9. The data extraction error self-correction method based on external text semantic parsing according to claim 1 or 2, characterized in that, Perform statistical aggregation on the deviation to calculate correction parameters, perform coordinate translation transformation on the initial data point set based on the correction parameters, and output the corrected data set; including: Calculate the global correction parameter based on the set of deviations for each matching pair; The correction parameters are applied to the initial data point set to perform a coordinate translation transformation, resulting in corrected data; Calculate the offset statistics or error assessment indicators before and after the correction to evaluate the effect of the correction.

10. The data extraction error self-correction method based on external text semantic parsing according to claim 2, characterized in that, Generate a visualization of the before-and-after correction results, overlay the key data points as an external text reference benchmark, and output a statistical report on the correction effect; including: Generate multi-layered visual charts; intuitively present the relative positional relationship between the original data and the corrected data and external text reference points through spatial comparison; Differentiate and label each data layer; use differentiated visual styles to distinguish original data points, corrected data points, key data points and their matching points, with key data points highlighted. Output structured correction results; generate structured correction results containing correction parameters, matching details and effect indicators, used to record the main correction process and results.