System and method for extracting and validating multi-modal data

GB2704265APending Publication Date: 2026-08-26GIST ADVISORY PTE LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
GB2025020592
Authority / Receiving Office
GB · GB
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-11-27
Filing Date
2025-12-02
Publication Date
2026-08-26

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

A computer-implemented data-extraction method is disclosed, comprising ingesting multi-modal data from a respective plurality of data sources 102. The ingested multi-modal data is integrated by retrie
Need to check novelty before this filing date? Find Prior Art

Description

Field of Invention The present invention relates to artificial intelligence (Al) systems for data extraction and validation, and more particularly to a system and method for multi-agent, multi-modal Al extraction and validation of complex metrics from unstructured data sources. Technical Background Modern organizations generate and rely upon vast quantities of data from diverse sources to make informed decisions, for example those regarding sustainability, environmental impact, and regulatory compliance. This data is fragmented and typically exists in multiple formats including textual documents, tabular data, geospatial information, images, and various structured and unstructured datasets. Traditional data processing systems often employ single-modal approaches that analyse one type of data at a time, requiring separate processing pipelines for different data formats. Data is typically extracted using basic retrieval-augmented generation (RAG) techniques, or simple parsing algorithms, to identify and extract relevant information from these disparate sources. Existing data extraction systems face significant challenges when processing the volume and variety of information required for comprehensive analysis. Organisations typically track hundreds of metrics, and taking sustainability data as an example, approximately 70% of such metrics lack standard definitions across different reporting frameworks. As a consequence of both the lack of standard definitions and the volume of data to be processed, current systems struggle to integrate information from multiple data modalities simultaneously, often resulting in incomplete or inconsistent metric extraction. The variety of reporting standards and frameworks that must be used creates further complexity as data must be mapped and validated against multiple regulatory requirements. In addition to the complexities associated with integration of multi-modal information, traditional approaches are known to suffer from limited accuracy and reliability. Systematic errors are common, with a company’s assets being incorrectly classified or tagged. Inaccuracies in location information can lead to a remote industrial facility such as a coal mine being incorrectly identified as a commercial office building. When considering the millions of corporate assets that might be held in a particular dataset, and the trillions of dollars of investment decisions that may be associated with such assets, it is clear that the impact of such systematic errors is extremely significant. The technical challenges inherent in existing multi-modal data retrieval and extraction systems is such that conventional systems typically operate on annual data cycles, allowing sufficient time to collect and integrate information from different datasets. The temporal lag between data collection, processing, and availability makes real-time decision-making impossible, while it can also be difficult to trace the provenance of extracted metrics back to their original sources or understand the reasoning behind specific classifications or validations. Current validation methodologies lack the sophistication to cross-reference multiple data sources and apply contextual analysis to verify the accuracy of extracted information. In summary, current data processing architectures struggle with scalability when handling large-scale, global datasets spanning thousands of companies and millions of assets. The integration of historical context and forward-looking analysis remains limited in existing systems, reducing their effectiveness for comprehensive risk assessment and regulatory compliance. Furthermore, traditional approaches provide limited customization capabilities, making it challenging to adapt extraction processes to specific organizational requirements or emerging regulatory frameworks. Embodiments of the present invention are developed in this context, and the methods and systems that are described aim to address at least some of the challenges set out above. Summary of Invention According to an aspect of the present invention, there is provided a computer-implemented data-extraction method, comprising: ingesting multi-modal data from a respective plurality of data sources; integrating the ingested multi-modal data by retrieving context-specific information from the ingested multi-modal data, and extracting context-specific quantitative and / or qualitative metrics from the retrieved context-specific information; verifying the extracted metrics against one or more predetermined confidence thresholds; generating an integrated report of extracted metrics meeting the one or more predetermined confidence thresholds, and outputting the integrated report. The ingestion of multi-modal data may be performed by integrating the functionalities of a plurality of specialised agents, working in combination to ensure that errors in data from one source are captured using reference data from a different source. For example, errors in text can be identified based on comparison with maps or images. In embodiments, the ingesting step further comprises accessing a historical database containing data to provide historical context for the multi-modal data, wherein the verifying steps uses the data in the historical database to determine whether the extracted metrics meet the one or more predetermined confidence thresholds. In embodiments, the ingesting step comprises identifying optimal data sources based on probabilistic analysis of data source relevance to a topic associated with the integrated report, and ingesting the multi-modal data comprises ingesting only data from the identified optimal data sources. In this way, processing of redundant information can be avoided. In embodiments, the ingesting step comprises identifying optimal data sources based on anticipated processing complexity associated with information from each data source, and ingesting the multi-modal data comprises ingesting only data from the identified optimal data sources. Such embodiments present an alternative way in which processing of redundant information can be avoided. In embodiments, verifying the extracted metrics comprises identifying inconsistencies and / or applying contextual analysis to verify accuracy against industry benchmarks and historical patterns. In embodiments, the method further comprises generating a provenance trail that links each validated metric to source evidence and reasoning. In this way, the output report does not provide only answers to questions, but also shows the specific evidence it used to derive the answer, so that it is possible to trace the origins of each result. In embodiments, the method comprises receiving instructions specifying a report topic; one or more indicators associated with the report; quantitative definitions of qualitative metrics; and an integrated report format. ln embodiments, the method further comprises outputting one or more updates to the instructions to reduce an amount of metrics failing to meet the one or more predetermined confidence thresholds. In embodiments, the method comprises receiving the one or more indicators from a database of indicators associated with a group of predetermined topics. In embodiments, verifying the extracted metrics comprises accessing one or more previously-generated integrated reports from a library and comparing the extracted metrics with the content of the previously-generated integrated reports, wherein the integrated report defines changes relative to the one or more previously-generated integrated reports. According to a further aspect of the present invention, there is provided a computer program which, when executed by a processor, is arranged to perform the above method. According to a further aspect of the present invention, there is provided a system comprising storage means and one or more processors, the storage means arranged to store one or more sets of computer-executable instructions which, when executed by the one or more processors, is arranged to perform the above method. Brief Description of Drawings Embodiments of the present invention will be described by way of example only, with reference to the accompanying drawings, of which: Figure 1 illustrates a data processing system according to embodiments of the present invention; Figure 2 illustrates a method of processing data according to embodiments of the present invention; and Figure 3 illustrates a schematic of the data flow of the process of quantitative data extraction, according to embodiments of the present invention. Detailed Description Figure 1 illustrates a data processing system 100 according to embodiments of the present invention. In terms of functionality and data flow, the system 100 represents a plurality of architectural layers or stages, including an ingestion layer 114, processing modules 122, and an output layer 134. The system interfaces with a data source layer 102 which contains data which passes through the system 100 in order to be processed into a format to be output by the output layer 134. In particular, the system 100 processes multi-modal data through a multi-modal architecture to generate validated outputs with associated confidence measures and traceability information. The system 100 interfaces with a historical database 132 for the purpose of data validation, in a manner to be described in more detail below. The data source layer 102 comprises multiple data sources 110 including, but not limited to, documents 104, geospatial data 106, Internet of Things (loT) sensors 108, government and public datasets 110, and visual data (images, heatmaps, videos) data 112. These data sources provide diverse input formats to the system 100 for processing and analysis. The data may be user-owned, third-party or partner-owned, or associated with a combination of ownership models. For example, the documents 104 may provide unstructured textual content from corporate reports, regulatory filings, and disclosure documents, including, but not limited to, text documents, portable document format (.pdf) files, spreadsheets, presentations, and so on. The geospatial data source 106 may supply location-based datasets, mapping information, and spatial data layers. The loT sensors 108 may collect real-time data streams, such as data representing energy consumption, biodiversity parameters, temperature, air quality, and so on. The government datasets 110 may provide regulatory frameworks, compliance requirements, and official disclosure standards. This information typically includes publicly available datasets from regulatory agencies, environmental monitoring organizations, and policy enforcement bodies. The system 100 processes government datasets 110 to provide regulatory context and compliance validation for extracted metrics. The image data source 112 may include photographs, satellite imagery, graphical representations, motion patterns, and video data. Further data sources (not shown) may also be envisaged, such as audio data sources, capturing audio of animals for, for example, generation of biodiversity reports. The ingestion layer 114 acts to integrate data from the different sources in the data source layer 102, so that it can be correlated, validated, and synthesized. Multi-modal integration allows the system 100 to leverage complementary information across data formats to improve accuracy, completeness and robustness of extracted metrics compared to single-modal processing approaches. For example, cross-referencing of data from different sources, based on retrieval and conversion of data into uniform "tokens" where images and text, etc. can be consumed together, enables processing of high-quality data. In particular, the fusion of multi-modal data by the ingestion layer 114 is achieved in the embodiment illustrated in Figure 1 by means of three types of specialized agents: classifier agents 116, extractor agents 118, and quality-check agents 120, arranged as a sequential pipeline. The classifier agents 116 receive input from all data sources in the data source layer 102, and are responsible for identifying and categorizing information from the relevant data sources, determining data type and format, and routing data to appropriate processing pathways, thus enabling selective handling. For example, in a case where images and a text description are retrieved, the classifier agents 116 my determine that both information sources relate to the same information. Classification may be performed using any suitable analysis technique in the art, taking into account the content and format of the information extracted. The extractor agents 118 apply domain-specific extraction algorithms to retrieve quantitative and qualitative information. The extraction process produces structured output data from unstructured or semi-structured input sources. The quality-check agents 120 validate outputs through systematic verification procedures, apply validation rules and anomaly detection, and may generate confidence scores, although as described below these may also be generated in the output layer 134. The quality-check agents 120 may flag low-confidence results for additional review. A quality-checking framework of the type provided by the quality-check agents 120 enables addressing of technical challenges in data reliability and completeness through systematic validation processes and automated quality control mechanisms. Traditional data processing systems face limitations in handling incomplete datasets and ensuring consistent data quality across diverse data sources. The quality assurance framework provides comprehensive validation capabilities that maintain data integrity throughout the multi-modal processing pipeline. The processing layer 122 is responsible for taking the data ingested by the ingestion layer 114 and arranging and structuring it into a form in which it can be prepared for output by the output layer 134, and verified for accuracy, The processing layer 122 includes specialized processing components. In the embodiment illustrated in Figure 1, such processing components include geospatial Al 124, context-tuned LLMs 126, data verification 128, and scenario analysis 130. The geospatial Al 124 processes geospatial data from multiple sources including satellite imagery providers, aerial survey datasets, and ground-based observation systems. Spatial coordinates are maintained, together with asset type classifications, and verification status information for each tracked asset. This large-scale tracking capability enables comprehensive monitoring of corporate assets, infrastructure facilities, and industrial operations across different geographic regions. Context-tuned LLMs 126 process natural language content and apply domain-specific language understanding to the input data. In some embodiments, the LLMs 126 are finetuned on sustainability, environmental, and regulatory datasets to understand specialized terminology, reporting frameworks, and industry-specific language patterns, although other contexts may also be employed where required. In this way, ambiguity can be resolved, and implicit relationships between data elements can be understood. The models may incorporate knowledge of multiple regulatory standards such as banking authority guidelines, sustainability reporting frameworks, and compliance requirements. The use of context-tuned LLMs 126 provides an advantage over more general-purpose algorithms or Al modules integrated into web browsers, for example, which may attempt to fit particular template language to retrieved data, rather than generating language appropriate to a particular context. As such, the context-tuned LLMs 126 are deterministic and avoid hallucinations, ensuring consistency and accuracy of output. Data verification component 128 cross-references multiple data sources and applies contextual analysis to verify accuracy. For example, the data verification component 128 may evaluate factors such as industry benchmarks, company size, geographic location, and operational characteristics to identify potentially erroneous values. In particular, extracted metrics are compared against historical patterns stored in historical database 132 to detect anomalies or unusual variations. The data verification component 128 may flag metrics that deviate significantly from established company-specific or industry-wide trends for additional review. A further function of the data verification component 128, in some embodiments, is the evaluation of the quality and reliability of source documents, considering factors such as document completeness, formatting consistency, and information clarity. The component may assign lower confidence scores to metrics extracted from poor-quality sources. Scenario analysis component 130 generates projections and predicts potential future outcomes. In particular, the scenario analysis component 130 extrapolates historical trends from historical database 132 and incorporates external factors to predict multiple future scenarios. In some embodiments, the scenario projections are time-dependent, to show how different factors may evolve over various time horizons, supporting both shortterm and long-term planning requirements. The extractor agents 118 of the ingestion layer 114 connect to both the geospatial Al 124 and context-tuned LLMs 126. The quality-check agents 120 connect to the data verification 128, while the geospatial Al 124 connects to the scenario analysis 130. The historical database 132 provides data storage and retrieval capabilities, connecting (not shown) to both the extractor agents 118 and quality-check agents 120 to support processing operations with historical context and validation information. In particular, the historical database 132 maintains structured records of extracted metrics, validation results, and contextual information across multiple years, supporting temporal analysis and trend identification. In one example, the historical database 132 may store records of floods in a particular region, in terms of dates, water depths, and weather conditions in both the short and longer-term periods leading up to the flood, and the financial and constructional damage associated with each flood. In the embodiment illustrated in Figure 1, the output layer 134 contains three output components, responsible for validated metrics 136, confidence scores 138, and a provenance trail 140. The validated metrics component 136 provides verified quantitative and qualitative measurements, which comply with standardized formats, regulatory requirements and industry frameworks. The confidence scores component 138 quantifies reliability and uncertainty levels for each metric. In embodiments, at least some of the functionality of the data verification component 128 may be implemented in the confidence scores component 138. In other embodiments, the confidence scores component 138 provides additional functions such as comparing results across multiple processing attempts and validation checks. The component 138 may thus identify variations in extracted values and reduce confidence scores when inconsistencies are detected across different extraction runs. The provenance trail component 140 links each metric to source evidence, validation steps, and decision rationale. The links may include including specific page numbers, table references, or image coordinates. The provenance trail component 140 records all validation decisions made by the data verification component 128, including which validation rules were applied, what cross-references were checked, and why specific validation outcomes were reached. The provenance trail component 140 may document both successful validations and failed validation attempts with explanatory rationale. The quality-check agents 120 connect to both validated metrics 136 and confidence scores 138 components (not shown), while the data verification component 128 connects to the provenance trail component 140. The distributed processing illustrated in Figure 1 allows parallel operation of multiple agents across different data sources simultaneously. Agent specialization enables optimized processing algorithms tailored to specific data types and extraction requirements. The modular architecture supports scalable processing capabilities where additional agents can be deployed to handle increased data volumes or new data source types. Figure 2 illustrates a worked example of the operation of the system 100 of Figure 1, according to an embodiment of the present invention. The example that is presented is described in the context of a task to generate a report on deforestation levels in a particular area, attributed to activities of a particular organisation or group of companies. The report is described herein as an integrated report, presented in a context-specific format and integrating qualitative and / or quantitative metrics. In step S201, the system 100 initiates deforestation analysis by accessing standardized definitions and indicators associated with the topic ‘deforestation’ through government datasets 110, retrieving regulatory frameworks, forest classification criteria, and sustainable finance taxonomy definitions. Context-tuned LLMs 126 process these multiple regulatory frameworks to establish physical deforestation thresholds to which observations can be compared. For example, such thresholds may incorporate minimum forest area thresholds of 0.5 hectares, tree height requirements of 5 meters at maturity, canopy cover percentages with 10% minimum coverage, and temporal analysis periods for change detection. In step S202, classifier agents 116 identify and categorize relevant data sources across multiple modalities. Geospatial data component 106 provides satellite imagery from multiple temporal periods, while document component 104 supplies corporate sustainability reports and land use permits. Government datasets 110 contribute official forest inventory records and protected area boundaries, and loT sensors 108 deliver ground-based forest monitoring station data and biodiversity sensor information. In step S203, geospatial Al component 124 processes the satellite imagery to identify forest boundaries using spectral analysis techniques, calculating baseline forest coverage for the specified region and detecting temporal changes in forest cover between time periods. The system classifies land use changes including agricultural conversion, urban development, and logging activities, while generating confidence maps that show certainty levels for detected changes. In step S204, which may be performed in parallel with S203 in some embodiments, extractor agents 118 process corporate reports and regulatory filings to extract disclosed deforestation metrics and commitments, identify supply chain relationships with deforestation-risk commodities, and retrieve land acquisition and development project information. Context-tuned LLMs 126 interpret sustainability language to distinguish between deforestation prevention policies and actual performance data, extracting quantitative targets and achievement information while identifying the geographic scope of corporate operations. In step S205, data verification component 128 performs comprehensive cross-modal validation by comparing satellite-detected forest loss with corporate disclosures, crossreferencing geospatial changes with reported land use activities, and validating corporate claims against government forest inventory data. The system checks consistency between different satellite data sources and temporal periods to ensure accuracy across multiple information streams. In step S206, historical database 132 provides temporal context by comparing current deforestation rates with historical regional patterns, identifying seasonal variations and natural forest cycle patterns, and assessing long-term deforestation trends over multiple years. The system flags unusual deforestation spikes or anomalous patterns that deviate from established baselines. Additionally, undisclosed data gaps in data sources can be filled. In step S207, quality-check agents 120 apply validation frameworks to verify satellite imagery quality and cloud cover limitations, assess temporal alignment between different data sources, and evaluate completeness of corporate disclosure coverage. Confidence scores component 138 generates reliability metrics by assigning confidence levels based on satellite image resolution and clarity, weighting scores according to corporate disclosure completeness, and factoring in validation consistency across multiple data sources. In step S208, validated metrics component 136 compiles deforestation measurements for output, including total hectares of forest loss within the specified region and time period, deforestation rate calculations expressed as hectares per year, attribution analysis linking forest loss to specific activities or entities, and comparison metrics against regulatory thresholds and targets. These validated outputs provide quantified assessments suitable for regulatory reporting and investment decision-making. In step S209, provenance trail component 140 creates a comprehensive audit trail by linking each deforestation metric to specific satellite images and coordinates, documenting which corporate reports contributed to attribution analysis, recording validation decisions and confidence assessment rationale, and maintaining references to regulatory definitions and standards applied throughout the analysis process. In step S210, scenario analysis component 130 generates forward-looking projections by modelling potential future deforestation under different policy scenarios, assessing climate impact implications of detected forest loss, and generating risk indicators for continued deforestation in the region. These projections support both regulatory compliance planning and investment risk assessment. In step S211, a comprehensive deforestation report is output, that delivers quantified forest loss metrics with confidence intervals, attribution analysis linking deforestation to specific causes and entities such as deforestation-linked suppliers, forward-looking risk projections and scenario analysis, and a complete audit trail enabling independent verification of all findings. By including references to source data, it is possible to ensure that reasoning behind report content is possible. This integrated approach ensures that deforestation reporting meets both regulatory requirements and investment decisionmaking needs while maintaining full transparency and traceability throughout the analytical process. In embodiments, reports may be summarised, and in combination with the traceability described above, operational simplicity in quality assurance and delivery is ensured. The embodiment illustrated in Figure 2 sets out an example of the level of sophistication that can be achieved in a report generated using the system 100 of Figure 1. It will also be appreciated that the system 100 can be customised and configured in a number of different ways, depending on the requirements of a particular task, such as the topic, topic indicators, data to be included in a particular report, the sources of data to be consulted, and thresholds to be considered. In the first instance, the customisation is achieved by means of a report definition. In the example of Figure 2, the report definition can be derived from a government dataset 110 holding a particular definition. However, it will be appreciated that there are a number of circumstances in which such a definition either does not exist, or is not appropriate. For example, an organisation may wish to generate a report on financial performance that includes proprietary indices used in industry benchmarking. In this circumstance, a report definition may be selected from a series of predefined templates that the organisation frequently accesses, via a report format database which may be comprised within or accessible to, the ingestion layer 114, and controlled via a user interface. Such an interface to the ingestion layer 114 may allow for the generation of new report types on the fly. Depending on the nature of the report, some of data source layer 102 may not be consulted, or additional data source layer 102 not shown in Figure 1 may need to be consulted. More generally, it will be appreciated that reports may be generated on a variety of topics including, but not limited to, human rights, diversity, corporate governance, taxonomy eligibility, tracked across products, suppliers, targets, and so on. Each report is characterised by one or more key performance indicators (KPIs) extracted from the data source layer 102. Further configuration can be applied to the output layer 134 in order to specify parameters such as the presentation format, file type and data storage format, and encryption schemes and access privileges to be applied to the report that is generated. Such configuration may be performed via a user interface to the output layer, through which parameters may be selected or defined from menus, drop-down lists, check-boxes and the like. Similarly, the components of the processing layer 122 may be configured via a user input, or using machine-learning algorithms. For example, the context-tuned LLMs 126 may be tuned using data input from previous reports, and labels, so that the LLMs may develop its ability to interpret data correctly. Reports may be customised in terms of the quantitative and qualitative metrics which are included, particular topic flags or keywords, the nature of explainability or reasoning (for example, links to data sources), categorisation, and flexible metric weighting. In embodiments, the ingestion layer 114, and particularly the classifier agents 116, are configured to identify the best data sources in the data source layer 102 for interrogation. For this purpose, the ingestion layer 114 may include an intelligence module 119 which takes, as its input, the definition of report parameters in the manner described above, and automatically identifies both the type of data that may be most likely to contain pertinent information, and the locations within the data sources (for example, page numbers) of that information, fora particular topic. In doing so, the ingestion layer 114 is able to ensure that resources are used efficiently in processing data from the data source layer 102, avoiding the need to crawl vast amount of redundant information (that might otherwise occur if, for example, a full web search were to be performed). For example, satellite imagery need not be consulted for a report on financial performance. Text documents with particular phrases such as ‘estimated’, ‘projected’ or ‘draft’ may be excluded from consideration where historical performance records with particular numbers required. Articles from particular journalists may be preferred to those from others. Websites which are known to be regularly updated may be preferred to those which are dormant or known to contain out-of-date information. The intelligence module 119 executes one or more machine learning algorithms which determine, on a dynamic and probabilistic basis, the components of the data source layer 102 which are to most likely to provide appropriate results, based on a priori knowledge associated with a particular format of information, and a posteriori knowledge derived from previous analysis. The intelligence module 119 may alternatively, or additionally, take into account the anticipated processing complexity associated with a particular information source. For example, a lengthy report may consume large processing resources, whereas a consolidated report of equivalent data, such as a report by a journalist extracting key points or data, may be more efficient to process. Reports that are produced by the output layer 134 may be stored in the data source layer 102 and / or the historical database 132 in a library accessible in future comparisons with newly generated reports. This can simplify the process of report generation, enabling analysis of changes relative to previous reports, rather than repeating parts of the analysis unnecessarily. Reports may be verified, edited, or annotated by a user and fed back into the output layer 134 as feedback to improve future cycles of report generation. In embodiments, one or more aspects of the customisations described above may be achieved by means of prompts which initiate the process of preparing a particular report. In embodiments, the prompts are provided to the system 100 via the ingestion layer 114 but may alternatively be provided to a combination of the ingestion layer 114 with the processing layer 122 and / or the output layer 134. The following gives an example of the user inputs which may be used to configure the system 100 to produce a report for a sustainability analyst based on particular KPIs. KPI Definition: TOTAL GREENHOUSE GAS (GHG) EMISSIONS (DIRECT AND INDIRECT) defined as: sum of a company's greenhouse gas emissions from both direct sources (Scope 1) and indirect sources (Scope 2). Keywords for searching: Emission Total, Total Direct and Indirect Emissions, Combined Emissions, Total Carbon Footprint, Organizational Greenhouse Gas Emissions, Total Operational Emissions, GHG Emissions, Total Corporate Emissions, Company-Wide Emissions, Total C02e Emissions, Aggregate GHG Emissions, Fossil Fuel and Purchased Energy Emissions, Overall Carbon Emissions, GHG Protocol Total, SASB Reporting, CDP Total Emissions Disclosure. The keywords may be configured in accordance with the Scope 1 and Scope 2 emission sources identified in the KPI definition. Units: 100 Million tonnes, Hundred MTCO2e, MTCO2E, Metric tonnes CO2e, Million MTCO2e, Ten Thousand MTCO2e, Thousand MTCO2e, Tonnes CH4, Tonnes N2O, kg CO2e, lbs CO2e. What were the TOTAL GHG EMISSIONS DIRECT AND INDIRECT EMISSIONS for this company, corresponding to the KPI-CODE GHG-279 for the year 2023. Historical Data: The previous years data for this company for this KPI was this: 2020: 2038 Thousand tonnes CO2e, 2021: 2322 Thousand tonnes CO2e, 2022: 2201 Thousand tonnes CO2e. Task Prompt “Your task is to extract KPI values from the provided documents using KPI, all values are present in these documents, if multiple potential values are present send each one as separate. Some tables may contain data over multiple years, strictly extract only for year 2023”. Specified Output Format The extracted data should follow this format: {\n \"KPI_CODE\": \"GHG-2\",\n \"value\": \"6381250.0\",\n \"units\": \"MTCO2e\",\n \"year\": \"2022\",\n \"report_name\": \"SUZANO_ON_2022_Air_pollutant_PAGE.pdf\",\n V'Page no\": \"22\"\n}. If any entry is missing, return 'null' forthat key". Figure 3 illustrates a schematic of the data flow of the process of quantitative data extraction, performed using a system operating according to the principles described using the embodiments of Figures 1 and 2, according an embodiment of the present invention. In the example of Figure 3, the input prompts are directed towards the extraction of information pertaining to a topic from a group of KPIs 301 for a predictive report, such as future estimations of GHG emissions, flooding risk, air pollution, water consumption, water and land pollution, waste generation and disposal, fuel or human capital. The arrangement of Figure 3 is shown not in terms of physical architectural connections, but in terms of the logical flow of data from input to output. In particular, the embodiment of Figure 3 illustrates a process by which prompts to the ingestion layer 114 of system 100 of Figure 1 can be refined as a consequence of the data extraction process performed by the ingestion layer. A series of initial prompts 302 are prepared, in a similar manner to the example described above. The prompts reflect the term definition performed in step S201 of Figure 2, namely prompts based on those previously used, or stored in conjunction with a particular KPI selected from the KPI group 301 (for example, definitions associated with GHG emissions), or prompts which are created manually by user for the first instance of the generation of a particular KPI report. The prompts cause the ingestion of data and classification in a manner described above in relation to Figure 1, and in this regard, the KPI prompts are shown as causing an output which connects to context-tuned LLMs 126, although it will be appreciated that in practice the KPIs drive processes by the ingestion layer 114 and the processing layer 122 as a whole. As described in relation to Figure 1, the context-tuned LLMs 126 receive data from a variety of sources, such as a document database 104, or a database of company reports 105 which may be included in the data source layer 102 in embodiments, in order to tune their ability to interpret context. The output of the processing layer 122 which is provided to the output layer 134 is represented in Figure 3 by the connection between the context-tuned LLMs 126 and an accuracy measurement module 303, which represents functionality provided by, or in addition to the validated metrics component 136 and the confidence scores component 138 of the system 100 of Figure 1. The accuracy measurement module accesses a ground truth component 304, which may, in embodiments, correspond to the historical database 132 of Figure 1, but may be a standalone component that crawls data as a background operation in order to build up a repository of information for use in accuracy determination or verification. The accuracy measurement module 303 performs a comparison of information which is output by the processing layer 122 and information available from the ground truth component 304, and deduces whether the processing layer 122 has a) been able to predict information associated that the KPI selected from the KPI group 301, or has not been able to predict the required information and b) if it has been able to predict information, whether the prediction is correct 306 or incorrect 307. The three possible results of ‘not predicted’ 305, ‘predicted and correct’ 306 and ‘predicted and incorrect’ 307 are provided to an output component 308 as the result of the process performed by the accuracy measurement module 303, together with the information obtained from the processing layer 122 and information from the KPI prompts 302, so than an analysis module 308 can identify one or more causes of a failed prediction or an inaccurate prediction, or can extract learnings from a correct prediction in order to label the KPI prompt accordingly. Using either such learnings, or by suggesting modifications to the prompts, the analysis module 309 outputs one or more updated prompts 310 for use in future tasks. The analysis module 309 may perform any of a plurality of sophisticated routines in order to assess the success of the prediction process. One such routine is the assessment of the suitability of a particular qualitative parameter to generate meaningful quantitative data. The analysis module 309 may build up a profile for each KPI. Where the output of the analysis module 309 is provided to a user via a display, user interface, or via electronic communication, natural language generation may be employed to assist with user with understanding any problems associated with the data collection, and enable the user to contribute to the evolution, or the reformulation of the prompts. For example, the analysis module 309 may determine the following: ‘Waste KPIs exhibited the highest error rates, indicating that they are the most challenging category for the Al model to accurately capture. The primary contributing factors include’. a. The inherent complexity associated with the reporting of Waste KPIs, b. The complexity of calculations required for Waste KPI tagging". In another example, the analysis module 309 may determine the following: “Fuel KPIs also has high contribution in mismatches and the reason for it could be: a. Percentage reporting of fuel KPIs which involves calculations b. Reporting in non-standard format.". From the output of the analysis module 309, it may be determined whether a different KPI should be considered, or whether an acceptable error rate can be tolerated that an indication can be included in a particular report to qualify its accuracy. Additionally, in embodiments, the output of the analysis module 309 may be input to the intelligence module 119 of Figure 1 to update the algorithm on the basis of which optimum data sources as selected and consulted. Embodiments of the present invention achieve significant technical improvements in data processing accuracy and reliability through its multi-modal Al architecture. The multimodal fusion capabilities of the embodiments address fundamental technical limitations in conventional data integration approaches. Traditional systems often struggle to correlate information across different data formats, leading to incomplete or inconsistent analysis results. The system's ability to simultaneously process, for example, textual, geospatial, tabular, and video data through coordinated agent workflows enables comprehensive analysis that leverages complementary information across all available data sources. Systems of embodiments of the present invention therefore address fundamental technical challenges in handling incomplete datasets by implementing sophisticated gap-filling algorithms that maintain data integrity while providing comprehensive metric coverage across diverse reporting frameworks. The multi-agent architecture provides substantial computational efficiency improvements over conventional sequential processing approaches. By deploying parallel processing configurations, the system may reduce overall processing time compared to traditional linear data processing pipelines. The distributed agent architecture enables dynamic load balancing and resource optimization, allowing the system to scale processing capabilities based on data volume and complexity requirements without degrading performance quality. Embodiments deliver enhanced data validation capabilities through systematic crossreferencing of multiple information sources. Conventional validation systems typically rely on single-source verification methods that may miss systematic errors or inconsistencies. Validation may be performed across a variety of dimensions, enabling detection of validation errors that would remain unidentified in traditional processing approaches. This multi-dimensional validation approach may reduce false positive rates while maintaining high sensitivity for genuine data anomalies. The system provides significant improvements in processing scalability and throughput capacity. Traditional data extraction systems often experience exponential performance degradation when handling large-scale datasets spanning thousands of companies and millions of assets. The present invention's modular architecture enables linear scaling characteristics, where processing capacity may be increased proportionally by deploying additional agent instances without architectural modifications. This scalability improvement enables real-time processing of global datasets that would require days or weeks using conventional approaches. Embodiments may achieve superior temporal analysis capabilities through accessing historical database 132 with real-time processing workflows. Conventional systems typically operate on static data snapshots that lack temporal context for anomaly detection and trend analysis. In this way, embodiments of the present invention maintain continuous historical context that enables identification of subtle patterns and deviations that may indicate data quality issues or emerging trends. Embodiments of the present invention may provide enhanced transparency and explainability through comprehensive provenance tracking capabilities. Conventional systems often operate as "black boxes" that provide limited insight into decision-making processes, creating challenges for regulatory compliance and audit requirements. The provenance trail component 140, for example, maintains detailed records of all processing steps, validation decisions, and source attributions, enabling complete traceability from raw input data to final validated metrics. This transparency capability supports regulatory compliance requirements while enabling continuous system improvement through detailed performance analysis. Alternative system architectures and methodological approaches to those illustrated in Figures 1-3 may be implemented while maintaining the core technical advantages of multi-modal data processing and validation of embodiments of the present invention. For example, the multi-agent architecture illustrated in Figure 1 may be configured with different agent specializations, such as domain-specific agents focused on particular industry sectors or regulatory frameworks, or hybrid agents that combine classification and extraction functionalities within single processing units. Sequential processing pipelines may alternatively be implemented as a parallel processing network where multiple agent types operate simultaneously on the same data sources, or as a hierarchical processing structure with multiple validation layers. The processing layer 122 may be reconfigured to emphasize different analytical capabilities, such as prioritizing real-time streaming analysis over batch processing, or incorporating additional specialized components such as blockchain verification modules or quantum computing elements for enhanced security and processing power. The output layer 134 may be configured to generate different output formats including real-time dashboards, API endpoints for system integration, or standardized regulatory reporting formats that automatically adapt to different jurisdictional requirements. These alternative implementations may provide equivalent technical benefits while accommodating different operational requirements, regulatory environments, or technological constraints without departing from the fundamental multi-modal validation approach that characterizes embodiments of the present invention. The components and modules described herein may be implemented using software, hardware, firmware, or any combination thereof, depending on specific operational requirements and technological constraints. For example, the ingestion layer 114, processing layer 122, and output layer 134 may be realized as software applications executing on general-purpose computing hardware, specialized hardware implementations such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs), or hybrid configurations that combine software algorithms with dedicated processing hardware for optimized performance. The classifier agents 116, extractor agents 118, and quality-check agents 120 may be implemented as software processes, microservices, containerized applications, or hardware-accelerated processing units, while maintaining their functional characteristics regardless of the underlying implementation approach. Similarly, components such as geospatial Al 124, context-tuned LLMs 126, data verification 128, and scenario analysis 130 may be understood as defining functional components, rather than mandatory physical configurations, where each component may be distributed across multiple computing resources, consolidated within single processing units, or implemented through cloudbased services that provide equivalent analytical capabilities. The modular architecture enables flexible deployment strategies where individual components may be scaled, replaced, or reconfigured independently while preserving overall system functionality and performance characteristics. This implementation flexibility allows the embodiments to adapt to diverse technological environments and operational requirements without compromising the fundamental multi-modal data processing and validation capabilities that define the system's technical advantages. Features of any of the examples or embodiments outlined above may be combined to create additional examples or embodiments without losing the intended effect. It should be understood that the description of an embodiment or example provided above is by way of example only, and various modifications could be made by one skilled in the art. Furthermore, one skilled in the art will recognise that numerous further modifications and combinations of various aspects are possible. Accordingly, the described aspects are intended to encompass all such alterations, modifications, and variations that fall within the scope of the appended claims. Although the examples above a primarily described in the context of environmental, social and governance (ESG), the operating principles that are described are applicable to any other data analysis in non-ESG context, such as risk analysis, the monitoring of the technical status of a complex systems such as 5 industrial facilities or distributed technological centres such as server rooms or communication networks.

Claims

1. A computer-implemented data-extraction method, comprising:ingesting multi-modal data from a respective plurality of data sources;integrating the ingested multi-modal data by retrieving context-specific information from the ingested multi-modal data, and extracting context-specific quantitative and / or qualitative metrics from the retrieved context-specific information;verifying the extracted metrics against one or more predetermined confidence thresholds;generating an integrated report of extracted metrics meeting the one or more predetermined confidence thresholds, and outputting the integrated report.

2. The method of claim 1, wherein the ingesting step further comprises accessing a historical database containing data to provide historical context for the multi-modal data, wherein the verifying steps uses the data in the historical database to determine whether the extracted metrics meet the one or more predetermined confidence thresholds.

3. The method of claim 1 or claim 2, wherein the ingesting step comprises identifying optimal data sources based on probabilistic analysis of data source relevance to a topic associated with the integrated report and ingesting the multi-modal data comprises ingesting only data from the identified optimal data sources.

4. The method of claim 1 or claim 2, wherein the ingesting step comprises identifying optimal data sources based on anticipated processing complexity associated with information from each data source and ingesting the multi-modal data comprises ingesting only data from the identified optimal data sources.

5. The method of any one of the preceding claims, wherein verifying the extracted metrics comprises identifying inconsistencies and / or applying contextual analysis to verify accuracy against industry benchmarks and historical patterns.

6. The method of any one of the preceding claims, further comprising generating a provenance trail that links each validated metric to source evidence and reasoning.

7. The method of any one of the preceding claims, comprising receiving instructions specifying:a report topic;one or more indicators associated with the report;quantitative definitions of qualitative metrics; and an integrated report format.

8. The method of claim 7, further comprising outputting one or more updates to the instructions to reduce an amount of metrics failing to meet the one or more predetermined confidence thresholds.

9. The method of claim 7 or claim 8, comprising receiving the one or more indicators from a database of indicators associated with a group of predetermined topics.

10. The method of any one of the preceding claims, wherein verifying the extracted metrics comprises accessing one or more previously-generated integrated reports from a library and comparing the extracted metrics with the content of the previously-generated integrated reports,wherein the integrated report defines changes relative to the one or more previously-generated integrated reports.

11. A computer program which, when executed by a processor, is arranged to perform the method of any one of the preceding claims.

12. A system comprising storage means and one or more processors, the storage means arranged to store one or more sets of computer-executable instructions which, when executed by the one or more processors, is arranged to perform the method of any one of claims 1 to 10.A

Citation Information

Patent Citations

  • ViewUS2010/0114899A1onEspacenetopensinnewtab