Real-Time ESG Analytics via Point-in-Time Data Reprocessing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data processing techniques struggle with efficiently generating real-time analytics from large unstructured data sets, particularly for Environmental, Social, and Governance (ESG) signals, due to resource constraints and the inability to reprocess historical data with new methodologies, leading to rigid outputs that cannot be retroactively updated.
Innovation Solution
A multi-pipeline architecture that ingests and processes data from various sources, applies different processing rules to tag entities and events with ESG scores, and uses point-in-time aliasing to reprocess historical data, allowing for flexible and retroactive assignment of responsibilities and behaviors, enabling real-time analytics and backward compatibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If human analysts process news data manually, then evaluation accuracy and judgment flexibility are improved, but processing speed and real-time capability deteriorate
Solution Approach 1:
The patent replaces manual human analysis with automated machine learning models that process news data automatically. The system uses NLP models to extract entities, events, and sentiments from news articles, eliminating the need for human analysts to manually review each story while maintaining high evaluation accuracy through trained algorithms.
Solution Approach 2:
The system enables self-service processing where the machine learning model autonomously evaluates news data without human intervention. The automated pipeline independently performs data ingestion, entity extraction, event detection, and score generation, allowing the system to process data in real-time without requiring human analysts to revisit historical documents.
2Quantity of substance
If data is processed once and stored with fixed analysis, then storage efficiency is improved, but adaptability to reprocess with new methodologies deteriorates
Solution Approach 1:
The patent implements a dynamic processing architecture where the same stored data can be reprocessed multiple times with different machine learning models or evaluation criteria. The system maintains historical data in its original format and applies new processing methodologies as needed, allowing the analysis to adapt to changing requirements without requiring data recollection.
Solution Approach 2:
The system performs preliminary data collection and storage of raw news data in a centralized repository. This pre-stored data can then be repeatedly processed with different machine learning models or evaluation frameworks as methodologies evolve, eliminating the need to collect data again when new analysis approaches are needed.
3Measurement precision
If new evaluation criteria are applied to historical data, then analysis relevance and accuracy are improved, but processing time and resource consumption increase
Solution Approach 1:
The patent segments the data processing into distinct stages: data collection, data storage, and multiple rounds of analysis with different models. By separating these functions, the system can efficiently store data once and then apply different machine learning models in parallel to historical data without time penalty, as the heavy computational work is distributed across multiple processing layers.
Solution Approach 2:
The system changes processing parameters by applying different machine learning models or evaluation criteria to the same stored data without requiring data recollection. The architecture allows parameter changes in the analysis stage while maintaining efficient processing by leveraging pre-stored data and parallel computation capabilities.
4Speed
If real-time analytics are generated from large datasets, then responsiveness to current events is improved, but computational resource consumption increases
Solution Approach 1:
The patent extracts and stores only the essential data elements (entities, events, dates, locations) from the full news articles during the initial processing stage. This extracted data is stored in a structured format that can be quickly queried and reprocessed, reducing the computational burden during real-time analysis while maintaining responsiveness to current events.
Solution Approach 2:
The system transforms unstructured news articles into structured data representations with standardized fields. This dimensional transformation from raw text to structured entities enables efficient real-time processing by reducing the complexity of data manipulation and allowing parallel processing across multiple dimensions (entities, events, sentiments) simultaneously.
Data Source
AI summary
Systems of the present disclosure may ingest content from a plurality of data sources with the content including ingested documents referencing entities and events relevant to the ESG signals. The content may be stored in a content database. The System may also identify metadata and a body of text associated with each document to produce a set of preprocessed documents. An entity may be tagged to a first preprocessed document from the set of preprocessed documents, and the document may include a first document identifier. The System may generate an event score related to a first ESG signal including a direction and a magnitude associated with an event identified in the body of text. The event score may be tagged to the document. The system may write to an unstructured data set the document identifier in association with the tagged entity and the tagged event score for delivery.


