Real-Time ESG Analytics via Point-in-Time Data Reprocessing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data processing techniques struggle with efficiently generating real-time analytics from large unstructured data sets, particularly for Environmental, Social, and Governance (ESG) signals, due to resource constraints and the inability to reprocess historical data with new methodologies, leading to rigid outputs that cannot be retroactively updated.

Innovation Solution

A multi-pipeline architecture that ingests and processes data from various sources, applies different processing rules to tag entities and events with ESG scores, and uses point-in-time aliasing to reprocess historical data, allowing for flexible and retroactive assignment of responsibilities and behaviors, enabling real-time analytics and backward compatibility.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If human analysts process news data manually, then evaluation accuracy and judgment flexibility are improved, but processing speed and real-time capability deteriorate

Engineering Contradiction:
Improveevaluation accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSSpeed

Solution Approach 1:

The patent replaces manual human analysis with automated machine learning models that process news data automatically. The system uses NLP models to extract entities, events, and sentiments from news articles, eliminating the need for human analysts to manually review each story while maintaining high evaluation accuracy through trained algorithms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service processing where the machine learning model autonomously evaluates news data without human intervention. The automated pipeline independently performs data ingestion, entity extraction, event detection, and score generation, allowing the system to process data in real-time without requiring human analysts to revisit historical documents.

Inventive Principle:
Principle #25Self-service

2Quantity of substance

If data is processed once and stored with fixed analysis, then storage efficiency is improved, but adaptability to reprocess with new methodologies deteriorates

Engineering Contradiction:
Improvestorage efficiencyVSAvoidreprocess capability
Core Design Contradiction:
Quantity of substanceVSAdaptability or versatility

Solution Approach 1:

The patent implements a dynamic processing architecture where the same stored data can be reprocessed multiple times with different machine learning models or evaluation criteria. The system maintains historical data in its original format and applies new processing methodologies as needed, allowing the analysis to adapt to changing requirements without requiring data recollection.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system performs preliminary data collection and storage of raw news data in a centralized repository. This pre-stored data can then be repeatedly processed with different machine learning models or evaluation frameworks as methodologies evolve, eliminating the need to collect data again when new analysis approaches are needed.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If new evaluation criteria are applied to historical data, then analysis relevance and accuracy are improved, but processing time and resource consumption increase

Engineering Contradiction:
Improveanalysis accuracyVSAvoidreprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the data processing into distinct stages: data collection, data storage, and multiple rounds of analysis with different models. By separating these functions, the system can efficiently store data once and then apply different machine learning models in parallel to historical data without time penalty, as the heavy computational work is distributed across multiple processing layers.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes processing parameters by applying different machine learning models or evaluation criteria to the same stored data without requiring data recollection. The architecture allows parameter changes in the analysis stage while maintaining efficient processing by leveraging pre-stored data and parallel computation capabilities.

Inventive Principle:
Principle #35Parameter changes

4Speed

If real-time analytics are generated from large datasets, then responsiveness to current events is improved, but computational resource consumption increases

Engineering Contradiction:
Improvereal-time responsivenessVSAvoidcomputational resources
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The patent extracts and stores only the essential data elements (entities, events, dates, locations) from the full news articles during the initial processing stage. This extracted data is stored in a structured format that can be quickly queried and reprocessed, reducing the computational burden during real-time analysis while maintaining responsiveness to current events.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system transforms unstructured news articles into structured data representations with standardized fields. This dimensional transformation from raw text to structured entities enables efficient real-time processing by reducing the complexity of data manipulation and allowing parallel processing across multiple dimensions (entities, events, sentiments) simultaneously.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11928143B2Systems, methods, and devices for generating real-time analytics
Publication Date: 2024.03.12 TRUVALUE LABS
  • US11928143B2 patent drawing
  • US11928143B2 patent drawing
  • US11928143B2 patent drawing

AI summary

Systems of the present disclosure may ingest content from a plurality of data sources with the content including ingested documents referencing entities and events relevant to the ESG signals. The content may be stored in a content database. The System may also identify metadata and a body of text associated with each document to produce a set of preprocessed documents. An entity may be tagged to a first preprocessed document from the set of preprocessed documents, and the document may include a first document identifier. The System may generate an event score related to a first ESG signal including a direction and a magnitude associated with an event identified in the body of text. The event score may be tagged to the document. The system may write to an unstructured data set the document identifier in association with the tagged entity and the tagged event score for delivery.