Surprisingness Score Algorithm for Geoscience Text Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods fail to compute a surprisingness score for geoscience sentences using theory-guided Natural Language Processing (NLP) and Machine Learning (ML), limiting the ability to identify and surface insightful information in search results for geoscientists, who often face overwhelming amounts of data.

Innovation Solution

A system and method that compute a surprisingness score for geoscience sentences by utilizing informative features, named entities, domain relevance, and noun phrases, combined with an exponential weighting algorithm, to rank sentences within search results, facilitating serendipitous encounters and learning opportunities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If unsupervised statistical techniques are used to compute surprisingness scores, then the system can process large volumes of text data, but the results lack domain-specific insight and theoretical guidance

Engineering Contradiction:
Improvetext processing capacityVSAvoiddomain-specific measurement accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent introduces theory-guided NLP components as intermediaries between raw text data and surprisingness scoring. These components include domain-specific lexicons, named entity recognition systems, and grammatical analysis modules that mediate the processing to ensure both scalability and domain accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system segments the text processing into distinct analytical layers: statistical frequency analysis, domain-specific entity recognition, grammatical structure analysis, and contextual meaning evaluation. Each layer processes specific aspects independently before integrating results for the final surprisingness score.

Inventive Principle:
Principle #1Segmentation

2Reliability

If search results are optimized for relevance to specific tasks, then the system provides accurate information retrieval, but it fails to facilitate serendipitous encounters and learning opportunities

Engineering Contradiction:
Improveinformation retrieval accuracyVSAvoidserendipity facilitation capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent implements a dynamic search result ranking system that adjusts between relevance-based and surprisingness-based ordering based on user context, query type, and interaction patterns. This allows the system to switch between providing reliable task-specific information and facilitating serendipitous discoveries.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system applies different quality criteria to different portions of search results. High-relevance results are optimized for task completion, while a separate 'surprising findings' section presents domain-insightful content that may have lower direct relevance but higher potential for serendipitous learning.

Inventive Principle:
Principle #3Local quality

3Adaptability or versatility

If collaborative filtering techniques are used to generate serendipitous information encounters, then the system can personalize recommendations, but suggestions are limited by previous activity and usage data

Engineering Contradiction:
Improvepersonalized recommendation capabilityVSAvoidavailable usage data
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent pre-computes domain knowledge graphs, entity relationships, and theoretical frameworks before they are needed for recommendation. This preliminary structuring of domain knowledge allows the system to generate serendipitous recommendations based on theoretical insights rather than waiting for sufficient usage data to emerge.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12106063B2Method and system for generating a surprisingness score for sentences within geoscience text
Publication Date: 2024.10.01 EXXONMOBIL TECHNOLOGY & ENGINEERING CO
  • US12106063B2 patent drawing
  • US12106063B2 patent drawing
  • US12106063B2 patent drawing

AI summary

The invention is a data processing method and system for suggesting insightful and surprising sentences to geoscientists from unstructured text. The data processing system makes the necessary calculations to assign a surprisingness score to detect sentences containing several signals which when combined exponentially, have tendencies to give rise to surprise. In particularly, the data processing system operates on any digital unstructured text derived from academic literature, company reports, web pages and other sources. Detected sentences can be used to stimulate ideation and learning events for geoscientists in industries such as oil and gas, economic mining, space exploration and Geo-health.