Surprisingness Scoring for Geoscience Text Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current search engines fail to effectively surface surprising or insightful information relevant to geoscientists, as they rely on relevance and popularity rather than serendipitous encounters, and existing methods for computing surprisingness scores in geoscience sentences are limited by their reliance on unsupervised statistical techniques that do not account for domain-specific user expectations or informative features.
Innovation Solution
A method and system using theory-guided natural language processing and machine learning to compute a surprisingness score for geoscience sentences, incorporating informative features, specificity, domain relevance, and noun phrase ratios, with an exponential weighting algorithm to rank sentences and surface the most surprising ones in search results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If search engines rely on relevance and popularity to rank results, then the most commonly accessed information is surfaced, but surprising or insightful information relevant to geoscientists remains hidden
Solution Approach 1:
The patent introduces a new parameter (surprisingness score) to complement existing ranking parameters like relevance and popularity. By computing this score based on domain-specific features (geoscience terminology, sentence structure, information density) and combining it with traditional ranking signals, the system can surface information that is both relevant and surprising, resolving the contradiction between reliability and information loss
Solution Approach 2:
The patent introduces an intermediary component (the surprisingness scoring module) that acts as a mediator between the search engine's traditional ranking system and the final result presentation. This intermediary computes domain-specific scores and integrates them with existing ranking signals, enabling the system to balance relevance and surprise without completely replacing established ranking mechanisms
2Productivity
If unsupervised statistical techniques are used to compute surprisingness scores, then general patterns can be detected, but domain-specific user expectations and informative features are not accounted for
Solution Approach 1:
The patent applies preliminary action by pre-computing and storing domain-specific resources (geoscience terminology databases, predefined informative feature templates, domain knowledge graphs) before the actual search and scoring process. This preparation enables the system to quickly evaluate domain-specific features during query processing without sacrificing computational efficiency, thereby improving measurement precision while maintaining productivity
Solution Approach 2:
The patent segments the surprisingness scoring process into distinct components: domain-specific feature extraction (using pre-defined geoscience features), informative feature detection (using templates and patterns), and score computation. This segmentation allows each component to be optimized independently, maintaining computational efficiency while improving overall measurement precision through specialized processing
3Quantity of substance
If too many search results are returned, then comprehensive coverage is achieved, but geoscientists cannot read through all results and potential knowledge remains hidden
Solution Approach 1:
The patent extracts and highlights the most surprising sentences from search results using the computed surprisingness scores. Instead of requiring users to review entire documents, the system extracts key surprising information and presents it prominently, reducing the time needed to review results while maintaining comprehensive coverage through the underlying full-text search capability
Solution Approach 2:
The patent creates condensed copies of search results in the form of surprising sentence excerpts and summaries. These copies capture the essential surprising information without requiring users to read full documents, effectively reducing review time while preserving the core valuable knowledge from the comprehensive result set
Data Source
AI summary
The invention is a data processing method and system for suggesting insightful and surprising sentences to geoscientists from unstructured text. The data processing system makes the necessary calculations to assign a surprisingness score to detect sentences containing several signals which when combined exponentially, have tendencies to give rise to surprise. In particular, the data processing system operates on any digital unstructured text derived from academic literature, company reports, web pages and other sources. Detected sentences can be used to stimulate ideation and learning events for geoscientists in industries such as oil and gas, economic mining, space exploration and Geo-health.


