NLP Pipeline for Extracting Self-Reported Activities
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual analysis of self-reported activities from freestyle narrative text is time-consuming and prone to human bias, making it inefficient for extracting insights over time.
Innovation Solution
A processor-implemented method using natural language processing and a supervised classification model to automatically extract self-reported activities by generating candidate activity phrases, identifying context tokens, and classifying them as self-reported or referred activities based on predefined grammar patterns.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual analysis of freestyle narrative text is performed to extract self-reported activities, then the extraction can be done with human comprehension and judgment, but the process becomes time-consuming and prone to human bias
Solution Approach 1:
The patent replaces the manual mechanical analysis process with an automated computational system. The system uses natural language processing techniques, including tokenization, part-of-speech tagging, and supervised machine learning classifiers, to automatically extract self-reported activities from freestyle narrative text. This substitution eliminates human time investment while maintaining extraction accuracy through algorithmic processing.
Solution Approach 2:
The system enables self-service extraction by allowing the text processing pipeline to automatically identify and classify activities without human intervention. The supervised learning model is trained on annotated data and then autonomously processes new text inputs, performing the extraction task independently while consistently applying the learned classification rules.
2Adaptability or versatility
If manual analysis is used to extract self-reported activities, then human comprehension can distinguish self-reported from referred activities, but the process becomes complex and requires significant time investment for long lists of texts
Solution Approach 1:
The patent segments the complex analysis task into distinct processing stages: text tokenization, part-of-speech tagging, candidate activity phrase identification, context window extraction, and classification. Each stage handles a specific aspect of the analysis, breaking down the overall complexity into manageable, automated components that can be processed systematically.
Solution Approach 2:
The system changes the parameters of analysis by using computational metrics and algorithmic rules instead of human judgment. The supervised learning model transforms qualitative human comprehension into quantitative classification based on trained parameters, enabling the system to distinguish self-reported from referred activities through learned patterns rather than manual evaluation.
3Reliability
If manual listing and maintaining of self-reported activities is performed, then the activities can be accurately identified, but the process is time-consuming and requires continuous human effort
Solution Approach 1:
The patent replaces the manual listing and maintaining process with automated computational extraction. The system continuously processes incoming freestyle narrative text through the NLP pipeline, automatically identifying and extracting self-reported activities without human intervention. This maintains reliability through consistent algorithmic application while dramatically improving productivity by eliminating continuous human effort.
Solution Approach 2:
The system enables continuous extraction of self-reported activities from ongoing text inputs. The automated pipeline can process multiple texts in sequence without interruption, maintaining continuous useful action rather than requiring periodic manual analysis. This ensures both reliability through consistent processing and productivity through uninterrupted operation.
Data Source
AI summary
This disclosure relates generally to methods and systems for automatic extraction of self-reported activities of an individual from a freestyle narrative text. Manual extraction of such self-reported activities of the individual from the freestyle narrative text over the period of time is a complex task and consume a significant amount of time. The present systems and methods utilize a predefined grammar pattern and a natural language processing technique to generate one or more candidate activity phrases, from the pre-processed input text posted by the individual. A deep learning based supervised classification model is utilized to automatically extract the one or more self-reported activities of the individual, from the one or more candidate activity phrases. Manual intervention and efforts of analyzing the freestyle narrative text to extract the self-reported activities are avoided. Longitudinal assessment of the self-reported activities may reveal routines and behavior of the individual.


