Cross-Study Data Joining via Variable Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for combining research studies are limited to studies with similar collection methods, respondents, questions, and variable types, making it difficult to analyze studies with different characteristics.
Innovation Solution
A processor-implemented method and system that identifies matched variables in statistically representative samples across multiple research studies, establishes schemas of segments, determines statistical representativeness, and creates a joint study by combining responses within each segment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional meta-analysis or comparative analysis methods are used to combine research studies, then studies with similar characteristics can be analyzed together, but studies with different collection methods, respondents, questions, or variable types cannot be combined
Solution Approach 1:
The patent segments the research studies into discrete variables and creates segmented datasets where each variable is independently processed. The system identifies matched variables across studies, segments the data accordingly, and then recombines them into a unified structured format. This segmentation approach enables studies with different characteristics to be joined by treating each variable independently rather than requiring entire studies to be identical.
Solution Approach 2:
The patent transforms unstructured or semi-structured research data into a standardized structured format by changing the parameters of data organization. It applies variable matching algorithms that compare variable names, descriptions, and metadata across studies, then restructures the data to align matched variables. This parameter transformation allows studies with different collection methods and question formats to be combined through standardized variable mapping.
2Loss of information
If research studies with different characteristics are combined, then more diverse insights can be extracted, but erroneous data and statistical representativeness issues arise
Solution Approach 1:
The patent performs preliminary variable matching and data segmentation before combining studies. It pre-identifies matched variables across studies using algorithms that compare variable names, descriptions, and metadata, and pre-segments the data into comparable units. This preliminary processing ensures that only statistically compatible data points are combined, maintaining reliability while preserving complete response data.
Solution Approach 2:
The system incorporates feedback mechanisms that evaluate the quality and representativeness of matched variables during the joining process. It assesses whether variables from different studies are truly comparable based on multiple criteria including data types, measurement scales, and sampling characteristics. This feedback loop ensures that erroneous matches are identified and corrected, maintaining statistical representativeness while combining diverse studies.
Data Source
AI summary
A processor-implemented method and a system for user-selected selected joining two or more research studies to extract analytical insights for enabling cross-study analysis is disclosed. The method includes identifying, using an identification module, at least one matched variable in a statistically representative sample in each of the two or more research studies. The method further includes establishing, using a segment establishing module, for each combination of matched variables, a schema of segments, from the two or more research studies. The method furthermore includes determining, using the segment establishing module if each of the established schema of segments is statistically representative. The method furthermore includes creating a joint study, using a joint study creation module, by combining responses from each of the two or more research studies within each segment of the statistically representative schema of segments.


