Dynamic Infectious Disease Keyword Generation via Word Embedding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for predicting infectious diseases rely on empirical human knowledge for selecting keywords, which lacks objective quantification of similarity to the disease and fails to account for temporal variability in search words.

Innovation Solution

A method for automatically generating infectious disease prediction keywords using word embedding and online search volume data, which adapts to temporal changes without requiring expert experience. This involves converting words into embedding vectors, calculating similarities, extracting relevant vectors, and generating keywords based on correlation coefficients with confirmed case data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If keywords are selected based on empirical human knowledge, then the selection process is simple and requires less computational resources, but the similarity between the selected keyword and the infectious disease cannot be objectively quantified and temporal variability of search words cannot be accounted for

Engineering Contradiction:
Improveobjective quantification of keyword similarityVSAvoidcomplexity of keyword generation system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces the manual/empirical keyword selection process with an automated computational system using word embedding and similarity calculation algorithms. This substitution enables objective quantification of keyword-disease similarity through mathematical computations rather than human expertise, directly resolving the contradiction between measurement precision and system complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent transforms static keyword selection into dynamic parameter-based selection by calculating similarity scores and temporal correlation coefficients. The system adjusts keyword selection based on changing parameters such as search volume trends and correlation with confirmed cases, enabling objective measurement while managing complexity through algorithmic parameter optimization.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If static keywords are used for disease prediction, then the prediction model is simple and easy to implement, but it cannot reflect temporal changes in public interests and search behavior

Engineering Contradiction:
Improvetemporal adaptability of prediction keywordsVSAvoidcomplexity of temporal adaptation mechanism
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic keyword selection by continuously updating the keyword list based on temporal patterns in search data. The system calculates correlation coefficients between search volumes and confirmed cases over time, automatically adjusting keywords to reflect changing public interests. This dynamic approach enables temporal adaptability while managing complexity through systematic algorithmic processes.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent incorporates feedback mechanisms where the system monitors the correlation between search keyword volumes and actual confirmed case numbers. Based on this feedback, the system automatically adjusts and refines the keyword list over time, ensuring the prediction model adapts to temporal changes in disease patterns and public search behavior without requiring manual intervention.

Inventive Principle:
Principle #23Feedback

3Reliability

If automated keyword generation using word embedding is implemented, then temporal variability of search words can be accounted for and objective quantification is achieved, but the computational resources and processing time increase

Engineering Contradiction:
Improveaccuracy of infectious disease predictionVSAvoidcomputational energy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent applies partial action by selectively processing only the most relevant portions of the data. Instead of analyzing all possible words, the system focuses on calculating similarities for words that meet certain criteria (e.g., minimum similarity threshold, sufficient search volume). This selective approach maintains prediction accuracy while reducing computational energy consumption compared to exhaustive analysis.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent optimizes computational resources by dynamically adjusting parameters such as the number of top similar words to analyze, the similarity threshold, and the time window for correlation calculations. The system can modify these parameters based on available computational resources and data characteristics, achieving reliable predictions while managing energy consumption through parameter-based optimization.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250079024A1Method for generating infectious disease prediction keyword that changes over time based on word embedding and apparatus performing the same
Publication Date: 2025.03.06 THE CATHOLIC UNIV OF KOREA IND ACADEMIC COOP FOUND
  • US20250079024A1 patent drawing
  • US20250079024A1 patent drawing
  • US20250079024A1 patent drawing

AI summary

The method comprises obtaining a document including a target infectious disease as a corpus for each of a plurality of time sections, converting a plurality of words included in the obtained corpus into embedding vectors, calculating similarities between the respective converted embedding vectors and an embedding vector indicating the target infectious disease, extracting an embedding vector of which the calculated similarity is higher than a predetermined first threshold value, for each time section, obtaining first time series data indicating a search volume, over time, of a word corresponding to the embedding vector extracted for each time section, calculating a correlation coefficient between the obtained first time series data and second time series data, and generating a word corresponding to the first time series data of which the calculated correlation coefficient is higher than a predetermined second threshold value as a keyword for each time section.