Network Data Collection for ML Training via Geographic Perimeter Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Acquiring vast and computationally expensive data for machine learning model training is time-consuming and inefficient, particularly in identifying patterns from network traffic associated with mobile communications devices.

Innovation Solution

A network data collection system identifies a target data source location by establishing a perimeter around mobile communications network base stations, collects network packets, and filters out packets from outside the perimeter to create a training data set for machine learning models, leveraging common events like disabling airplane mode to reduce data collection time.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data is collected from all base stations in a network, then the quantity of training data increases, but the time and computational resources required for data processing increase

Engineering Contradiction:
Improvequantity of training dataVSAvoiddata collection time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent segments the network data collection process by establishing geographic perimeters around target data source locations and selectively collecting data only from base stations within those perimeters. This segmentation allows the system to divide the vast network into manageable geographic zones, collecting relevant training data while excluding unnecessary data from other areas, thereby reducing overall data collection time and processing requirements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by focusing data collection efforts on specific geographic locations where target events are likely to occur, rather than uniformly collecting data across the entire network. By concentrating resources on high-value locations and filtering for locally relevant data patterns, the system achieves efficient training data acquisition with reduced time and computational overhead.

Inventive Principle:
Principle #3Local quality

2Quantity of substance

If data is collected from base stations outside the target perimeter, then more comprehensive data is obtained, but data filtering requirements and processing complexity increase

Engineering Contradiction:
Improvecomprehensiveness of dataVSAvoiddata filtering complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent implements preliminary action by pre-establishing geographic perimeters around target data source locations before data collection begins. These perimeters serve as pre-defined filtering criteria that automatically determine which base station data to include or exclude. This preliminary geographic boundary setting simplifies the filtering process by providing clear inclusion/exclusion rules upfront, reducing the complexity of real-time data filtering while maintaining comprehensive coverage of relevant areas.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If traditional data collection methods are used without targeting specific events, then diverse network patterns are captured, but the time required to identify relevant training patterns increases

Engineering Contradiction:
Improvediversity of network patternsVSAvoidpattern identification time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-identifying target data source locations and expected event types before data collection begins. This advance planning allows the system to focus on capturing specific network patterns associated with predetermined events, rather than searching through all possible network traffic patterns. The preliminary identification of what to look for significantly reduces the time required to find relevant training patterns while maintaining adaptability to various event types.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11386294B2Data harvesting for machine learning model training
Publication Date: 2022.07.12 AT&T INTELLECTUAL PROPERTY I L P
  • US11386294B2 patent drawing
  • US11386294B2 patent drawing
  • US11386294B2 patent drawing

AI summary

Concepts and technologies disclosed herein are directed to data harvesting for machine learning model training. According to one aspect of the concepts and technologies disclosed herein, a network data collection system can identify a target data source location from which to harvest data for a machine learning system to utilize during a machine learning model training process. The data can be associated with a plurality of mobile communications devices operating in communication with at least one base station of a mobile communications network that serves the target data source location. The network data collection system can collect the data and provide the data to the machine learning system. The machine learning system, in turn, can create a training data set for use during the machine learning model training process based, at least in part, upon the data.