Automated URL Identification via Fuzzy Matching Validation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manual identification of URLs from large volumes of data is costly and time-consuming, especially in fields like retail, where merchant URLs are often not provided and need to be monitored for violations or agreement enforcement.

Innovation Solution

A system and method for automatic URL identification using fuzzy and exact matching algorithms, which processes input data from various sources, validates URLs through non-URL data items like business names and addresses, and generates reports for electronic delivery to e-commerce platforms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual identification of URLs is performed from large volumes of data, then accuracy can be maintained through human review, but time consumption and costs increase significantly

Engineering Contradiction:
ImproveURL identification accuracyVSAvoidTime for URL identification
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the URL identification process into multiple independent stages: data collection from multiple sources, pattern matching algorithms, validation rules, and verification steps. Each stage processes specific aspects of URL identification independently, allowing parallel processing and reducing overall time while maintaining accuracy through cumulative validation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediary validation mechanisms including pattern matching algorithms that act as filters between raw data and final URL identification. These intermediaries (validation rules, pattern matchers) automatically verify URL legitimacy without requiring manual human review, thus reducing time loss while preserving accuracy through multiple layers of verification.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If manual URL identification is performed to ensure high precision, then accuracy is maintained, but operational complexity and costs increase

Engineering Contradiction:
ImproveURL identification precisionVSAvoidSystem complexity for URL identification
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs self-service through automated validation rules and pattern matching algorithms that independently verify URL legitimacy without external human intervention. The validation engine automatically cross-references collected data against predefined patterns and rules, enabling the system to maintain high precision while reducing operational complexity by eliminating manual processes.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent transforms the identification process by changing parameters from manual human judgment to automated algorithmic validation. By converting qualitative human review into quantitative pattern matching and validation rules, the system maintains precision through structured criteria while reducing complexity through standardized automated processes.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If automated URL identification is implemented to reduce time and costs, then productivity increases, but measurement precision may deteriorate without proper validation

Engineering Contradiction:
ImproveURL identification speedVSAvoidURL identification accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent ensures continuous validation through multiple sequential stages: data collection, pattern matching, validation rule application, and verification. Each stage continuously processes data without interruption, maintaining high productivity through automated flow while preserving precision through uninterrupted multi-layer validation that catches errors at each step.

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system implements feedback mechanisms where validation results from each stage feed into the next stage. Pattern matching outcomes feedback to validation rules, which in turn feedback to final verification. This continuous feedback loop enables automated high-speed processing while maintaining accuracy through iterative validation that corrects errors progressively.

Inventive Principle:
Principle #23Feedback

4Quantity of substance

If multiple data sources are processed to improve URL identification coverage, then completeness increases, but data processing complexity increases

Engineering Contradiction:
ImproveVolume of processed dataVSAvoidData processing system complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent implements a universal data collection framework that handles multiple data sources (websites, databases, files, APIs) through a single standardized interface. The collection engine performs multi-functional operations including web scraping, database queries, file parsing, and API calls using unified code, reducing processing complexity while maintaining comprehensive data coverage from diverse sources.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20230161831A1Systems and Methods for Automatic URL Identification From Data
Publication Date: 2023.05.25 INSURANCE SERVICES OFFICE INC
  • US20230161831A1 patent drawing
  • US20230161831A1 patent drawing
  • US20230161831A1 patent drawing

AI summary

Systems and methods for automatic URL identification from data are provided. The system receives and processes one or more sources of data, such as merchant data, and processes the input data to identify one or more URLs present in the data. The identified URLs are automatically validated by the system using one or more fuzzy and/or exact matching algorithms. The validation could be performed by matching one or more non-URL data items, such as a business name, address, e-mail, country, zip code, or any other suitable non-URL data item, to ensure that only valid URLs are identified. Once the URLs are validated, a report is generated by the system.