Multi-Language Database Creation via Machine Translation and Noise Removal

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing database creation apparatuses are limited to creating databases from text information in a single language, restricting their data collection range and resulting in lower search usefulness.

Innovation Solution

A database creation apparatus that acquires and combines text information in multiple languages by translating non-target language text into a primary language, removing noise, and associating sensitivity information with the mixed text information, allowing for a broader data range and improved search results.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If text information in multiple languages is acquired and combined, then the data collection range and search usefulness are improved, but the processing complexity and noise removal difficulty increase

Engineering Contradiction:
Improvedata collection rangeVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the multi-language text processing into distinct functional modules: acquisition unit for collecting text, translation unit for converting foreign language text to target language, mixing unit for combining translated text with original target language text, and noise removal unit for filtering unwanted content. This segmentation allows each module to handle specific tasks independently, managing complexity while achieving multi-language database creation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a translation unit as an intermediary component that converts foreign language text into the target language before combining it with original target language text. This intermediary translation step enables the system to process multi-language input while maintaining a unified target language output, resolving the complexity of directly handling multiple languages without requiring separate processing pipelines for each language.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If translated text information is combined with original text information, then the information range is expanded, but the noise information increases

Engineering Contradiction:
Improveinformation rangeVSAvoidnoise information
Core Design Contradiction:
Quantity of substanceVSObject-generated harmful factors

Solution Approach 1:

The patent employs a noise removal unit that extracts and removes unwanted noise information from the mixed text data after combining translated text with original target language text. This extraction approach allows the system to retain the expanded information range from multi-language sources while actively removing translation artifacts, inconsistencies, and other noise elements that degrade data quality.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent converts the potential harm of noise information introduced by translation into a benefit by implementing a dedicated noise removal unit that processes the mixed text. The translation process, while potentially introducing noise, also enables access to foreign language information; the noise removal unit then systematically identifies and removes translation artifacts, converting the harmful noise into an opportunity to refine and enhance the overall data quality through targeted cleaning.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

3Measurement precision

If noise removal processing is executed, then the search accuracy is improved, but the processing time increases

Engineering Contradiction:
Improvesearch accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements noise removal processing as a preliminary step before database creation and search operations. By removing noise information from the mixed text data in advance, the system ensures that the database is populated with clean, high-quality data, thereby improving search accuracy without adding time pressure during actual search operations. The noise removal is performed once during data preparation rather than repeatedly during searches.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11436278B2Database creation apparatus and search system
Publication Date: 2022.09.06 HONDA MOTOR CO LTD
  • US11436278B2 patent drawing
  • US11436278B2 patent drawing
  • US11436278B2 patent drawing

AI summary

To provide a database creation apparatus and the like capable of creating a database with its usefulness increased. A data processing server 2 acquires Japanese language data and foreign language data from external servers 6, creates machine-translated data by translating the foreign language data into data written in the Japanese language using machine translation, creates mixed data by combining the machine-translated data as an additional part of the Japanese language data, and creates retained data using the mixed data.