Multi-Language Database Creation via Machine Translation and Noise Removal
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing database creation apparatuses are limited to creating databases from text information in a single language, restricting their data collection range and resulting in lower search usefulness.
Innovation Solution
A database creation apparatus that acquires and combines text information in multiple languages by translating non-target language text into a primary language, removing noise, and associating sensitivity information with the mixed text information, allowing for a broader data range and improved search results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If text information in multiple languages is acquired and combined, then the data collection range and search usefulness are improved, but the processing complexity and noise removal difficulty increase
Solution Approach 1:
The patent segments the multi-language text processing into distinct functional modules: acquisition unit for collecting text, translation unit for converting foreign language text to target language, mixing unit for combining translated text with original target language text, and noise removal unit for filtering unwanted content. This segmentation allows each module to handle specific tasks independently, managing complexity while achieving multi-language database creation.
Solution Approach 2:
The patent introduces a translation unit as an intermediary component that converts foreign language text into the target language before combining it with original target language text. This intermediary translation step enables the system to process multi-language input while maintaining a unified target language output, resolving the complexity of directly handling multiple languages without requiring separate processing pipelines for each language.
2Quantity of substance
If translated text information is combined with original text information, then the information range is expanded, but the noise information increases
Solution Approach 1:
The patent employs a noise removal unit that extracts and removes unwanted noise information from the mixed text data after combining translated text with original target language text. This extraction approach allows the system to retain the expanded information range from multi-language sources while actively removing translation artifacts, inconsistencies, and other noise elements that degrade data quality.
Solution Approach 2:
The patent converts the potential harm of noise information introduced by translation into a benefit by implementing a dedicated noise removal unit that processes the mixed text. The translation process, while potentially introducing noise, also enables access to foreign language information; the noise removal unit then systematically identifies and removes translation artifacts, converting the harmful noise into an opportunity to refine and enhance the overall data quality through targeted cleaning.
3Measurement precision
If noise removal processing is executed, then the search accuracy is improved, but the processing time increases
Solution Approach 1:
The patent implements noise removal processing as a preliminary step before database creation and search operations. By removing noise information from the mixed text data in advance, the system ensures that the database is populated with clean, high-quality data, thereby improving search accuracy without adding time pressure during actual search operations. The noise removal is performed once during data preparation rather than repeatedly during searches.
Data Source
AI summary
To provide a database creation apparatus and the like capable of creating a database with its usefulness increased. A data processing server 2 acquires Japanese language data and foreign language data from external servers 6, creates machine-translated data by translating the foreign language data into data written in the Japanese language using machine translation, creates mixed data by combining the machine-translated data as an additional part of the Japanese language data, and creates retained data using the mixed data.


