Multi-Language Email Character Code Embedding via OCR Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing email systems face garbling issues when sending emails that contain text information in different character codes, particularly when languages like Japanese, Korean, and Chinese are involved, due to mismatched character codes between senders and receivers, leading to distorted text display.

Innovation Solution

A communication apparatus equipped with input, recognition, embedding, and sending means to identify and specify the character code type of extracted text information using OCR, embedding this information with a corresponding identifier in the email, ensuring correct character code usage for each language type.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If text information extracted by OCR is embedded into e-mail text, then text information from scanned documents can be utilized, but garbled display occurs when character codes of different languages are mixed

Engineering Contradiction:
Improveability to process multi-language documentsVSAvoidtext display accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments the e-mail text into multiple portions, each corresponding to a specific language type. By dividing the text stream and associating each segment with its corresponding character code table, the system prevents garbled display while maintaining the ability to process multi-language documents. The segmentation is achieved by inserting delimiters between different language sections and applying appropriate character codes to each segment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different character code tables to different portions of the e-mail text based on the language type of each segment. Instead of using a single uniform character code for the entire message, the system dynamically selects and applies the appropriate character code table (e.g., Japanese, Korean, Chinese) to each local segment, ensuring accurate display of each language while maintaining overall document versatility.

Inventive Principle:
Principle #3Local quality

2Device complexity

If a single character code is used for the entire e-mail text, then processing is simple, but text information in different languages becomes garbled

Engineering Contradiction:
Improvecharacter code processing complexityVSAvoidtext display accuracy
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent introduces dynamic character code switching capability into the e-mail processing system. The character code table is no longer static but changes dynamically based on the language type of each text segment. The system automatically switches between different character code tables (Japanese, Korean, Chinese, etc.) according to the delimiters and language identifiers present in the text, maintaining simplicity of operation while ensuring accurate display of multiple languages.

Inventive Principle:
Principle #15Dynamics

3Loss of information

If OCR extraction is applied to scanned documents, then text information can be embedded in e-mail, but character code mismatch causes garbled display

Engineering Contradiction:
Improvetext information retentionVSAvoidcharacter code compatibility
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The patent performs preliminary identification of the language type and character code table before embedding the OCR-extracted text information into the e-mail. By pre-determining the appropriate character code table based on the source document's language and pre-configuring the text segment with the correct character code association, the system prevents character code mismatch and garbled display while retaining the extracted text information.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10305836B2Communication apparatus, information processing method, program, and storage medium
Publication Date: 2019.05.28 CANON KK
  • US10305836B2 patent drawing
  • US10305836B2 patent drawing
  • US10305836B2 patent drawing

AI summary

This invention has as its object to avoid occurrence of garble even when an e-mail message to be created includes text information described in character codes of different kinds of language. To achieve this object, a communication apparatus according to this invention includes an input unit which inputs image information, a recognition unit which extracts text information included in the image information input by the input unit, and recognizing a type of character code of the extracted text information, an embedding unit which embeds the extracted text information in a text of e-mail using character codes of the type recognized by the recognition unit, and describing the recognized type (510, 516) of character code and an identifier (509, 515, 526) indicating a description range of the extracted text information in the text of e-mail, and a sending unit which sends e-mail data embedded by the embedding unit.