Super-Character Ideogram Matrix for Latin Text Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine learning technologies face challenges in effectively learning the meaning of written Latin-alphabet based languages, as they are not well-equipped to handle the complex combinations of ideograms present in these languages, limiting their ability to understand and process the nuanced meanings conveyed by multiple characters.

Innovation Solution

A multi-layer two-dimensional symbol is created, where a string of Latin-alphabet based language texts is transformed into a matrix of pixels representing a super-character, divided into sub-matrices that represent individual ideograms, allowing image processing techniques like convolutional neural networks to classify and learn the combined meaning of these ideograms.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional machine learning approaches are used to process Latin-alphabet based languages, then the system can handle individual characters, but it cannot effectively learn the combined meaning of multiple ideograms

Engineering Contradiction:
Improveunderstanding of combined meaningVSAvoidhandling of complex ideogram combinations
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent segments a Latin-alphabet word into multiple ideogram components, where each ideogram is represented as a sub-matrix within a larger 2-D symbol. This segmentation allows the machine learning system to process and understand the combined meaning of multiple ideograms by treating them as distinct yet related units within the super-character structure.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a nested structure where multiple ideogram sub-matrices are contained within a larger 2-D symbol (super-character). Each sub-matrix represents an individual ideogram, and these are nested within the overall 2-D symbol structure, allowing the system to learn both individual ideogram meanings and their combined meanings simultaneously.

Inventive Principle:
Principle #7Nested doll (Nesting)

2Loss of information

If Latin-alphabet text is processed as traditional text strings, then processing is straightforward, but the system cannot capture nuanced meanings conveyed by multiple characters

Engineering Contradiction:
Improvenuanced meaningVSAvoidtext representation structure
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent transforms traditional 1-D text string representation into a 2-D symbol structure. By arranging ideogram sub-matrices in a two-dimensional grid format, the system can preserve and convey nuanced meanings that are lost in linear text representation, while the structured 2-D format actually simplifies the processing compared to handling complex combinatorial relationships in 1-D strings.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Measurement precision

If the system uses detailed 2-D symbol representation with multiple sub-matrices, then it can learn combined meaning of ideograms, but the processing complexity increases

Engineering Contradiction:
Improveclassification accuracyVSAvoidimage processing structure
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates a universal 2-D symbol structure that can represent multiple ideograms in a standardized format. This multi-functional representation allows the same image processing techniques and machine learning algorithms to handle various Latin-alphabet words and languages uniformly, reducing the need for specialized processing logic despite the increased representational detail.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10192148B1Machine learning of written Latin-alphabet based languages via super-character
Publication Date: 2019.01.29 GYRFALCON TECHNOLOGY INC
  • US10192148B1 patent drawing
  • US10192148B1 patent drawing
  • US10192148B1 patent drawing

AI summary

A string of Latin-alphabet based language texts is received and formed a multi-layer 2-D symbol in a computing system. The received string contains at least one word with each word containing at least one letter of the Latin-alphabet based language. 2-D symbol comprises a matrix of N×N pixels of data representing a super-character. The matrix is divided into M×M sub-matrices. Each sub-matrix represents one ideogram formed from the at least one letter contained in a corresponding word in the received string. Ideogram has a square format with a dimension EL letters by EL letters (i.e., row and column). EL is determined from the total number of letters (LL) contained in the corresponding word. EL, LL, N and M are positive integers. Super-character represents a meaning formed from a specific combination of at least one ideogram. Meaning of the super-character is learned with image classification of the 2-D symbol.