Multi-lingual Action Identification via Universal Format Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing natural language processing systems face challenges in accurately identifying and categorizing actionable items across multiple languages due to differences in syntax and positioning of action verbs, leading to inaccurate translations and loss of nuances.

Innovation Solution

A multi-lingual action identification system that includes a training manager, language manager, converter, evaluator, and inference manager to train a machine learning model, which identifies action tokens and linguistic features, maps them to a universal format, and predicts action token classes, enabling effective categorization of communications across languages.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional translation methods are used to process multi-lingual communications, then translations can be produced, but accuracy is lost due to differences in syntax and positioning of action verbs

Engineering Contradiction:
Improvetranslation accuracyVSAvoidloss of contextual nuances
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The system segments the communication into discrete action tokens and linguistic features, allowing precise identification and mapping of each component across languages without losing contextual information through holistic translation approaches

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A universal format serves as an intermediary between different languages, enabling accurate translation by mapping action tokens and linguistic features from source languages to the universal format and back, thereby preserving contextual nuances

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If action verbs are identified in their original language format, then linguistic accuracy is maintained, but cross-language comparability is reduced due to syntax differences

Engineering Contradiction:
Improvelinguistic accuracyVSAvoidcross-language comparability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system creates a universal format that can represent action tokens and linguistic features from multiple languages, enabling the same structured representation to be used across different languages while maintaining linguistic accuracy through precise mapping

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If a machine learning model is trained to identify action tokens across multiple languages, then multi-lingual action identification accuracy is improved, but system complexity increases

Engineering Contradiction:
Improveaction token identification accuracyVSAvoidsystem architecture complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system divides the complex task of multi-lingual action identification into separate manageable components: language identification, action token identification, linguistic feature extraction, and mapping to universal format, each handled by specialized modules that simplify the overall architecture

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The universal format acts as an intermediary layer between different language-specific processing modules and the machine learning model, reducing system complexity by providing a standardized interface that simplifies data transformation and model training

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11354504B2Multi-lingual action identification
Publication Date: 2022.06.07 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11354504B2 patent drawing
  • US11354504B2 patent drawing
  • US11354504B2 patent drawing

AI summary

Embodiments relate to an intelligent computer platform to identify and process communications across multiple languages. An originating communication is identified, including identification of language, action tokens, and linguistic features. A first map of the identified action tokens and linguistic features from the originating language to a second format is created and populated into identifying machine learning model (MLM). A second communication is identified, including the originating language, action tokens, and linguistic features, and a second map is created of the identified action tokens and linguistic features of the second communication. The second map and the MLM are leveraged to identify and return a predicted action token class of the identified action tokens in the second communication.