Token marking method for large language model input data and cognitive system

By using a four-dimensional discrete spatiotemporal system as a token for large language models, the problems of confusion and logical jumps in multi-turn dialogues and IoT control are solved, enabling more accurate dialogue understanding and device control, and constructing a unified cognitive system.

CN122334486APending Publication Date: 2026-07-03黄宝明
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
黄宝明
Filing Date
2026-04-06
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing large language models lack explicit context differentiation, logical progression indicators, physical world associations, and a unified tokenization framework when handling multi-turn dialogues and IoT control. This leads to model confusion, logical jumps, and inaccurate device control in complex tasks.

Method used

The token is a four-dimensional discrete spatiotemporal system used as a large language model. It includes coordinates of four dimensions: time (T), context (C), logic (L), and device (D). These coordinates explicitly identify the token's time, dialogue order, logical position, and device identity, thus constructing a structured cognitive system.

Benefits of technology

It improves the model's dialogue understanding ability, enhances the coherence of logical reasoning, supports time synchronization and device control in the Internet of Things scenario, is compatible with existing model architecture, and has low deployment cost and good scalability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The application discloses a Token marking method and a cognitive system for large language model input data, and belongs to the technical field of artificial intelligence. The application generates corresponding discrete coordinates for each Token input into the large language model, including: a time coordinate for providing a unified time reference; a context coordinate for identifying a dialogue time sequence; a logic coordinate for identifying the sequential position of the Token in a text or instruction, and the logic coordinate is independently numbered within the range defined by the time coordinate or the context coordinate; and a device coordinate for identifying an Internet of Things device, a role or a priority. The coordinates jointly constitute a unified discrete space-time cognitive framework, and provide a structured worldview for the large language model, so that the large language model can perform structured understanding based on multi-dimensional information such as time, context, logical sequence and identity. The application does not change the core architecture of the model, has strong compatibility, and can be widely applied to the fields of multi-round dialogue, intelligent customer service, industrial automation, smart home, Internet of Vehicles and the like.
Need to check novelty before this filing date? Find Prior Art