Long text information processing method and system based on large model, medium, product and terminal

By performing preset threshold segmentation and differential processing on long texts, compressed semantic vectors and text semantic vector representations are generated, solving the problems of computational power consumption and context integrity in long text processing of large language models, and realizing efficient and lossless long text information processing.

CN121562565BActive Publication Date: 2026-06-02SHANGHAI GUANGYU XINCHEN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI GUANGYU XINCHEN TECHNOLOGY CO LTD
Filing Date
2026-01-20
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing large language models suffer from high computational costs and easy loss of contextual integrity when processing long text information. Furthermore, solutions that modify the model architecture or rely on external knowledge bases have high development costs, poor compatibility, or the risk of information loss.

Method used

By dividing a long text into a first sub-text segment and a second sub-text segment using a preset threshold, image conversion and encoding compression and lexicalization are applied respectively to generate compressed semantic vectors and text semantic vector representations, which are then input into a large language model through splicing and adaptation.

Benefits of technology

Without modifying the large language model architecture or relying on external knowledge bases, it achieves efficient processing of long texts while preserving the complete semantic logic, reducing computational complexity and resource consumption, and ensuring processing efficiency and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121562565B_ABST
    Figure CN121562565B_ABST
Patent Text Reader

Abstract

The application provides a long text information processing method, system, medium, product and terminal based on a large model. The long text is divided into a first subtext segment and a second subtext segment through text division based on a preset threshold and a differentiated processing strategy. The corresponding semantic vector representation is obtained through first path processing and second path processing, respectively. Finally, semantic lossless fusion is realized through ordered splicing and feature space projection. The application realizes high compression of long text and reduction of computational complexity under the premise of ensuring that the performance is basically consistent with the original performance of the large language model. The application solves the problems of high computational resource occupation, high processing delay and the like caused by high attention calculation cost and semantic loss in the traditional long text processing method.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Method and device for training large language model based on long text

    CN119004107A

  • Fault-tolerant processing method and system for super-long text in large model service, and storage medium

    CN121144496A