An ai semantic analysis and data processing method based on a large language model

CN121901426BActive Publication Date: 2026-05-29ZHONGNAN INFORMATION TECH (SHENZHEN) CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHONGNAN INFORMATION TECH (SHENZHEN) CO LTD
Filing Date
2026-03-23
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing semantic analysis solutions cannot accurately understand the true meaning of technical terms in different contexts when processing professional domain texts, leading to errors in structured data entry and confusion in medical record archiving, thus reducing the accuracy of data processing.

Method used

We employ an AI semantic analysis method based on a large language model. By constructing contextual fragments through word segmentation, term extraction, and local context windows, and combining hidden layer output vectors and a three-state semantic working mode state machine, we dynamically adjust word frequency weights, identify the semantic stability level of professional terms, and accurately restore the weight of terms in specific contexts through co-occurrence word matching and vector distance determination.

Benefits of technology

It improves the accuracy and recall of knowledge graph construction and intelligent question answering systems in professional fields, avoids misjudgment of professional terms and retrieval omissions, and ensures the stability and accuracy of document extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121901426B_ABST
    Figure CN121901426B_ABST
Patent Text Reader

Abstract

The application relates to the field of data processing, in particular to an AI semantic analysis and data processing method based on a large language model. First, professional terms and context windows in vertical field texts are extracted, hidden layer vectors are extracted by using the large language model, and spatial dispersion is calculated, and the terms are divided into three-state working modes according to semantic stability. For polysemous terms, independent semantic branches are identified through vector clustering, co-occurrence word matching and vector distance determination are used to realize accurate semantic attribution in a new context, and dynamic segmented adjustment of static term frequency (TF-IDF) weight is carried out accordingly, and finally the improved weight is injected into the downstream analysis process. The application gives the same term different weight expressions in different contexts, and the low-overhead cascade disambiguation mechanism efficiently solves the polysemy problem of terms, avoids confusion of structured data and error clustering of documents, and significantly improves the accuracy of text matching, information extraction and document classification and the like.
Need to check novelty before this filing date? Find Prior Art