A language cognitive coding analysis method and system based on topological difference and a storage medium

CN122819152APending Publication Date: 2026-09-25马寅方
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611007050.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-08
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

1. 语义依赖性强:需大量标注数据训练,输出受训练数据分布影响,无法捕捉文本背后的真实意图与逻辑结构;

Benefits of technology

1. 结构解析深度:突破传统NLP语义表层分析,引入拓扑学视角识别四类逻辑结构,揭示语言对思维的模具作用;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The application discloses a language cognitive coding analysis method and system based on topology difference and a storage medium, and relates to the technical field of artificial intelligence and natural language processing. The method constructs L0-L3 ontology architecture, performs name-real mutual test triple search, calculates macroscopic EII overflow value Delta eii and implied meaning weight probability P-Weight, quantifies confidence by mu(x)=C_topo / (K_conflict+F_ambig), and realizes non-semantic analysis through a Base / Fuzzy dual-mode kit. The application can accurately identify the interest driving and cost transfer path behind the implied meaning, solves the defects that the prior art cannot quantify the interest conflict and is interfered by the physical layer, and is suitable for financial compliance, public opinion monitoring and the like.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Title: Method, System and Storage Medium for Linguistic CognitiveEncoding Analysis Based on Topological Differencing

[0002] Technical Field: The field of interdisciplinary technology of artificial intelligence and natural language processing (NLP), particularly a non-semantic analysis method targeting the coding layer of human cognition, applicable to scenarios such as financial compliance review, public opinion monitoring and early warning, corporate due diligence and macroeconomic policy analysis. Background Technology

[0003] Current technologies for processing language text primarily rely on deep learning-based NLP models (such as BERT and GPT), which have the following inherent drawbacks: 1. Strong semantic dependence: It requires a large amount of labeled data for training, and the output is affected by the distribution of training data, making it unable to capture the true intent and logical structure behind the text; 2. Lack of structural analysis ability: Treating text as a linear sequence of words, unable to identify the underlying topological structure of language (tree / circular / dome / root logic), and finding it difficult to analyze the shaping effect of different logical structures on cognition; 3. Lack of Quantification of Conflicts of Interest: When analyzing financial statements, policy documents, and other texts involving conflicts of interest, it is impossible to quantify the difference between "nominal claims" and "actual interests," making it difficult to identify subtext, cost-shifting paths, and institutional biases. 4. Physical layer interference: Some technologies mix L0 physical data such as EEG and heart rate with L1 language data, resulting in analysis results that are strongly correlated with physical state and lose the objectivity of cognitive analysis.

[0004] Therefore, there is an urgent need for a language cognitive coding analysis method that can strip away semantic representations, analyze underlying topological structures, quantify interest differences, and strictly isolate physical layer interference. Summary of the Invention

[0005] (a) Purpose of the invention This invention aims to provide a language cognitive encoding analysis method, system, and storage medium based on topological difference, addressing the aforementioned problems in existing technologies. By constructing an L0-L3 ontology architecture, the analysis is strictly limited to the L1 human cognitive encoding layer. Utilizing topological difference operations and quaternion operation logic, high-precision analysis of deep text structure, subtext, and interest flow is achieved.

[0006] (II) Technical Solution The method provided by this invention includes the following steps: 1. Ontological Architecture Construction and Input Gate Verification: A four-level architecture is constructed, comprising an L0 physical information layer, an L1 human cognitive encoding layer, an L2 topological differential operation layer, and an L3 technological incarnation layer. Only L1 layer data is processed. A triple-check retrieval of name and reality is performed on the input L1 material. * Name → Reality Verification: Verify whether core proper nouns (entity name, brand name) correspond to real entities in compliant public information sources (business registration, official directories). If the verification fails, the analysis will be terminated. * Real-name triple verification: For the core claimed attributes of an entity (business hours, capital background, etc.), obtain cross-verification from at least 3 independent and compliant sources. If the 3 sources are inconsistent, it is marked as "mismatch between name and reality". If there are fewer than 3 sources, it is marked as "information missing". * Interest Anchor Extraction: Anchors are extracted from the compliantly disclosed L0 layer economic interest data (revenue, costs, financing, fines, sanctions records) as the basis for L1 layer analysis.

[0007] 2. Survival-Subject-Stance Anchoring (STEP 0) Anchoring is performed after passing through the input gate: identifying subject type (natural person / legal person / state apparatus subsystem / cluster / media machine), text proxy layer (official stance / direct leader's voice / institutionalized writing / filtering layer), authentication root (personal will / capital / state authorization / algorithmic feedback), anchoring survival drive (fundamental survival pressure) and stance (maintained interest boundaries), establishing a data volume baseline (Level 1-4) and setting initial confidence levels. If economic interest tracing reveals significant penalty risks or sudden changes in the weights of the macro interest matrix, this step is forcibly restarted.

[0008] 3. Topological Positioning (STEP 1) Determine the cognitive geometry of the object of analysis: Identify the native language base (Indo-European / Sino-Tibetan / Persian / Asiatic language family features), determine the structural complexity (single structure / double structure / triple structure / quadruple structure, corresponding to different environmental complexities), identify differential features (strong rule type, contextual type, pervasive type, anchoring type), and analyze the topological deformation driven by interests (such as the tendency of circular structures to converge into tree structures under punishment pressure).

[0009] 4. Syntactic Anatomy and Core Operations (STEP 2): Deconstructing language features and performing core operations: * Lexical closure system analysis: Determine the high cohesion, high inclusiveness, or discretization characteristics of term usage; * Argument chain feature analysis: Identifying linear derivation, contextual resonance, authoritative citation, or semantic derivation patterns; * Subtext Three Interface Analysis: Identifying Channel Separation (Key Information Hidden Channel and Public Channel), Defining Right-of-Way Barriers (Attribution of Right to Define Core Terms and Exclusion Items), and Cost Transfer Paths (Explicit / Implicit Cost Bearers and Transfer Paths); * Institutional Segregation Index (EII) Calculation: For collective subjects, calculate the degree of contamination of the decision-making system by emotions / personality (high segregation / low segregation / mask type). * Calculation of macro EII spillover value (Δeii): Δeii = sum (W_module} times D_module), where W_module represents the weights of the four macro modules (policy compliance W_p, fiscal performance W_f, public opinion stability W_s, and technological survival W_t), and D_module represents the difference between an individual's wording and the optimal solution for that module. The criteria are: Δeii > 0.2 indicates institutional isolation failure (individual recklessness), and Δeii < -0.2 indicates institutional repression (individual conservatism). * Subtext weight probability (P-Weight) calibration: Based on δ_{eii}, determine the driving module and individual status (powerful faction / mouthpiece / technocrat / radical faction).

[0010] 5. Cognitive Computational Power Distribution Inference (STEP 3) Based on text features, infer the activity level of cognitive dimensions: executive control (fluency of register switching, metacognitive monitoring), crystallized intelligence (terminology expertise, cross-domain integration ability), and fluid reasoning (cross-domain modeling ability, theoretical construction ability), output comprehensive cognitive features (such as high control-high aggregation type), and supplement with benefit pressure correction terms (executive control is passively improved under high penalty pressure).

[0011] 6. Core Consciousness Architecture Determination (STEP 4) Comprehensive determination of consciousness structure: assess logical carrying capacity, field adaptability, value anchoring power, semantic penetration power, reality mapping ability (L1→L2→L3 level implementation ability) and cross-scale mapping potential (isomorphism between micro-individual cognition and macro-civilization structure), output a comprehensive feature summary, and predict future actions based on Delta_{eii} and P-Weight.

[0012] 7. Multitrack Differential Confidence Measurement: The confidence level is calculated using the statutory formula: \mu(x) = \frac{C_{topo}}{K_{conflict} + F_{ambig}}, where C_{topo} is the topological overlap, K_{conflict} is the logical conflict, and F_{ambig} is the semantic ambiguity. If \Delta_{eii} exceeds the normal range, the confidence weight of \mu(x) is reduced proportionally.

[0013] 8. Dual-mode kit computation (L3 layer) * Base module (basic topology engine): sequentially passes through the authentication module (outputs material confidence weight w_{auth}), the routing module (outputs topology type label T_{type} and complexity score S_{comp}), the execution module (outputs differential confidence score \mu(x) and topology closure rating R_{close}), and the closure module (outputs topology breakpoint marker P_{break}). * Fuzzy module: Only starts when \mu(x) < 0.7 or R_{close} \neq is complete, performs dynamic membership assignment, cost transfer path quantization, threshold adaptation, and three-level source chain binding (quantization result → topology node → text basis). * Extended submodules: The Benefit Differential Quantization submodule calculates the Delta_{eii} value and the differential source module; the P-Weight Calibration submodule outputs the latent weight probability P and the influence type label.

[0014] (III) Beneficial Effects Compared with the prior art, the present invention has the following advantages: 1. Depth of structural analysis: Breaking through the traditional semantic surface analysis of NLP, it introduces a topological perspective to identify four types of logical structures, revealing the molding effect of language on thinking; 2. Ability to quantify benefits: Pioneering the Delta_{eii} and P-Weight models to quantify the difference between the macro-benefit matrix and individual will, accurately identify the benefit drivers and cost transfer paths behind the subtext, and solve the problem of "inconsistency between words and deeds" in compliance review; 3. Physical layer isolation: Strictly define the boundary between L0 and L1 layers, and use L0 only as an empirical anchor point rather than a computational source to avoid interference from noise such as biological signals and ensure the scientific nature of cognitive analysis; 4. Dual-mode robustness: The Base / Fuzzy dual-mode architecture balances rule interpretability with quantization accuracy in complex scenarios; 5. Wide range of applications: It can cover high-value scenarios such as KYB, AML, public opinion early warning, executive background checks, and geopolitical policy analysis, and has extremely high commercial and strategic value. Attached Figure Description

[0015] * Figure 1 : A flowchart illustrating the language cognitive coding analysis method based on topological difference in this embodiment of the invention. * Figure 2 : A schematic diagram of the hierarchical relationship of the ontological architecture (L0-L3) in this embodiment of the invention * Figure 3 : A schematic diagram of the I / O logic of the LTM Base (basic topology engine) quad module in this embodiment of the invention. * Figure 4 The calculation logic of \Delta_{eii} and the P-Weight calibration flowchart in this embodiment of the invention. * Figure 5 : Workflow diagram of the dual-mode kit (Base / Fuzzy) in this embodiment of the invention * Figure 6 : A schematic diagram of the name-real-name mutual verification triple retrieval process in an embodiment of the present invention * Figure 7 : A schematic diagram of the core consciousness architecture determination logic in this embodiment of the invention Detailed Implementation

[0016] Special Note: The core improvement of this invention lies in constructing an L0-L3 ontology architecture and performing differential operations on the L1 human cognitive coding layer. The following embodiments are merely specific application examples of this invention in the field of compliant text analysis and do not constitute a limitation on the technical solution itself. Those skilled in the art can implement cognitive coding analysis in other fields based on the content disclosed in this invention.

[0017] The present invention will be further described in detail below with reference to specific embodiments. The embodiments are only used to explain the present invention and do not constitute a limitation on the scope of protection.

[0018] Example 1: Cognitive Encoding Analysis in a General Abstract Scenario Input: A publicly released policy text from an organization (L1 material, approximately 2000 words, including 5 subject claim attributes).

[0019] step: 1. Input gate verification: Perform a triple verification of name and reality. Name → Reality verification confirms that the publishing entity is a real entity in a compliant public information source (such as a government-published directory); Reality → Name triple verification obtains cross-verification from at least 3 independent information sources for 5 core claimed attributes. If one attribute (claiming "consensus among all members") is inconsistent with the publicized signature coverage rate (60%), it is marked as "name and reality do not match".

[0020] 2. Anchoring: The identified subject type is a state apparatus subsystem, the text proxy layer is official statements, the authentication root is state authorization, the survival driver is the continuation of the system, and the stance is maintaining the legitimacy of the system. Establish a data volume baseline Level 3.

[0021] 3. Topological localization: The native language base is identified as having Indo-European characteristics (subject-verb-object linear logic), with a structural complexity of a triple structure (linear logic + contextual logic + value anchoring), and the difference features are strongly rule-based.

[0022] 4. Syntactic Analysis: Analyzing the three interfaces of subtext: channel separation (key information hidden in unpublished attachments); definition of power checkpoints ("consensus" definition excludes opposing samples); cost transfer path (implicit costs are borne by the grassroots). Calculate the macro EII spillover value \Delta_{eii}, setting policy compliance weight W_p=0.3, public opinion stability weight W_s=0.3, and fiscal performance weight W_f=0.4 (example assignment, actual calibration based on scenario). The difference between individual wording and the optimal solution of W_s is D=0.25, resulting in \Delta_{eii}=-0.225, triggering an institutional repression alarm. Calibrate the subtext weight probability P-Weight, determining the driving module as W_s and the individual status as a mouthpiece.

[0023] 5. Dual-mode operation: Calculate the multi-track differential confidence level mu(x) = 0.38. Since mu(x) < 0.7, activate the Fuzzy module to quantify the probability of the public opinion guidance valve (the specific value is based on scenario calibration and is not limited in this invention).

[0024] 6. Output: Generate a comprehensive analysis report, marking the text as a mouthpiece statement under institutional repression, indicating the pressure on grassroots implementation, and suggesting the addition of original samples for verification.

[0025] Example 2: Enterprise Supplier KYB (Know Your Business) Screening Input: Supplier list (Excel format) submitted by the purchaser and related publicly available business / financial information (L1 materials).

[0026] step: 1. Input gate verification: Verify the entities of 3 candidate suppliers (compare business registration information); conduct triple verification on the attributes of "registered capital / number of employees / business scope", and find that 1 supplier claims to be "technology development" but its main business in its annual report is "trade transit", which is marked as "mismatch between name and reality" and PR filter material.

[0027] 2. Anchoring: The main type of the purchaser is a legal person, and the survival driver is supply chain compliance and cost reduction; supplier cluster topology characteristics are evident.

[0028] 3. Topological positioning: The purchasing party is dominated by tree-like logic (contracts / rules); the suppliers exhibit root-and-stem topological characteristics (multi-layered nesting, non-linear relationships).

[0029] 4. Syntactic Analysis: Analyzing the three interfaces of subtext: Channel separation (the "related party" column in the procurement list is blank, and the business registration shows cross-shareholding among shareholders); Definition of rights checkpoints (the definition of "compliant supplier" excludes the "trade transit" subcategory); Cost transfer path (explicit cost is the procurement premium of 5%, implicit cost is the compliance investigation time and potential associated risks). Calculate \Delta_{eii}, set the weight of W_f (financial performance) to 0.5 and the weight of W_p (compliance) to 0.5. A significant difference between the supplier's individual wording and the optimal solution of W_p triggers the anchoring restart mechanism.

[0030] 5. Dual-mode operation: The Base module marks the topological break point P_{break} (the "related party" definition weight checkpoint fails); the Fuzzy module constructs the transfer path map G_{cost} and quantifies the risk transfer probability p_{trans} (the specific value is based on scenario calibration and is not limited in this invention).

[0031] 6. Output: \mu(x) = 0.42 (medium confidence level, labeled difference), generating a screening report indicating that one supplier has a high risk of "misrepresentation + related penetration", and recommends supplementing on-site verification for compliance personnel's reference.

[0032] Example 3: Construction of Core Awareness Profiles for Senior Executives (Anonymized) Input: Public speeches, social media posts, and regulatory compliance penalty records of a CEO over the past three years (L1 materials).

[0033] step: 1. Input gate verification: Verify CEO identity information; extract publicly available compliance penalty records as a benefit anchor.

[0034] 2. Anchoring: The subject type is a natural person, and the text proxy layer changes with the scenario (public speech / social media).

[0035] 3. Topological localization: Identified as a quadruple structure, possessing cross-scale mapping potential.

[0036] 4. Syntactic analysis: PR filter features were detected (public speeches showed "verb reduction" and "risk point removal"); the EBI index was calculated to show that the public mood was isolated; the Delta_{eii} was calculated to show that the difference between the public statement and the company's actual performance (W_f) has been maintained above 0.25 for a long time, which is to identify the person as a radical and powerful figure.

[0037] 5. Dual-mode operation: The Base module establishes the baseline topology; the Fuzzy module handles the uncertainties brought about by PR filters and calibrates the probability of true intent.

[0038] 6. Output: Construct a dynamic core consciousness profile, marking its aggressive and adventurous personality and high-control executive characteristics. Disclaimer: This analysis is based on publicly available text data and does not constitute a qualitative evaluation of the individual; it is for institutional decision-making reference only.

Claims

1. A language cognitive coding analysis method based on topological difference, characterized in that, Includes the following steps: Ontological architecture construction steps: Construct a four-level architecture including the L0 physical information layer, the L1 human cognitive encoding layer, the L2 topological differential operation layer, and the L3 technological incarnation layer, and set it to process only the data of the L1 human cognitive encoding layer; Input gate verification steps: Perform name-real mutual verification triple retrieval on the input L1 materials, including name → reality verification, reality → name triple verification, and interest anchor extraction. The interest anchor extraction step includes extracting revenue, cost, financing, fines, and sanctions records from the compliantly disclosed L0 layer economic interest data as analysis anchors for the L1 layer; Anchoring steps: Identify the subject type, text proxy layer, authentication root, survival drive, and stance of L1 material, and establish a baseline for data volume; Topology positioning steps: Identify the native language base characteristics, structural complexity, and differential characteristics of L1 materials, and analyze the topological deformation driven by interests; Syntactic dissection steps: Deconstruct the lexical closure system and argument chain characteristics of L1 materials, analyze the three interfaces of the subtext including channel separation, definition of the right checkpoint, and cost transfer path, and calculate the institutional isolation index (EII); Differential calculation steps: Calculate the macroscopic EII overflow value \Delta_{eii} = \sum (W_{module} \timesD_{module}), where W_{module} is the weight of the four macroscopic modules, and D_{module} is the difference between the individual wording and the optimal solution of the module; Calibration steps: Based on \Delta_{eii}, calibrate the subtext weight probability (P-Weight) to determine the driving module and individual status; Confidence quantification step: The multitrack differential confidence score is calculated using the formula mu(x) = ...

2. The method according to claim 1, characterized in that, The triple retrieval of name-real-object mutual verification in the input gate verification step specifically includes: verifying whether the core proper noun corresponds to the real entity in the compliant public information source; obtaining cross-verification from at least 3 independent compliant information sources for the core claimed attributes of the entity; if the information from the 3 information sources is inconsistent, it is marked as name-real-object mismatch; if 3 independent information sources cannot be obtained, it is marked as information missing.

3. The method according to claim 1, characterized in that, The Base module operations include: the authentication module inputs L1 material and outputs material confidence weight w_{auth}; the routing module inputs L1 material with w_{auth} and outputs topology type label T_{type} and complexity score S_{comp}; the execution module inputs T_{type}, S_{comp} and L1 material and outputs differential confidence level \mu(x) and topology closure rating R_{close}; the closure module inputs \mu(x), R_{close} and L1 material and outputs topology breakpoint marker P_{break}.

4. The method according to claim 1, characterized in that, The Fuzzy module operations include: dynamic membership assignment, calculating membership degree m based on logical conflict degree K_{conflict} and semantic ambiguity F_{ambig}; cost transfer path quantization, extracting the rights and responsibilities allocation and benefit flow text in L1 materials, constructing a transfer path map G_{cost} and calculating the transfer probability p_{trans}; threshold adaptation, automatically matching the quantization threshold based on L1 material type labels; and source tracing binding, generating a three-level source tracing chain, including quantization results, topological nodes, and textual evidence.

5. The method according to claim 1, characterized in that, It also includes a benefit difference quantization submodule and a P-Weight calibration submodule, which are used to calculate the δDelta_{eii} value and the difference source module, respectively, and output the latent weight probability P and the influence type label.

6. The method according to claim 1, characterized in that, After the confidence quantification step, the method further includes: if \Delta_{eii} > 0.2 or \Delta_{eii} < -0.2, then the confidence weight of \mu(x) is reduced proportionally.

7. A language cognitive coding analysis system based on topological difference, characterized in that, include: The input gate module is used to perform triple verification of identity and reality and extraction of benefit anchor points; The anchoring module is used to perform survival-subject-position anchoring and establish a baseline for data volume; the topology localization module is used to identify native language base features, structural complexity, and differential features. The syntactic anatomy module is used to perform subtext parsing, calculate the institutional segregation index (EII) and the macro EII spillover value A calibration module is used to calibrate the subtext weight probability (P-Weight) based on Confidence quantification module is used to calculate the multi-track difference confidence level A dual-mode computation engine includes a Base module and a Fuzzy module. The Base module is used to handle steady-state analysis, and the Fuzzy module is used to start when the confidence level is below a threshold or the topological closure is incomplete, performing dynamic membership assignment, cost transfer path quantization, and source binding.

8. The system according to claim 7, characterized in that, The dual-mode computing engine also includes a benefit difference quantization submodule and a P-Weight calibration submodule, which are used to calculate the Delta_{eii} value and the difference source module, and output the subtext weight probability P and the influence type label, respectively.

9. The system according to claim 7, characterized in that, The system is configured to perform the method as described in any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 6.