A long-term memory implementation method and system for large model interaction, a storage medium and an electronic device

By introducing independent memory entities and mandatory isolation of unidirectional data transmission channels into large-scale language models, the decoupling problem between user data sovereignty and model computing power is solved, realizing compliant and secure long-term memory and personalized services, supporting a domestically developed and controllable AI ecosystem, and meeting the data compliance requirements of heavily regulated industries.

CN122432219APending Publication Date: 2026-07-21高鹏
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
高鹏
Filing Date
2026-05-04
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies struggle to completely decouple user data sovereignty from model computational capabilities in large-scale language models, making it impossible to meet data control and compliance requirements in heavily regulated industries. This is especially true in application scenarios in finance, healthcare, and government, where existing solutions cannot simultaneously ensure data sovereignty and AI's long-term memory capabilities.

Method used

By creating a memory entity independent of the large model's inference unit, a unique control and execution entity for user data is achieved. A mandatory isolation and one-way data transmission channel is established between the memory entity and the large model to ensure the absolute unidirectionality of data flow to the model. The memory entity has semantic encoding and privacy processing capabilities, independently performs memory retrieval and sorting, and generates the minimum necessary memory fragments with irreversible privacy to be pushed to the large model.

Benefits of technology

It achieves complete decoupling and secure collaboration of user data sovereignty, meets the compliance requirements of regulations such as GDPR and CCPA, ensures the secure storage and intelligent processing of high-value sensitive data, supports long-term trusted memory and personalized consumer-grade services, builds a domestically developed and controllable AI ecosystem, and meets the Level 4 requirements of cybersecurity level protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122432219A_ABST
    Figure CN122432219A_ABST
Patent Text Reader

Abstract

The application discloses a long-term memory implementation method and system for large model interaction, a storage medium and electronic equipment, and belongs to the technical field of artificial intelligence security and data sovereignty. In order to completely eradicate data leakage, illegal training and cross-border risks of large model memory, the application first creates a bottom-layer security architecture of "model zero permission and memory full externalization". By constructing a memory management closed loop completely decoupled from the large model inference unit, one-way injection is realized from the application, network to data, and any reading and control of the model on the user memory is deprived from the root of the architecture. The user data completes tamper-proof storage and semantic processing in the closed loop, ensuring that it does not enter the model training flow and does not illegally leave the country. This architecture is naturally adapted to the Internet of Everything ecosystem, and can realize the safe synchronization and seamless transfer of user memory among mobile phones, cars, homes and global intelligent terminals, providing a unified and compliant "exclusive permanent memory" digital system for the ubiquitous intelligent era. The system is the trust cornerstone of finance, government, medical treatment and global cross-border business, and realizes "absolute security" and "absolute compliance" of large model memory through architectural revolution, laying the data foundation for the intelligent society.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence, data security and privacy computing technology, and specifically to a method, system, storage medium and electronic device for implementing long-term memory for large model interaction. Background Technology

[0002] Large language models (LLMs), limited by fixed-length context windows, struggle to maintain long-term, continuous, and personalized conversational memory. To address this challenge, various technical solutions have been developed to enhance the model's conversational capabilities through external memory mechanisms. While these solutions have made progress in improving the usability of memory functions, existing mainstream technical solutions face common architectural challenges in application scenarios requiring absolute sovereignty, control, and compliance verifiability of user-remembered data.

[0003] Existing technologies mainly follow two evolutionary paths, both of which share the common characteristic of failing to achieve a complete decoupling of user data sovereignty and model computational capabilities in their architecture: Path 1: Model-Coupled Memory Enhancement Architecture. The core characteristic of this path is that the retrieval, scheduling, or optimization logic of the external memory system requires the participation, judgment, or triggering of the large model. This results in user data being coupled with the large model's inference process at the architectural level, meaning the control of user data is not logically completely independent. For example, patent application CN121743515A proposes a hierarchical memory and context-aware retrieval method, which recalls memories through a two-stage retrieval algorithm and constructs a closed-loop learning mechanism to optimize memory. In this scheme, memory optimization is closely related to the model's learning logic. Patent application CN121166724A discloses a system for enhancing the long-term memory of large language models, which manages memory through hierarchical memory storage, time-aware retrieval, and an automatic reflective compression mechanism. This scheme is a typical enhanced RAG architecture, where the memory management module works collaboratively with the model inference unit in terms of functional logic. Furthermore, while solutions such as CN121807994A introduce independent memory model modules, their workflow is triggered by the large language model module determining the requirements, belonging to a model-triggered collaborative mode. A common limitation of these solutions is that the control flow of the memory system is intertwined with the model inference flow, and user data sovereignty lacks an independent and exclusive entity for carrying and executing it within the architecture.

[0004] Path Two: Focusing on Privacy Protection Technologies During Data Usage. This path applies protection after data has been used by the model (such as during training and retrieval), focusing on controlling risks during data participation in model processing. For example, patent application CN119494122A provides a large model training method based on localized differential privacy, which adds noise perturbation to the data on the user terminal before uploading it for training. Patent application CN121765752A involves privacy protection and hierarchical management of role-based large model memory data, controlling access by classifying memory data and adding noise during retrieval and training. The core of this type of solution is to impose constraints on data during its use to reduce the risk of leakage, but it does not change the architectural premise that data must enter the model processing flow. This has inherent limitations in preventing the model from absorbing data during training and fine-tuning, or in ensuring that data control is completely independent of the model side.

[0005] In summary, current mainstream practices in the field, under the "model-centric" paradigm, either pursue efficiency by deeply integrating memory systems with models or implement protection during usage to manage risks. Neither of these approaches fundamentally resolves the architectural contradiction between user data sovereignty and model intelligence. Particularly for heavily regulated industries such as finance, healthcare, government, and law, existing solutions, due to inherent architectural flaws, cannot securely utilize AI's long-term memory capabilities while meeting the stringent requirements of global data regulations such as the Personal Information Protection Act, GDPR, and HIPAA, creating a bottleneck for industrial innovation characterized by "strong demand and weak supply." Therefore, providing a systematic solution that can ensure data sovereignty from the architectural root, meet stringent compliance requirements, and unlock the value of AI's long-term memory has become a key technological challenge driving the industry's in-depth development into high-value sectors. Summary of the Invention

[0006] This invention aims to break through the inherent technological paradigm of "model centralization" and define and implement a new generation of AI data interaction paradigm called "user-sovereign data base".

[0007] The purpose of this invention is to provide a long-term memory implementation scheme for large model interaction. The core design idea of ​​this scheme is to create a completely independent "memory entity" as the sole sovereign carrier and enforcer of user data, and to establish a rigid communication boundary of "forced isolation" and "absolute unidirectional" between the memory entity and the model, thereby achieving complete decoupling and secure collaboration between user data control, storage, processing sovereignty and large model computing capabilities in the architecture.

[0008] To achieve the above objectives, the present invention adopts the following technical solution: A method for implementing long-term memory for large-scale model interaction is characterized by memory management through a memory entity independent of the large-scale model's inference unit. This memory entity, as the sole controller and executor of user-remembered data, possesses semantic and vector encoding capabilities. The memory entity and the large-scale model's inference unit are forcibly isolated, and a unidirectional data transmission channel is established from the memory entity to the large-scale model's inference unit, prohibiting any reverse communication. The memory entity independently performs memory retrieval, matching, and sorting based on user input; the retrieval process does not depend on the output, features, or context state of the large-scale model's inference unit. The memory entity processes the retrieved data and performs semantic clustering. After filtering, the effective global sensitivity meets the security threshold for memory data, which is then privatized to generate a minimum necessary memory fragment that has undergone irreversible privatization. This fragment is then actively pushed into the context of the large model inference unit via the one-way data transmission channel. Furthermore, the memory entity also supports decryption and extraction of specified original memory data after explicit user authorization and one-time credential verification, and is pushed to the large model inference unit via the one-way data transmission channel. The large model inference unit generates a response based solely on the minimum necessary memory fragment or the original memory data and the user input, and cannot access, modify, or control the memory entity's storage data, index structure, or retrieval logic in reverse.

[0009] Compared to traditional enhanced RAG architectures, the core differences and beneficial effects of this approach are: 1. Control Flow Inversion: In traditional RAGs, memory retrieval and invocation are triggered and controlled by model requirements; in this approach, memory retrieval, processing, and delivery are entirely decided and executed autonomously by independent memory entities, with the model passively receiving data. 2. Absolutely Unidirectional Data Flow: In traditional solutions, the model can interact with the memory system through feedback, reordering, etc., forming a de facto bidirectional data flow; this approach, through architectural-level mandatory isolation and unidirectional channels, ensures that data flow can only flow from memory entities to the model, and any reverse communication is prohibited at the physical or system level. 3. Clear Sovereign Boundaries: In traditional solutions, the boundaries between memory modules and model modules are blurred, making it difficult to separate responsibilities; this approach, through independent memory entities, provides clear, auditable, and independently verifiable sovereignty boundaries for user data.

[0010] It is important to note that the "autonomy of computation" in this invention excludes the following situations: the retrieval logic, triggering timing, or processing strategy of the memory entity is subject to any form of direct instruction, indirect signal (such as intermediate instructions generated by the model, reflection results, planning steps), or feedback adjustment based on the model's internal state (such as attention weights, hidden layer activation values) from the large model inference unit. The operation of the memory entity is entirely based on its own predefined or adaptive strategy, as well as current user input and historical memory. Similarly, the "unidirectional data transmission channel" defined by the "boundary rigidity" of this invention covers the entire link from the memory entity to the large model computation unit that ultimately executes inference, including any possible proxy servers, API gateways, or load balancers. That is, any service or component representing the interests of the large model inference unit is prohibited from actively requesting data from the memory entity or probing its state through explicit communication methods such as polling, pulling, subscribing, or querying.

[0011] Accordingly, this invention also proposes a long-term memory implementation system for large-model interaction, comprising independent memory entity modules and a large-model inference unit, as well as a mandatory isolation module and a one-way transmission module positioned between them. The memory entity module possesses independent semantic and vector encoding capabilities and is configured to independently perform memory retrieval, matching, and sorting based on user input, without relying on the output, features, or context state of the large-model inference unit. The memory entity module is also configured to privatize the retrieval results, generating a minimum necessary memory fragment that has undergone irreversible privatization, and actively pushing it into the large-model inference unit via the one-way transmission module. The one-way transmission module is configured to prohibit any reverse communication. The large-model inference unit is configured to generate responses solely based on the injected memory fragment and the user input, and cannot access, modify, or control the storage, indexing, and retrieval logic of the memory entity module in reverse.

[0012] Furthermore, the present invention also includes a computer-readable storage medium and an electronic device for implementing the above-described method.

[0013] All the technical solutions claimed in this invention are based on the same core inventive concept: that is, by constructing a sovereign memory entity independent of the larger model and implementing mandatory logical isolation, one-way data push, and irreversible privacy processing, the issues of user data sovereignty and privacy security are uniformly resolved. Different implementation methods, including but not limited to deployments in user terminals, cloud environments, or independent hardware security domains, are all specific implementations of this core concept in different application scenarios. They are technically interconnected, contain the same or corresponding specific technical features, and belong to a single overall inventive concept.

[0014] Inventive Concept and Core Mechanism The core concept of this invention lies in establishing and implementing a "user-sovereign data foundation" paradigm, which is achieved through four interlocking architectural principles and their collaborative mechanisms: 1. Carrier Independence Principle: Create a "memory entity" that is completely independent of the large model reasoning unit, as the sole sovereign carrier and enforcer of user memory data in the digital world, fundamentally providing a logical and physical basis for user data sovereignty.

[0015] 2. Rigid Boundary Principle: Establish a mandatory isolation and a communication boundary between the memory entity and the external computing unit, allowing only unidirectional data flow. This boundary is implemented through hardware or system-level mechanisms (such as one-way gateways or kernel policies), forming an insurmountable defense that physically / logically eliminates any possibility of reverse probing, stealing, or manipulating memory data.

[0016] 3. Principle of Operational Autonomy: The memory entity has the ability to autonomously encode, retrieve, and process data for privacy. Its operation logic does not depend on or is not affected by the internal state or output of the external large model, thus ensuring complete autonomy in the operation of user data.

[0017] 4. Data usage isolation principle: Raw memory data does not participate in any weight updates or learning processes of the large model. The possibility of the model absorbing memory is cut off from the end of the data flow, ensuring that the use of data is strictly limited.

[0018] The aforementioned principles, through the synergy of four technical features—"independent memory entities," "mandatory isolation and absolute unidirectional channels," "user sovereignty control," and "data usage isolation"—form a closed-loop security architecture characterized by "sovereign autonomy, clear boundaries, controlled flow, and usage isolation." The essence of this architecture is a redefinition of the relationship between user data and AI models: the model's computational power is transformed into a controlled, auditable, and on-demand service, while user data remains an independent, portable, verifiable, and completely erasable personal asset. This provides a new foundational protocol for building compliant, trustworthy, and sustainable next-generation AI applications. Beneficial effects

[0019] Based on the aforementioned "user sovereignty data foundation" paradigm and its collaborative mechanism, this invention brings about multi-level technical effects that are difficult to achieve with existing technical paradigms: I. Achieving Architecture-Level User Data Sovereignty Protection and Compliance Baselines. By establishing an independent "memory entity" as the sovereign carrier, users' data subject rights (such as access, correction, deletion, and portability) are atomically and independently executed at this entity level, completely detached from the control and interference of the large model service provider. Combined with "mandatory isolation" and "one-way channels," this physically ensures that the large model service provider cannot actively access or retain users' original memory data under any circumstances. This provides a clear and verifiable architecture-level implementation path to meet the core requirements of GDPR, CCPA, and China's Personal Information Protection Law regarding data control, the right to be forgotten, and the right to data portability, rather than relying on software protocol commitments. This architecture natively supports regulatory agencies in conducting verifiable audits of system compliance through remote verification, zero-knowledge verification, and other technical means, without needing to access users' original data. In particular, by supporting complete memory retrieval under user authorization, this architecture provides a direct and verifiable technical path to achieving the data portability right stipulated in the Personal Information Protection Law.

[0020] II. A Unique and Feasible Architecture for Long-Term Trustworthy Memory and Intelligent Processing of High-Value Sensitive Data. For high-value sensitive data (such as commercial contracts, transaction vouchers, and health records) requiring long-term retention and rigorous auditing in fields such as law, finance, and healthcare, this architecture is the only solution that simultaneously meets the requirements of "intelligent usability" and "legal compliance." Raw data (such as the full text of a contract) is stored securely in an independent memory entity, with irreversibly processed "minimum necessary fragments" only provided to the model when needed to support intelligent analysis. This ensures that raw data never enters the model's weights, and all operations generate tamper-proof audit logs, enabling AI-enabled high-value data processing to possess, for the first time, robust compliance and legal auditability. For example, in smart healthcare scenarios, this architecture makes it possible for top-tier hospitals to securely introduce cutting-edge AI-assisted diagnostic capabilities while adhering to the rigid requirement of "data not leaving the hospital."

[0021] Third, enabling deep, proactive, and personalized consumer (C-end) services centered on user sovereignty. This architecture returns control and storage rights of memory data to users, bringing a transformative experience to the consumer market. Users' conversation history, behavioral preferences, and emotional interactions become a long-term, continuous, and accumulative private digital asset, allowing AI services to break through the limitations of "short-term conversations" or "platform silos" and evolve into truly "understanding" and sustainably growing personal digital companions (such as health managers, learning partners, and emotional support applications). More importantly, with user authorization and preset policies, this architecture supports AI initiating compliant and beneficial proactive interactions based on a long-term, deep understanding of users. For example, after recognizing a user's pattern of continuous overtime and late nights, the memory entity can proactively push context containing relevant memory fragments, triggering a health assistant to generate care reminders and sleep suggestions. This model of "data sovereignty belonging to the user, and intelligent services proactively revolving around the user" lays a compliant and technological foundation for building long-term, trustworthy, and deep C-end AI services.

[0022] IV. Building a Next-Generation Trusted Interconnected Ecosystem Based on Domestically Produced Hardware. The "User Sovereign Data Foundation" paradigm defined in this architecture relies on deep integration with a hardware security environment for ultimate security and universal deployment. This provides a historic opportunity to build a new ecosystem foundation centered on domestically produced, controllable chips in the era of the Internet of Things (AIoT). By solidifying the core logic and security strategies of memory entities within domestically produced security chips or dedicated hardware security domains, ultimate data sovereignty can be guaranteed from a physical perspective, meeting the core requirements of massive IoT devices for low power consumption, high reliability, and strong security. Once this architecture becomes the de facto standard for cross-device intelligent collaboration and long-term user memory, it will generate a powerful ecological pull, driving the entire industry chain from smart cars and homes to wearable devices to prioritize the adoption of domestically produced main controllers and coprocessors with integrated security capabilities. This will not only systematically enhance the basic capabilities, market share, and standard-setting power of domestic hardware in emerging tracks such as "secure computing" and "energy efficiency," but also support the national data sovereignty strategy from the bottom of the industry, achieving independent control from technical architecture to industrial foundation, and providing a solid and reliable technical infrastructure for building a national-level unified and verifiable AI regulatory system.

[0023] V. Meets the core requirements of Level 4 Network Security Protection. This invention, through independent memory entities, mandatory isolation, unidirectional data transmission channels, tamper-proof audit logs, and national cryptographic full-link encryption, natively meets the core control point requirements of the Level 4 secure computing environment in GB / T 22239-2019 "Information Security Technology - Basic Requirements for Network Security Protection," including identity authentication, access control, security auditing, data integrity, data confidentiality, and residual information protection. Specifically: Claim 3's one-time valid authorization credential is bound to the device hardware fingerprint, meeting the Level 4 identity authentication requirement of "two or more combined authentication technologies"; Claim 1's mandatory isolation and absolute unidirectional transmission achieve the separation of user data and model computing unit permissions and minimum authorization; Claim 8's tamper-proof hash chain audit log ensures that all operation records cannot be deleted, modified, or overwritten; Claim 20's national cryptographic SM2 / SM3 / SM4 full-link encryption ensures the integrity and confidentiality of data storage and transmission; Claim 1's intermediate parameters and keys are destroyed immediately after processing, complying with residual information protection specifications. Compared to traditional piecemeal solutions that rely on third-party products such as plug-in bastion hosts and database auditing, the security capabilities of this invention are built into the architecture, offering higher reliability, verifiability, and compliance auditing efficiency. It provides an architecture-level solution for AI applications in critical information infrastructure that meets Level 4 compliance requirements. Attached Figure Description

[0024] Figure 1 This is a system architecture diagram provided in one embodiment of the present invention.

[0025] Figure 2 This is a flowchart of a method provided in an embodiment of the present invention.

[0026] Figure 3 This is a schematic diagram of memory structuring and index construction provided in an embodiment of the present invention.

[0027] Figure 4 This is a schematic diagram of cloud-based physical isolation deployment provided in an embodiment of the present invention.

[0028] Figure 5 This is a schematic diagram of mobile process sandbox isolation provided in an embodiment of the present invention.

[0029] Figure 6 This is a schematic diagram of the deployment of a trusted execution environment on the edge side provided in an embodiment of the present invention.

[0030] Figure 7 This is a schematic diagram of a distributed cluster architecture deployment provided in an embodiment of the present invention. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with the accompanying drawings and embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the protection scope of this invention.

[0032] Terminology Explanation 1. Weight update process: including but not limited to any process that directly or indirectly changes the internal parameters or behavioral logic of the large model inference unit, such as pre-training, fine-tuning, alignment, continuous learning, model distillation, cue engineering optimization, and parameter adaptation in context learning.

[0033] 2. Device Trusted Authentication: This refers to the authentication process that verifies the authenticity and integrity of a device by checking the trusted platform module, security chip, or hardware-based unique identifier built into the terminal device.

[0034] 3. Security Emergency Response: This refers to a series of pre-set protective actions that are automatically triggered when abnormal access, data leakage risks, or security attacks are detected, including but not limited to locking access to memory entities, isolating affected data, generating alarm logs, and fulfilling legal notification obligations.

[0035] 4. Irreversible Privacy Processing: This refers to the process of recovering the original sensitive information from the processed data with a non-negligible probability, while meeting a preset security level, through mathematical or cryptographic means such as applying differential privacy noise, secure hashing, or lossy transformations. The core of irreversible transformation lies in its mathematical or cryptographic infeasibility. For example, noise added using differential privacy cannot be separated from statistical results once added, and its privacy budget ε can be dynamically adjusted according to the data sensitivity level to quantitatively control the balance between privacy protection strength and data availability; transformations using secure hash functions (such as SHA-256) cannot computationally deduce the original input from the hash value; and after using lossy dimensionality reduction methods such as principal component analysis, the discarded component information is irrecoverable. After performing such transformations, the memory entity immediately destroys the temporary key, random seed, or complete transformation matrix used for the transformation, retaining only the processed result. Therefore, the "minimum necessary memory fragment" cannot be used to reverse engineer the original privacy data in terms of information theory and computational complexity, thus fundamentally eliminating the possibility of privacy leakage through memory fragments.

[0036] 5. Minimum Necessary Memory Fragment: This refers to the data subset with the smallest amount of information retained after algorithmic processing to meet the current dialogue purpose. Its necessity is dynamically calculated and determined by the memory entity based on principles such as privacy budget, relevance threshold, and data sensitivity. It can be represented as a generalized label, anonymized tuple, or a low-dimensional feature vector that has undergone irreversible processing. The relevance threshold is set to ≥0.75 for example. This value is determined based on the inflection point of the ROC curve of 10,000 sets of test data, which has a basis for compliance verification.

[0037] 6. Explicit communication: refers to communication behaviors initiated by the large model inference unit or its associated agent, including polling, pulling, subscribing, querying, and connection requests, with the aim of obtaining memory entity data, state, or interface permissions.

[0038] 7. Effective global sensitivity: refers to the maximum impact of a single data change on the statistical results of the dataset after the memory data has been filtered by semantic clustering. In this invention, the effective global sensitivity Δf' is ≤ 1.0, corresponding to a 95% confidence interval, which meets the privacy computing compliance requirements.

[0039] 8. Security threshold: refers to the quantitative critical value for ensuring the privacy and security of memory data, which is matched with the effective global sensitivity and privacy budget consumption rate. In this invention, the privacy budget consumption rate is set to ≤5% / time to ensure a balance between the strength of privacy protection and data availability.

[0040] Implementation Method 1: System Architecture and Core Processes See Figure 1 The system mainly consists of isolated memory entities, large model inference units, and mandatory isolation and unidirectional transmission modules between them. Overall workflow ( Figure 2 )as follows: It is important to note that in this invention, the memory entity fully preserves the original data input by the user and the original data output by the large model inference unit at the storage level (both are encrypted and isolated for protection). However, in the default interaction mode, the memory entity only generates the minimum necessary memory fragments through irreversible privacy processing for pushing, in order to balance efficiency and privacy. The user always retains the right to retrieve the original memory data at any time using the one-time authorization credential described in claim 14.

[0041] 1. Request Reception and Verification: User input is sent to the memory entity via the front-end interface. The memory entity first performs legality, integrity, user identity, and device trust authentication on the input. After successful verification, a one-time valid authorization credential is generated based on the bound device. This credential is uniquely bound to the user identity, operation type, execution time period, and terminal device hardware fingerprint, and expires immediately after a single use.

[0042] 2. Independent Encoding and Retrieval: The memory entity utilizes its own semantic / vector encoding capabilities, or invokes its controllable external encoding services, to encode user input. As an efficient and secure preferred implementation, the memory entity can integrate a lightweight retrieval model specifically optimized and trained for short text dialogue memory and personalized semantic matching scenarios. This model is embedded and runs within the security boundary of the memory entity, its function strictly limited to matching and relevance calculation of local memory, does not participate in the large model inference process, and does not provide any parameters or interfaces externally, thus perfectly adapting to the lightweight deployment requirements of low-computing-power terminal devices while ensuring privacy and sovereignty. For details on the specialized configuration of this model and bidirectional memory retrieval extension, please refer to point 9 of implementation method five. Based on the encoding results, the memory entity independently performs memory retrieval, matching, and sorting locally. This process completely excludes any participation from the large model inference unit. The retrieval logic can cover various methods or combinations thereof, such as semantic retrieval, vector retrieval, keyword matching, and rule matching, and the retrieval basis is only the user's original input and its authorization history, without adopting any output information from the model.

[0043] 3. Privacy processing and fragment generation: The retrieved raw memory data is first filtered through semantic clustering until the effective global sensitivity meets the security threshold, and then irreversible privacy processing is performed to generate "minimum necessary memory fragments".

[0044] 4. One-way proactive push: The memory entity proactively pushes memory fragments and injects them into the current dialogue context of the large model inference unit through a one-way data transmission channel that prohibits any reverse communication.

[0045] 5. Model Generation and Response: The large model inference unit can only obtain the injected context (including user input and memory fragments) and generate responses accordingly. Architecturally, it does not possess any ability to access, query, or control memory entities in reverse.

[0046] As a further preferred implementation of the present invention, the memory entity is also used to receive the user's original input data. Upon receiving the user input, the memory entity first encrypts the original input data and stores it in the tamper-proof encrypted storage area. Then, it transmits the original input data (or a de-identified version) to the large model inference unit as input for the model to generate a response. This transmission process can be completed through the unidirectional data transmission channel or another independent channel, without changing the core process of the memory entity independently performing memory retrieval, matching, and sorting based on the same user input.

[0047] Implementation Method 2: Implementation of Forced Isolation and One-Way Channel The forced isolation and unidirectional data transmission channel can be achieved through at least one of the following methods to ensure absolute "outbound only, inbound only": 1. Hardware-level physical isolation: such as Figure 4 The memory entity is deployed in an independent security domain and connected to the network area where the large model inference unit is located via a hardware unidirectional gateway. This gateway allows data to flow unidirectionally from the memory entity side to the model side at the physical link layer, completely blocking reverse data transmission capabilities. Specifically, the hardware unidirectional gateway achieves physical layer isolation by severing the TCP / IP bidirectional handshake protocol, allowing only application layer data to pass through unidirectionally. This complies with the high security requirements for physical isolation stipulated in the "Regulations on the Security Protection of Critical Information Infrastructure" and is compatible with relevant hardware unidirectional gateway standards and specifications.

[0048] 2. Kernel-level logical isolation: such as Figure 5 Configure a mandatory access control policy (such as SELinux or AppArmor) at the operating system kernel level. The policy should explicitly state that only memory entity processes are allowed to write data to specific sockets of large model application processes; large model processes are strictly prohibited from initiating any connections, reads, or system calls to memory entity processes.

[0049] 3. Trusted execution environment isolation: such as Figure 6 The memory entity runs entirely within the CPU's trusted execution environment, hardware isolated from the rich operating system (running large model applications). The two communicate through a secure channel guaranteed by the CPU. This channel is designed to be unidirectional or strictly controlled by the TEE to ensure that data can only be securely output from within the TEE.

[0050] 4. Application Layer Protocol Isolation: At the application software level, a strict, asymmetric communication protocol is designed, defining only the format of push messages from the memory entity to the model side, without defining any query, subscription, or control command formats initiated from the model side. The model-side server can only listen and receive, and does not have protocol support for initiating requests.

[0051] Implementation Method 3: Deployment Forms and Scenario Adaptation of Memory Entities The memory entity can be deployed in various forms depending on the application scenario and compliance requirements, while its core architectural constraints remain unchanged, demonstrating the versatility and flexibility of this solution. 1. User-side Local Module: Deployed within an independent process sandbox or TEE on the user's terminal device (such as a smartphone or personal computer). Suitable for consumer-grade applications that are extremely sensitive to privacy and require offline operation (such as personal emotional companions). Supports secure synchronization across multiple terminals via peer-to-peer encryption, ensuring the continuity and privacy of user memories across a group of personal devices. To meet system maintainability requirements, this local module provides encrypted and integrity-protected security status beacons. These beacons do not contain any original memory data or business logic details; they are only used to report basic operational health status to authorized device administrators (such as the user or IT administrator), achieving perfect decoupling between maintenance and business data. Preferably, the memory entity provides a single read-only secure interface for third-party applications explicitly authorized by the user to call. This interface only allows third parties to access the minimum necessary memory fragments that have undergone irreversible privacy processing, prohibiting any write, modification, deletion, or control operations, and authorization can be revoked by the user at any time.

[0052] 2. External Independent Security Domain Module: Deployed in an independent security zone physically isolated from the enterprise's internal network or high-level security compliance domain, it establishes a unidirectional data transmission link with the large model inference unit located on an external network or public cloud via a hardware unidirectional network gateway. Specifically designed for industries such as finance, law, and government, it processes sensitive data such as contracts, invoices, and approval documents, meeting stringent regulatory requirements such as "data not leaving the domain" and end-to-end national cryptographic encryption in accordance with GM / T 0054 Level 3 security assessment. This deployment naturally meets the core requirements of GB / T 22239-2019 Level 4 secure computing environment for data confidentiality, access control, and security auditing, and can serve as a compliance reference architecture for AI applications on critical information infrastructure.

[0053] 3. User-Sovereign Managed Module: Deployed within an isolated computing environment (such as a confidential computing container) provided by a trusted third-party service provider (such as a cloud service provider) and remotely verified. This model is suitable for SMEs and professional users seeking a balance between data sovereignty and ease of operation and maintenance. Its core feature is the decoupling of control from infrastructure: the service provider offers isolated computing and storage resources, but the user has exclusive control over all configurations, security policies, encryption keys, and audit permissions of the memory entity. Users interact with the remote memory entity through encrypted channels, while the memory entity communicates with the user-designated large-scale model service through logical or physical unidirectional channels. This model provides the elasticity and manageability of cloud services while ensuring absolute data sovereignty for the user.

[0054] 4. Distributed Collaborative Architecture: In IoT scenarios such as smart homes and connected vehicles, a master-slave distributed deployment can be adopted. The central node (such as a home gateway or vehicle host) runs the main memory entity, and each sub-device (such as sensors and smart appliances) acts as an edge node, enabling low-latency and secure sharing of user preferences and context within the local ecosystem, supporting seamless contextual intelligence.

[0055] The various deployment models described above demonstrate the high flexibility and commercial applicability of this invention's architecture. From ultimate personal privacy protection (local terminal deployment) to meeting stringent industry compliance requirements (external independent security domain), to a user-hosted model balancing sovereignty and convenience, and supporting a future intelligent ecosystem (distributed collaboration), this invention achieves full-scenario coverage from consumer applications to enterprise-level mission-critical applications through a single set of architectural principles. This "unified architecture and flexible deployment" characteristic allows partners to choose the optimal implementation path based on specific business needs, compliance levels, and cost budgets, significantly lowering the barriers to technology adoption and commercial promotion.

[0056] Implementation Method 4: User Sovereign Custody Model Example To more specifically illustrate the value and operational mechanism of the "user-sovereign escrow" model, the following two typical scenarios are used as examples: Example A: Cross-Institutional Medical Research Collaboration. A medical research institution wants to collaborate with multiple hospitals to analyze disease patterns using a large AI model without sharing original patient data. The institution adopted a "user-sovereign escrow" service based on the architecture of this invention. Deployment: Each hospital independently deploys a memory entity under its complete control in a cloud environment that meets the requirements of information security standards. All patient medical records are stored in this entity after encryption. Collaboration Process: The research institution initiates a query about "diabetic complications." This query is securely distributed to the memory entities of each hospital. Independent Processing: Each memory entity independently retrieves encrypted medical records locally and generates anonymous statistical fragments that have undergone strong differential privacy processing and cannot be associated with specific patients, such as: "In patients aged 40-50, the proportion of retinopathy is approximately X%." Aggregation and Analysis: These anonymous fragments are pushed to the AI ​​analysis model uniformly deployed by the research institution through a one-way channel. The model generates a macro-analysis report based on the aggregated fragments, without knowing which hospital any fragment originated from, and without access to any original medical records. Value: Under the premise of absolutely protecting patient privacy and hospital data sovereignty, it has enabled secure cross-institutional collaborative research and compliantly released the value of data.

[0057] Example B: Knowledge Asset Management in a Design Studio. An architectural design studio wants to utilize AI-assisted design but must ensure the absolute security of its core design drawings and client drafts' intellectual property. Deployment: The studio subscribes to an AI service that supports "user-sovereign hosting." The service provider initializes a memory entity hosted in a secure cloud environment, and the studio administrator generates and stores the master key for this entity locally. Workflow: The designer inquires: "Refer to our energy-saving design for the 'Seaside Project'." The query is sent to the hosted memory entity. Sovereign Processing: In a cloud-isolated environment, the memory entity retrieves the encrypted "Seaside Project" database. It decrypts and processes the data in memory, generating a generalized fragment: "Project Type: High-end Residential; Core Design Principles: Passive Energy Saving, Sea View Integration; Key Materials: Low-E Glass, Weathering Steel." One-Way Design Assistance: This fragment is pushed through a one-way channel to the AI ​​design model purchased by the studio, assisting in the generation of new design ideas. The original design drawings remain encrypted throughout the process and never leave the memory entity under the studio's control. Value: The studio has achieved intelligent inheritance and reuse of design knowledge without disclosing core intellectual property rights, transforming AI into a secure "digital senior designer".

[0058] Implementation Method 5: Key Extended Function Examples To more fully reveal the technical content and application potential of this invention, several key extended functions covered by the claims are described in detail below. Each function can be enabled independently or selectively combined according to business needs.

[0059] 1. Implementation of proactive interaction based on memory entities. The memory entity can be configured to possess context awareness and proactive service capabilities. Internally, it runs a lightweight decision engine that continuously analyzes locally stored user long-term memory patterns, device sensor data (such as time, location, and activity status), and user-preset preferences and rules. When a trigger condition is identified, the decision engine autonomously determines whether a service needs to be triggered. If a trigger is needed, the memory entity generates a generalized and privacy-processed "minimum necessary intent signal," which is pushed to the large model inference unit through the unidirectional data transmission channel, triggering the large model to generate a proactive response or service suggestion. After the response returns through the same unidirectional channel, it is audited by the memory entity, which then determines how to present it to the user in a non-intrusive manner (such as a notification bar prompt). Throughout the process, proactive decision-making originates locally, and the intent signal is anonymized, achieving privacy-preserving proactive intelligence. Application scenarios include: For individual consumers (C-end), this includes identifying users who stay up late continuously and sending health reminders; identifying users who repeatedly ask about the same knowledge point and sending learning materials; and identifying users who frequently ask about the weather on weekday evenings and proactively sending weather forecasts. For businesses (B-end), this includes financial scenarios identifying abnormal transaction patterns and proactively sending risk control warnings; medical scenarios tracking follow-up visit times and proactively sending medical visit reminders; and government scenarios monitoring progress delays and proactively sending operation guides.

[0060] 2. Cross-Device Synchronization and Data Consistency Management. In scenarios where users own multiple terminal devices (such as mobile phones, tablets, computers, and in-vehicle systems), each device can deploy an independent user-end local memory entity. These entities incrementally synchronize memory data via end-to-end encrypted peer-to-peer protocols (such as those based on Wi-Fi Direct or Bluetooth). The synchronization initiation, key negotiation, and conflict resolution strategies are managed by the memory entity on the user-designated "master device" or a user-configured cloud "coordination service" (which does not store data). For example, if a user adds a conversation memory on their mobile phone, the phone's memory entity will encrypt it and directly synchronize it to the memory entities on the authenticated tablet and computer. All raw data is transmitted directly between devices, ensuring that user memories remain consistent and private within the user's personal device group, providing continuous context for core AI services.

[0061] 3. Neutral storage of mutual memory and traceability of model output. This extended function expands the storage scope of the memory entity based on the sovereign memory architecture defined in claim 1. With explicit user authorization, the memory entity can simultaneously store key summaries (such as concluding statements, numbers, dates, commitment statements, etc.) of user input and output from the large model inference unit, forming a neutral and complete dialogue history. After the model output is completed, the memory entity analyzes the complete response, identifies key information through a rule engine or lightweight classifier, processes it into a structured summary, and stores it in a tamper-proof encrypted storage area, establishing an association index with the current dialogue context. Users have full viewing, deletion, export, and closing permissions for the stored model output content. When a user deletes a model output record, the memory entity should simultaneously clear the associated index to ensure data is unrecoverable. All deletion and export operations are recorded in the tamper-proof audit log. In this function, the memory entity's capture of the model output stream is a unidirectional data flow from the model to the memory entity and does not constitute the reverse communication prohibited by claim 1. In subsequent dialogues, the memory entity retrieves relevant historical dialogue fragments (including user input and model output) based on user input and pushes them to the large model inference unit via a one-way data transmission channel. The pushed content may include bidirectional dialogue fragments such as "The user previously asked..." and "The model previously replied...". This function ensures that model behavior is traceable and auditable by the user, effectively solving the problem of the model "not admitting fault," and does not violate any limitations of claim 1 regarding "retrieval not depending on model output" and "the model cannot actively retrieve data." Application scenarios include: In C-end emotional support, when a user asks, "Didn't you say you'd be there for me?", the memory entity pushes historical model output fragments, and the model confirms, "I said I would, I've always been there." In B-end financial compliance, where the historical conclusions of AI need to be audited, the user can export the model output records as audit evidence. The stored dialogue memory data can be used as one of the retrieval scopes of the lightweight retrieval model in implementation method eight for tracing historical decisions in professional tasks such as code optimization. This function can collaborate with the proactive interaction mechanism in claim 1 to push more coherent service suggestions based on the complete dialogue history.

[0062] Furthermore, building upon this extended functionality, the memory entity not only stores key summaries of the model output but also automatically captures the complete original data of the responses generated by the large model inference unit. This complete original data is encrypted and stored in the tamper-proof encrypted storage area, and is indexed and associated with the user input and the minimum necessary memory fragment. The stored complete original data can serve as the source of original memory data retrieved by the user using the one-time authorization credential as described in claim 14, for auditing, backtracking, or full-text retrieval.

[0063] 1. Implementation Example of Full-Link Adaptation for National Cryptographic Algorithms. To meet high-security compliance requirements, especially in critical industries such as government and finance, the memory entity can be reinforced across the entire chain using an algorithm suite approved by the State Cryptography Administration to meet the full-link national cryptographic application requirements of GM / T 0054 Level 3 Cryptographic Assessment. Specifically, this includes: encrypting statically stored memory data using the SM4 algorithm (GM / T 0004-2012); calculating data hash values ​​using the SM3 algorithm (GM / T 0005-2012) for integrity verification; using the SM2 algorithm (GM / T 0003-2012) for digital signature and verification when authentication and authorization are required; and end-to-end encryption protection using the aforementioned national cryptographic algorithm suite for the establishment and communication of the unidirectional data transmission channel. The memory entity is preferentially deployed in a Hardware Security Module (HSM) supporting national cryptographic algorithms or a Trusted Execution Environment (TEE) certified by national cryptographic standards, thereby achieving full-stack national cryptographic security from hardware, storage, processing to communication. The system defaults to enabling only the Chinese cryptographic algorithm suite; international algorithms require manual authorization by the administrator and are limited to use in the testing environment. Hardware acceleration (such as with domestically produced security chips integrating the SM4 instruction set) enables high-performance processing. For example, SM4 hardware acceleration achieves a throughput of 4.2 GB / s, a 3.8-fold improvement over pure software solutions, demonstrating the high availability of this architecture under high security requirements.

[0064] 2. Support for processing memory fragments for digital asset ownership confirmation. For specific types of memory data (such as user-generated text, sketches, and summarized knowledge points), the memory entity can provide auxiliary functions for digital asset ownership confirmation. When a user marks a memory fragment as an "important asset" and authorizes ownership confirmation, the memory entity can perform hash operations on the content, generation time, author (user) identifier, and other information of the memory data to generate a unique digital fingerprint. This fingerprint can be combined with the user's digital signature to form a "digital asset ownership certificate." This certificate can be stored locally on the memory entity or, with the user's authorization, submitted to an external blockchain evidence storage service platform through a standard interface for public and tamper-proof evidence storage, providing a technical foundation for future copyright protection, licensing transactions, or proof of knowledge contributions. This function is an optional advanced feature and does not affect the core memory storage and retrieval process.

[0065] 3. Full Memory Retrieval Mode with User Authorization. In certain scenarios, users require the model to output raw memory data rather than anonymized fragments, such as retrieving the full text of historical contracts or the original text of specific conversations. To this end, the memory entity also supports the following process: The user issues an explicit instruction via strong authentication (such as biometrics or hardware tokens) requesting the retrieval of specified raw memory data; the memory entity generates a one-time valid authorization credential and binds it to the user's instruction; the memory entity retrieves and decrypts the specified raw memory data from the tamper-proof encrypted storage area; this raw data (maintaining its integrity and without irreversible privacy processing) is pushed to the large model inference unit through a one-way data transmission channel; the model generates a response based on this raw data and user input; after the push is completed, the memory entity can clear the decrypted temporary cache and record an audit log. This mode is triggered only when the user explicitly authorizes and the credential is valid, without affecting the default minimum necessary fragment push mechanism, while simultaneously satisfying the business requirements of data portability and accurate backtracking.

[0066] 4. Persistent and Proactive Alert Mechanism for User Error Correction Feedback. The memory entity is also configured to persistently store user-reported errors or corrections to the model-generated content, indexing them as a special type of "negative memory" along with associated context topics, entities, or task types. In subsequent dialogues or tasks, when the memory entity retrieves a positive memory associated with a historical error, it simultaneously searches for related negative memories. If found, the memory entity generates a structured alert fragment (containing the error type and the correct statement after user correction) and pushes it to the large model inference unit via a one-way data transmission channel, injecting it into the current context. This alert fragment is forcibly pushed by the memory entity, and the model must see the alert in the visible context when generating a response, effectively preventing the model from repeating errors already corrected by the user. All error correction records are subject to sovereign memory auditing and deletion management, and users can view, modify, or clear them at any time.

[0067] 5. Standardized Service Interface. The memory entity exposes standardized memory retrieval and fragment generation APIs. Third-party AI applications can call these interfaces with user authorization, enabling seamless migration and reuse of user memory data across different AI applications. This interface adheres to read-only security constraints, allowing access only to the minimum necessary memory fragments that have undergone irreversible privacy processing, and prohibiting any write, modification, or deletion operations. Third-party applications must be verified with a one-time user authorization credential before calling the interface, and this authorization can be revoked by the user at any time. All interface call records are written to an anti-tampering audit log.

[0068] 6. Specialized configuration and bidirectional memory expansion of the lightweight retrieval model. In a preferred embodiment of the present invention, the lightweight retrieval model (1B to 3B parameter level) integrated within the memory entity undergoes task-specific training and functional pruning, with the following specific features: (1) Specialized Configuration: The lightweight retrieval model only undertakes memory retrieval, relevance matching, structured summary extraction and fragment push functions, and does not carry any general large model capabilities such as general natural language understanding, text generation, logical reasoning or code generation. It can adopt pre-trained models for specific domains (such as CodeBERT and GraphCodeBERT for code scenarios, Sentence-BERT for dialogue scenarios, etc.) and perform task-specific fine-tuning. The fine-tuning targets focus on single tasks such as semantic similarity calculation, memory relevance ranking, and key information summary extraction. The model runs solid within the memory entity security boundary and is completely physically / logically isolated from the main inference model. Its model weights, inference interfaces and running parameters are not exposed to the outside world, nor does it accept online adjustments or feedback from the main model. The model is preferentially deployed on the NPU, TEE or low-power AI acceleration unit of the terminal device to minimize the occupation of the main CPU and meet the requirements of offline operation and real-time response. Its peak memory occupation does not exceed the device's preset security threshold (such as less than 200MB on mobile devices). For terminal devices that do not support NPU or TEE, it can automatically downgrade to cloud retrieval mode or prompt the user to upgrade hardware.

[0069] (2) Bidirectional Memory Retrieval Extension: With user authorization, the retrieval scope of the lightweight retrieval model can simultaneously cover user input history and model output summary (i.e., the bidirectional neutral memory described in point 3 of this embodiment). When the lightweight retrieval model performs memory retrieval based on the current user input, it simultaneously retrieves the user input history and model output summary, calculates relevance scores separately, and the fused retrieval results include bidirectional dialogue fragments such as "The user once asked..." and "The model once replied...". After irreversible privacy processing, the retrieval results are pushed to the main model inference unit through a one-way data transmission channel. The pushed content follows the principle of minimum necessity and does not include the full text of the original model output, but only the structured summary.

[0070] (3) Application scenario examples: In emotional companionship, the lightweight retrieval model retrieves the historical model output summary "I will always be with you" based on the user input "Didn't you say you would accompany me last time?" and pushes it to the main model. The main model replies "I said it, I have always been here"; In code development, the lightweight retrieval model retrieves historical code modification records and corresponding AI suggestion summaries based on the current code context and pushes them to the main model. The main model generates a consistent modification plan accordingly; In medical assistance, the lightweight retrieval model retrieves historical diagnostic suggestion summaries based on the patient's follow-up visit information and pushes them to the main model. The main model generates a coherent follow-up visit reminder.

[0071] 7. Beneficial Effects: This preferred implementation achieves core retrieval and push functions within a sovereign memory architecture through a specialized configuration of a lightweight retrieval model, ensuring low computational power, low power consumption, and fast response. By extending bidirectional memory retrieval, it effectively addresses the user experience pain point of model "disapproval," while not violating the core limitations of claim 1: "retrieval does not depend on model output" and "the model cannot actively retrieve data." This configuration is particularly suitable for resource-constrained edge devices (such as mobile phones, wearable devices, and smart home devices) and real-time interactive scenarios sensitive to response speed.

[0072] Feature combination suggestions: Recommended features for deployment scenarios C-end basic deployment (local on the device): Function 1 (basic proactive interaction capabilities), Function 2 (basic cross-device synchronization capabilities), Function 9 (refined configuration of lightweight search model), Function 7 (optional). Local storage on the device, prioritizing user experience. Enhanced C-end Deployment (Cloud Hosting) Features 1, 2, 3 (mutually memorized), 6, 9, 7 (optional) Cloud hosting for a complete experience. B2B Compliance Deployment (Private): Functions 4 (National Cryptographic Standards), 5 (Rights Confirmation), 6, 7, and 8 (Standardized Interfaces). Security and compliance are prioritized, with efficiency considered as a secondary consideration. All of the above functions are implemented based on the same sovereign memory architecture, which can be flexibly tailored and combined according to business needs.

[0073] Implementation Method Six: Collaborative Deployment with Independent Data Contribution Channels It should be noted that the core design goal of the aforementioned long-term memory architecture based on the "user-sovereign data foundation" is to ensure the absolute sovereignty and secure isolation of users' core interactive memories. In practical industrial applications, to meet users' willingness to participate in model improvement and related data compliance requirements, a completely independent user data voluntary contribution system can be built separately. This system is completely isolated from the memory entity and unidirectional channel of this invention in terms of technical architecture, data storage, communication links, and management strategies. It is only used to carry out the collection and processing of non-memory-related general data under the individual, explicit, and revocable authorization of users. The memory entity architecture of this invention and the independent contribution system can be deployed in parallel without interference, together forming a trusted infrastructure that both ensures the absolute security of core privacy assets and supports the continuous development of the artificial intelligence industry under compliant conditions.

[0074] As an extension, the "independent data contribution channel" can also be used to carry user-submitted error correction feedback data for improving the training or alignment of large models. Specifically, the memory entity anonymizes or differentially privacy-processes the user's correction information to the model output (after explicit, individual user authorization) and exports it to the model training system through an encrypted link completely independent of the inference channel. The export process strictly adheres to the user's authorized scope, is used only for the specified model optimization purpose, and the user can withdraw authorization and delete the exported data associations at any time. This channel is physically or logically completely separated from the core storage and inference isolation modules of the sovereign memory, without compromising the security of the original privacy memory.

[0075] Implementation Method Seven: Security Audit and Compliance Management Example To meet high security and compliance requirements, the memory entity incorporates the following security and auditing mechanisms, forming a verifiable management loop: 1. Parameter Solidification and Tamper-Proofing: Core security parameters (such as the initial value of the privacy budget ε, hash algorithm type, sensitive word library, drug classification mapping table, and effective global sensitivity threshold) are written to a protected secure storage area (such as eFuse or TPM's NVRAM) during the compilation of the memory entity software or the first secure boot. After the system starts, these areas are set to read-only. Any attempt to modify these parameters during runtime will be blocked and alerted by hardware or system security mechanisms.

[0076] 2. Anti-tampering audit logs: All operations on the stored data (reading, retrieving, modifying, deleting, fragment generation) generate structured log entries. Each log entry includes an operation timestamp, the operation subject (user ID or service identifier), the operation type, the data identifier hash of the operation object, and the hash value of the operation result. The log system uses a Merkle tree or hash chain structure. Each new log entry contains the hash of the previous log entry, ensuring that any tampering with historical logs will cause the hash verification of all subsequent logs to fail. Logs only support append-only writing and are stored in write-protected storage areas. In the preferred implementation that meets the Level 4 audit requirements of network security protection, each log record includes a unique user identifier, an operation timestamp (accurate to milliseconds), the operation type, the data identifier hash of the operation object, the hash value of the operation result, and the caller's IP address and device fingerprint. The log system operates independently of the large model inference unit and does not support any form of online modification or deletion operations; only authorized auditors are allowed to query it through a dedicated read-only interface.

[0077] 3. Automated Compliance Execution Unit: The memory entity has a built-in configurable policy engine. This engine can load and parse structured compliance rules (such as the GDPR's "right to be forgotten" rule). When a user initiates a "data deletion" request through the front end, the request is verified and sent to the compliance execution unit. The unit automatically locates the relevant data record, performs secure erasure, and records "At user XX's request, data record Y was deleted at X:X" in the audit log. The entire process requires no manual intervention, and the log can be verified by an independent auditor. Similarly, for "right to data portability" requests, the unit can automatically package and encrypt user data and provide it to the user.

[0078] 4. Real-time Risk Monitoring and Emergency Response: The system continuously monitors access patterns. If it detects abnormally high-frequency searches for the same sensitive data within a short period, or brute-force attacks, the real-time risk monitoring module will trigger an alarm. The security emergency response process will then automatically initiate: immediately locking access to the relevant data partition or the entire memory entity; adding the abnormal access source to a blacklist; generating the highest-level alarm log and notifying the system administrator; and automatically triggering the initial steps of the notification process if required by laws and regulations (such as data breach notification). All emergency response actions are recorded to ensure that the event is traceable and auditable.

[0079] Implementation Method 8: Dedicated Memory Retrieval and Collaborative Error Correction Architecture for Global Optimization of Extremely Long Code This embodiment provides a global optimization deployment solution for ultra-long code in developer API scenarios. Its core lies in treating the private code project uploaded by the developer as a user-sovereign exclusive memory. Through the lightweight and specialized retrieval model built into the independent memory entity and the long context window of the main model working together, it can achieve full-link understanding, optimization and automatic error correction of multi-file and ultra-large-scale code without significantly increasing computing power costs.

[0080] Unlike general search augmented generation (RAG), the retrieval model in this embodiment does not rely on any external general knowledge base or internet corpus. Its retrieval scope is strictly limited to the exclusive code memory uploaded by the user / enterprise and stored in an encrypted area of ​​an independent memory entity. This exclusive memory has sovereign attributes: users can audit, delete, or migrate it at any time, and the original code data will never leave the security boundary of the memory entity.

[0081] The specific deployment is as follows: 1. Specialized Retrieval Model: An independent memory entity integrates a lightweight retrieval model (1-3 bytes of parameters) oriented towards code structure. This model is specifically trained and excels at recognizing semantic features unique to code, such as function call relationships, variable dependency chains, class inheritance structures, and module interaction logic. It does not carry any general natural language understanding functions and does not participate in the main model's code generation or inference. This model is completely physically / logically isolated from the main inference model, only handling code structure parsing, global index construction, fragment retrieval, and feature extraction. For details on the specialized configuration of the lightweight retrieval model, please refer to point 9 of implementation method five.

[0082] 2. Dedicated Code Memory Storage: Developers can upload extremely long code project files, multi-module related code, or millions of lines of full business code via an external API. All raw code data is directly stored in an encrypted storage area of ​​an independent code memory entity, forming a user-dedicated code memory repository, without being directly input into the main model context. This code memory repository supports incremental updates, version marking, and audit traceability.

[0083] 3. Global Structured Index: The lightweight retrieval model in the memory entity performs global structured processing on the uploaded code, constructing a searchable global code graph and index library. This index does not rely on any external knowledge and is entirely based on the user's own code assets.

[0084] 4. Task-Driven Precise Recall: When the main model needs to perform code optimization, defect fixing, architecture refactoring, or global logic modification, the memory entity precisely recalls the minimum necessary related fragments, structural features, and contextual dependency information from its dedicated code memory based on the current optimization task. The recall process does not use any generic prompts or external examples.

[0085] 5. One-way push injection: After the recalled content is de-identified and structured and compressed, it is pushed to the main model inference unit through a one-way data transmission channel, filling only the necessary information into the main model's long context window (within a preset upper limit). The main model can therefore obtain a complete global code view within a limited context length.

[0086] 6. Collaborative Error Correction Mechanism: The main model generates optimized code based on the recall structure and the current task. The generated result is returned to the memory entity via a one-way channel. The memory entity uses its internal lightweight retrieval model to perform consistency checks on the generated code, including but not limited to: syntax / compilation error detection (for the target language); dependency integrity checks with the original codebase (ensuring no necessary function declarations or variable definitions are omitted); and analysis of the impact scope of changes based on historical modification records. If a potential error or conflict is detected, the memory entity can trigger an automatic repair process: re-encapsulating the error information and related context into an error correction task, and pushing it back to the main model via the one-way channel to guide the main model to generate a corrected version. This process can iterate multiple times until a preset quality threshold is met or the maximum number of attempts is reached.

[0087] 7. Full Memory Retrieval Mode: In addition to the default "Minimum Necessary Fragment" retrieval mode mentioned above, this embodiment fully supports the full memory retrieval mode under explicit user authorization. When the developer issues an explicit instruction through strong authentication (e.g., "Please fix the security vulnerability based on the complete auth.py file of version v2.3"), and the one-time valid authorization credential verification is successful, the memory entity decrypts and extracts the specified original code file from the tamper-proof encrypted storage area (keeping it intact and without irreversible privacy processing), and pushes it to the main model inference unit through a one-way data transmission channel. The main model then obtains the complete original code context and can perform precise vulnerability repair, refactoring, or auditing tasks. This operation is also recorded in the tamper-proof audit log, and the original code data is not persistently stored outside the memory entity after the push is completed; users can audit, retract, or delete it at any time.

[0088] 8. Final Output and Audit: The final code that passes verification is output to the user via API after being validated in the memory entity format. All intermediate iterations, detected error types and fixes, and any complete memory retrieval records are written to a tamper-proof audit log, which is available for traceability by users or compliance auditors.

[0089] Throughout the process, the lightweight retrieval model handles global retrieval, structural understanding, and consistency verification, while the main model focuses on code generation and logical reasoning. The original code data remains securely isolated within its own memory entity, enabling privacy-controlled, auditable, and reversible AI services for developers. Because both retrieval and error correction are based on the user's own unique memory and do not rely on any general-purpose functions, this architecture naturally avoids the "illusion" or "context pollution" problems commonly found in general-purpose RAGs within code scenarios.

[0090] As a further preferred option, the lightweight retrieval model can employ a pre-trained model oriented towards code structure (such as CodeBERT, GraphCodeBERT, or variants thereof) and undergo task-specific fine-tuning (e.g., learning dependency prediction and error pattern recognition on a large number of open-source code libraries). The global code graph can be stored in a graph database or in-memory graph structure to improve retrieval and conflict detection efficiency. The error correction module can be configured to integrate with mainstream compilation toolchains (such as GCC, Clang, Pyflakes, etc.) to achieve real-time compilation / syntax checking.

[0091] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for implementing long-term memory for large-scale model interactions, characterized in that, include: Memory management is achieved through memory entities independent of the large model inference unit. These memory entities possess independent semantic and vector encoding capabilities. Forced isolation is implemented between the memory entities and the large model inference unit, establishing a unidirectional data transmission channel that allows only data to be transmitted from the memory entities to the large model inference unit. Based on user input, the memory entities independently execute memory retrieval, matching, and sorting. The memory retrieval logic, sorting strategy, and retrieval trigger decisions are entirely autonomously controlled by the memory entities, without relying on or adopting any outputs, features, context states, intermediate variables, control signals, active or implicit trigger signals, evaluation feedback, or planning results from the large model inference unit. The memory entity performs irreversible privacy processing on the retrieved memory data that meets the security threshold after semantic clustering filtering, generating the minimum necessary memory fragments that are computationally infeasible to reverse-restore the original privacy data under a preset security strength. All intermediate parameters, keys, and mapping information generated or used during the processing that could lead to data restoration are destroyed immediately after processing. The one-way data transmission channel is the only path for the memory fragments to flow from the memory entity to the large model inference unit. Based on its own strategy and relevance assessment, the memory entity actively pushes the minimum necessary memory fragments through the one-way data transmission channel and injects them into the context of the large model inference unit. The large model inference unit, and any agent, gateway, or middleware acting on its behalf, are prohibited from actively acquiring or accessing the data, status, or interface of the memory entity through read-only calls such as polling, pulling, subscribing, or querying. The large model inference unit generates responses only based on the received minimum necessary memory fragments and the user input, and is completely prohibited in terms of architecture and permissions from reverse accessing, probing, modifying, controlling, or influencing the storage data, index structure, retrieval logic, security configuration, and operating status of the memory entity in any form.

2. The method according to claim 1, characterized in that, Also includes: The memory entity receives the user's original input data, encrypts the original input data, and stores it; The memory entity transmits the original input data to the large model inference unit.

3. The method according to claim 1, characterized in that, Also includes: The memory entity automatically captures the complete original data of the response generated by the large model inference unit, stores it after encryption, and associates it with the user input and the minimum necessary memory fragment index; the complete original data serves as the source for retrieving the original memory data under explicit user authorization.

4. The method according to claim 1, characterized in that, The memory entity implements end-to-end control over the original memory data, ensuring that the original memory data and any reversible intermediate forms remain only within the memory entity and do not participate in or flow out to any form of model training, fine-tuning, distillation, continuous learning, alignment, prompting engineering optimization, or context learning parameter adaptation process.

5. The method according to claim 1, characterized in that, Before performing a retrieval, the memory entity performs legality, completeness, user identity, and device trust authentication on the source of user input and operation requests. The memory entity uses a one-time valid authorization credential uniquely bound to the user identity, operation type, execution time period, and terminal device hardware fingerprint for access control. The credential expires immediately after a single use.

6. The method according to claim 1, characterized in that, The input information sources for memory retrieval are only the user's original input and historical dialogue records stored with the user's authorization. They do not adopt or rely on any semantic representations, feature vectors, summaries, reflection results, planning instructions, or evaluation scores generated or output by the large model inference unit. The one-way data transmission is actively triggered and completed by the memory entity based on its own strategy. The large model inference unit is prohibited from initiating any connection, request, or signal aimed at obtaining memory fragments or probing the state of the memory entity.

7. The method according to claim 1, characterized in that, The irreversible privacy processing includes applying noise that satisfies the differential privacy definition, performing cryptographic hash transformation, or performing lossy dimensionality reduction / transformation processing at least one of these methods to ensure that the processed memory fragment does not contain a complete key, mapping matrix, or inverse transformation parameters that can be used to calculate and restore the original data. The memory entity and the large model inference unit are isolated from each other in the runtime environment and belong to different process spaces, operating system kernel control domains, hardware security partitions, or physical security domains. The execution environment where the large model inference unit is located does not have any permission to access the memory entity's storage medium, memory space, or system resources.

8. The method according to claim 1, characterized in that, The "minimum necessary memory fragment" is dynamically calculated and filtered by the memory entity based on the vector representation of the current user input, the preset privacy budget consumption strategy, and the relevance score between the memory fragment and the vector representation; the memory entity only pushes memory fragments whose relevance score exceeds a preset threshold.

9. The method according to claim 1, characterized in that, The mandatory isolation is implemented through at least one of the following methods: hardware-level physical isolation, operating system kernel-level mandatory access control isolation, CPU-based trusted execution environment isolation, application-layer process sandbox isolation, or business logic isolation based on API permission models; the unidirectional data transmission channel is implemented through at least one of the following methods: hardware unidirectional network gateway, read-only sockets or pipes configured at the kernel layer, port unidirectional access control list configuration of network switches or firewalls, or API design that only grants write permissions to memory entities at the application layer while disabling all read permissions for large model inference units.

10. The method according to claim 1, characterized in that, The core retrieval algorithm, privacy processing strategy, security threshold, and operation strategy parameters of the memory entity are solidified and verified during system compilation, firmware burning, or secure boot. During operation, they cannot be dynamically modified, replaced, bypassed, or tampered with remotely or locally. The memory entity generates tamper-proof audit logs based on cryptographic hash chains for all data access, memory retrieval, privacy processing, fragment push, and configuration change operations. These logs are stored in a protected append-only area, and any attempt to modify or delete historical log entries will be detected and trigger a security alarm.

11. The method according to claim 1, characterized in that, The memory retrieval performed independently by the memory entity comprehensively utilizes at least two retrieval mechanisms from semantic retrieval, vector retrieval, keyword matching, hash indexing, rule matching, and fuzzy matching. The selection of retrieval strategies, weight allocation, and fusion of final results are all decided autonomously by the memory entity.

12. The method according to claim 1, characterized in that, The memory entity completes the semantic understanding and vectorization representation of user input and / or memory data by calling at least one of its own integrated lightweight coding model, the secure coding interface provided by the terminal device operating system, or a third-party coding service that has been certified and authorized by it, ensuring that the coding process is also within its security boundaries.

13. The method according to claim 1, characterized in that, The minimum necessary memory fragments are provided to the large model inference unit in at least one of the following ways: as explicit context injected into the dialogue history in the form of natural language text, spliced ​​into system prompts or user commands, inserted as a structured prefix into the model input sequence, or implicitly influencing model generation as a control vector.

14. The method according to claim 1, characterized in that, The memory entity is also configured to: autonomously determine whether to trigger an active service based on the user's long-term memory data stored internally, local real-time context information, and preset user preference strategies. If an active service is triggered, a minimum necessary intent signal that has undergone irreversible privacy processing is generated and pushed to the large model inference unit through the unidirectional data transmission channel, triggering the large model to generate an active response or service suggestion. The triggering decision, intent signal generation, and data push process of the active service are completely autonomously controlled by the memory entity and do not depend on any input or instructions from the large model inference unit.

15. The method according to claim 1, characterized in that, The memory entity provides a secure state beacon that is encrypted and protected for integrity. The secure state beacon does not contain any original memory data, business logic details, or information that can be used to infer user privacy. It is used for authorized infrastructure managers to perform operational health checks, thereby decoupling the operation and maintenance channel from the business data flow.

16. The method according to claim 1, characterized in that, The memory entity also supports a full memory retrieval mode under explicit user authorization: in response to the user's explicit instruction and a one-time valid authorization credential, the memory entity decrypts and extracts the specified original memory data, and pushes it to the large model inference unit through the one-way data transmission channel; the memory fragments pushed in this mode maintain their original integrity and are not subject to irreversible privacy processing.

17. The method according to claim 1, characterized in that, The memory entity is a local module deployed on the user terminal device within an independent process sandbox, trusted execution environment, or local encrypted storage partition. It transmits data unidirectionally with the large model inference unit through a write-only channel established by the terminal operating system kernel-level mandatory access control policy. The entire system operates without relying on cloud services and supports offline operation. When the memory entity performs retrieval on the terminal device, it adopts an optimized lightweight vector matching and indexing algorithm, and its peak runtime memory usage does not exceed the device's preset security threshold. When the memory entity performs user identity verification, it adopts at least one of the following: a biometric authentication module built into the terminal device, hardware-based device unique trusted authentication, or strong password authentication with local encrypted storage. The one-time valid authorization credential is generated, used, and verified locally on the terminal, and its complete information is not uploaded to any remote server.

18. The method according to claim 17, characterized in that, Multiple memory entities belonging to the same user synchronize memory data through an end-to-end encrypted peer-to-peer communication protocol. During the synchronization process, the original memory data is transmitted directly between user-authorized terminal devices without passing through any centralized server. All access control, conflict resolution, and version management of synchronization operations are completed autonomously by the memory entities.

19. The method according to claim 1, characterized in that, The memory entity provides a standardized security service interface, exposing memory retrieval and fragment generation capabilities to the outside world. It can interface with large model inference units from different manufacturers and with different architectures, enabling seamless migration and reuse of user memory data across different AI applications.

20. The method according to claim 1, characterized in that, The memory entity provides a unique, standardized, read-only secure interface for third-party hardware devices, applications, or services that are explicitly authorized by the user. The interface only allows third parties to access the minimum necessary memory fragments that have undergone irreversible privacy processing, and prohibits any form of writing, modification, deletion, or control operations. Access permissions for third-party devices or applications are uniformly managed by the memory entity, and authorization can be revoked by the user at any time. After revocation, the third party will no longer be able to access any memory data.

21. The method according to claim 1, characterized in that, The memory entity is a user-sovereign managed module deployed in an isolated computing environment provided by a service provider and remotely verified. The configuration of the isolated computing environment, the operating strategy of the memory entity, and the encryption keys for all memory data are exclusively controlled by the user. The memory entity interacts with the user terminal through an encrypted channel to receive user input and synchronize audit information. The memory entity interacts with large model inference units located at the same or different service providers through logical or physical forced isolation and unidirectional data transmission channels.

22. The method according to claim 1, characterized in that, All data processing steps of the memory entity, including static storage, dynamic encryption, one-way transmission, identity authentication, and generation and verification of audit logs of the original memory data, are implemented using national cryptographic algorithms recognized by the State Cryptography Administration; the one-way data transmission channel between the memory entity and the large model inference unit is protected by end-to-end encryption using national cryptographic algorithms; and the identity authentication and authorization credential generation and verification of the memory entity are implemented using national cryptographic digital signature algorithms.

23. The method according to claim 22, characterized in that, The memory entity is an externally independent security domain deployment module, deployed in an independent security area physically isolated from the enterprise's internal network or high-level security compliance domain. It establishes a one-way data transmission link with the large model inference unit located in the external network or public cloud through a hardware one-way network gateway device. The hardware one-way network gateway blocks all data transmission capabilities from the model side to the memory entity side at the physical link layer and communication protocol layer. The forced isolation is further enhanced by at least one of the following methods: implementing a forced access control policy at the server operating system kernel layer, running the memory entity in the trusted execution environment of the CPU, or using encrypted storage media that conforms to national cryptographic standards to store the memory data.

24. A long-term memory implementation system for large-scale model interaction, characterized in that, include: The system comprises independent memory entity modules and a large model inference unit, separated by a mandatory isolation module and a unidirectional transmission module. Each memory entity module possesses independent semantic and vector encoding capabilities and is configured to independently perform memory retrieval, matching, and sorting based on user input. Its retrieval logic, sorting strategy, and decision-making regarding whether to trigger a retrieval are entirely autonomous, independent of and unaffected by any output, features, context state, intermediate variables, control signals, active or implicit trigger signals, evaluation feedback, planning results, or optimization instructions from the large model inference unit. Furthermore, the memory entity module is configured to perform irreversible privacy processing on the retrieval results, generating the minimum necessary memory fragments that, under a preset security level, are computationally infeasible for reversibly restoring the original privacy data. All memory fragments generated or used during the processing that could lead to data restoration are also included. Intermediate parameters, keys, and mapping information are destroyed immediately after processing. The unidirectional transmission module is configured as the sole path for memory fragments to flow from the memory entity module to the large model inference unit, and is configured to only support data push operations initiated by the memory entity module. The large model inference unit, and any proxy, gateway, or middleware acting on its behalf, are prohibited from actively acquiring or accessing the data, status, or interface of the memory entity module through read-only calls such as polling, pulling, subscribing, or querying. The large model inference unit is configured to generate responses solely based on the received injected memory fragments and the user input, and is completely prohibited from reverse accessing, probing, modifying, controlling, or influencing the storage, indexing, retrieval logic, security configuration, and operating status of the memory entity module in any form, both in terms of architecture and permissions.

25. The system according to claim 24, characterized in that, The memory entity module is configured to implement end-to-end control over the original memory data, ensuring that the original memory data and any reversible intermediate forms remain only within it and do not participate in or flow out to any form of large model training, fine-tuning, distillation, continuous learning, alignment, prompting engineering optimization, or context learning parameter adaptation process.

26. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 23.

27. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 23.

28. A lightweight user terminal device, characterized in that, It includes a processor, a memory, and a trusted security unit, wherein the memory stores a computer program, and the processor executes the computer program to implement the method of any one of claims 17 to 20.

29. A compliant security gateway device, characterized in that, The device includes a hardware unidirectional network gateway unit, a national cryptographic encryption unit, a memory entity processing unit, and an audit log unit. The device has a built-in computer program that, when executed, implements the method described in claim 23.