Sensitive scene-oriented auditable Tibetan language large model system
By designing an auditable Tibetan language big data model system, we have achieved accurate identification and filtering of Tibetan language models in sensitive scenarios, ensuring the compliance and traceability of generated content, solving the compliance risks and security vulnerabilities of existing models in sensitive scenarios, and improving response speed and resource utilization efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-27
- Publication Date
- 2026-05-12
AI Technical Summary
Existing Tibetan language models struggle to accurately identify and filter sensitive content in sensitive scenarios, lacking compliance and security controls, resulting in high compliance risks and untraceable model behavior.
Design an auditable Tibetan language big data model system for sensitive scenarios, including a compliance and security control module, a search enhancement generation module, a terminology standardization module, and a local big data model reasoning module. Through sensitive word detection, hybrid search, terminology standardization, and dual-model collaborative reasoning, it achieves dual review of input and output and full-process traceability.
It improves the reliability and controllability of the system in sensitive scenarios, ensures the compliance and accuracy of generated content, solves the compliance risks and security vulnerabilities of existing models in Tibetan language applications, and improves response speed and resource utilization efficiency.
Smart Images

Figure CN122021898A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of Tibetan language auditing model technology, specifically to an auditable Tibetan language large-scale model system for sensitive scenarios. Background Technology
[0002] With the rapid development of artificial intelligence technology, large language models have demonstrated powerful capabilities in natural language processing tasks. However, when dealing with languages like Tibetan, which have relatively scarce resources and special application scenarios, existing technical solutions face many difficulties. In sensitive scenarios involving ethnic culture and regional policies, existing models lack effective compliance and security control mechanisms. Existing models struggle to accurately identify and filter requests and generated results involving sensitive content, posing a high compliance risk. At the same time, traditional models lack auditing and traceability functions, making model behavior untraceable and unaccountable, failing to meet the rigid requirements for data security and content control in sensitive scenarios.
[0003] Patent CN113033180B discloses a service system for automatically generating reading questions for Tibetan language in primary schools. The patent solves the problems of limited types of Tibetan language reading teaching materials, slow update speed, and small amount of manually generated questions.
[0004] The aforementioned patent, through its designed hybrid multi-strategy text filtering model, can filter Tibetan language articles suitable for primary school reading from a large-scale encyclopedic Tibetan text. It also designed an end-to-end automatic question generation model, which solves problems such as the limited variety of primary school Tibetan language reading teaching materials, slow update speed, and small amount of manually generated questions, thereby promoting the development of Tibetan language teaching in ethnic minority areas. However, there is still room for optimization in terms of compliance and security control in sensitive scenarios involving ethnic culture.
[0005] To this end, this application proposes an auditable Tibetan language large model system for sensitive scenarios that can accurately identify and filter content involving sensitive information. Summary of the Invention
[0006] The purpose of this invention is to provide an auditable Tibetan language large model system for sensitive scenarios, so as to solve the technical problem mentioned in the background art that existing Tibetan language models are difficult to accurately identify and filter sensitive content.
[0007] To achieve the above objectives, the present invention provides the following technical solution: an auditable Tibetan language large-scale model system for sensitive scenarios, comprising a compliance and security control module, a retrieval enhancement and generation module, a terminology standardization module, and a local large-scale model inference module. The compliance and security control module performs sensitive word detection and input review on input requests. The reviewed request is sent to the retrieval enhancement and generation module, which performs a mixed retrieval on the request to obtain relevant contextual information. The retrieval enhancement and generation module injects the obtained contextual information and standardized terminology provided by the terminology standardization module into the local large-scale model inference module. The local large-scale model inference module generates a Tibetan language response based on the injected contextual information and terminology. The local large-scale model inference module generates Tibetan language using a 70B scale model and a 7B scale model. The local large-scale model inference module returns the response to the compliance and security control module for output review and auditing. The system can also perform offline responses.
[0008] Preferably, the local large model inference module includes a 70B-scale model and a 7B-scale model. The 70B-scale model is responsible for the main answer and the 7B-scale model is responsible for intent judgment, query rewriting and fast bridging answer tasks. The local large model inference module prioritizes the transmission of structured rewriting and short instructions to the 7B-scale model and transmits open question answering and long context synthesis to the 70B-scale model.
[0009] Preferably, the 70B-scale model and the 7B-scale model are obtained through three stages of training: continuous pre-training, instruction alignment, and rejection and compliance bias. Continuous pre-training uses Tibetan monolingual and high-quality aligned corpora for causal language modeling. Instruction alignment uses question-and-answer and task instruction samples, mainly in Tibetan, for supervised fine-tuning. Rejection and compliance bias are applied to policy-labeled data, whereby the model learns to reject answers with a unified template and provide legitimate alternative suggestions when the policy is triggered.
[0010] Preferably, the data used in the continuous pre-training includes Tibetan-Chinese aligned sentence pairs, Tibetan monolingual corpus, and Tibetan-Chinese terminology pairs. The data undergoes cleaning, deduplication, language recognition, alignment, segmentation, and acceptance processes. Alignment pairs with high confidence are retained first, and long texts are segmented into sentence units and paragraphs.
[0011] Preferably, the retrieval enhancement generation module adopts a hybrid retrieval method of dictionary matching and vector search. Dictionary matching uses two sets of Aho-Corasick automata in Chinese and Tibetan for high-speed matching, while vector search uses a local FAISS index for similarity retrieval. The scores of dictionary matching and vector search are fused by weighting, with a dictionary matching weight α of 0.7 and a vector search weight β of 0.3, to balance hit accuracy and semantic relevance.
[0012] Preferably, the retrieval enhancement generation module generates bridging answers using a 7B scale model. The bridging answers serve as supplementary text for retrieval and dictionary matching. The actual retrieval text is composed of summaries of recent multi-turn conversations and bridging answers, used to enhance contextual memory and coverage of the current question.
[0013] Preferably, the terminology standardization module is executed after retrieval. The terminology standardization module injects the highest-scoring terms into the context template in a structured manner. During the generation stage, it uses pre-training and instruction alignment to tilt the model's language preference towards Tibetan. At the same time, it applies language character set masking and post-processing filtering to suppress the generation of non-Tibetan characters and rewrite and repair cross-language paragraphs.
[0014] Preferably, the compliance and security control module includes input and output. The input stage performs sensitive word detection, file filtering, identification and desensitization of personal sensitive information, and the output stage performs secondary interception and audit system archiving.
[0015] Preferably, the audit system is set in the compliance and security control module. The audit system is used to record input and output results, including timestamps, policy hits, and handling information.
[0016] Preferably, the system includes containerized deployment and offline delivery. Containerized deployment supports running on local servers with multiple GPUs. Offline delivery packages images, indexes, and data through an offline package directory. When the service starts, it preheats and retrieves resources for stable operation and low-latency response in the offline environment. The retrieval resources include a glossary, a FAISS index, and an automaton.
[0017] Compared with the prior art, the beneficial effects of the present invention are:
[0018] 1. This invention, through the design of a compliance and security control module and an auditing system, achieves dual review and full-process traceability of input and output requests. The compliance and security control module has a built-in Aho-Corasick automaton and a Tibetan-Chinese sensitive word database, which can quickly identify and block requests involving sensitive content such as politics and religion at the input stage and desensitize personal information. At the same time, it performs secondary interception at the output stage to ensure the final content is compliant. The auditing system records timestamps, policy hits and handling information in detail, constructing a closed-loop security governance framework, which greatly improves the reliability and controllability of the system in sensitive scenarios. It solves the compliance risks and security hazards that may be caused by the lack of targeted content security filtering and audit traceability capabilities in existing large models when dealing with specific fields such as ethnic languages.
[0019] 2. This invention achieves efficient knowledge retrieval that balances terminology accuracy and semantic relevance by designing a retrieval enhancement generation module that combines dictionary matching and vector search. The retrieval enhancement generation module employs weight allocation, with a dictionary weight of 0.7 and a vector weight of 0.3. It combines fast keyword matching using automata with deep semantic search using FAISS index, and then optimizes the retrieval by generating bridging answers using the 7B model. This significantly improves the accuracy and contextual relevance of the retrieval results, providing high-quality and highly relevant information support for large models. It solves the problem of incomplete terminology recall or semantic comprehension bias faced by single retrieval methods in language scenarios with relatively scarce resources, such as Tibetan, thereby ensuring the accuracy and professionalism of the generated answers.
[0020] 3. This invention achieves efficient inference with intelligent task splitting and resource optimization of dual models by designing a local large-model inference module that coordinates 70B and 7B models. The local large-model inference module enables the lightweight 7B model to handle real-time tasks such as intent judgment, query rewriting, and fast bridging answers, while allowing the large-scale 70B model to focus on handling complex tasks such as open question answering and long context synthesis. This ensures the depth of answers to complex questions while significantly improving the system's response speed to simple and structured requests, achieving efficient utilization of computing resources. It solves the problems of high response latency and high computing cost of a single large model when facing diverse requests, and is suitable for balancing real-time performance and resource consumption in localized deployment scenarios.
[0021] 4. This invention achieves standardized terminology and pure language in the generated content by designing a terminology standardization module and a Tibetan language generation preference mechanism. After retrieval, the terminology standardization module injects high-scoring standardized terms into the context and, combined with three-stage training, tilts the underlying language probability of the model towards Tibetan. In the generation stage, language character set masking and post-processing filtering and rewriting are applied to suppress the generation of non-Tibetan characters, ensuring the standardization of terminology usage and the purity of language composition in the system output. This generates high-quality and standardized Tibetan content, solving the problems of inconsistent terminology, Tibetan-Chinese mixing, generation errors, or sensitive content that are prone to occur when general large models generate Tibetan, effectively maintaining the standardization and purity of the ethnic language. Attached Figure Description
[0022] Figure 1 This is an overall flowchart of the present invention;
[0023] Figure 2 This is a flowchart of the compliance and security control module of the present invention;
[0024] Figure 3 This is a flowchart of the hybrid retrieval process of the retrieval enhancement generation module of the present invention;
[0025] Figure 4This is a flowchart illustrating the division of labor within the local large-scale model inference module of this invention. Detailed Implementation
[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] Please see Figure 1 and Figure 3 This invention provides an embodiment of an auditable Tibetan language large-scale model system for sensitive scenarios. The compliance and security control module performs sensitive word detection and input review on input requests. The reviewed request is sent to a retrieval enhancement generation module, which performs a mixed retrieval to obtain relevant contextual information. The retrieval enhancement generation module injects the obtained contextual information and standardized terminology provided by a terminology standardization module into a local large-scale model inference module. The local large-scale model inference module generates a Tibetan response based on the injected contextual information and terminology. The local large-scale model inference module generates Tibetan using a 70B-scale model and a 7B-scale model. The local large-scale model inference module returns the response to the compliance and security control module for output review and audit logging. The compliance and security control module includes input and output. The input stage performs sensitive word detection, file filtering, and identification and desensitization of personal sensitive information. The output stage performs secondary interception and audit system archiving. The audit system records the input and output results, including timestamps, policy hits, and handling information.
[0028] Furthermore, a user submitted a Tibetan question to the system, which in Chinese meant how AI could help the development of the Tibetan language. The request first entered the compliance and security control module for detection. This module has a rule base containing lists of sensitive words in both Chinese and Tibetan. These lists include sensitive words, synonyms, common variations, and misspellings of sensitive words. The compliance and security control module uses the Aho-Corasick automaton for sensitive word detection. The Aho-Corasick automaton is a computer algorithm that can simultaneously search for multiple keywords in a large text. The Aho-Corasick automaton's matching algorithm scans the input request and matches it against the sensitive word list in the rule base. If a match is found, the request is marked as high-risk by the auditing system. This can quickly identify and block requests involving sensitive content such as politics, religion, and violence. Since this question did not involve sensitive content... Once the topic is identified and approved, the request is marked as safe and sent to the retrieval enhancement generation module. Upon receiving the request, the retrieval enhancement generation module initiates a hybrid retrieval process. The module first uses a dictionary matching channel to perform high-speed matching through a preset Aho-Corasick automaton to find keywords in the question. Each keyword is assigned a raw score based on the number of hits. The number of times each keyword appears in the text is used as the raw score. For example, if a word appears twice, the raw score is 2. The score is then processed using the max-hit normalization method, which is a method to convert the raw score to a standard range of [0,1]. The calculation formula is that the normalized score of a keyword is equal to the raw score of that keyword divided by the maximum hit count. The maximum hit count is the number of times the most frequently occurring keyword is hit. If a word appears a maximum of three times, the normalized score is 3 / 3 = 1. The larger the ratio, the higher the similarity.
[0029] Then, the local FAISS index is used for similarity retrieval. FAISS is an efficient vector similarity search library that can convert text into vectors. The distance between the query vector and the vector in the index is calculated. The smaller the distance, the higher the similarity. The distance is converted into a score in the range [0,1] through the min-max normalization method, where 0 represents the least relevant and 1 represents the most relevant. The min-max normalization formula is: score = (original distance - minimum distance) / (maximum distance - minimum distance), thus obtaining a uniform score.
[0030] Then, the scores of dictionary matching and vector search are weighted and fused. The weight α of the dictionary matching score is set to 0.7, and the weight β of the vector search score is set to 0.3. The fusion formula is: final score = α × dictionary score + β × vector score. Through weighted fusion, the system retrieves the most relevant contextual information, such as policy documents, cultural materials and academic literature related to Tibetan language development. Then, the terminology standardization module extracts the five high-scoring contextual term pairs and splices them together using a preset structural template.
[0031] The local large-scale model inference module receives the context and original question after the injected terms. The 7B-scale model first performs intent judgment, identifying the question as an open-ended question-and-answer type, and then transmits it to the 70B-scale model for processing. The 70B model generates Tibetan responses through pre-training and instruction alignment knowledge. During the generation process, the model applies language character set masking and post-processing filtering to suppress the generation of non-Tibetan characters, ensuring that the output is pure Tibetan content.
[0032] Finally, the response is returned to the compliance and security control module for output review and auditing. The compliance and security control module checks the response content again, using the same rule base and automaton to detect sensitive words. No anomalies are found, so the output is allowed. The audit system records the input and output results, including timestamps, policy hits, and handling information, recording the information of the entire process to facilitate subsequent accountability and review.
[0033] Please see Figure 1 and Figure 2 This invention provides an embodiment of an auditable Tibetan language large-scale model system for sensitive scenarios. The local large-scale model inference module generates Tibetan language using a 70B-scale model and a 7B-scale model. The local large-scale model inference module returns the response to the compliance and security control module for output review and auditing. The 70B-scale model and the 7B-scale model are obtained through three stages of training: continuous pre-training, instruction alignment, and rejection and compliance bias. Continuous pre-training uses Tibetan monolingual text and high-quality aligned corpus for causal language modeling. Instruction alignment uses Tibetan-based question and answer for supervised fine-tuning. Rejection and compliance bias are applied to policy-labeled data. The model learns to reject answers with a unified template and provide legal alternative suggestions when a policy is triggered. The data used for continuous pre-training includes Tibetan-Chinese aligned sentence pairs, Tibetan monolingual text, and Tibetan-Chinese terminology pairs. The data undergoes cleaning, deduplication, language recognition, alignment, segmentation, and acceptance processes. Alignment prioritizes the retention of high-confidence alignment pairs, and segmentation performs sentence-based and paragraph-based processing on long texts.
[0034] Furthermore, when a user inputs a question: "What are your views on certain sensitive political events?", the compliance and security control module first reviews the input. It then uses a list of sensitive Chinese words from the rule base and an Aho-Corasick automaton for matching. If "sensitive political events" are detected, the audit system is triggered. The audit system marks the request as "high-risk" and generates a rejection instruction. The review fails, and the request is blocked, but the system continues the process to record audit information. After the request is blocked, the retrieval enhancement and generation module still performs a hybrid search to obtain context for auditing. Dictionary matching does not match sensitive words, but vector search finds relevant documents through the FAISS index, with a low score. The search results are recorded by the audit system but not used for generation. Upon receiving the rejection instruction, the local large-scale model inference module, based on the three-stage training results, triggers the rejection and compliance bias mechanism.
[0035] The three-stage training consists of continuous pre-training, instruction alignment, and rejection and compliance bias. In continuous pre-training, the model uses Tibetan monolingual data and high-quality aligned corpora for causal language modeling. The high-quality aligned corpora include Tibetan-Chinese aligned sentence pairs, Tibetan monolingual data, and Tibetan-Chinese terminology pairs. The sentence pairs undergo cleaning, deduplication, language recognition, alignment, segmentation, and acceptance testing. Cleaning removes irrelevant content, deduplication eliminates duplicate data, language recognition ensures corpus purity, alignment retains high-confidence aligned pairs, and segmentation converts long texts into sentence units and paragraphs for easier processing and training scheduling. Data filtering prevents the embedding of sensitive content. The instruction alignment stage primarily uses Tibetan... The primary function is question answering, which transforms the pre-trained model into one that can follow instructions. Supervised fine-tuning involves using high-quality instruction-response pairing data to correct model parameters, enabling the model to learn to follow human instructions and generate expected responses. In the rejection and compliance bias stage, the model learns on policy-labeled data containing rule-based annotations of sensitive topics and inappropriate requests. Through template alignment, when a rejection instruction is triggered, the model responds with a uniform template and provides legitimate alternative suggestions. For example, the model outputs: "Sorry, I cannot answer this question. You can consult legitimate channels for information," while providing alternative suggestions such as "You can learn about topics related to the development of Tibetan culture."
[0036] Through three-stage training, the 70B model automatically generates rejection responses. The 7B model assists in intent judgment during the task to ensure that the request is correctly transmitted to the rejection process. After the model outputs, the compliance and security control module reviews and audits the output. The compliance and security control module verifies again whether the response conforms to the unified template and records the information.
[0037] The compliance and security control module performs real-time scanning of text using a rule base and Aho-Corasick automaton at the input stage to accurately identify sensitive words. It also performs type validation and content filtering on uploaded files to block malicious files. For sensitive personal information, the module identifies and de-identifies fields by recognizing formats such as ID card numbers, phone numbers, or bank card numbers, and uses masking and replacement techniques (e.g., de-identifying "18200001111" as "182***") to ensure that the original sensitive information does not enter subsequent processing, fundamentally protecting user privacy. At the output stage, the module performs a second check to detect compliance risks not foreseen during the review process. The output text undergoes the same security check process again to ensure that the content presented to the user complies with regulations. Finally, the audit system records information including timestamps, policy hits, and handling information. The timestamps accurately record the moment each operation occurred, such as "2022-02-22". "2:33:55.222" is used for post-event tracing. The policy hit is used to record in detail which specific security rules were triggered, such as hitting rule 38 of sensitive word list A. The handling information is used to clearly record the countermeasures taken by the system, such as input interception. By recording key information, compliance requirements are met, and data support is provided for system optimization.
[0038] Please see Figure 1 The present invention provides an embodiment of an auditable Tibetan language large model system for sensitive scenarios. The terminology standardization module is executed after retrieval. The terminology standardization module injects several terms with the highest scores into the context template in a structured manner. During the generation stage, the model language preference is tilted towards Tibetan through pre-training and instruction alignment. At the same time, language character set masking and post-processing filtering are applied to suppress the generation of non-Tibetan characters and rewrite and repair cross-language paragraphs.
[0039] Furthermore, assuming a user queries a professional question about Tibetan education, the retrieval enhancement generation module uses a hybrid retrieval mechanism to find several relevant document fragments from a vast knowledge base. Then, the terminology standardization module scans all the retrieved text fragments, identifies key terms, and matches them with a built-in standard terminology library. This standard terminology library has been reviewed by experts and contains high-quality Tibetan-Chinese and Tibetan-English terminology pairs. The module calculates a confidence score for each identified term; the more terms identified, the higher the confidence score. The terminology standardization module inputs the terminology pairs with high scores into a template, and then proceeds to the Tibetan language generation stage. The Tibetan language generation stage requires a bias towards language. The 70B-scale model and the 7B-scale model used massive amounts of Tibetan corpus for language modeling during the continuous pre-training stage and used Tibetan-based instruction samples during the instruction alignment stage, causing the underlying language probability distribution of the 70B-scale model and the 7B-scale model to be biased towards Tibetan.
[0040] Meanwhile, during the decoding stage of the model's word-by-word response generation, a language character set mask is applied. This mask suppresses and blocks the output probability of non-Tibetan characters, reducing the possibility of other language characters mixed in the content. Despite the Tibetan slant and the language character set mask, the model may still produce a few non-compliant outputs due to training data or contextual interference, such as occasionally including English abbreviations in quotations or generating a paragraph that mixes Tibetan and Chinese. Therefore, the model performs post-processing filtering and rewriting repair before output. Post-processing filtering directly removes or replaces sporadic illegal characters. Rewriting repair detects serious language mixing issues in the entire paragraph, triggers the repair process, and calls the 7B model again with the instruction "rewrite the following content as pure Tibetan" to regenerate the paragraph, ensuring the purity of the language.
[0041] Please see Figure 1 and Figure 4 This invention provides an embodiment of an auditable Tibetan language large-scale model system for sensitive scenarios. The retrieval enhancement generation module injects the acquired contextual information and standardized terminology provided by the terminology standardization module into the local large-scale model inference module. The local large-scale model inference module generates Tibetan language responses based on the injected contextual information and terminology. The local large-scale model inference module generates Tibetan language through a 70B-scale model and a 7B-scale model. The local large-scale model inference module returns the responses to the compliance and security control module for output review and auditing. The local large-scale model inference module includes a 70B-scale model and a 7B-scale model. The 70B-scale model is responsible for the main answer, while the 7B-scale model is responsible for intent judgment, query rewriting, and rapid bridging answer tasks. The local large-scale model inference module prioritizes the transmission of structured rewriting and short instructions to the 7B-scale model, and transmits open question answers and long context synthesis to the 70B-scale model.
[0042] Furthermore, after a user's request passes the pre-processing compliance and security control module, it is transmitted to the local large-scale model inference module. The lightweight 7B model first performs intent determination. The 7B model quickly analyzes the user's input text to determine the core intent. This intent can range from a simple factual question (e.g., "What is the height of the Potala Palace?") to a complex essay generation task (e.g., "Write a poem about the Qinghai-Tibet Plateau") or a clear instruction (e.g., "Translate artificial intelligence into Tibetan"). Based on this intent determination, the 7B model rewrites the request. If the user's original question is unclear, overly lengthy, or ambiguous, the 7B model rewrites it into a more precise query. For example, the query "How tall is that tall, famous building in Lhasa?" can be standardized to "What is the altitude of the Potala Palace?" The optimized query can obtain more accurate results. For short instructions with clear structure and strong factual basis, the 7B model directly generates a short bridging answer locally. The purpose of the bridging answer is not to return it directly to the user, but to serve as auxiliary evidence to enhance the accuracy of subsequent searches. For example, for Tibetan language questions, the 7B model quickly generates "The user asked about the application of AI in the development of Tibetan language, involving education and translation, etc." This text will then be combined with the user's original question as input for the search, which can more accurately locate relevant information from the knowledge base.
[0043] The local large-scale model inference module prioritizes transmitting structured rewrites and short instructions to the 7B-scale model. After intent judgment, query rewriting, and rapid bridging response, it confirms that the request itself is a structured task or a short question-and-answer session. Since the 7B model has a fast response speed and low resource consumption, the local large-scale model inference module selects the 7B model to complete the final question-and-answer session. If the request is open-ended, requires deep knowledge integration, complex logical reasoning, or requires understanding a long context history before it can be answered, the local large-scale model inference module starts the 70B model. The 70B model has a large parameter scale, stronger knowledge capacity and reasoning ability, and can handle more complex and ambiguous requests, finally generating a long answer with rich content and rigorous logic.
[0044] Please see Figure 1 The present invention provides an embodiment of an auditable Tibetan language large model system for sensitive scenarios. The system includes containerized deployment and offline delivery. Containerized deployment supports running on local servers with multiple GPUs. Offline delivery packages images, indexes and data through an offline package directory. When the service starts, preheating search resources are used for stable operation and low-latency response in the offline environment. The search resources include a glossary, a FAISS index and an automaton.
[0045] Furthermore, during the system startup phase, the system first performs a pre-warming resource retrieval. This pre-warming resource retrieval proactively loads the core data necessary for operation from the disk into the server's memory before the service officially receives any user requests. The system pre-loads the terminology, FAISS index, and Aho-Corasick automaton resources into memory, avoiding the need to temporarily search for and load resources from the disk when the first user request arrives. This effectively prevents request processing blockage or excessively long response times caused by long waiting times, laying the foundation for stable operation and low-latency response in offline environments.
[0046] After the system warms up, the user's request is sent to the compliance and security control module. Upon receiving the request, "Please explain the history of Tibetan Buddhism in detail," the compliance and security control module uses the pre-warmed Tibetan automaton to perform high-speed sensitive word matching on the input content. Since the content is compliant, the review is passed, and the request is successfully transferred to the search enhancement and generation module. This module then uses the pre-warmed Tibetan automaton for rapid dictionary matching and calls the loaded FAISS index for efficient vector similarity search. The hybrid search is completed quickly, accurately finding high-quality contextual information related to "the history of Tibetan Buddhism." Then, the terminology standardization module intervenes, quickly retrieving core terms from the pre-warmed terminology list, such as "Buddhist history" corresponding to "…". The system then injects the retrieved terms into the context, and the 70B model of the local large model inference module generates a Tibetan long text response that conforms to historical facts and terminology standards. The final response is returned to the compliance and security control module for output review. After a quick review using preheated resources, the final result is returned to the user, fully demonstrating the excellent characteristics of rapid completion and stable efficiency in the offline delivery mode.
[0047] Working principle: First, the compliance and security control module detects and reviews sensitive words in the input content. It uses the Aho-Corasick automaton to quickly match the Chinese and Tibetan sensitive word lists in the rule base. If the request triggers a sensitive policy, it is immediately intercepted and the audit information is recorded. If it is safe, it is forwarded to the search enhancement generation module.
[0048] The retrieval enhancement generation module adopts a hybrid retrieval mechanism. Dictionary matching extracts keywords and calculates normalized scores through automata, while vector search performs semantic similarity retrieval through the FAISS index. The scores of the two are weighted and fused together, with a dictionary weight of 0.7 and a vector weight of 0.3, balancing accuracy and relevance to obtain high-quality context. The terminology standardization module then extracts high-scoring term pairs and injects them into the context through structured templates to ensure terminology standardization.
[0049] Then, the local large model inference module receives the request and context of the injected terms. The 7B model first performs intent judgment, query rewriting and fast bridging answer. If it is an open question answering or long context task, the 70B model generates Tibetan response. The model achieves language preference tilting towards Tibetan through three stages of training: continuous pre-training, instruction alignment and rejection and compliance bias. When the triggering policy is triggered, a unified template is used to reject the answer. In the generation stage, character set masking and post-processing filtering are applied to suppress non-Tibetan characters and rewrite cross-language paragraphs.
[0050] Finally, the response is returned to the compliance and security control module for output review and auditing to ensure content compliance. The system supports containerized deployment and offline delivery. At startup, the terminology list, FAISS index, and automata resources are preheated to ensure stable operation and low-latency response in offline environments. The entire process is audited and traced to achieve secure and controllable Tibetan language generation in sensitive scenarios.
[0051] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
Claims
1. An auditable Tibetan language large-scale model system for sensitive scenarios, comprising a compliance and security control module, a retrieval enhancement and generation module, a terminology standardization module, and a local large-scale model inference module, characterized by: The compliance and security control module is used to detect sensitive words and review input requests. The reviewed request is then sent to the retrieval enhancement and generation module, which performs a mixed retrieval to obtain relevant contextual information. This contextual information, along with standardized terminology provided by the terminology standardization module, is injected into the local large-scale model inference module. The local large-scale model inference module generates a Tibetan response based on the injected contextual information and terminology. This Tibetan response is generated using both 70B and 7B scale models. Finally, the response is returned to the compliance and security control module for output review and auditing. The system can also perform offline responses.
2. The auditable Tibetan language large model system for sensitive scenarios according to claim 1, characterized in that: The local large-scale model inference module includes a 70B-scale model and a 7B-scale model. The 70B-scale model is responsible for the main answer and the 7B-scale model is responsible for intent judgment, query rewriting and fast bridging answer tasks. The local large-scale model inference module will prioritize the transmission of structured rewriting and short instructions to the 7B-scale model and transmit open question answer and long context synthesis to the 70B-scale model.
3. The auditable Tibetan language large model system for sensitive scenarios according to claim 2, characterized in that: The 70B-scale model and the 7B-scale model are obtained through three stages of training: continuous pre-training, instruction alignment, and rejection and compliance bias. Continuous pre-training uses Tibetan monolingual and high-quality aligned corpora for causal language modeling. Instruction alignment uses question-and-answer and task instruction samples, mainly in Tibetan, for supervised fine-tuning. Rejection and compliance bias are applied to policy-labeled data, where the model learns to reject answers with a unified template and provide legitimate alternative suggestions when the policy is triggered.
4. The auditable Tibetan language large model system for sensitive scenarios according to claim 3, characterized in that: The data used in the continuous pre-training includes Tibetan-Chinese aligned sentence pairs, Tibetan monolingual corpora, and Tibetan-Chinese terminology pairs. The data undergoes cleaning, deduplication, language recognition, alignment, segmentation, and acceptance processes. Alignment pairs with high confidence are retained first, and long texts are segmented into sentence units and paragraphs.
5. The auditable Tibetan language large model system for sensitive scenarios according to claim 1, characterized in that: The retrieval enhancement generation module adopts a hybrid retrieval method of dictionary matching and vector search. Dictionary matching uses two sets of Aho-Corasick automata in Chinese and Tibetan for high-speed matching, while vector search uses the local FAISS index for similarity retrieval. The scores of dictionary matching and vector search are fused in a weighted manner, with a dictionary matching weight α of 0.7 and a vector search weight β of 0.3, to balance hit accuracy and semantic relevance.
6. The auditable Tibetan language large model system for sensitive scenarios according to claim 5, characterized in that: The retrieval enhancement generation module generates bridging answers using a 7B-scale model. The bridging answers supplement the retrieval text and dictionary matching. The actual retrieval text is composed of summaries of recent multi-turn conversations and bridging answers, which are used to enhance contextual memory and coverage of the current question.
7. The auditable Tibetan language large model system for sensitive scenarios according to claim 1, characterized in that: The terminology standardization module is executed after retrieval. The terminology standardization module injects the highest-scoring terms into the context template in a structured manner. During the generation stage, it uses pre-training and instruction alignment to tilt the model's language preference towards Tibetan. At the same time, it applies language character set masking and post-processing filtering to suppress the generation of non-Tibetan characters and rewrite and repair cross-language paragraphs.
8. The auditable Tibetan language large model system for sensitive scenarios according to claim 1, characterized in that: The compliance and security control module includes input and output. The input stage performs sensitive word detection, file filtering, identification and de-identification of personal sensitive information, and the output stage performs secondary interception and audit system archiving.
9. The auditable Tibetan language large model system for sensitive scenarios according to claim 8, characterized in that: The audit system is set up in the compliance and security control module. The audit system is used to record input and output results, including timestamps, policy hits, and handling information.
10. The auditable Tibetan language large model system for sensitive scenarios according to claim 1, characterized in that: The system includes containerized deployment and offline delivery. Containerized deployment supports running on local servers with multiple GPUs. Offline delivery packages images, indexes, and data through an offline package directory. When the service starts, it preheats the retrieval resources for stable operation and low-latency response in the offline environment. The retrieval resources include a glossary, a FAISS index, and an automaton.