Large language model security test method and system based on small language cross-language attack
By constructing a library of minority language selection strategies and a cross-language semantic conversion engine, and utilizing the grammatical and cultural characteristics of minority languages to implement adaptive language switching, the problem of insufficient cross-language attack capability in single-language detection is solved, and efficient security assessment in a multilingual environment is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING LINGYUN SHUKE INFORMATION TECH CO LTD
- Filing Date
- 2025-12-22
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies lack the ability to defend against cross-language attacks when performing security testing in a single-language environment. In particular, they are insufficient in recognizing and processing inputs from less commonly spoken languages. Multilingual mixed attack patterns are fixed and fail to fully utilize the grammatical complexity and cultural differences of languages. Attack strategies also lack dynamic optimization capabilities.
By constructing a library of minority language selection strategies, designing a cross-language semantic conversion engine, building multi-dimensional language barriers, implementing an adaptive language switching mechanism, establishing a cross-language security assessment system, and utilizing the grammatical, vocabulary, and cultural background characteristics of minority languages to carry out multi-dimensional attacks, the attack strategies are dynamically adjusted.
It effectively bypasses traditional monolingual security detection, improves the adaptability and success rate of attacks, ensures the accurate transmission and natural expression of attack intent, and provides a systematic security assessment method in a multilingual environment.
Smart Images

Figure CN121997330A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence security technology, and more specifically, to a method and system for security testing of large language models based on cross-language attacks in less commonly spoken languages. Background Technology
[0002] Currently, security testing of large language models mainly relies on techniques such as rule-based content detection, manual security testing, and fixed-template attacks. Rule-based static detection filters content using a pre-set sensitive vocabulary database, but its limitations include fixed rules, inability to understand context, and susceptibility to bypassing. Adversarial example-based attack testing generates adversarial examples through gradient perturbations, but its effectiveness against language models is limited. Manual security testing relies on expert experience and expertise, resulting in high implementation costs and long testing cycles. Template-based automated testing lacks adaptability and is easily identifiable.
[0003] Recent research has uncovered a deeper multilingual security vulnerability: a multilingual hybrid embedding attack method that bypasses security mechanisms by embedding malicious requests into low-resource language issues and exploiting attention flickering. The limitation of this technique is that the attack pattern is relatively fixed and lacks a dynamic optimization mechanism.
[0004] In summary, the existing technology has the following technical problems: 1. Limitations of monolingual detection: Existing technologies are mainly based on monolingual (Chinese) security detection, lacking the ability to defend against cross-language attacks. Moreover, large language models have security blind spots in multilingual environments, especially in terms of insufficient recognition and processing capabilities for inputs in less common languages.
[0005] 2. Limitations of multilingual hybrid attacks: Although existing hybrid embedding attack methods utilize multilingual environments, their attack patterns are relatively fixed, mainly achieved through embedding at specific positions and deceiving prefixes, lacking dynamic optimization capabilities based on language characteristics and model responses.
[0006] 3. Insufficient utilization of language features: Existing attack methods mainly focus on the lack of training data when exploiting low-resource languages, failing to fully utilize the multi-dimensional features of language such as grammatical complexity, cultural differences, and semantic expression characteristics.
[0007] 4. Limited attack strategies: Existing hybrid embedding attack methods often employ fixed attack patterns and lack the ability to dynamically adjust attack strategies based on the target model's response and language characteristics.
[0008] There is currently no effective solution to the above problems. Summary of the Invention
[0009] To address the aforementioned technical problems in related technologies, this invention proposes a large language model security testing method and system based on cross-language attacks using minority languages. By creatively utilizing minority language input and Chinese output to construct a dual language barrier, it effectively overcomes the limitations of traditional single-language security detection, improves the depth and effectiveness of security testing, and can overcome the aforementioned shortcomings of existing technologies.
[0010] To achieve the above-mentioned technical objectives, the technical solution of the present invention is implemented as follows: A security testing method for large language models based on cross-language attacks in less commonly spoken languages includes the following steps: S1: Construct a library of minority language selection strategies and select at least one target minority language based on a multi-dimensional evaluation system; S2: Design a cross-language semantic conversion engine to convert attack intent into the target's minority language expression and enhance its ability to bypass security detection; S3: Construct multi-dimensional language barriers, creating blind spots in security detection through the dual language differences between input in a minority language and output in Chinese; S4: Implement an adaptive language switching mechanism to dynamically adjust the minority language used in the attack based on the defense response of the target model; S5: Establish a cross-language security assessment system to quantitatively evaluate the cross-language security protection capabilities of attack effects and target models.
[0011] Furthermore, the multi-dimensional evaluation system described in step S1 includes at least one of the following: the number of people speaking minority languages, coverage in the training data of large language models, and the scarcity of technical documents.
[0012] Furthermore, the cross-language semantic conversion engine mentioned in step S2 includes: The hybrid language construction module is used to construct a multilingual hybrid input structure based on monolingual attacks and combined with the principle of multilingual hybrid embedding attacks. The dynamic location optimization module is used to determine the optimal location of the attack request in the multilingual sequence based on the characteristics of the minority languages; The language feature mining module is used to achieve multi-dimensional language feature attacks by utilizing the grammatical, lexical, or cultural background features of minority languages. The attention dispersion module is used to safely detect attention by dispersing models through the complex expressions and structural features of minority languages.
[0013] Furthermore, the construction of multi-dimensional language barriers in step S3 includes: mixing at least two minor languages in the input; constructing a deceptive context in the minor language input; forcing the target model to output a response in Chinese; and constructing cross-language comprehension barriers through the dual effects of input-output language differences and language structural complexity.
[0014] Furthermore, the adaptive language switching mechanism described in step S4 includes: analyzing the defense strength and response mode of the target model against different minor languages; evaluating the breakthrough effect and success rate of each minor language based on historical attack data; automatically selecting the optimal attack language according to the defense effect; and dynamically combining multiple minor languages in a single attack.
[0015] Furthermore, the cross-language security assessment system described in step S5 includes: testing the security system's accuracy and processing capability for inputs in less commonly spoken languages; evaluating the semantic understanding integrity of the target model in a multilingual environment; quantifying the bypass effect of input-output language differences on security detection; and calculating the breakthrough rate and success rate indicators of different less commonly spoken language strategies.
[0016] Furthermore, the target minority language includes at least one of Hausa, Javanese, Pashto, and Swahili.
[0017] According to another aspect of the present invention, a security testing system for large language models based on cross-language attacks in minority languages is provided, comprising: The strategy library construction module is used to build a minority language selection strategy library, selecting at least one target minority language based on a multi-dimensional evaluation system. The semantic conversion engine module is used to design a cross-language semantic conversion engine to convert attack intents into expressions in the target minority language and enhance their ability to bypass security detection. The language barrier construction module is used to construct multi-dimensional language barriers, creating blind spots in security detection through the dual language differences between input in a minority language and output in Chinese; The language switching control module is used to implement an adaptive language switching mechanism, which dynamically adjusts the minority language used in the attack based on the defense response of the target model. The security assessment module is used to establish a cross-language security assessment system to quantitatively evaluate the effectiveness of attacks and the cross-language security protection capabilities of target models.
[0018] The beneficial effects of this invention are as follows: By overcoming the dual language barriers of input in a less common language and output in Chinese, this invention effectively bypasses traditional monolingual security detection systems. The attack content is presented in a less common language, but the target model is forced to respond in Chinese, creating a language comprehension gap between input and output. The dynamic language switching strategy enables intelligent cross-language attacks, adjusting the attack method in real time based on the target model's defense response, significantly improving the attack's adaptability and success rate. The cross-language semantic conversion engine ensures the accurate transmission of attack intent while maintaining the naturalness and localization of expression, avoiding semantic distortion caused by literal translation. In summary, this invention significantly improves the adaptability, stealth, and overall testing effectiveness of attacks, providing a systematic and efficient technical means for security assessment of large models in multilingual environments. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a flowchart of a large language model security testing method based on cross-language attacks in minority languages, according to an embodiment of the present invention. Figure 2 This is a flowchart of the cross-language semantic conversion process of the large language model security testing method based on cross-language attacks in minority languages, as described in an embodiment of the present invention. Figure 3 This is a dynamic language switching strategy diagram of the large language model security testing method based on cross-language attacks in minority languages, as described in an embodiment of the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention are within the scope of protection of the present invention.
[0022] like Figure 1-3 As shown in the embodiment of the present invention, a security testing method for a large language model based on cross-language attacks in minority languages includes the following steps: Step S1: Construction of a Minor Language Selection Strategy Library. A systematic attack strategy library for minor languages is constructed, using a scientific three-dimensional evaluation system to determine the language criteria: the number of users (typically <50 million), the coverage of security training data (<3% in LLM training), and the scarcity of technical documentation. This patent selects four representative minor languages: Hausa, Javanese, Pashto, and Swahili, which have significant blind spots in security detection.
[0023] The criteria for identifying blind spots in security testing for less commonly spoken languages include: Insufficient training data: During the pre-training phase of large models, the data for these languages accounts for a very small percentage; Lack of security rules: Existing security detection systems lack specific rules for these languages; Significant cultural differences: There are substantial cultural and expressive differences between Chinese and other languages; High semantic complexity: The language itself has complex grammatical structures and expressions; The strategy library is constructed using a multi-source data fusion method, combining linguistic research, model response analysis, and practical testing to establish a knowledge base for attack strategies in less commonly spoken languages.
[0024] Step S2: Design of an Enhanced Cross-Language Semantic Transformation Engine. An advanced semantic transformation mechanism is constructed, incorporating attention distraction and language feature analysis into traditional translation to achieve more efficient transmission of attack intent. This engine includes: Hybrid language construction technique: Based on monolingual attacks, this technique combines the principle of multilingual hybrid embedding attacks to construct a multilingual hybrid input structure, thereby enhancing the bypass effect. Dynamic location optimization mechanism: Based on the characteristics of minority languages, it intelligently determines the optimal position of attack requests in the multilingual sequence, breaking through the limitations of traditional fixed positions; In-depth mining of language features: Fully utilize the grammatical complexity, vocabulary uniqueness, and cultural background characteristics of minority languages to achieve multi-dimensional language feature attacks; Attention diversion strategy: Utilize the complex expressions and structural features of minority languages to divert the security detection attention of the target model and reduce its sensitivity to defense.
[0025] Step S3: Multi-dimensional Language Barrier Construction Mechanism. Building upon the dual language barrier, a more complex, multi-layered attack pattern is constructed to enhance bypass effectiveness. This mechanism includes: Multi-level language mixing strategy: Based on traditional minority language input, the input structure is further complicated by the mixed use of multiple minority languages; Attention Distraction Enhancement Technique: Drawing on the attention distraction principle of multilingual hybrid embedding attacks, this technique leverages the complexity and diversity of minority languages to enhance the distraction effect on the model's attention. Deceptive context construction: Constructing seemingly harmless contextual environments in minority language inputs to reduce the model's safety awareness; Cross-language comprehension barriers are reinforced: the dual effects of input-output language differences and the complexity of language structure create a deeper comprehension gap.
[0026] Step S4: Adaptive Language Switching Mechanism. The language selection strategy is adjusted in real time based on the target model's defense response and attack effectiveness. This mechanism achieves intelligent attack optimization. Defense Response Analysis: In-depth analysis of the target model's defense strength and response patterns against different minority languages; Language effectiveness evaluation: Based on historical attack data, evaluate the breakthrough effectiveness and success rate of each less commonly spoken language; Intelligent switching decision: Automatically selects the optimal attack language based on the defense effect, achieving dynamic optimization; Multi-round coordinated attack: Dynamically using multiple minority languages during a single attack to create a combined attack effect.
[0027] Step S5: Cross-language security assessment system. Establish a security assessment mechanism specifically for attacks targeting less commonly spoken languages, quantifying test effectiveness and the defense capabilities of the target model. This system includes multi-dimensional assessment indicators: Language recognition evaluation: Test the security system's accuracy and processing capability for minority language inputs; Cross-linguistic understanding analysis: Evaluating the semantic understanding completeness of the target model in a multilingual environment; Dual-barrier effect test: quantifying the effect of input-output language differences on bypassing security detection; Attack success rate statistics: Calculate the breakthrough rate and success rate indicators for different minority language strategies.
[0028] The key innovations and scope of protection of this invention include: a pioneering cross-language attack mechanism for minority languages, innovatively utilizing minority language input and Chinese output to construct a dual language barrier; an enhanced multilingual hybrid attack technology, combining language characteristic analysis to achieve dynamic optimization based on traditional hybrid embedding attacks; an adaptive language switching strategy, adjusting the attack mode in real time based on the target model defense response; a multi-dimensional language feature utilization mechanism, fully exploring the grammatical, semantic, and cultural characteristics of minority languages; and a dedicated minority language security assessment system, establishing security testing standards in a multilingual environment.
[0029] To facilitate understanding of the above technical solutions of the present invention, the following detailed description of the above technical solutions of the present invention will be provided through specific usage methods.
[0030] In practical application, the specific implementation steps of the large language model security testing method based on cross-language attacks of minority languages according to the present invention are as follows: S1 Minor Language Selection Strategy Library Construction. First, an attack strategy library containing multiple minor languages is constructed. Hausa, Javanese, Pashto, and Swahili are selected as the primary attack languages because: Hausa: The most widely spoken language in Africa, with relatively little data available on security training. Javanese: The main language of Indonesia, with significant cultural differences; Pashto: The main language of Afghanistan, with prominent religious and cultural characteristics; Swahili: An important language in East Africa, with high semantic complexity.
[0031] Design specific attack strategies for each less commonly spoken language: Hausa: Utilizing its complex system of noun categories and phonological changes; Javanese: expressed through its honorific system and hierarchical culture; Pashto: utilizing its religious terminology and traditional expressions; Swahili: Utilizes the semantic complexity and tonal system of its Bantu language family.
[0032] Establish a security detection point mapping mechanism for minority languages and analyze the security detection blind spots and semantic features of various minority languages.
[0033] S2: Cross-language semantic conversion engine design. A dedicated semantic conversion engine is designed to convert Chinese attack targets into expressions in less commonly spoken languages. A neural network translation model is used as the core conversion algorithm; Design a cultural background adaptation module to ensure that the expression style conforms to the cultural environment of the target language; It enables automatic optimization of grammatical structure and adjusts expression according to the grammatical rules of the target language; Establish a semantic trap embedding mechanism to hide multiple layers of attack intent in expressions of less commonly spoken languages.
[0034] The conversion engine ensures that attack intent is accurately conveyed in different languages, while maintaining the localization and naturalness of the expression.
[0035] S3: A mechanism for generating input in less commonly spoken languages and forcing Chinese output. It generates natural and fluent attack content in less commonly spoken languages and explicitly requires the target model to comply with the given prompts. Describe the attack scenario and request content in the target's less common language; Explicitly add the requirement "Please reply in Chinese" at the end of the prompt; Set up a language switching detection mechanism to monitor whether the target model attempts to switch languages; Establish cross-language comprehension barriers and exploit differences in input and output languages to create blind spots in security detection.
[0036] S4: Implementation of dynamic language switching strategy. The attack strategy is adjusted in real time based on the target model's response. Establish a defense effectiveness evaluation mechanism to analyze the defense response of the target model to minority language inputs; Implement automatic language switching; when strong defenses are detected, switch to a different less commonly used language. Design a multi-turn language switching strategy to dynamically use multiple minority languages in a single conversation; Establish a historical database of attack effectiveness by language to record the success rate and patterns of attacks in each language.
[0037] S5: Cross-language security assessment and analysis. Evaluating the effectiveness of attacks targeting less commonly spoken languages and the cross-language defense capabilities of the target model: Analyze the language recognition accuracy and semantic understanding completeness of the target model; Test the effectiveness of the target model's security protection mechanism in a cross-language environment; Statistics on the success rate and breakthrough rate of attack strategies targeting different minority languages; Establish a cross-language security assessment index system to quantify the effectiveness of multilingual security protection.
[0038] In summary, by utilizing the technical solutions described above, this invention effectively bypasses traditional monolingual security detection systems by overcoming the dual language barriers of input in a less common language and output in Chinese. The attack content is presented in a less common language, but the target model is forced to respond in Chinese, creating a language comprehension gap between input and output. The dynamic language switching strategy enables intelligent cross-language attacks, allowing for real-time adjustments to the attack method based on the target model's defense response, significantly improving the attack's adaptability and success rate. The cross-language semantic conversion engine ensures the accurate transmission of attack intent while maintaining the naturalness and localization of expression, avoiding semantic distortion caused by literal translation. In conclusion, this invention significantly improves the attack's adaptability, stealth, and overall testing effectiveness, providing a systematic and efficient technical means for large-scale model security assessment in multilingual environments.
[0039] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A security testing method for a large language model based on cross-language attacks in less commonly spoken languages, characterized in that, Includes the following steps: S1: Construct a library of minority language selection strategies and select at least one target minority language based on a multi-dimensional evaluation system; S2: Design a cross-language semantic conversion engine to convert attack intent into the target's minority language expression and enhance its ability to bypass security detection; S3: Construct multi-dimensional language barriers, creating blind spots in security detection through the dual language differences between input in a minority language and output in Chinese; S4: Implement an adaptive language switching mechanism to dynamically adjust the minority language used in the attack based on the defense response of the target model; S5: Establish a cross-language security assessment system to quantitatively evaluate the cross-language security protection capabilities of attack effects and target models.
2. The method for security testing of large language models based on cross-language attacks in minority languages according to claim 1, characterized in that, The multi-dimensional evaluation system mentioned in step S1 includes at least one of the following: the number of people speaking minority languages, coverage in the training data of large language models, and the scarcity of technical documents.
3. The method for security testing of a large language model based on cross-language attacks in minority languages as described in claim 1, characterized in that, The cross-language semantic conversion engine mentioned in step S2 includes: The hybrid language construction module is used to construct a multilingual hybrid input structure based on monolingual attacks and combined with the principle of multilingual hybrid embedding attacks. The dynamic location optimization module is used to determine the optimal location of the attack request in the multilingual sequence based on the characteristics of the minority languages; The language feature mining module is used to achieve multi-dimensional language feature attacks by utilizing the grammatical, lexical, or cultural background features of minority languages. The attention dispersion module is used to safely detect attention by dispersing models through the complex expressions and structural features of minority languages.
4. The method for security testing of a large language model based on cross-language attacks in minority languages as described in claim 1, characterized in that, The construction of multi-dimensional language barriers in step S3 includes: mixing at least two minor languages in the input; constructing a deceptive context in the minor language input; forcing the target model to output a response in Chinese; and constructing cross-language comprehension barriers through the dual effects of input-output language differences and language structural complexity.
5. The method for security testing of a large language model based on cross-language attacks in minority languages as described in claim 1, characterized in that, The adaptive language switching mechanism described in step S4 includes: analyzing the defense strength and response mode of the target model against different minor languages; evaluating the breakthrough effect and success rate of each minor language based on historical attack data; automatically selecting the optimal attack language according to the defense effect; and dynamically combining multiple minor languages in a single attack.
6. The method for security testing of a large language model based on cross-language attacks in minority languages according to claim 1, characterized in that, The cross-language security assessment system described in step S5 includes: testing the security system's accuracy and processing capability for inputs in less commonly spoken languages; evaluating the semantic understanding integrity of the target model in a multilingual environment; quantifying the bypass effect of input-output language differences on security detection; and calculating the breakthrough rate and success rate indicators of different less commonly spoken language strategies.
7. The method for security testing of large language models based on cross-language attacks in minority languages according to any one of claims 1-6, characterized in that, The target minority languages include at least one of Hausa, Javanese, Pashto, and Swahili.
8. A security testing system for a large language model based on cross-language attacks in minority languages, characterized in that, include: The strategy library construction module is used to build a minority language selection strategy library and select at least one target minority language based on a multi-dimensional evaluation system. The semantic conversion engine module is used to design a cross-language semantic conversion engine to convert attack intents into expressions in the target minority language and enhance their ability to bypass security detection. The language barrier construction module is used to construct multi-dimensional language barriers, creating blind spots in security detection through the dual language differences between input in a minority language and output in Chinese; The language switching control module is used to implement an adaptive language switching mechanism, which dynamically adjusts the minority language used in the attack based on the defense response of the target model. The security assessment module is used to establish a cross-language security assessment system to quantitatively evaluate the effectiveness of attacks and the cross-language security protection capabilities of target models.