Safety protection method and device based on LLM safety protection system

Through an LLM-based security protection system, the to-process text sets are processed and analyzed, and the error or misleading content is identified and deleted or intercepted, which solves the problem that the prior art is difficult to adapt to the complexity of LLM output, and improves security protection performance and user experience.

CN120181065AActive Publication Date: 2025-06-20GUANGDONG PLANNING & DESIGNING INST OF TELECOMM

Patent Information

Application Number
CN202510645893.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-20
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

Existing security guardrail products are difficult to adapt to the diversity and complexity of large language model (LLM) output, resulting in the generation of inaccurate or misleading content, affecting the user's user experience.

Method used

A security protection method based on an LLM security protection system is provided, identifying and deleting or intercepting erroneous or misleading content by obtaining a set of pending texts and processing and analyzing them according to user processing requirements parameters.

Benefits of technology

Reduces errors or misleading content in text collection after processing, improves the security protection performance of LLM, and thus improves the security of user information viewing and use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181065A_ABST
    Figure CN120181065A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information security, and discloses a security protection method and device based on an LLM security protection system, and the method comprises the steps: carrying out the processing operation of an obtained to-be-processed text set according to a user processing demand parameter, and obtaining a processed text set; according to the processed text set, analyzing the processed text set to obtain an analysis result of the processed text set; and according to an analysis result of the processed text set, judging whether target operation needs to be performed on the processed text set, and if so, performing the target operation on the processed text set. Visibly, by implementing the method and the device, the to-be-processed text set can be processed and analyzed based on the user processing demand parameters, the analysis result of the processed text set is obtained, and the processed text set is deleted and / or intercepted, so that wrong or misleading contents in the processed text set are reduced, the safety protection performance of LLM is improved, and the user experience is improved. Therefore, the information viewing / using safety of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information security technology, and in particular, to a security protection method and device based on an LLM security protection system. Background Art

[0002] In recent years, LLM (Large Language Model) technology has achieved in-depth understanding and efficient processing of natural language through deep learning and big data training. This model can capture the complexity and diversity of language, thus performing excellently in various NLP (Natural Language Processing) tasks. However, with the wide application of LLM technology, the uncertainty of its output and potential security risks have become increasingly prominent: for example, most existing security fence products use rule- or template-based methods for content filtering and correction, which are difficult to adapt to the diversity and complexity of LLM output, and thus are prone to generating inaccurate or misleading content, affecting the user experience. It can be seen that it is particularly important to provide a method that can improve the security protection performance of LLM. Summary of the Invention

[0003] The present invention provides a security protection method and device based on an LLM security protection system, which reduces the incorrect or misleading content in the processed text set, thereby improving the security protection performance of LLM, and thus improving the security of users' information viewing / usage.

[0004] To solve the above technical problems, in the first aspect of the present invention, a security protection method based on an LLM security protection system is disclosed, and the method includes: Obtain a text set to be processed, and perform a processing operation on the text set to be processed according to preset user processing requirement parameters to obtain a processed text set; the user processing requirement parameters include at least one of a text deduplication requirement parameter, a data visualization requirement parameter, a text service requirement parameter, and a text analysis accuracy requirement parameter; Perform an analysis operation on the processed text set according to the processed text set to obtain an analysis result of the processed text set; the analysis operation includes a star prediction operation and / or a risk conduction analysis operation; Judge whether a target operation needs to be performed on the processed text set according to the analysis result of the processed text set. If so, perform the target operation on the processed text set; the target operation includes a text deletion operation or a text interception operation.

[0005] As an optional implementation manner, in the first aspect of the present invention, the performing a processing operation on the text set to be processed according to preset user processing requirement parameters to obtain a processed text set includes: Obtain the text parameters of the text set to be processed; the text parameters of the text set to be processed include at least one of the text source parameter, text time parameter, text type parameter, and text content parameter of each text to be processed in the text set to be processed; According to the text parameters of the text set to be processed and the preset user processing requirement parameters, perform a preprocessing operation on the text set to be processed to obtain a preprocessed text set; the preprocessing operation includes at least one of a text format conversion operation, a text format normalization operation, a text word segmentation and sentence splitting operation, and a word vector conversion operation; According to the preprocessed text set and the user processing requirement parameters, perform a similarity calculation operation on every two preprocessed texts in the preprocessed text set to obtain the similarity parameters corresponding to every two preprocessed texts; According to the similarity parameters corresponding to every two preprocessed texts, screen out every two preprocessed texts in the preprocessed text set whose similarity parameters are less than or equal to the preset similarity threshold as the processed text set.

[0006] As an optional implementation manner, in the first aspect of the present invention, the step of performing a similarity calculation operation on every two preprocessed texts in the preprocessed text set according to the preprocessed text set and the user processing requirement parameters to obtain the similarity parameters corresponding to every two preprocessed texts includes: According to the text parameters of the preprocessed text set, perform an inverted index operation on historical data of the preprocessed text set to obtain the keywords of each preprocessed text in the preprocessed text set; Perform a sentence-level prefix tree construction operation on each preprocessed text to obtain the prefix tree of each preprocessed text; According to the keywords of each preprocessed text and the corresponding prefix tree, determine the feature information of each preprocessed text, and according to the feature information of all the preprocessed texts, determine the cosine similarity parameter and / or Hamming distance parameter of every two preprocessed texts as the similarity parameters corresponding to every two preprocessed texts; the feature information includes at least one of text vector feature information, string feature information, and binary feature information.

[0007] As an optional implementation manner, in the first aspect of the present invention, the step of performing an analysis operation on the processed text set according to the processed text set to obtain the analysis result of the processed text set includes: According to the processed text set, determine the target information of the processed text set; the target information includes at least one of text sentiment tendency information, user interest point information, text label information, and target object attribute information; Determine the star - rating prediction parameters of the processed text set according to the target information of the processed text set and a preset target library; the target library includes at least one of a knowledge base, a corpus, and a thesaurus, and the star - rating prediction parameters of the processed text set include the basic star - rating prediction parameters of each processed text; For each of the processed texts, determine the prediction frequency parameter that matches the processed text from the historical star - rating prediction data; Determine the target star - rating prediction parameters of the processed text set based on the basic star - rating prediction parameters of all the processed texts and the corresponding prediction frequency parameters, and use them as the analysis result of the processed text set.

[0008] As an alternative implementation manner, in the first aspect of the present invention, the step of analyzing the processed text set according to the processed text set to obtain the analysis result of the processed text set further includes: Determine the target entity information of the processed text set according to the processed text set; the target entity information includes the entity information of each processed text and the entity relationship information among all the processed texts; Construct a knowledge graph of the processed text set according to the target entity information of the processed text set; Determine the target risk conduction parameters of the processed text set according to the knowledge graph of the processed text set, and use them as the analysis result of the processed text set.

[0009] As an alternative implementation manner, in the first aspect of the present invention, the step of determining the target risk conduction parameters of the processed text set according to the knowledge graph of the processed text set includes: Determine the risk conduction type parameter and the risk conduction path parameter of each processed text according to the knowledge graph of the processed text set; the risk conduction path parameter includes a risk source parameter, a risk receptor parameter, and a risk conduction direction parameter; Determine the risk conduction parameter of each processed text according to the risk conduction type parameter and the risk conduction path parameter of each processed text; the risk conduction parameter includes at least one of a risk conduction speed parameter, a risk influence range parameter, and a loss type parameter; Determine the target risk conduction parameters of the processed text set according to the risk conduction parameters of all the processed texts.

[0010] As an alternative implementation manner, in the first aspect of the present invention, the step of determining whether to perform a target operation on the processed text set according to the analysis result of the processed text set includes: Determine the target scenario information of the processed text set; the target scenario information includes viewing scenario information and / or usage scenario information; According to the target scenario information, determine the decision parameter threshold of the processed text set; the decision parameter threshold includes a star rating prediction parameter threshold and / or a risk conduction parameter threshold; According to the target parameter of the processed text set and the decision parameter threshold, determine whether the target parameter is greater than or equal to the decision parameter threshold. If so, determine that a target operation needs to be performed on the processed text set; the target parameter includes the target star rating prediction parameter of the processed text set and / or the target risk conduction parameter of the processed text set.

[0011] A second aspect of the present invention discloses a security protection device based on an LLM security protection system, the device includes: An acquisition module, configured to acquire a text set to be processed; A processing module, configured to perform a processing operation on the text set to be processed according to preset user processing requirement parameters, and obtain a processed text set; the user processing requirement parameters include at least one of a text deduplication requirement parameter, a data visualization requirement parameter, a text service requirement parameter, and a text analysis accuracy requirement parameter; An analysis module, configured to perform an analysis operation on the processed text set according to the processed text set, and obtain an analysis result of the processed text set; the analysis operation includes a star rating prediction operation and / or a risk conduction analysis operation; A judgment module, configured to judge whether a target operation needs to be performed on the processed text set according to the analysis result of the processed text set; A target operation module, configured to perform a target operation on the processed text set when the judgment result of the judgment module is yes; the target operation includes a text deletion operation or a text interception operation.

[0012] As an optional implementation manner, in the second aspect of the present invention, the manner in which the processing module performs a processing operation on the text set to be processed according to preset user processing requirement parameters to obtain a processed text set specifically includes: Obtain the text parameters of the text set to be processed; the text parameters of the text set to be processed include at least one of the text source parameter, text time parameter, text type parameter, and text content parameter of each text to be processed in the text set to be processed; According to the text parameters of the text set to be processed and the preset user processing requirement parameters, perform a preprocessing operation on the text set to be processed to obtain a preprocessed text set; the preprocessing operation includes at least one of a text format conversion operation, a text format normalization operation, a text word segmentation and sentence splitting operation, and a word vector conversion operation; According to the preprocessed text set and the user processing requirement parameters, perform a similarity calculation operation on every two preprocessed texts in the preprocessed text set to obtain the similarity parameters corresponding to every two preprocessed texts; According to the similarity parameters corresponding to every two preprocessed texts, screen out every two preprocessed texts in the preprocessed text set whose similarity parameters are less than or equal to a preset similarity threshold as the processed text set.

[0013] As an optional implementation manner, in the second aspect of the present invention, the manner in which the processing module performs a similarity calculation operation on every two preprocessed texts in the preprocessed text set according to the preprocessed text set and the user processing requirement parameters to obtain the similarity parameters corresponding to every two preprocessed texts specifically includes: Perform an inverted index operation on historical data of the preprocessed text set according to the text parameters of the preprocessed text set to obtain the keywords of each preprocessed text in the preprocessed text set; Perform a sentence-level prefix tree construction operation on each preprocessed text to obtain the prefix tree of each preprocessed text; Determine the feature information of each preprocessed text according to the keywords and the corresponding prefix tree of each preprocessed text, and determine the cosine similarity parameter and / or Hamming distance parameter of every two preprocessed texts according to the feature information of all the preprocessed texts as the similarity parameters corresponding to every two preprocessed texts; the feature information includes at least one of text vector feature information, string feature information, and binary feature information.

[0014] As an optional implementation manner, in the second aspect of the present invention, the manner in which the analysis module performs an analysis operation on the processed text set according to the processed text set to obtain the analysis result of the processed text set specifically includes: Determine the target information of the processed text set according to the processed text set; the target information includes at least one of text sentiment tendency information, user interest point information, text label information, and target object attribute information; Determine the star rating prediction parameter of the processed text set according to the target information of the processed text set and a preset target library; the target library includes at least one of a knowledge library, a corpus, and a word library, and the star rating prediction parameter of the processed text set includes the basic star rating prediction parameter of each processed text; For each processed text, determine the prediction frequency parameter matching the processed text from the historical star rating prediction data; Based on the basic star - rating prediction parameters and the corresponding prediction frequency parameters of all the processed texts, determine the target star - rating prediction parameter of the processed text set as the analysis result of the processed text set.

[0015] As an alternative implementation, in the second aspect of the present invention, the way that the analysis module analyzes the processed text set according to the processed text set to obtain the analysis result of the processed text set specifically further includes: According to the processed text set, determine the target entity information of the processed text set; the target entity information includes the entity information of each processed text and the entity relationship information between all the processed texts; Construct a knowledge graph of the processed text set according to the target entity information of the processed text set; Determine the target risk conduction parameter of the processed text set according to the knowledge graph of the processed text set as the analysis result of the processed text set.

[0016] As an alternative implementation, in the second aspect of the present invention, the way that the analysis module determines the target risk conduction parameter of the processed text set according to the knowledge graph of the processed text set specifically includes: According to the knowledge graph of the processed text set, determine the risk conduction type parameter and the risk conduction path parameter of each processed text; the risk conduction path parameter includes a risk source parameter, a risk receptor parameter, and a risk conduction direction parameter; According to the risk conduction type parameter and the risk conduction path parameter of each processed text, determine the risk conduction parameter of each processed text; the risk conduction parameter includes at least one of a risk conduction speed parameter, a risk influence range parameter, and a loss type parameter; Determine the target risk conduction parameter of the processed text set according to the risk conduction parameters of all the processed texts.

[0017] As an alternative implementation, in the second aspect of the present invention, the way that the judgment module determines whether to perform a target operation on the processed text set according to the analysis result of the processed text set specifically includes: Determine the target scenario information of the processed text set; the target scenario information includes viewing scenario information and / or usage scenario information; According to the target scenario information, determine the decision parameter threshold of the processed text set; the decision parameter threshold includes a star - rating prediction parameter threshold and / or a risk conduction parameter threshold; Based on the target parameter of the processed text set and the determination parameter threshold, determine whether the target parameter is greater than or equal to the determination parameter threshold. If so, determine that a target operation needs to be performed on the processed text set; the target parameter includes the target star rating prediction parameter of the processed text set and / or the target risk conduction parameter of the processed text set.

[0018] The third aspect of the present invention discloses another security protection device based on an LLM security protection system. The device includes: A memory storing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory and executes the security protection method based on the LLM security protection system disclosed in the first aspect of the present invention.

[0019] The fourth aspect of the present invention discloses a computer storage medium. The computer storage medium stores computer instructions, which are used to execute the security protection method based on the LLM security protection system disclosed in the first aspect of the present invention when the computer instructions are called.

[0020] Compared with the prior art, the embodiments of the present invention have the following beneficial effects: In the embodiments of the present invention, according to the user processing requirement parameters, a processing operation is performed on the obtained text set to be processed to obtain a processed text set; based on the processed text set, an analysis is performed on the processed text set to obtain an analysis result of the processed text set; according to the analysis result of the processed text set, it is determined whether a target operation needs to be performed on the processed text set. If so, a target operation is performed on the processed text set. It can be seen that implementing the present invention can perform processing and analysis operations on the text set to be processed based on the user processing requirement parameters, obtain the analysis result of the processed text set, and perform deletion and / or interception operations on the processed text set. In this way, the incorrect or misleading content in the processed text set is reduced, thereby improving the security protection performance of the LLM and thus improving the security of the user's information viewing / usage. Description of the Drawings

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a flowchart of a security protection method based on an LLM security protection system disclosed in an embodiment of the present invention; Figure 2 It is a schematic flowchart of another security protection method based on the LLM security protection system disclosed in the embodiments of the present invention; Figure 3 It is a schematic structural diagram of a security protection device based on the LLM security protection system disclosed in the embodiments of the present invention; Figure 4 It is a schematic structural diagram of another security protection device based on the LLM security protection system disclosed in the embodiments of the present invention. Detailed implementation manners

[0023] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0024] The terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or terminal including a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or terminals.

[0025] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment can be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0026] The present invention discloses a security protection method and device based on an LLM security protection system, which reduces the incorrect or misleading content in the processed text set, thereby improving the security protection performance of the LLM, and thus improving the security of users' information viewing / using.

[0027] Embodiment 1 Please refer to Figure 1 , Figure 1 It is a schematic flowchart of a security protection method based on the LLM security protection system disclosed in the embodiments of the present invention. Among them, Figure 1The described security protection method based on the LLM security protection system can be applied to the fields of software code security, public security, enterprise security, etc., and the embodiments of the present invention are not limited thereto. Optionally, this method can be implemented by a security protection device, which can be integrated in a security protection device (such as a smart computer), or can be a local server or a cloud server used to process the LLM security protection device process, etc., and the embodiments of the present invention are not limited thereto. As Figure 1 shown, the security protection method based on the LLM security protection system may include the following operations: 101. Obtain the text set to be processed, and perform a processing operation on the text set to be processed according to the preset user processing requirement parameters to obtain the processed text set.

[0028] In the embodiments of the present invention, this step can be implemented by the data layer of the LLM security protection system. Additionally, a storage operation can also be performed on the text set to be processed. Optionally, the user processing requirement parameters include at least one of text deduplication requirement parameters (such as text repetition rate), data visualization requirement parameters (such as whether the data chart used is two-dimensional or three-dimensional), text business requirement parameters, and text analysis accuracy requirement parameters.

[0029] 102. Analyze the processed text set according to the processed text set to obtain the analysis result of the processed text set.

[0030] In the embodiments of the present invention, this step can be implemented by the data layer of the LLM security protection system. Optionally, the analysis operation includes a star rating prediction operation (which includes processes such as sentiment analysis and keyword extraction) and / or a risk conduction analysis operation (which includes processes such as knowledge graph and risk conduction analysis).

[0031] 103. Determine whether it is necessary to perform a target operation on the processed text set according to the analysis result of the processed text set. If so, perform a target operation on the processed text set.

[0032] In the embodiments of the present invention, this step can be implemented by the service layer and application layer of the LLM security protection system. Optionally, the target operation includes a text deletion operation or a text interception operation. Further, user parameters, such as user role parameters and user permission scope parameters, can also be combined to determine whether it is necessary to perform a target operation on the processed text set.

[0033] It can be seen that implementing the embodiments of the present invention can process and analyze the text set to be processed based on the user processing requirement parameters, obtain the analysis result of the processed text set, and perform deletion and / or interception operations on the processed text set. In this way, the incorrect or misleading content in the processed text set is reduced, thereby improving the security protection performance of the LLM, and thus improving the security of users' information viewing / usage.

[0034] Embodiment 2 Please refer to Figure 2 , Figure 2 which is a schematic flowchart of another security protection method based on the LLM security protection system disclosed in the embodiments of the present invention. Among them, Figure 2 the described security protection method based on the LLM security protection system can be applied to the fields of software code security, public security, enterprise security, etc., and the embodiments of the present invention do not make limitations. Optionally, this method can be implemented by a security protection device, and this security protection device can be integrated in a security protection device (such as a smart computer), or can be a local server or a cloud server used to process the LLM security protection device process, etc., and the embodiments of the present invention do not make limitations. As Figure 2 shown, the security protection method based on the LLM security protection system may include the following operations: 201. Obtain a text set to be processed and obtain the text parameters of the text set to be processed.

[0035] In the embodiments of the present invention, optionally, the text parameters of the text set to be processed include at least one of the text source parameter, text time parameter, text type parameter, and text content parameter of each text to be processed in the text set to be processed.

[0036] 202. Perform a preprocessing operation on the text set to be processed according to the text parameters of the text set to be processed and the preset user processing requirement parameters to obtain a preprocessed text set.

[0037] In the embodiments of the present invention, optionally, the preprocessing operation includes at least one of a text format conversion operation, a text format normalization operation, a text tokenization and sentence splitting operation, and a word vector conversion operation.

[0038] 203. Perform a similarity calculation operation on the pairwise preprocessed texts in the preprocessed text set according to the preprocessed text set and the user processing requirement parameters to obtain the similarity parameters corresponding to the pairwise preprocessed texts.

[0039] In the embodiments of the present invention, among them, the similarity calculation operation may include a cosine similarity calculation operation and / or a Hamming distance parameter calculation operation.

[0040] 204. Screen out the pairwise preprocessed texts with similarity parameters less than or equal to the preset similarity threshold from the preprocessed text set as the processed text set.

[0041] In an embodiment of the present invention, first, duplicate removal operation is performed on the text set to be processed to reduce the computational amount and meet the user's visualization requirements. In addition, operations such as stop word filtering, spelling correction, and text noise reduction can also be performed on the processed text set to update the processed text set.

[0042] 205. Analyze the processed text set according to the processed text set to obtain the analysis result of the processed text set.

[0043] 206. Determine whether it is necessary to perform a target operation on the processed text set according to the analysis result of the processed text set. If so, perform the target operation on the processed text set.

[0044] In an embodiment of the present invention, for other descriptions of steps 205 and 206, please refer to the detailed descriptions of steps 102 and 103 in Embodiment 1, and the embodiments of the present invention will not be elaborated here.

[0045] It can be seen that implementing the embodiments of the present invention can process the text set to be processed by obtaining the text parameters of the text set to be processed and combining the specific processing requirements of the user, such as text format conversion and normalization, text duplicate removal, etc., to obtain the processed text set. In this way, the processing reliability and accuracy of the text set to be processed are improved, and further the analysis reliability and accuracy of the processed text set are improved, which is conducive to performing accurate deletion and interception operations on the processed text set, and realizing the high security protection performance of the LLM security protection system.

[0046] In an alternative embodiment, the operation of calculating the similarity between two preprocessed texts in the preprocessed text set according to the preprocessed text set and the user processing requirement parameters in step 203 above to obtain the similarity parameters corresponding to the two preprocessed texts includes: Perform an inverted index operation on historical data of the preprocessed text set according to the text parameters of the preprocessed text set to obtain the keywords of each preprocessed text in the preprocessed text set; Perform a sentence-level prefix tree construction operation on each preprocessed text to obtain the prefix tree of each preprocessed text; Determine the feature information of each preprocessed text according to the keywords and the corresponding prefix tree of each preprocessed text, and determine the cosine similarity parameter and / or Hamming distance parameter of the two preprocessed texts according to the feature information of all the preprocessed texts as the similarity parameters corresponding to the two preprocessed texts.

[0047] In this optional embodiment, optionally, the feature information includes at least one of text vector feature information, string feature information, and binary feature information. Further, when constructing the inverted index, the weight of each keyword in the text can be calculated, such as using methods like TF-IDF (Term Frequency-Inverse Document Frequency), which helps to more accurately evaluate the importance of keywords in subsequent similarity calculations. Further, in order to improve the accuracy and efficiency of text similarity calculation, algorithms such as cosine similarity and Hamming distance can be optimized. For example, approximate nearest neighbor search algorithms (such as LSH, HNSW, etc.) can be used to accelerate the process of finding similar texts, etc.

[0048] It can be seen that this optional embodiment can, by combining the historical data inverted index and the prefix tree construction technology, achieve the similarity calculation of pairwise preprocessed texts in the preprocessed text set, obtain the similarity parameters corresponding to the pairwise preprocessed texts. In this way, the efficiency and accuracy of text similarity calculation can be improved, the text processing time can be shortened, and further, the analysis accuracy and efficiency of the subsequent processed text set can be improved, thereby providing more reliable and accurate text processing results for users.

[0049] In another optional embodiment, for the analysis operation on the processed text set according to the processed text set in step 205 above, the analysis result of the processed text set includes: Determine the target information of the processed text set according to the processed text set; Determine the star rating prediction parameter of the processed text set according to the target information of the processed text set and the preset target library; For each processed text, determine the prediction frequency parameter that matches the processed text from the historical star rating prediction data; Determine the target star rating prediction parameter of the processed text set according to the basic star rating prediction parameters of all processed texts and the corresponding prediction frequency parameters as the analysis result of the processed text set.

[0050] In this optional embodiment, among them, the star rating prediction parameter of the processed text set includes the basic star rating prediction parameter of each processed text. Optionally, the target information includes at least one of text sentiment tendency information, user interest point information, text label information, and target object attribute information (such as service attribute, commodity attribute, activity attribute, etc.). Further optionally, the target library includes at least one of a knowledge base, a corpus, and a thesaurus.

[0051] Further, in the star rating prediction parameter, dynamic adjustment factors can also be introduced, and these dynamic adjustment factors can be adjusted in real time according to external factors such as the latest market trends, user feedback, and competitor situations, making the star rating prediction closer to reality.

[0052] Furthermore, for the determination of the prediction frequency parameter, it can not only rely on the historical star rating prediction data, but also consider factors such as the popularity and timeliness of the currently processed text. For example, texts with high popularity may have a higher prediction frequency parameter, while texts with a longer timeliness may have a lower prediction frequency parameter, etc.

[0053] It can be seen that this optional embodiment can accurately determine the star rating prediction parameters of the processed text set through multiple dimensions of information of the processed text set (such as text sentiment tendency, user interest points, text tags, target object attributes, etc.), and in combination with the preset target library and historical star rating prediction data. At the same time, by introducing the dynamic adjustment factor and optimizing the prediction frequency parameter, the star rating prediction is made closer to the actual situation, further improving the accuracy and timeliness of the star rating prediction parameters, so as to enhance the readability and practicality of the analysis results of the subsequent processed text set, and provide users with more intuitive and comprehensive analysis results.

[0054] In another optional embodiment, the above-mentioned step of analyzing the processed text set according to the processed text set to obtain the analysis result of the processed text set further includes: Determine the target entity information of the processed text set according to the processed text set; Construct a knowledge graph of the processed text set according to the target entity information of the processed text set; Determine the target risk conduction parameter of the processed text set according to the knowledge graph of the processed text set as the analysis result of the processed text set.

[0055] In this optional embodiment, the target entity information includes the entity information of each processed text (such as person names, place names, organization names, etc.) and the entity relationship information between all processed texts.

[0056] Furthermore, as an optional implementation method, determining the target risk conduction parameter of the processed text set according to the knowledge graph of the processed text set includes: Determine the risk conduction type parameter and risk conduction path parameter of each processed text according to the knowledge graph of the processed text set; Determine the risk conduction parameter of each processed text according to the risk conduction type parameter and risk conduction path parameter of each processed text; Determine the target risk conduction parameter of the processed text set according to the risk conduction parameters of all processed texts.

[0057] In this optional embodiment, the risk conduction path parameters include risk source parameters, risk receptor parameters, and risk conduction direction parameters. Optionally, the risk conduction parameters include at least one of a risk conduction speed parameter, a risk influence range parameter, and a loss type parameter.

[0058] It can be seen that this optional embodiment can analyze and mine the processed text set through the knowledge graph and risk conduction parameters of the processed text set. In this way, it not only reveals the complex relationships between entities but also provides an intuitive and visual risk conduction path, which helps users understand more clearly the propagation mode and influence range of risks in the entity network, and provides coping strategies for users to develop risk warning systems.

[0059] In another optional embodiment, determining whether to perform a target operation on the processed text set according to the analysis result in step 206 above includes: Determine the target scenario information of the processed text set; According to the target scenario information, determine the decision parameter threshold of the processed text set; the decision parameter threshold includes a star rating prediction parameter threshold and / or a risk conduction parameter threshold; According to the target parameter of the processed text set and the decision parameter threshold, determine whether the target parameter is greater than or equal to the decision parameter threshold. If so, determine that a target operation needs to be performed on the processed text set.

[0060] In this optional embodiment, optionally, the target scenario information includes viewing scenario information and / or usage scenario information, such as review scenarios, fact-checking scenarios, user monitoring scenarios, etc. Further optionally, the target parameter includes the target star rating prediction parameter of the processed text set and / or the target risk conduction parameter of the processed text set.

[0061] It can be seen that this optional embodiment can achieve more accurate and flexible risk judgment and management of the processed text set by introducing a dynamic adjustment of the decision parameter threshold, a user feedback mechanism, and diverse target operation methods, making the text processing results more in line with the actual needs of users, thereby effectively reducing the potential losses caused by text risks, and thus enhancing the practical value and competitiveness of the LLM security protection system in the fields of risk management and decision support.

[0062] Embodiment III Please refer to Figure 3 , Figure 3 which is a schematic structural diagram of a security protection device based on the LLM security protection system disclosed in the embodiments of the present invention. As Figure 3 shown, the security protection device based on the LLM security protection system may include: An acquisition module 301, configured to acquire a text set to be processed; A processing module 302, configured to perform a processing operation on a text set to be processed according to preset user processing requirement parameters, so as to obtain a processed text set; An analysis module 303, configured to perform an analysis operation on the processed text set according to the processed text set, so as to obtain an analysis result of the processed text set; A judgment module 304, configured to judge whether a target operation needs to be performed on the processed text set according to the analysis result of the processed text set; A target operation module 305, configured to perform a target operation on the processed text set when the judgment result of the judgment module 304 is yes; the target operation includes a text deletion operation or a text interception operation.

[0063] In an embodiment of the present invention, the user processing requirement parameters include at least one of a text deduplication requirement parameter, a data visualization requirement parameter, a text service requirement parameter, and a text analysis accuracy requirement parameter; the analysis operation includes a star prediction operation and / or a risk conduction analysis operation.

[0064] It can be seen that implementing Figure 3 The described security protection device based on the LLM security protection system can perform processing and analysis operations on the text set to be processed based on user processing requirement parameters, obtain the analysis result of the processed text set, and perform deletion and / or interception operations on the processed text set. In this way, the incorrect or misleading content in the processed text set is reduced, thereby improving the security protection performance of the LLM, and thus improving the security of the user's information viewing / using.

[0065] In an alternative embodiment, the manner in which the processing module 302 performs a processing operation on the text set to be processed according to preset user processing requirement parameters to obtain a processed text set specifically includes: Obtain the text parameters of the text set to be processed; According to the text parameters of the text set to be processed and the preset user processing requirement parameters, perform a preprocessing operation on the text set to be processed to obtain a preprocessed text set; According to the preprocessed text set and the user processing requirement parameters, perform a similarity calculation operation on the pairwise preprocessed texts in the preprocessed text set to obtain similarity parameters corresponding to the pairwise preprocessed texts; According to the similarity parameters corresponding to the pairwise preprocessed texts, screen out the pairwise preprocessed texts in the preprocessed text set whose similarity parameters are less than or equal to a preset similarity threshold as the processed text set.

[0066] In this optional embodiment, the text parameters of the text set to be processed include at least one of the text source parameter, text time parameter, text type parameter, and text content parameter of each text to be processed in the text set to be processed; the preprocessing operations include at least one of text format conversion operation, text format normalization operation, text word segmentation and sentence splitting operation, and word vector conversion operation.

[0067] It can be seen that implementing Figure 3 The security protection device of the LLM-based security protection system described above can process the text set to be processed by obtaining the text parameters of the text set to be processed and combining the specific processing requirements of the user, such as text format conversion and normalization, text deduplication, etc., to obtain the processed text set. In this way, the processing reliability and accuracy of the text set to be processed are improved, and then the analysis reliability and accuracy of the processed text set are improved, which is conducive to accurate deletion and interception operations on the processed text set, realizing the high security protection performance of the LLM-based security protection system.

[0068] In another optional embodiment, the processing module 302 calculates the similarity between every two preprocessed texts in the preprocessed text set according to the preprocessed text set and the user processing requirement parameters to obtain the similarity parameters corresponding to every two preprocessed texts. The specific method includes: Perform an inverted index operation on historical data for the preprocessed text set according to the text parameters of the preprocessed text set to obtain the keywords of each preprocessed text in the preprocessed text set; Perform a sentence-level prefix tree construction operation on each preprocessed text to obtain the prefix tree of each preprocessed text; Determine the feature information of each preprocessed text according to the keywords and the corresponding prefix tree of each preprocessed text, and determine the cosine similarity parameter and / or Hamming distance parameter of every two preprocessed texts according to the feature information of all preprocessed texts as the similarity parameters corresponding to every two preprocessed texts.

[0069] In this optional embodiment, the feature information includes at least one of text vector feature information, string feature information, and binary feature information.

[0070] It can be seen that implementing Figure 3 The security protection device of the LLM-based security protection system described above can implement the similarity calculation of every two preprocessed texts in the preprocessed text set by combining the inverted index of historical data and the prefix tree construction technology, and obtain the similarity parameters corresponding to every two preprocessed texts. In this way, the efficiency and accuracy of text similarity calculation can be improved, the text processing time can be shortened, and then the analysis accuracy and efficiency of the processed text set can be improved, thus providing more reliable and accurate text processing results for users.

[0071] In yet another alternative embodiment, the analysis module 303 analyzes the processed text set according to the processed text set to obtain the analysis result of the processed text set. The specific methods include: Determine the target information of the processed text set according to the processed text set; Determine the star prediction parameters of the processed text set according to the target information of the processed text set and a preset target library; For each processed text, determine the prediction frequency parameter matching the processed text from the historical star prediction data; Determine the target star prediction parameter of the processed text set according to the basic star prediction parameters of all processed texts and the corresponding prediction frequency parameters, and use it as the analysis result of the processed text set.

[0072] In this alternative embodiment, the target information includes at least one of text sentiment tendency information, user interest point information, text label information, and target object attribute information; the target library includes at least one of a knowledge base, a corpus, and a word library, and the star prediction parameters of the processed text set include the basic star prediction parameters of each processed text.

[0073] It can be seen that the implementation Figure 3 The described security protection device based on the LLM security protection system can, through multiple dimensional information of the processed text set (such as text sentiment tendency, user interest point, text label, target object attribute, etc.), and in combination with a preset target library and historical star prediction data, achieve the accurate determination of the star prediction parameters of the processed text set. At the same time, by introducing the optimization of dynamic adjustment factors and prediction frequency parameters, the star prediction is made closer to the actual situation, further improving the accuracy and timeliness of the star prediction parameters, so as to enhance the readability and practicality of the subsequent analysis results of the processed text set and provide users with more intuitive and comprehensive analysis results.

[0074] In yet another alternative embodiment, the analysis module 303 analyzes the processed text set according to the processed text set to obtain the analysis result of the processed text set. The specific methods further include: Determine the target entity information of the processed text set according to the processed text set; Construct a knowledge graph of the processed text set according to the target entity information of the processed text set; Determine the target risk conduction parameter of the processed text set according to the knowledge graph of the processed text set, and use it as the analysis result of the processed text set.

[0075] In this optional embodiment, the target entity information includes the entity information of each processed text and the entity relationship information between all processed texts.

[0076] Further, as an optional implementation manner, the specific way for the analysis module 303 to determine the target risk conduction parameter of the processed text set according to the knowledge graph of the processed text set includes: Determine the risk conduction type parameter and the risk conduction path parameter of each processed text according to the knowledge graph of the processed text set; Determine the risk conduction parameter of each processed text according to the risk conduction type parameter and the risk conduction path parameter of each processed text; Determine the target risk conduction parameter of the processed text set according to the risk conduction parameters of all processed texts.

[0077] In this optional embodiment, the risk conduction path parameter includes a risk source parameter, a risk receptor parameter, and a risk conduction direction parameter; the risk conduction parameter includes at least one of a risk conduction speed parameter, a risk influence range parameter, and a loss type parameter.

[0078] It can be seen that implementing Figure 3 The described security protection device based on the LLM security protection system can analyze and mine the processed text set through the knowledge graph and risk conduction parameters of the processed text set. In this way, not only the complex relationships between entities are revealed, but also an intuitive and visual risk conduction path is provided, which helps users understand more clearly the propagation mode and influence range of risks in the entity network, and provides coping strategies for users to develop risk warning systems.

[0079] In another optional embodiment, the specific way for the judgment module 304 to judge whether a target operation needs to be performed on the processed text set according to the analysis result of the processed text set includes: Determine the target scenario information of the processed text set; Determine the judgment parameter threshold of the processed text set according to the target scenario information; the judgment parameter threshold includes a star prediction parameter threshold and / or a risk conduction parameter threshold; Judge whether the target parameter is greater than or equal to the judgment parameter threshold according to the target parameter of the processed text set and the judgment parameter threshold. If so, determine that a target operation needs to be performed on the processed text set.

[0080] In this optional embodiment, the target scenario information includes viewing scenario information and / or usage scenario information; the target parameter includes the target star prediction parameter of the processed text set and / or the target risk conduction parameter of the processed text set.

[0081] It can be seen that implementing Figure 3The described security protection device of the LLM-based security protection system can achieve more accurate and flexible risk judgment and management of the processed text set by introducing dynamic adjustment of the determination parameter threshold, user feedback mechanism, and diverse target operation methods, making the text processing results more in line with the actual needs of users, thereby effectively reducing the potential losses caused by text risks, and thus enhancing the practical value and competitiveness of the LLM-based security protection system in the fields of risk management and decision support.

[0082] Embodiment 4 Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of another security protection device of the LLM-based security protection system disclosed in the embodiments of the present invention. As Figure 4 shown, the security protection device of the LLM-based security protection system may include: A memory 401 storing executable program code; A processor 402 coupled to the memory 401; The processor 402 invokes the executable program code stored in the memory 401 and executes the steps in the security protection method of the LLM-based security protection system described in Embodiment 1 or Embodiment 2 of the present invention.

[0083] Embodiment 5 The embodiments of the present invention disclose a computer storage medium, which stores computer instructions that, when invoked, are used to execute the steps in the security protection method of the LLM-based security protection system described in Embodiment 1 or Embodiment 2 of the present invention.

[0084] Embodiment 6 The embodiments of the present invention disclose a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute the steps in the security protection method of the LLM-based security protection system described in Embodiment 1 or Embodiment 2.

[0085] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0086] Through the specific descriptions of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, and the storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disk memories, tape memories, or any other medium that can be used to carry or store data and is computer-readable.

[0087] Finally, it should be noted that: The security protection method and device based on the LLM security protection system disclosed in the embodiments of the present invention only disclose the preferred embodiments of the present invention, and are only used to illustrate the technical solutions of the present invention, rather than limiting them; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A security protection method based on the LLM security protection system, characterized in that: The method comprises: Obtaining a text set to be processed, and performing a processing operation on the text set to be processed according to preset user processing requirement parameters to obtain a processed text set; the user processing requirement parameters include at least one of a text deduplication requirement parameter, a data visualization requirement parameter, a text business requirement parameter, and a text analysis accuracy requirement parameter; According to the processed text set, an analysis operation is performed on the processed text set to obtain an analysis result of the processed text set; the analysis operation includes a star rating prediction operation and / or a risk conduction analysis operation; According to the analysis result of the processed text set, it is determined whether a target operation needs to be performed on the processed text set. If so, the target operation is performed on the processed text set; the target operation includes a text deletion operation or a text interception operation.

2. The security protection method based on the LLM security protection system according to claim 1 is characterized in that: The processing operation is performed on the to-be-processed text set according to preset user processing requirement parameters to obtain a processed text set, including: Acquire text parameters of the to-be-processed text set; the text parameters of the to-be-processed text set include at least one of a text source parameter, a text time parameter, a text type parameter, and a text content parameter of each to-be-processed text in the to-be-processed text set; According to the text parameters of the to-be-processed text set and the preset user processing requirement parameters, the to-be-processed text set is preprocessed to obtain a preprocessed text set; the preprocessing operation includes at least one of a text format conversion operation, a text format normalization operation, a text word segmentation operation, and a word vector conversion operation; According to the preprocessed text set and the user processing requirement parameter, a similarity calculation operation is performed on each pair of preprocessed texts in the preprocessed text set to obtain similarity parameters corresponding to each pair of preprocessed texts; According to the similarity parameters corresponding to the preprocessed texts in pairs, the preprocessed texts in pairs whose similarity parameters are less than or equal to a preset similarity threshold are screened out from the preprocessed text set as the processed text set.

3. The security protection method based on the LLM security protection system according to claim 2 is characterized in that: The performing of similarity calculation operations on the preprocessed texts in the preprocessed text set according to the preprocessed text set and the user processing requirement parameters to obtain similarity parameters corresponding to the preprocessed texts in pairs comprises: According to the text parameters of the preprocessed text set, a historical data inverted index operation is performed on the preprocessed text set to obtain a keyword for each preprocessed text in the preprocessed text set; Performing a sentence-level prefix tree construction operation on each of the preprocessed texts to obtain a prefix tree for each of the preprocessed texts; According to the keywords of each preprocessed text and the corresponding prefix tree, the feature information of each preprocessed text is determined, and according to the feature information of all the preprocessed texts, the cosine similarity parameters and / or Hamming distance parameters of the preprocessed texts are determined pairwise as the similarity parameters corresponding to the preprocessed texts pairwise; the feature information includes at least one of text vector feature information, character string feature information and binary feature information.

4. The security protection method based on the LLM security protection system according to any one of claims 1 to 3, characterized in that: The step of performing an analysis operation on the processed text set according to the processed text set to obtain an analysis result of the processed text set includes: Determining target information of the processed text set according to the processed text set; the target information includes at least one of text sentiment tendency information, user interest point information, text label information and target object attribute information; Determine the star rating prediction parameters of the processed text set according to the target information of the processed text set and a preset target library; the target library includes at least one of a knowledge base, a corpus and a thesaurus, and the star rating prediction parameters of the processed text set include basic star rating prediction parameters of each processed text; For each of the processed texts, determining a prediction frequency parameter that matches the processed text from historical star rating prediction data; According to the basic star rating prediction parameters of all the processed texts and the corresponding prediction frequency parameters, the target star rating prediction parameters of the processed text set are determined as the analysis result of the processed text set.

5. The security protection method based on the LLM security protection system according to claim 4 is characterized in that: The step of performing an analysis operation on the processed text set according to the processed text set to obtain an analysis result of the processed text set further includes: Determining target entity information of the processed text set according to the processed text set; the target entity information includes entity information of each processed text and entity relationship information between all processed texts; Constructing a knowledge graph of the processed text set according to target entity information of the processed text set; According to the knowledge graph of the processed text set, a target risk conduction parameter of the processed text set is determined as an analysis result of the processed text set.

6. The security protection method based on the LLM security protection system according to claim 5 is characterized in that: Determining the target risk transmission parameter of the processed text set according to the knowledge graph of the processed text set includes: Determine the risk transmission type parameter and the risk transmission path parameter of each processed text according to the knowledge graph of the processed text set; the risk transmission path parameter includes the risk source parameter, the risk receptor parameter and the risk transmission direction parameter; Determine the risk transmission parameter of each processed text according to the risk transmission type parameter and the risk transmission path parameter of each processed text; the risk transmission parameter includes at least one of a risk transmission speed parameter, a risk impact range parameter and a loss type parameter; According to the risk conduction parameters of all the processed texts, the target risk conduction parameters of the processed text set are determined.

7. The security protection method based on the LLM security protection system according to claim 6 is characterized in that: The step of judging whether it is necessary to perform a target operation on the processed text set according to the analysis result of the processed text set includes: Determining target scenario information of the processed text set; the target scenario information includes viewing scenario information and / or usage scenario information; Determining a judgment parameter threshold of the processed text set according to the target scenario information; the judgment parameter threshold includes a star prediction parameter threshold and / or a risk conduction parameter threshold; According to the target parameter of the processed text set and the judgment parameter threshold, it is determined whether the target parameter is greater than or equal to the judgment parameter threshold. If so, it is determined that a target operation needs to be performed on the processed text set; the target parameter includes a target star prediction parameter of the processed text set and / or a target risk conduction parameter of the processed text set.

8. A safety protection device based on the LLM safety protection system, characterized in that: The device comprises: An acquisition module is used to acquire a text set to be processed; A processing module, used to perform a processing operation on the to-be-processed text set according to preset user processing requirement parameters to obtain a processed text set; the user processing requirement parameters include at least one of a text deduplication requirement parameter, a data visualization requirement parameter, a text business requirement parameter, and a text analysis accuracy requirement parameter; An analysis module, configured to perform an analysis operation on the processed text set according to the processed text set to obtain an analysis result of the processed text set; the analysis operation includes a star rating prediction operation and / or a risk conduction analysis operation; A judgment module, used for judging whether a target operation needs to be performed on the processed text set according to the analysis result of the processed text set; The target operation module is used to perform a target operation on the processed text set when the judgment result of the judgment module is yes; the target operation includes a text deletion operation or a text interception operation.

9. A safety protection device based on the LLM safety protection system, characterized in that: The device comprises: A memory storing executable program code; a processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the security protection method based on the LLM security protection system as described in any one of claims 1-7.

10. A computer storage medium, characterized in that: The computer storage medium stores computer instructions, which, when called, are used to execute the security protection method based on the LLM security protection system as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Risk early warning method, system and equipment based on natural language processing and medium

    CN113094476A

  • LLM-driven adaptive industrial network security protection method and firewall device

    CN118138362A

  • Text information intelligent auditing method, device and equipment and storage medium

    CN119006005A

  • Method for identifying harmful information in short message content based on big data

    CN119907005A

  • Large language models firewall

    US20240388551A1

Cited By

  • Interaction monitoring method and system based on LLM security protection system

    CN120781348A