Information grading identification method based on rule engine and LLM cooperation and related device

CN122594873APending Publication Date: 2026-08-18GUOSEN SECURITIES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610225358.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-25
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

缺点:灵活性差,无法理解上下文语义,容易产生大量误报和漏报

Benefits of technology

[0020] This invention embodiment performs rule matching on the original data to be identified according to a preset text information recognition rule base to obtain a matching result; calculates a confidence score corresponding to the matching result according to a preset confidence algorithm; when the confidence score is lower than a first confidence threshold but higher than or equal to a second confidence threshold, inputs the original data to be identified into a lightweight model for verification to obtain a first verification result; when the first verification result is consistent with the matching result, the matching result is taken as the text information recognition result; when the confidence score is lower than a second confidence interval or the lightweight model fails to recognize the matching result... When the verification results are inconsistent, the original data to be identified is input into a large model for analysis and identification to obtain the large model identification result. This large model identification result is used as the text information identification result. This provides a highly efficient and accurate sensitive information identification solution: by introducing an intelligent routing mechanism, simple and clear identification tasks are handled by an efficient rule engine, while only complex and ambiguous identification tasks are handled by a large language model. This ensures that the overall identification accuracy is close to that of a pure LLM solution, while improving processing efficiency to a level close to that of a pure rule engine solution, and significantly reducing the cost of calling LLM. A dynamic identification system with self-learning and self-evolution capabilities is constructed: by establishing a two-way feedback channel between rules and models, the system can automatically learn new knowledge and correct old errors from daily operation, continuously enriching the rule base and optimizing model performance, thereby continuously improving the ability to identify new and variant sensitive information and solving the problem of poor adaptability of traditional static systems. The overall operating cost of sensitive information identification is significantly reduced: by minimizing reliance on expensive large language models, high-precision sensitive information identification technology can be deployed and applied on a large scale at a lower cost, improving the data security protection capabilities of enterprises.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122594873A_ABST
    Figure CN122594873A_ABST
Patent Text Reader

Abstract

The embodiment of the present application relates to the technical field of computer information security and artificial intelligence, and discloses a sensitive information automatic identification method combining a rule engine and a natural language processing technology and related devices, which comprises the following steps: receiving original data to be identified; performing rule matching on the original data to be identified to obtain a matching result; calculating a confidence score corresponding to the matching result; when the confidence score is lower than a first confidence threshold and higher than or equal to a second confidence threshold, inputting the matching result into a lightweight model for verification to obtain a first verification result; when the first verification result is consistent with the matching result, taking the matching result as a text information identification result; when the confidence score is lower than the second confidence threshold or the verification result of the lightweight model on the matching result is inconsistent, inputting the matching result into a large model for analysis and identification to obtain a large model identification result, and taking the large model identification result as the text information identification result. The embodiment of the present application realizes the balance between the text information identification efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network management technology, specifically to a text information hierarchical recognition method based on rule engine and LLM collaboration, a text information hierarchical recognition device based on rule engine and LLM collaboration, a computer device, and a computer-readable storage medium. Background Technology

[0002] With the deepening of digital transformation, the amount of data within enterprises and organizations is exploding, including a large amount of sensitive information such as personal identification information, trade secrets, financial data, and intellectual property. How to effectively identify, classify, and protect this sensitive information and prevent data breaches has become a core challenge in the field of information security.

[0003] Currently, mainstream sensitive information identification technologies are mainly divided into two categories: Category 1: Recognition methods based on rule engines.

[0004] This method identifies sensitive information by predefining a series of precise matching rules (such as regular expressions, keyword lists, data formats, etc.). Advantages: fast processing speed, high efficiency, low resource consumption, and strong interpretability of results (clearly knowing which rule matched). Disadvantages: poor flexibility, inability to understand contextual semantics, and a high risk of false positives and false negatives. For example, the rules cannot distinguish whether an 18-digit number is an ID card number or a product serial number, nor can they understand the semantic difference between "Please send a copy of my ID card to Zhang San" and "This invoice has an 18-digit serial number." Static rule bases are insufficient to cover new, variant, or deliberately disguised sensitive information.

[0005] The second category: recognition methods based on large language models.

[0006] This method leverages the powerful natural language understanding and reasoning capabilities of large language models to perform deep semantic analysis on text, thereby identifying complex and implicit sensitive information. Advantages: High recognition accuracy, ability to understand context, strong generalization ability, and effective handling of novel and variant sensitive information, reducing false positives and false negatives. Disadvantages: Extremely high computational cost, slow inference speed, and high consumption of GPU / CPU resources, making it difficult to meet the demands of processing massive amounts of real-time data. Furthermore, its decision-making process is similar to a "black box," resulting in poor interpretability.

[0007] Therefore, the existing methods have the following problems: 1. The trade-off between efficiency and accuracy: The two approaches in the background technology present a stark "binary opposition." Choosing a rule-based engine sacrifices recognition accuracy and adaptability; choosing a large language model sacrifices processing efficiency and cost-effectiveness. In real-world scenarios requiring simultaneous processing of massive amounts of data and high-precision recognition, existing technologies cannot provide an ideal balance. The fundamental reason is that existing technologies have failed to establish an intelligent decision-making mechanism to dynamically allocate the most appropriate computing resources based on the complexity and determinism of the content to be recognized.

[0008] 2. The Staticity and Lag of the System: Both the rule base of the rule engine and the initial knowledge of the large language model are relatively static. Once deployed, the system struggles to update and evolve itself. When new sensitive information leakage patterns emerge (such as new internet slang or novel data masquerading methods), the rule engine cannot identify them, and the large language model may fail because it has never encountered similar cases. The fundamental reason is that existing technologies lack an effective closed-loop feedback learning mechanism, failing to systematically feed new knowledge and erroneous experiences discovered during the identification process back into the system itself, thus hindering continuous capability enhancement.

[0009] 3. High operating costs: For scenarios requiring high-precision recognition, relying entirely on large language models results in continuous API call fees or the hardware and electricity costs of building a self-constructed cluster, which are prohibitive for most enterprises. This limits the widespread application of high-precision sensitive information recognition technology. The fundamental reason is that existing technologies have failed to achieve refined, on-demand scheduling and use of expensive large model resources. Summary of the Invention

[0010] In view of the above problems, embodiments of the present invention provide a text information hierarchical recognition method, related apparatus and computing device based on rule engine and LLM collaboration, which overcomes the above problems or at least partially solves the above problems.

[0011] According to one aspect of the present invention, a text information hierarchical recognition method based on the collaboration of a rule engine and an LLM is provided, the method comprising: Receive the raw data to be identified; According to the preset text information recognition rule base, the original data to be recognized is matched with the rules to obtain the matching result; The confidence score corresponding to the matching result is calculated according to the preset confidence algorithm; When the confidence score is lower than the first confidence threshold but higher than or equal to the second confidence threshold, the original data to be identified is input into the lightweight model for verification to obtain the first verification result. When the first verification result matches the matching result, the matching result is taken as the text information recognition result; When the confidence score is lower than the second confidence interval or the verification results of the lightweight model for the matching results are inconsistent, the original data to be identified is input into the large model for analysis and identification to obtain the large model identification result, and the large model identification result is used as the text information identification result.

[0012] In one optional approach, calculating the confidence score corresponding to the matching result according to a preset confidence algorithm includes: Obtain the matching results of the rule engine for the text to be identified, the matching results including the matched candidate sensitive information and the type identifier of the candidate sensitive information; Calculate the base confidence component, which is used to characterize the intrinsic credibility of the rule matching result; Calculate the context confidence component, which is used to characterize the semantic fit between the candidate sensitive information and the context in which the candidate sensitive information is located; Calculate the historical calibration confidence component, which is used to calibrate and adjust the current confidence based on historical identification data; The final comprehensive confidence level is calculated using a dynamic weighted fusion method based on the basic confidence level component, the context confidence level component, and the historical calibration confidence level component.

[0013] In an optional embodiment, after calculating the confidence score corresponding to the matching result according to a preset confidence algorithm, the method further includes: When the confidence score is higher than or equal to the first confidence threshold, the matching result is taken as the text information recognition result.

[0014] In an optional approach, when the confidence score is lower than the second confidence interval or the lightweight model's verification results for the matching results are inconsistent, the original data to be identified is input into a large model for analysis and identification to obtain the large model's identification result. After using the large model's identification result as the text information identification result, the method further includes: Monitor the recognition results of the large model; Based on the recognition results of the large model, when determining whether the large model has recognized information that does not exist in the text information recognition rule base, candidate rules are generated according to the reasoning process and recognition results of the large model.

[0015] In one embodiment, after inputting the original data to be identified into a lightweight model for verification and obtaining a first verification result when the confidence score is lower than a first confidence threshold but higher than or equal to a second confidence threshold, the method further includes: Statistically analyze the historical false alarm data and historical missed alarm data in the text information recognition results; Cluster analysis is performed on the historical false alarm data and historical false alarm data to extract feature information; The lightweight model is fine-tuned based on the feature information to obtain an optimized lightweight model.

[0016] In one embodiment, receiving the raw data to be identified includes: CPU and GPU resources can be scheduled using multithreading or an event loop. When the raw data to be identified flows in, the CPU prioritizes rule matching and routing decisions; When a large model needs to be invoked, the program uses the GPU to accelerate inference.

[0017] In one embodiment, the text information recognition rule base is a sensitive text information recognition rule base; the original data to be recognized is routed to a lightweight model or a large model through dynamic confidence routing.

[0018] Furthermore, embodiments of the present invention also provide a text information hierarchical recognition device based on the collaboration of a rule engine and LLM, the device comprising: The receiving module is used to receive the raw data to be identified; The rule matching module is used to perform rule matching on the original data to be identified based on a preset text information recognition rule base to obtain the matching result; The confidence calculation module is used to calculate the confidence score corresponding to the matching result according to a preset confidence algorithm; The lightweight model module is used to input the original data to be identified into the lightweight model for verification when the confidence score is lower than a first confidence threshold and higher than or equal to a second confidence threshold, to obtain a first verification result; when the first verification result is consistent with the matching result, the matching result is used as the text information recognition result. The large model module is used to input the original data to be identified into the large model for analysis and identification when the confidence score is lower than the second confidence interval or the verification results of the lightweight model for the matching results are inconsistent, so as to obtain the large model identification result and use the large model identification result as the text information identification result.

[0019] In one embodiment, the device further includes: The monitoring module is used to monitor the recognition results of the large model; The first determining module is used to determine whether the large model has identified information that does not exist in the text information recognition rule base based on the recognition result of the large model, and then generate candidate rules based on the reasoning process and recognition result of the large model.

[0020] This invention embodiment performs rule matching on the original data to be identified according to a preset text information recognition rule base to obtain a matching result; calculates a confidence score corresponding to the matching result according to a preset confidence algorithm; when the confidence score is lower than a first confidence threshold but higher than or equal to a second confidence threshold, inputs the original data to be identified into a lightweight model for verification to obtain a first verification result; when the first verification result is consistent with the matching result, the matching result is taken as the text information recognition result; when the confidence score is lower than a second confidence interval or the lightweight model fails to recognize the matching result... When the verification results are inconsistent, the original data to be identified is input into a large model for analysis and identification to obtain the large model identification result. This large model identification result is used as the text information identification result. This provides a highly efficient and accurate sensitive information identification solution: by introducing an intelligent routing mechanism, simple and clear identification tasks are handled by an efficient rule engine, while only complex and ambiguous identification tasks are handled by a large language model. This ensures that the overall identification accuracy is close to that of a pure LLM solution, while improving processing efficiency to a level close to that of a pure rule engine solution, and significantly reducing the cost of calling LLM. A dynamic identification system with self-learning and self-evolution capabilities is constructed: by establishing a two-way feedback channel between rules and models, the system can automatically learn new knowledge and correct old errors from daily operation, continuously enriching the rule base and optimizing model performance, thereby continuously improving the ability to identify new and variant sensitive information and solving the problem of poor adaptability of traditional static systems. The overall operating cost of sensitive information identification is significantly reduced: by minimizing reliance on expensive large language models, high-precision sensitive information identification technology can be deployed and applied on a large scale at a lower cost, improving the data security protection capabilities of enterprises.

[0021] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0022] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1A flowchart illustrating the information hierarchical identification method based on the collaboration of rule engine and LLM provided in an embodiment of the present invention is shown. Figure 2 A schematic diagram of the information hierarchical recognition process based on the collaboration of rule engine and LLM is shown in another embodiment of the present invention; Figure 3 This diagram illustrates the structure of an information hierarchical recognition device based on the collaboration of a rule engine and an LLM, as provided in an embodiment of the present invention. Figure 4 A schematic diagram of the structure of a computer device provided in an embodiment of the present invention is shown. Detailed Implementation

[0023] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein.

[0024] like Figure 1 As shown, this embodiment of the invention provides an information hierarchical recognition method based on the collaboration of a rule engine and LLM. This method is applied to computer devices, including desktop computers, tablets, smart terminals, wearable devices, distributed devices, etc., and this embodiment of the invention does not impose specific limitations. The method includes the following steps: Step 110: Receive the raw data to be identified.

[0025] In this embodiment of the invention, the original data to be identified is text data, such as text, documents, log streams, etc.; the content to be identified is sensitive information in the text data. For example, this embodiment of the invention takes a specific text recognition task as an example. The example text is: "Please send a copy of my ID card 230123XXXXXXX1234 to Zhang San, thank you." This text contains potentially sensitive information.

[0026] In this embodiment of the invention, the raw data to be identified is received through a network interface or file system, the data stream is received asynchronously, and the text is cached in memory.

[0027] Step 120: According to the preset text information recognition rule library, perform rule matching on the original data to be recognized to obtain the matching result.

[0028] The text information recognition rule base is a sensitive text information recognition rule base, used to perform rule matching on input data and output matching results. The matching results include information such as matching success / failure, the matched rule ID, the matched text fragment, and the matching degree. In this embodiment of the invention, the rules of the sensitive text information recognition rule base are stored in JSON or XML format, containing regular expressions, a keyword list, and rule metadata.

[0029] For example, the rule engine module scans text and applies ID number rules from the rule base (e.g., the regular expression \d{17}[\dXx] matches 18-digit numbers). In this embodiment, the rule successfully matches "230123XXXXXXX1234" and outputs the matching result (including rule ID, matching segment, and matching degree of 100%).

[0030] Step 130: Calculate the confidence score corresponding to the matching result according to the preset confidence algorithm.

[0031] In this embodiment of the invention, the preset confidence algorithm is a multi-dimensional confidence scoring model, which performs weighted calculations based on three dimensions: pattern completeness, contextual relevance, and historical accuracy, to obtain a confidence score. This confidence algorithm weights the scores from multiple dimensions to arrive at a comprehensive confidence score ranging from 0% to 100%.

[0032] Specifically, this is achieved in the following way: Step 1301: Obtain the matching result of the rule engine on the text to be identified, the matching result including the matched candidate sensitive information and the type identifier of the candidate sensitive information.

[0033] Step 1302: Calculate the basic confidence component, which is used to characterize the inherent credibility of the rule matching result.

[0034] First, calculate the confidence score of the pattern matching. The pattern matching confidence level is determined based on at least one of the following factors: Among them, the matching strength factor The match type is determined by the matching type. Match types include exact match, partial match, and fuzzy match.

[0035] Completeness factor It is determined based on the ratio of the matching length to the standard length.

[0036] Boundary quality factor The determination is based on whether the matching position has a word boundary.

[0037] The formula for calculating the pattern matching confidence level is as follows: ; in, The value range is [0, 1], with 1.0 for a complete match, 0.6 for a partial match, 0.3 for a fuzzy match, and 0 for no match; The value range is [0, 1]. When the matching length is within the preset standard length range, it is 1.0; otherwise, it is calculated according to the length deviation ratio. The value can be 0.5 or 1.0, with 1.0 for words and 0.5 otherwise.

[0038] Next, the format validation confidence score was calculated. The format verification confidence level is determined based on at least one of the following verification mechanisms: For ID card numbers: verify the validity of the region code, verify the reasonableness of the date of birth, and verify the correctness of the check digit; For mobile phone numbers: operator number segment verification, number length verification, and format standardization verification; For bank card numbers: Luhn algorithm verification, card number length verification, and bank card identification code (BIN) verification; For email addresses: domain validity verification, email format verification; The formula for calculating the confidence level of the format verification is as follows: ; in, For the first i The weighting coefficients of the item verification mechanism For the first i The pass / fail status of the item verification mechanism (1 for pass, 0 for fail).

[0039] Next, calculate the keyword matching confidence score. The keyword matching confidence level is determined based on the occurrence of preset keywords in the surrounding context of the candidate sensitive information.

[0040] The keywords are divided into three levels according to their indicative strength: strong indicative words, medium indicative words, and weak indicative words, and are assigned weight values ​​of 0.8-1.0, 0.4-0.7, and 0.1-0.3 respectively. The formula for calculating the keyword matching confidence score is as follows: ; in, The keyword influence coefficient ranges from [0.1, 0.2]. For the first i The weight of each keyword; For the first i The occurrence status of each keyword in the context (1 if it appears, 0 if it does not appear).

[0041] Calculate the base confidence level The base confidence level is the product of the three components mentioned above: ; Step 1303: Calculate the context confidence component. The context confidence component characterizes the semantic fit between the candidate sensitive information and its context.

[0042] First, calculate the semantic consistency confidence score. The semantic consistency confidence level is determined based on at least one of the following factors: The semantic distance factor is calculated based on the frequency of occurrence of semantic indicator words in the context; Detect whether there is a semantic pattern in the context that conflicts with this type of sensitive information, and calculate the semantic conflict factor; The topic relevance factor is calculated based on the frequency of occurrence of relevant terms in the context; The formula for calculating the semantic consistency confidence score is as follows: ; in, As a preset weighting coefficient, in one embodiment of the present invention, the values ​​can be 0.4, 0.3, and 0.3 respectively.

[0043] Calculate the confidence score of context relevance The context relevance confidence score is determined based on the relevance of text in different locations surrounding the candidate sensitive information.

[0044] The surrounding positions include: the immediate preceding position, the immediate following position, the preceding extended position (2-5 words before the matching position), and the following extended position (2-5 words after the matching position). The formula for calculating the context relevance confidence score is as follows: ; in, For the first i Relevance score of each location, For the first i The weight coefficients for each position are configured as follows: 0.4 for the immediate preceding position, 0.3 for the immediate following position, 0.2 for the preceding position, and 0.1 for the following position.

[0045] Calculate the confidence level of the reasonableness of the location distribution. The confidence level of the reasonableness of the location distribution is determined based on the location features of the candidate sensitive information in the document; The location features include: sentence location factor (calculated based on the sentence's sequence number in the document), paragraph location factor (calculated based on the type of the paragraph, including body text, list, table, and heading), and document type factor (calculated based on the document type, including form, article, code, etc.). The formula for calculating the confidence level of the rationality of the location distribution is as follows: ; Calculate context confidence The context confidence is a weighted sum of the three components mentioned above: ; in, As a preset weighting coefficient, in one embodiment of the present invention, the values ​​can be 0.4, 0.4, and 0.2 respectively.

[0046] Step 1304: Calculate the historical calibration confidence component, which is used to calibrate and adjust the current confidence based on historical identification data.

[0047] S4.1. Calculate the historical accuracy calibration factor The historical accuracy calibration factor is determined based on the deviation between the historical identification accuracy and the baseline accuracy of this type of sensitive information; The formula for calculating the historical accuracy calibration factor is as follows: ; in, The historical recognition accuracy of this type of sensitive information within a preset time window; This represents the baseline accuracy for this type of data. This is the calibration coefficient, with a value range of [0.2, 0.5].

[0048] S4.2. Calculation of type prior calibration factor The type of prior calibration factor is determined based on the prior distribution probability of the type of sensitive information in the training data or historical data; The formula for calculating the prior calibration factor of the aforementioned type is as follows: ; Where P(type) is the prior probability of the occurrence of this type of sensitive information; Pmax is the largest prior probability among all types.

[0049] S4.3. Calculate the aging attenuation calibration factor The time-degradation calibration factor is determined based on the interval between the update time of the rule base and the current time; The formula for calculating the aging attenuation calibration factor is as follows: ; in, This represents the time interval since the rule was last updated. This is the decay time constant, configured according to the type of sensitive information, with a value range of [90, 365] days; is the attenuation coefficient, with a value range of [0.1, 0.2].

[0050] S4.4. Calculate historical calibration confidence level The historical calibration confidence level is a weighted sum of the above three components: ; in, As a preset weighting coefficient, in one embodiment of the present invention, the values ​​can be 0.5, 0.3, and 0.2, respectively.

[0051] Step 1305: Calculate the final comprehensive confidence level using a dynamic weighted fusion method based on the basic confidence level component, the context confidence level component, and the historical calibration confidence level component.

[0052] ; in: W1, W2, and W3 are the weighting coefficients for the base confidence, contextual confidence, and historical calibration confidence, respectively, and their values ​​can be dynamically adjusted according to the type of sensitive information and the application scenario.

[0053] This is a confidence compensation factor used to adjust the confidence level in special cases.

[0054] The dynamic adjustment method for the weight coefficients W1, W2, and W3 includes: (1) Weight configuration based on sensitive information type: different weight configuration combinations are preset for different types of sensitive information; ID card number types: W1 ranges from [0.5, 0.7], W2 ranges from [0.15, 0.35], and W3 ranges from [0.10, 0.20]. Mobile phone number types: W1 has a value range of [0.45, 0.55], W2 has a value range of [0.30, 0.40], and W3 has a value range of [0.10, 0.20]. Email address types: W1 has a value range of [0.40, 0.50], W2 has a value range of [0.30, 0.40], and W3 has a value range of [0.15, 0.25]. Bank card number types: W1 has a value range of [0.65, 0.75], W2 has a value range of [0.10, 0.20], and W3 has a value range of [0.10, 0.20].

[0055] (2) Weight adjustment based on application scenario: The weights are adaptively adjusted according to the requirements of the application scenario; For high-precision scenarios: Increase the weights of W1 and W3, and decrease the weight of W2; In high-recall scenarios: increase the weight of W2 and appropriately decrease the weight of W1; In real-time processing scenarios: significantly increase the weight of W1 and decrease the weights of W2 and W3; In batch processing scenarios: appropriately reduce the weight of W1 and increase the weight of W3.

[0056] The judgment is made based on the preset three-level threshold (the first confidence threshold, also known as the high threshold H; the second confidence threshold, also known as the medium threshold M).

[0057] Specifically, when the confidence score is higher than or equal to the first confidence threshold, the matching result is used as the text information recognition result. In other cases, the original data to be recognized is routed to a lightweight model or a large model through dynamic confidence routing.

[0058] Step 140: When the confidence score is lower than the first confidence threshold but higher than or equal to the second confidence threshold, the original data to be identified is input into the lightweight model for verification to obtain the first verification result.

[0059] The lightweight model is a lightweight, dedicated validation model, such as a distillation model, a small neural network, or an efficient classifier. Its characteristics include extremely fast inference speed and computational cost far lower than large language models. It receives data from a medium-confidence path, performs rapid semantic verification, and returns the verification results (the results of the confirmation / correction / rejection rule engine) to the dynamic confidence routing, ultimately outputting the result. This approach aims to effectively improve the judgment accuracy of medium-difficulty tasks at a lower cost without significantly increasing costs.

[0060] Step 150: When the first verification result matches the matching result, the matching result is used as the text information recognition result.

[0061] When the first verification result is consistent with the matching result, it indicates that the lightweight model and the rule matching result are the same, the result is reliable, and therefore the matching result is output.

[0062] Step 160: When the confidence score is lower than the second confidence interval or the verification results of the lightweight model for the matching results are inconsistent, the original data to be identified is input into the large model for analysis and identification to obtain the large model identification result, and the large model identification result is used as the text information identification result.

[0063] The large model receives complex tasks from low-confidence paths, performs in-depth analysis using its powerful natural language understanding and reasoning capabilities, and outputs the final recognition results and classification.

[0064] For example: Suppose the input text is "code: 19801010abc123XYZ". The rule engine may match a date pattern but with low confidence. Rule engine matching: The rule matches "19801010" as a date format, but there are no explicit rules for sensitive information in the rule base, resulting in a low matching degree (e.g., 50%). Confidence calculation: Low pattern integrity score (40%), ambiguous context relevance (50%), average historical accuracy (70%), and a comprehensive score of only 53% < the mid-threshold TM (70%). Routing decision: Select the low-confidence path, routing the text and context to the large language model. LLM processing: LLM performs deep semantic analysis, identifying "19801010" as likely encoding rather than sensitive information. Combining this with the "code" context, it outputs the conclusion of "non-sensitive information".

[0065] In this embodiment of the invention, a bidirectional feedback enhancement mechanism is set up to improve the accuracy of both the large language model and the lightweight model. Specifically, the feedback channel from the large language model to the rules involves a method for structurally abstracting and extracting novel sensitive information patterns identified by the large language model that are not present in the rule base, and automatically generating candidate rules. This method includes parsing the LLM recognition results, pattern generalization, and the generation and submission review process of candidate rules. Specifically, the large model's recognition results are monitored; based on the large model's recognition results, if it determines whether the large model has identified information not present in the text information recognition rule base, candidate rules are generated based on the large model's reasoning process and recognition results.

[0066] The rule-to-model feedback channel involves clustering and analyzing high-frequency false positive cases generated by the rule engine to extract common features leading to these false positives (such as specific contexts and word collocations). These features are then transformed into structured negative knowledge or fine-tuning samples for continuous optimization of the large language model's knowledge base or for model fine-tuning. Specifically, historical false positive and false negative data in the text information recognition rule base are statistically analyzed; clustering analysis is performed on the historical false positive and false negative data to extract feature information; and the large model is fine-tuned based on this feature information to obtain an optimized large model.

[0067] The closed-loop system architecture, formed by the two feedback channels mentioned above, enables the rule base and model capabilities to promote each other, grow synergistically, and continuously adapt to new threats, thus achieving closed-loop self-evolution.

[0068] In this embodiment of the invention, to improve inference speed, CPU and GPU resources are scheduled through multi-threading or an event loop. When raw data to be identified flows in, the CPU prioritizes rule matching and routing decisions. When a large model needs to be invoked, the program calls the GPU to accelerate inference. The feedback module uses the CPU for batch data processing. The program achieves inter-module communication through memory sharing and message queues (such as Redis or Kafka) to ensure efficient parallelism.

[0069] This invention embodiment performs rule matching on the original data to be identified according to a preset text information recognition rule base to obtain a matching result; calculates a confidence score corresponding to the matching result according to a preset confidence algorithm; when the confidence score is lower than a first confidence threshold but higher than or equal to a second confidence threshold, inputs the original data to be identified into a lightweight model for verification to obtain a first verification result; when the first verification result is consistent with the matching result, the matching result is taken as the text information recognition result; when the confidence score is lower than a second confidence interval or the lightweight model fails to recognize the matching result... When the verification results are inconsistent, the original data to be identified is input into a large model for analysis and identification to obtain the large model identification result. This large model identification result is used as the text information identification result. This provides a highly efficient and accurate sensitive information identification solution: by introducing an intelligent routing mechanism, simple and clear identification tasks are handled by an efficient rule engine, while only complex and ambiguous identification tasks are handled by a large language model. This ensures that the overall identification accuracy is close to that of a pure LLM solution, while improving processing efficiency to a level close to that of a pure rule engine solution, and significantly reducing the cost of calling LLM. A dynamic identification system with self-learning and self-evolution capabilities is constructed: by establishing a two-way feedback channel between rules and models, the system can automatically learn new knowledge and correct old errors from daily operation, continuously enriching the rule base and optimizing model performance, thereby continuously improving the ability to identify new and variant sensitive information and solving the problem of poor adaptability of traditional static systems. The overall operating cost of sensitive information identification is significantly reduced: by minimizing reliance on expensive large language models, high-precision sensitive information identification technology can be deployed and applied on a large scale at a lower cost, improving the data security protection capabilities of enterprises.

[0070] like Figure 2As shown, a flowchart of a text information classification and recognition method based on the collaboration between a rule engine and an LLM according to another embodiment of the present invention is shown. Specifically, it includes the following steps: receiving the original data to be recognized; the rule engine performs rule matching on the original data to be recognized according to a preset text information recognition rule library to obtain a matching result; calculating a confidence score corresponding to the matching result according to a preset confidence algorithm. When the confidence score is higher than or equal to the first confidence threshold, the matching result is used as the text information recognition result. When the confidence score is lower than the first confidence threshold and higher than or equal to the second confidence threshold, the original data to be recognized is input into a lightweight model for verification to obtain a first verification result; when the first verification result is consistent with the matching result, the matching result is used as the text information recognition result; when the confidence score is lower than the second confidence interval or the verification result of the lightweight model for the matching result is inconsistent, the original data to be recognized is input into a large model for analysis and recognition to obtain a large model recognition result, and the large model recognition result is used as the text information recognition result. New patterns are extracted (LLM) according to the results of in-depth analysis and recognition by the large model, and after passing manual review, they are stored in the rule library to optimize the sensitive information recognition rule library. For the output recognition results, high-frequency misjudgment and false alarm cases are extracted through manual sampling to optimize the knowledge base or fine-tune the model of the lightweight model, thereby improving the performance of the lightweight model.

[0071] Among them, judgment is made according to a preset three-level threshold (the first confidence threshold, that is, the high threshold H; the second confidence threshold, that is, the medium threshold M): If the confidence score ≥ H (for example, 90%), the recognition result of the rule engine is directly output as the final result.

[0072] If M (for example, 70%) ≤ confidence score < H, the data is routed to the lightweight model for rapid secondary verification.

[0073] If the confidence score < M, it is determined that the complexity or ambiguity is high, and the data and complete context information are routed to the large language model module for in-depth analysis.

[0074] This invention embodiment performs rule matching on the original data to be identified according to a preset text information recognition rule base to obtain a matching result; calculates a confidence score corresponding to the matching result according to a preset confidence algorithm; when the confidence score is lower than a first confidence threshold but higher than or equal to a second confidence threshold, inputs the original data to be identified into a lightweight model for verification to obtain a first verification result; when the first verification result is consistent with the matching result, the matching result is taken as the text information recognition result; when the confidence score is lower than a second confidence interval or the lightweight model fails to recognize the matching result... When the verification results are inconsistent, the original data to be identified is input into a large model for analysis and identification to obtain the large model identification result. This large model identification result is used as the text information identification result. This provides a highly efficient and accurate sensitive information identification solution: by introducing an intelligent routing mechanism, simple and clear identification tasks are handled by an efficient rule engine, while only complex and ambiguous identification tasks are handled by a large language model. This ensures that the overall identification accuracy is close to that of a pure LLM solution, while improving processing efficiency to a level close to that of a pure rule engine solution, and significantly reducing the cost of calling LLM. A dynamic identification system with self-learning and self-evolution capabilities is constructed: by establishing a two-way feedback channel between rules and models, the system can automatically learn new knowledge and correct old errors from daily operation, continuously enriching the rule base and optimizing model performance, thereby continuously improving the ability to identify new and variant sensitive information and solving the problem of poor adaptability of traditional static systems. The overall operating cost of sensitive information identification is significantly reduced: by minimizing reliance on expensive large language models, high-precision sensitive information identification technology can be deployed and applied on a large scale at a lower cost, improving the data security protection capabilities of enterprises.

[0075] In one embodiment, for example, the confidence calculation for ID card number recognition: This embodiment uses ID card number recognition as an example to explain in detail the implementation process of the confidence calculation method.

[0076] Step S1: Obtain the matching results from the rule engine.

[0077] Suppose that the rule engine matches the candidate ID number "310101XXXXXXXX1234" with the type identifier "ID_CARD" in the text "Please provide your ID number: 310101XXXXXXXX1234 for real-name authentication".

[0078] Step S2: Calculate the base confidence components.

[0079] S2.1. Calculate the pattern matching confidence score Cpattern: Match strength factor: This match is a full regular expression match, Smatch = 1.0 Completeness factor: The standard length of an ID number is 18 characters, the matching length is 18 characters, Pfull = 1.0 Boundary quality factor: Matching positions with word boundaries such as spaces and colons, Qboundary = 1.0 The calculated value is: Cpattern = 1.0 × 1.0 × 1.0 = 1.0 S2.2. Calculate the format verification confidence score Cformat: Perform three checks: Area code validity check: The first 6 digits "310101" indicate Huangpu District, Shanghai, which is valid. W=0.3, V=1.

[0080] Birth date validity check: "XXXXXXXX" is XXXX year XX month XX day, which is valid, W=0.3, V=1.

[0081] Check bit correctness check: The last check bit is calculated correctly, W=0.4, V=1 The calculation yields: Cformat = (0.3×1 + 0.3×1 + 0.4×1) / (0.3+0.3+0.4) = 1.0 S2.3. Calculate the keyword matching confidence score Ckeyword: In the context "Please provide your ID number: ... for real-name authentication": "ID card" is a strong indicator with a weight of 1.0 and appears.

[0082] "Number" is a medium indicator word with a weight of 0.5, and it appears.

[0083] "Real-name authentication" is a strong indicator word with a weight of 0.9, and it appears.

[0084] Keyword weighted matching value = 1.0 + 0.5 + 0.9 = 2.4.

[0085] After normalization, the result is min(1.0, 2.4 / 3.0) = 0.8.

[0086] The calculation yields: C_keyword = 1 + 0.15 × 0.8 = 1.12.

[0087] S2.4. Calculate the basic confidence level: Cbase = 1.0 × 1.0 × 1.12 = 1.12.

[0088] Step S3: Calculate the context confidence component.

[0089] S3.1. Calculate the semantic consistency confidence score Csemantic: Semantic distance factor: When semantic indicator words such as "ID card" and "real-name authentication" appear in the context, F_semantic_distance = 0.9.

[0090] Semantic conflict factor: No conflicting words found, F_semantic_conflict = 1.0.

[0091] Topic relevance factor: The context topic is highly correlated with identity authentication, F_topic_relevance = 0.95.

[0092] The calculation yields: Csemantic = 0.4×0.9 + 0.3×1.0 + 0.3×0.95 = 0.945.

[0093] S3.2. Calculate the contextual relevance confidence score Ccontextual: Immediately preceding:: Correlation score 0.8.

[0094] Immediately following "used": Relevance score 0.7.

[0095] The preceding extension "Please provide your ID number": relevance score 0.9.

[0096] The subsequent expansion "for real-name authentication" has a relevance score of 0.85.

[0097] The calculation yields: C_contextual = 0.4×0.8 + 0.3×0.7 + 0.2×0.9 + 0.1×0.85 = 0.795.

[0098] S3.3. Calculate the confidence level C_position for the reasonableness of the location distribution: Sentence position factor: Located in the first sentence, P_sentence = 0.9.

[0099] Paragraph position factor: Located in the main text, P_paragraph = 1.0.

[0100] Document type factor: Form type document, P_document = 1.0.

[0101] The calculation yields: Cposition = 0.9 × 1.0 × 1.0 = 0.9.

[0102] S3.4. Calculate contextual confidence: Ccontext = 0.4×0.945 + 0.4×0.795 + 0.2×0.9 = 0.876.

[0103] Step S4: Calculate the historical calibration confidence components S4.1. Calculate the historical accuracy calibration factor: Assume the historical recognition accuracy rate for ID card number types is 0.95, and the baseline accuracy rate is 0.90.

[0104] β is set to 0.3. The calculated value is: Caccuracy = 1 + 0.3 × (0.95 - 0.90) = 1.015.

[0105] S4.2. Calculation of prior calibration factor: Assume that the prior probability of an ID number in the dataset is 0.30, and the maximum prior probability is 0.35.

[0106] The calculation yields: C_prior = √(0.30 / 0.35) = 0.926.

[0107] S4.3. Calculate the aging attenuation calibration factor: Assume the ID number rule was last updated 60 days ago, τ=180, γ=0.15.

[0108] The calculation yields: C_decay = 1 - 0.15 × (1 - e (-60 / 180) = 0.972.

[0109] S4.4. Calculate the historical calibration confidence level: Ccalibration = 0.5×1.015 + 0.3×0.926 + 0.2×0.972 = 0.985.

[0110] Step S5: Calculate the final overall confidence level.

[0111] For ID card number types, the following weighted configurations are used: W1=0.60, W2=0.25, W3=0.15 Without special compensation, Cboost = 0.

[0112] The calculation yields: Cfinal = 0.60×1.12 + 0.25×0.876 + 0.15×0.985 + 0= 0.672 + 0.219 +0.148= 1.039.

[0113] With the upper limit set to 1.0, we get: C_final = 1.0.

[0114] Step S6: Routing Decision The threshold values ​​used in the balanced mode are: Thigh=0.82, Tlow=0.45.

[0115] Since Cfinal=1.0 ≥ Thigh=0.82, it is determined to be of high confidence, and the result of the rule engine is directly used without calling LLM verification.

[0116] In another embodiment, confidence calculation for mobile phone number identification (low confidence case): This embodiment illustrates the confidence calculation and routing decision process in the low confidence case.

[0117] Step S1: Obtain the matching results from the rule engine.

[0118] Suppose the rule engine matches the candidate phone number "4001234567" with the type identifier "PHONE" in the text "The customer service phone number is 4001234567. Please contact us if you have any questions".

[0119] Steps S2-S5: Confidence Calculation: Calculations show that: Cpattern = 0.6 (partial match); Cformat = 0.7 (failed carrier number segment verification); Ckeyword = 1.05 (weak keyword); Cbase = 0.6 × 0.7 × 1.05 = 0.441; Csemantic = 0.6 (semantic relevance is average); Ccontextual = 0.5 (weak contextual relevance); Cposition = 0.8; Ccontext = 0.4×0.6 + 0.4×0.5 + 0.2×0.8 = 0.6; Caccuracy = 0.95 (historical accuracy); Cprior = 0.85 (low prior probability); Cdecay = 0.98; Ccalibration = 0.5×0.95 + 0.3×0.85 + 0.2×0.98 = 0.926; For mobile phone number types, the weighting configuration is: W1=0.50, W2=0.35, W3=0.15; The calculation yields: Cfinal = 0.50×0.441 + 0.35×0.6 + 0.15×0.926 = 0.221 + 0.21 + 0.139 = 0.57。

[0120] Step S6: Routing decision Since Tlow = 0.45 ≤ Cfinal = 0.57 < Thigh = 0.75, it is determined to be medium confidence, and the LLM needs to be called for verification.

[0121] In another specific implementation, parameter adaptive optimization is also performed. This embodiment illustrates the process of parameter adaptive optimization based on feedback data.

[0122] Step S7.1: Collect feedback data.

[0123] Suppose 1000 identified feedback data have been collected, including: confidence components (Cbase, Ccontext, Ccalibration), final confidence (Cfinal), decision results, LLM verification results, and manual annotation results.

[0124] Step S7.2: Update historical statistics. Calculate to obtain: Historical accuracy: 0.95 (increased by 0.02); Mean confidence of true positive samples: 0.88, standard deviation: 0.08; Mean confidence of false positive samples: 0.52, standard deviation: 0.15.

[0125] Step S7.3: Optimize weight coefficients.

[0126] Taking the maximization of the F1 score as the objective function, use the grid search method to search for the optimal solution in the weight space: Search range of W1: [0.5, 0.7], step size 0.02; Search range of W2: [0.15, 0.35], step size 0.02; Search range of W3: [0.10, 0.20], step size 0.01.

[0127] The optimal weight configuration is obtained by searching: W1 = 0.62 (original 0.60); W2 = 0.24 (original 0.25); W3 = 0.14 (original 0.15).

[0128] The F1 score after optimization is increased from 0.91 to 0.93.

[0129] Step S7.4: Optimize threshold parameters.

[0130] Calculate the new threshold based on the quantile method: Thigh = 10th percentile of confidence level for true positive samples = 0.84; Tlow = 75th percentile of confidence level for false positive samples = 0.48.

[0131] Update threshold configuration: Thigh: 0.82 → 0.84; Tlow: 0.45 → 0.48.

[0132] Figure 3 A schematic diagram of the structure of a text information hierarchical recognition device based on rule engine and LLM collaboration according to an embodiment of the present invention is shown. Figure 3 As shown, this text information hierarchical recognition device based on rule engine and LLM collaboration is applied to a computer device and includes: a receiving module 310, a rule matching module 320, a confidence calculation module 330, a lightweight model module 340, and a large model module 350. Wherein: The receiving module 310 is used to receive the raw data to be identified.

[0133] The rule matching module 320 is used to perform rule matching on the original data to be identified according to a preset text information recognition rule library to obtain a matching result.

[0134] The confidence calculation module 330 is used to calculate the confidence score corresponding to the matching result according to a preset confidence algorithm.

[0135] The lightweight model module 340 is used to input the original data to be identified into the lightweight model for verification when the confidence score is lower than a first confidence threshold and higher than or equal to a second confidence threshold, to obtain a first verification result; when the first verification result is consistent with the matching result, the matching result is used as the text information recognition result.

[0136] The large model module 350 is used to input the original data to be identified into the large model for analysis and identification when the confidence score is lower than the second confidence interval or the verification results of the lightweight model for the matching results are inconsistent, so as to obtain the large model identification result and use the large model identification result as the text information identification result.

[0137] The device further includes: The monitoring module is used to monitor the recognition results of the large model; The first determining module is used to determine whether the large model has identified information that does not exist in the text information recognition rule base based on the recognition result of the large model, and then generate candidate rules based on the reasoning process and recognition result of the large model.

[0138] The specific working process of this device is largely the same as the specific execution steps of the aforementioned method, and will not be repeated here.

[0139] This invention embodiment performs rule matching on the original data to be identified according to a preset text information recognition rule base to obtain a matching result; calculates a confidence score corresponding to the matching result according to a preset confidence algorithm; when the confidence score is lower than a first confidence threshold but higher than or equal to a second confidence threshold, inputs the original data to be identified into a lightweight model for verification to obtain a first verification result; when the first verification result is consistent with the matching result, the matching result is taken as the text information recognition result; when the confidence score is lower than a second confidence interval or the lightweight model fails to recognize the matching result... When the verification results are inconsistent, the original data to be identified is input into a large model for analysis and identification to obtain the large model identification result. This large model identification result is used as the text information identification result. This provides a highly efficient and accurate sensitive information identification solution: by introducing an intelligent routing mechanism, simple and clear identification tasks are handled by an efficient rule engine, while only complex and ambiguous identification tasks are handled by a large language model. This ensures that the overall identification accuracy is close to that of a pure LLM solution, while improving processing efficiency to a level close to that of a pure rule engine solution, and significantly reducing the cost of calling LLM. A dynamic identification system with self-learning and self-evolution capabilities is constructed: by establishing a two-way feedback channel between rules and models, the system can automatically learn new knowledge and correct old errors from daily operation, continuously enriching the rule base and optimizing model performance, thereby continuously improving the ability to identify new and variant sensitive information and solving the problem of poor adaptability of traditional static systems. The overall operating cost of sensitive information identification is significantly reduced: by minimizing reliance on expensive large language models, high-precision sensitive information identification technology can be deployed and applied on a large scale at a lower cost, improving the data security protection capabilities of enterprises.

[0140] Figure 4 The diagram shows a schematic of the structure of a computing device provided in an embodiment of the present invention. The specific embodiments of the present invention do not limit the specific implementation of the device.

[0141] like Figure 4 As shown, the computing device may include: a processor 402, a communications interface 404, a memory 406, and a communications bus 408.

[0142] The processor 402, communication interface 404, and memory 406 communicate with each other via communication bus 408. Communication interface 404 is used to communicate with other network elements such as clients or other servers. The processor 402 executes program 410, specifically performing the relevant steps in the above-described embodiment of the text information hierarchical recognition method based on rule engine and LLM collaboration.

[0143] Specifically, program 410 may include program code that includes computer operation instructions.

[0144] Processor 402 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The device may include one or more processors of the same type, such as one or more CPUs; or it may include processors of different types, such as one or more CPUs and one or more ASICs.

[0145] Memory 406 is used to store program 410. Memory 406 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0146] Specifically, program 410 can be used to cause processor 402 to perform the following operations: Receive the raw data to be identified; According to the preset text information recognition rule base, the original data to be recognized is matched with the rules to obtain the matching result; The confidence score corresponding to the matching result is calculated according to the preset confidence algorithm; When the confidence score is lower than the first confidence threshold but higher than or equal to the second confidence threshold, the original data to be identified is input into the lightweight model for verification to obtain the first verification result. When the first verification result matches the matching result, the matching result is taken as the text information recognition result; When the confidence score is lower than the second confidence interval or the verification results of the lightweight model for the matching results are inconsistent, the original data to be identified is input into the large model for analysis and identification to obtain the large model identification result, and the large model identification result is used as the text information identification result.

[0147] In an optional embodiment, after calculating the confidence score corresponding to the matching result according to a preset confidence algorithm, the method further includes: When the confidence score is higher than or equal to the first confidence threshold, the matching result is taken as the text information recognition result.

[0148] In an optional approach, when the confidence score is lower than the second confidence interval or the lightweight model's verification results for the matching results are inconsistent, the original data to be identified is input into a large model for analysis and identification to obtain the large model's identification result. After using the large model's identification result as the text information identification result, the method further includes: Monitor the recognition results of the large model; Based on the recognition results of the large model, when determining whether the large model has recognized information that does not exist in the text information recognition rule base, candidate rules are generated according to the reasoning process and recognition results of the large model.

[0149] In one embodiment, after inputting the original data to be identified into a lightweight model for verification and obtaining a first verification result when the confidence score is lower than a first confidence threshold but higher than or equal to a second confidence threshold, the method further includes: Statistically analyze the historical false alarm data and historical false alarm data in the text information recognition rule base; Cluster analysis is performed on the historical false alarm data and historical false alarm data to extract feature information; The large model is fine-tuned based on the feature information to obtain an optimized large model.

[0150] In one embodiment, receiving the raw data to be identified includes: CPU and GPU resources can be scheduled using multithreading or an event loop. When the raw data to be identified flows in, the CPU prioritizes rule matching and routing decisions; When a large model needs to be invoked, the program uses the GPU to accelerate inference.

[0151] In one embodiment, the text information recognition rule base is a sensitive text information recognition rule base; the original data to be recognized is routed to a lightweight model or a large model through dynamic confidence routing.

[0152] This invention embodiment performs rule matching on the original data to be identified according to a preset text information recognition rule base to obtain a matching result; calculates a confidence score corresponding to the matching result according to a preset confidence algorithm; when the confidence score is lower than a first confidence threshold but higher than or equal to a second confidence threshold, inputs the original data to be identified into a lightweight model for verification to obtain a first verification result; when the first verification result is consistent with the matching result, the matching result is taken as the text information recognition result; when the confidence score is lower than a second confidence interval or the lightweight model fails to recognize the matching result... When the verification results are inconsistent, the original data to be identified is input into a large model for analysis and identification to obtain the large model identification result. This large model identification result is used as the text information identification result. This provides a highly efficient and accurate sensitive information identification solution: by introducing an intelligent routing mechanism, simple and clear identification tasks are handled by an efficient rule engine, while only complex and ambiguous identification tasks are handled by a large language model. This ensures that the overall identification accuracy is close to that of a pure LLM solution, while improving processing efficiency to a level close to that of a pure rule engine solution, and significantly reducing the cost of calling LLM. A dynamic identification system with self-learning and self-evolution capabilities is constructed: by establishing a two-way feedback channel between rules and models, the system can automatically learn new knowledge and correct old errors from daily operation, continuously enriching the rule base and optimizing model performance, thereby continuously improving the ability to identify new and variant sensitive information and solving the problem of poor adaptability of traditional static systems. The overall operating cost of sensitive information identification is significantly reduced: by minimizing reliance on expensive large language models, high-precision sensitive information identification technology can be deployed and applied on a large scale at a lower cost, improving the data security protection capabilities of enterprises.

[0153] This invention provides a non-volatile computer storage medium storing at least one executable instruction that can execute the text information hierarchical recognition method based on rule engine and LLM collaboration in any of the above method embodiments.

[0154] Executable instructions can specifically be used to cause the processor to perform the following operations: Receive the raw data to be identified; According to the preset text information recognition rule base, the original data to be recognized is matched with the rules to obtain the matching result; The confidence score corresponding to the matching result is calculated according to the preset confidence algorithm; When the confidence score is lower than the first confidence threshold but higher than or equal to the second confidence threshold, the original data to be identified is input into the lightweight model for verification to obtain the first verification result. When the first verification result matches the matching result, the matching result is taken as the text information recognition result; When the confidence score is lower than the second confidence interval or the verification results of the lightweight model for the matching results are inconsistent, the original data to be identified is input into the large model for analysis and identification to obtain the large model identification result, and the large model identification result is used as the text information identification result.

[0155] In an optional embodiment, after calculating the confidence score corresponding to the matching result according to a preset confidence algorithm, the method further includes: When the confidence score is higher than or equal to the first confidence threshold, the matching result is taken as the text information recognition result.

[0156] In an optional approach, when the confidence score is lower than the second confidence interval or the lightweight model's verification results for the matching results are inconsistent, the original data to be identified is input into a large model for analysis and identification to obtain the large model's identification result. After using the large model's identification result as the text information identification result, the method further includes: Monitor the recognition results of the large model; Based on the recognition results of the large model, when determining whether the large model has recognized information that does not exist in the text information recognition rule base, candidate rules are generated according to the reasoning process and recognition results of the large model.

[0157] In one embodiment, after inputting the original data to be identified into a lightweight model for verification and obtaining a first verification result when the confidence score is lower than a first confidence threshold but higher than or equal to a second confidence threshold, the method further includes: Statistically analyze the historical false alarm data and historical false alarm data in the text information recognition rule base; Cluster analysis is performed on the historical false alarm data and historical false alarm data to extract feature information; The large model is fine-tuned based on the feature information to obtain an optimized large model.

[0158] In one embodiment, receiving the raw data to be identified includes: CPU and GPU resources can be scheduled using multithreading or an event loop. When the raw data to be identified flows in, the CPU prioritizes rule matching and routing decisions; When a large model needs to be invoked, the program uses the GPU to accelerate inference.

[0159] In one embodiment, the text information recognition rule base is a sensitive text information recognition rule base; the original data to be recognized is routed to a lightweight model or a large model through dynamic confidence routing.

[0160] This invention embodiment performs rule matching on the original data to be identified according to a preset text information recognition rule base to obtain a matching result; calculates a confidence score corresponding to the matching result according to a preset confidence algorithm; when the confidence score is lower than a first confidence threshold but higher than or equal to a second confidence threshold, inputs the original data to be identified into a lightweight model for verification to obtain a first verification result; when the first verification result is consistent with the matching result, the matching result is taken as the text information recognition result; when the confidence score is lower than a second confidence interval or the lightweight model fails to recognize the matching result... When the verification results are inconsistent, the original data to be identified is input into a large model for analysis and identification to obtain the large model identification result. This large model identification result is used as the text information identification result. This provides a highly efficient and accurate sensitive information identification solution: by introducing an intelligent routing mechanism, simple and clear identification tasks are handled by an efficient rule engine, while only complex and ambiguous identification tasks are handled by a large language model. This ensures that the overall identification accuracy is close to that of a pure LLM solution, while improving processing efficiency to a level close to that of a pure rule engine solution, and significantly reducing the cost of calling LLM. A dynamic identification system with self-learning and self-evolution capabilities is constructed: by establishing a two-way feedback channel between rules and models, the system can automatically learn new knowledge and correct old errors from daily operation, continuously enriching the rule base and optimizing model performance, thereby continuously improving the ability to identify new and variant sensitive information and solving the problem of poor adaptability of traditional static systems. The overall operating cost of sensitive information identification is significantly reduced: by minimizing reliance on expensive large language models, high-precision sensitive information identification technology can be deployed and applied on a large scale at a lower cost, improving the data security protection capabilities of enterprises.

[0161] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The required structure for constructing such systems is apparent from the above description. Furthermore, the embodiments of the present invention are not directed to any particular programming language. It should be understood that the content of the invention described herein can be implemented using various programming languages, and the above description of specific languages ​​is for the purpose of disclosing the best mode of implementation of the invention.

[0162] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0163] Similarly, it should be understood that, in order to streamline the invention and aid in understanding one or more of the various aspects of the invention, features of the embodiments of the invention are sometimes grouped together in a single embodiment, figure, or description thereof in the above description of exemplary embodiments of the invention. However, this disclosure should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim.

[0164] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from that embodiment. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and can be divided into multiple sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.

[0165] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.

Claims

1. A text information hierarchical recognition method based on the collaboration of rule engine and LLM, characterized in that, The method includes: Receive the raw data to be identified; According to the preset text information recognition rule base, the original data to be recognized is matched with the rules to obtain the matching result; The confidence score corresponding to the matching result is calculated according to the preset confidence algorithm; When the confidence score is lower than the first confidence threshold but higher than or equal to the second confidence threshold, the original data to be identified is input into the lightweight model for verification to obtain the first verification result. When the first verification result matches the matching result, the matching result is taken as the text information recognition result; When the confidence score is lower than the second confidence interval or the verification results of the lightweight model for the matching results are inconsistent, the original data to be identified is input into the large model for analysis and identification to obtain the large model identification result, and the large model identification result is used as the text information identification result.

2. The method according to claim 1, characterized in that, The step of calculating the confidence score corresponding to the matching result according to the preset confidence algorithm includes: Obtain the matching results of the rule engine for the text to be identified, the matching results including the matched candidate sensitive information and the type identifier of the candidate sensitive information; Calculate the base confidence component, which is used to characterize the intrinsic credibility of the rule matching result; Calculate the context confidence component, which is used to characterize the semantic fit between the candidate sensitive information and the context in which the candidate sensitive information is located; Calculate the historical calibration confidence component, which is used to calibrate and adjust the current confidence based on historical identification data; The final comprehensive confidence level is calculated using a dynamic weighted fusion method based on the basic confidence level component, the context confidence level component, and the historical calibration confidence level component.

3. The method according to claim 1, characterized in that, After calculating the confidence score corresponding to the matching result according to the preset confidence algorithm, the method further includes: When the confidence score is higher than or equal to the first confidence threshold, the matching result is taken as the text information recognition result.

4. The method according to claim 1, characterized in that, When the confidence score is lower than the second confidence interval or the verification results of the lightweight model for the matching results are inconsistent, the original data to be identified is input into the large model for analysis and identification to obtain the large model identification result. After using the large model identification result as the text information identification result, the method further includes: Monitor the recognition results of the large model; Based on the recognition results of the large model, when determining whether the large model has recognized information that does not exist in the text information recognition rule base, candidate rules are generated according to the reasoning process and recognition results of the large model.

5. The method according to claim 1, characterized in that, When the confidence score is lower than a first confidence threshold but higher than or equal to a second confidence threshold, the original data to be identified is input into a lightweight model for verification to obtain a first verification result. The method further includes: Statistically analyze the historical false alarm data and historical missed alarm data in the text information recognition results; Cluster analysis is performed on the historical false alarm data and historical false alarm data to extract feature information; The lightweight model is fine-tuned based on the feature information to obtain an optimized lightweight model.

6. The method according to claim 4, characterized in that, The receiving of the raw data to be identified includes: CPU and GPU resources can be scheduled using multithreading or an event loop. When the raw data to be identified flows in, the CPU prioritizes rule matching and routing decisions; When a large model needs to be invoked, the program uses the GPU to accelerate inference.

7. The method according to any one of claims 1-6, characterized in that, The text information recognition rule base is a sensitive text information recognition rule base; through dynamic confidence routing, the original data to be recognized is routed to a lightweight model or a large model.

8. A text information hierarchical recognition device based on rule engine and LLM collaboration, characterized in that, The device includes: The receiving module is used to receive the raw data to be identified; The rule matching module is used to perform rule matching on the original data to be identified based on a preset text information recognition rule base to obtain the matching result; The confidence calculation module is used to calculate the confidence score corresponding to the matching result according to a preset confidence algorithm; The lightweight model module is used to input the original data to be identified into the lightweight model for verification when the confidence score is lower than a first confidence threshold and higher than or equal to a second confidence threshold, to obtain a first verification result; when the first verification result is consistent with the matching result, the matching result is used as the text information recognition result. The large model module is used to input the original data to be identified into the large model for analysis and identification when the confidence score is lower than the second confidence interval or the verification results of the lightweight model for the matching results are inconsistent, so as to obtain the large model identification result and use the large model identification result as the text information identification result.

9. A computer device, characterized in that, include: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction, which causes the processor to perform the steps of the text information hierarchical recognition method based on rule engine and LLM collaboration according to any one of claims 1-7.

10. A computer storage medium storing at least one executable instruction that causes a processor to perform the steps of the text information hierarchical recognition method based on rule engine and LLM collaboration according to any one of claims 1-7.