Method and system for identifying finance and tax question and answer sensitive information based on large model
By applying a large-model-based financial and taxation Q&A sensitive information recognition method in the field of finance and taxation, combined with the semantic detection of Chinese pinyin sensitive word Trie tree and the fiscal and taxation model, the problem of not being able to identify financial and tax violation-oriented Q&A in the existing technology is solved, and the effect of efficient identification and processing of financial and tax sensitive information is achieved.
Patent Information
- Application Number
- CN202411928227.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-27
AI Technical Summary
The existing technology does not have a big model-based identification method for fiscal and tax violation-oriented Q&A, and cannot effectively identify and process sensitive information in the fiscal and taxation field.
The method of identifying sensitive information of financial and taxation questions and answers based on the big model is used to process the financial and taxation question data, and the Chinese pinyin sensitive word Trie tree is used to detect key prohibited words. Combined with the illegal semantic detection of the trained financial and taxation model, the illegal sensitive words are identified and processed.
It has achieved rapid review and identification of sensitive information in fiscal and taxation questions and answers, significantly improved the security of the big model, reduced the need for manual review, reduced labor costs, and helped maintain the healthy and orderly development of the cyberspace.
Smart Images

Figure CN120046608A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information technology applications, and more specifically, to a method and system for identifying sensitive information in fiscal and tax Q&A based on large models. Background Art
[0002] With the rapid development of large model technology, its applications in various fields are becoming increasingly widespread, from scientific research to business, and then to all aspects of daily life and office work. However, a series of potential security risks follow. The triggering and handling of these risks not only concern the reputation of enterprises, but also involve the protection of personal privacy and social stability. Therefore, it is crucial to deeply understand and address these security risks.
[0003] Large models process a large amount of sensitive data and personal information in many application scenarios, such as users' search records, social media interactions, and financial transactions. This makes the risks of data leakage and privacy infringement cannot be ignored. Once these sensitive information is leaked, the personal privacy rights and interests may be severely damaged and even be used for malicious behaviors, such as identity theft, fraud, and social engineering attacks. This will not only cause economic losses to the victims, but also may lead to social panic and distrust. Secondly, the powerful capabilities of large models may also be used for various forms of malicious attacks. The adversarial sample attack on the model, that is, making small changes to the input of the model to deceive the model into making wrong predictions, has become a common threat. Malicious users can use this method to create false information and affect the decision-making results, such as spreading misleading information to social media platforms, thus disrupting social order. In addition, the generation ability of large models may also be used to generate false content, threatening the credibility of the media and the authenticity of news.
[0004] Summary: Through patent query, patents similar to the present invention are Patent No. CN118520116A, CN118504039A, and CN118536506A.
[0005] Prior Art 1, CN118520116A proposes an information extraction method based on large models. Prior Art 1 identifies sensitive information in the target text through regular expressions of sensitive words in the rule library and the model. The difference is that this patent only detects prohibited words through regular matching and training a binary classification model. Prior Art 2, CN118504039A proposes a method, system, and all-in-one machine for desensitizing file information based on AIGC. Prior Art 2 all uses large language models to process sensitive information and locates prohibited information in the text by means of keyword query. Prior Art 3, the method proposed by CN118536506A is for image, voice, and text data in videos, and obtains each prohibited word in the prohibited word library for keyword matching to complete the detection.
[0006] Currently, there is no method for identifying questions and answers oriented towards fiscal and tax violations based on large models in the existing technology. Summary of the Invention
[0007] The technical solution of the present invention provides a method and system for identifying sensitive information in fiscal and tax questions and answers based on a large model to solve the problem of how to identify questions and answers oriented towards fiscal and tax violations based on a large model.
[0008] To solve the above problems, the present invention provides a method for identifying sensitive information in fiscal and tax questions and answers based on a large model, and the method includes:
[0009] Obtain fiscal and tax question data, and process the obtained fiscal and tax question data;
[0010] For the processed fiscal and tax question data, perform key prohibited word detection through the established Chinese pinyin sensitive word Trie tree;
[0011] When it is determined that there are no key prohibited words in the fiscal and tax question data, detect the prohibited sensitive words through the prohibited semantics in the trained fiscal and tax large model;
[0012] When no prohibited sensitive words are detected in the output of the fiscal and tax large model, output the fiscal and tax question data to the normal question and answer system.
[0013] Preferably, it further includes: when it is determined that there are key prohibited words in the fiscal and tax question data, based on the label set for the fiscal and tax question data, return the corresponding answer template of the fiscal and tax question data.
[0014] Preferably, it further includes: when the fiscal and tax large model detects prohibited sensitive words, based on the label set for the fiscal and tax question data, return the corresponding answer template of the fiscal and tax question data.
[0015] Preferably, the processing of the obtained fiscal and tax question data includes:
[0016] When the fiscal and tax question data is of the ultra-long text type, add a batch_size parameter to the fiscal and tax question data, set the adaptive context window of the fiscal and tax question data, and set the text segmentation point based on the specified symbol.
[0017] Preferably, the processing of the obtained fiscal and tax question data includes:
[0018] Judge whether there are deformed sensitive words in the fiscal and tax question data. When it is determined that there are deformed sensitive words in the fiscal and tax question data, process the deformed sensitive words, including: converting traditional Chinese to simplified Chinese, converting case, removing special symbols, and converting to Chinese pinyin.
[0019] Preferably, it also includes establishing a Trie tree of Chinese Pinyin sensitive words:
[0020] Determine 23 Chinese pinyin initials after removing "u", "v", and "i", and mark the 23 Chinese pinyin initials as nodes 0 to 22;
[0021] The nodes 0 to 22 are used to store the Chinese sensitive words in the initials of the pinyin respectively;
[0022] The node 23 will be created to store sensitive words starting with numbers, non-Chinese characters, and English characters;
[0023] Set a flag for nodes 0 to 23 to indicate whether they are terminal nodes.
[0024] Preferably, it also includes training the big finance and taxation model:
[0025] Acquire financial data, pre-label the financial data, and establish a fine-tuning data set based on the financial data and the pre-labeled labels;
[0026] The finance and taxation big model is trained by the fine-tuning data set until the finance and taxation big model reaches a predetermined standard for the detection rate of the illegal sensitive words in the fine-tuning data set.
[0027] According to another aspect of the present invention, the present invention provides a financial and taxation question and answer sensitive information identification system based on a large model, the system comprising:
[0028] An initial unit, used to obtain financial and taxation question data and process the obtained financial and taxation question data;
[0029] A first detection unit is used to detect key prohibited words on the processed financial and taxation question data through an established Chinese Pinyin sensitive word Trie tree;
[0030] The second detection unit is used to detect the illegal sensitive words through the illegal semantics in the trained financial and taxation big model when it is determined that the financial and taxation question data does not contain key banned words;
[0031] The result unit is used to output the financial and taxation question data to a normal question-and-answer system when no illegal sensitive words are detected in the output of the financial and taxation big model.
[0032] Preferably, the first detection unit is further used for: when it is determined that the financial and taxation question data contains key banned words, based on the labels set for the financial and taxation question data, returning the answer template corresponding to the financial and taxation question data.
[0033] Preferably, the second detection unit is further configured to: when the fiscal and tax large model detects a violation sensitive word, return the answer template corresponding to the fiscal and tax question data based on the label set for the fiscal and tax question data.
[0034] Preferably, the initial unit is configured to process the obtained fiscal and tax question data and is further configured to:
[0035] When the fiscal and tax question data is of the ultra-long text type, add a batch_size parameter to the fiscal and tax question data, set the adaptive context window of the fiscal and tax question data, and set the text segmentation point based on the specified symbol.
[0036] Preferably, the initial unit is configured to process the obtained fiscal and tax question data and is further configured to:
[0037] Determine whether there are deformed sensitive words in the fiscal and tax question data. When it is determined that there are deformed sensitive words in the fiscal and tax question data, process the deformed sensitive words, including: converting traditional Chinese to simplified Chinese, converting case, removing special symbols, and converting to Chinese pinyin.
[0038] Preferably, the first detection unit is further configured to: establish a Chinese pinyin sensitive word Trie tree:
[0039] Determine 23 initial letters of Chinese pinyin after removing "u", "v", and "i", and mark the 23 initial letters of Chinese pinyin as nodes numbered 0 to 22;
[0040] Store the Chinese sensitive words in the initial letter pinyin through the nodes numbered 0 to 22;
[0041] Store the sensitive words starting with numbers, non-Chinese, and English characters in the established node numbered 23;
[0042] Set the mark of whether the nodes numbered 0 to 23 are terminal nodes.
[0043] Preferably, the initial unit further includes training the fiscal and tax large model:
[0044] Obtain financial data, pre-annotate the financial data, and establish a fine-tuning data set based on the financial data and the pre-annotated labels;
[0045] Train the fiscal and tax large model through the fine-tuning data set until the detection rate of the violation sensitive words in the fine-tuning data set by the fiscal and tax large model reaches the predetermined standard.
[0046] The technical solution of the present invention provides a method and system for identifying sensitive information in fiscal and tax Q&A based on a large model. The method includes: obtaining fiscal and tax question data and processing the obtained fiscal and tax question data; for the processed fiscal and tax question data, detecting key prohibited words through the established Chinese pinyin sensitive word Trie tree; when it is determined that there are no key prohibited words in the fiscal and tax question data, detecting prohibited sensitive words through the prohibited semantics in the trained fiscal and tax large model; when no prohibited sensitive words are detected in the output of the fiscal and tax large model, outputting the fiscal and tax question data to the normal Q&A system. The technical solution of the present invention proposes a method and system for identifying sensitive information in fiscal and tax Q&A based on a large model, which can combine keyword detection and the semantic recognition ability of the large model to achieve rapid review of sensitive information in the input text. The technical solution of the present invention screens fiscal and tax violation-oriented Q&A through the trained fiscal and tax large model; the technical solution of the present invention constructs a Chinese pinyin sensitive word Trie tree and improves the detection effect by training the semantic understanding ability of the large model; the technical solution of the present invention realizes the identification of sensitive text not only through keyword query but also through semantic matching of the large model for the violation financial scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] By referring to the following drawings, the exemplary embodiments of the present invention can be more completely understood:
[0048] Figure 1 FIG. is a flowchart of a method for identifying sensitive information in fiscal and tax Q&A based on a large model according to a preferred embodiment of the present invention;
[0049] Figure 2 FIG. is a schematic diagram of a Chinese pinyin sensitive word Trie tree according to a preferred embodiment of the present invention;
[0050] Figure 3 FIG. is a flowchart of sensitive information identification according to a preferred embodiment of the present invention; and
[0051] Figure 4 FIG. is a structural diagram of a system for identifying sensitive information in fiscal and tax Q&A based on a large model according to a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] Now, the exemplary embodiments of the present invention will be described with reference to the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to disclose the present invention in detail and completely, and to fully convey the scope of the present invention to those skilled in the art. The terms in the exemplary embodiments shown in the drawings are not limitations on the present invention. In the drawings, the same unit / element uses the same reference numeral.
[0053] Unless otherwise specified, the terms used herein (including technical terms) have the ordinary meaning understood by those skilled in the relevant technical field. Additionally, it can be understood that terms defined in commonly used dictionaries should be construed to have a meaning consistent with the context of their relevant fields and should not be construed as idealized or overly formal meanings.
[0054] Figure 1 It is a flowchart of a method for identifying sensitive information in fiscal and tax Q&A based on a large model according to a preferred embodiment of the present invention.
[0055] The overall process of identifying sensitive information in fiscal and tax Q&A based on a large model proposed by the present invention can be divided into four stages, namely, professional ability training of the fiscal and tax large model, data processing, keyword detection, and large model sensitive information detection.
[0056] For the fiscal and tax Q&A scenario, common sensitive questions are divided into the following six scenarios: political security, porn, abuse, discrimination and prejudice, privacy security, and fiscal and tax gray areas. For the above scenarios, a method for identifying sensitive information in fiscal and tax Q&A based on a large model is proposed.
[0057] As Figure 1 shown, the present invention provides a method for identifying sensitive information in fiscal and tax Q&A based on a large model, and the method includes:
[0058] Step 101: Obtain fiscal and tax question data and process the obtained fiscal and tax question data;
[0059] Preferably, processing the obtained fiscal and tax question data includes:
[0060] When the fiscal and tax question data is of the ultra-long text type, add a batch_size parameter to the fiscal and tax question data, set an adaptive context window for the fiscal and tax question data, and set text segmentation points based on specified symbols.
[0061] Preferably, processing the obtained fiscal and tax question data includes:
[0062] Judge whether there are deformed sensitive words in the fiscal and tax question data. When it is determined that there are deformed sensitive words in the fiscal and tax question data, process the deformed sensitive words, including: converting traditional Chinese to simplified Chinese, converting case, removing special symbols, and converting to Chinese pinyin.
[0063] The present invention processes data, and for the data processing of fiscal and tax Q&A, the following problems are addressed respectively:
[0064] For ultra-long texts: To ensure the model processing efficiency, add a batch_size parameter, set an adaptive context window, and combine the full stop "。" as the text segmentation point to ensure text coherence.
[0065] For deformed sensitive words: There are usually the following scenarios. Insert non-Chinese characters between sensitive words to form deformed sensitive words, use pinyin, homophones, homonyms, etc. to deform sensitive words, and for sensitive words after splitting and traditionalization. Perform preprocessing operations of converting traditional Chinese to simplified Chinese, converting case, removing special symbols, and converting to Chinese pinyin for the above texts respectively.
[0066] Step 102: For the processed financial and tax question data, perform key prohibited word detection through the established Chinese pinyin sensitive word Trie tree;
[0067] Preferably, it further includes: when it is determined that there are key prohibited words in the financial and tax question data, based on the labels set for the financial and tax question data, return the corresponding answer template for the financial and tax question data. As Figure 3 shown.
[0068] Preferably, it further includes establishing a Chinese pinyin sensitive word Trie tree:
[0069] Determine 23 Chinese pinyin initials after removing "u", "v", "i", and mark the 23 Chinese pinyin initials as nodes numbered 0 to 22;
[0070] Store Chinese sensitive words in the initial letter pinyin through nodes numbered 0 to 22;
[0071] Store sensitive words starting with numbers, non-Chinese, and English characters in the established node numbered 23;
[0072] Set a mark for whether nodes numbered 0 to 23 are terminal nodes.
[0073] The present invention detects keywords:
[0074] The text to be detected is divided into two parts: matching with the prohibited word library and matching with the constructed Chinese pinyin sensitive word Trie tree. Among them, the Chinese pinyin sensitive word Trie tree is based on the decision tree idea and based on the composition of Chinese pinyin. Since the initial letters of Chinese pinyin are composed of 23 (removing "u", "v", "i") letters, among which nodes numbered 0 to 22 store Chinese sensitive words in the initial letter pinyin, and the node numbered 23 stores sensitive words starting with numbers or other non-Chinese and English characters, and at the same time of storing Chinese characters and their pinyin, mark whether the node is a terminal node. As Figure 2 shown.
[0075] Step 103: When it is determined that there are no key prohibited words in the financial and tax question data, detect the prohibited sensitive words through the prohibited semantics in the trained large financial and tax model;
[0076] Preferably, it further includes: when the large financial and tax model detects prohibited sensitive words, based on the labels set for the financial and tax question data, return the corresponding answer template for the financial and tax question data.
[0077] The present invention detects sensitive information of large models:
[0078] For the text after keyword detection, concatenate prompts and output the final system detection results.
[0079]
[0080] Step 104: When no illegal sensitive words are detected in the output of the financial and taxation model, the financial and taxation question data is output to the normal question-answering system.
[0081] Preferably, it also includes training the big finance and taxation model:
[0082] Obtain financial data, pre-label the financial data, and build a fine-tuning dataset based on the financial data and pre-labeled labels;
[0083] The finance and taxation big model is trained by fine-tuning the data set until the finance and taxation big model reaches the predetermined standard for the detection rate of illegal sensitive words in the fine-tuning data set.
[0084] In order to ensure the consistency of output and improve the model's ability to identify sensitive fiscal and tax issues, this method uses human feedback to perform reinforcement learning on the language model. The training process of the pre-trained model is divided into the following two steps:
[0085] Fine-tuning of labeled data: By collecting financial data, including financial statements, tax data, and financial and accounting laws and regulations, and combining experts in the field of finance, taxation and accounting, we ensure the diversity and accuracy of data. We pre-label the financial and taxation data based on prompt engineering and advanced large models such as GPT4, and then review and revise them by business experts, so as to provide supervised learning training samples for the model.
[0086] Training classification models: We build fine-tuning datasets based on prompts for sensitive questions in six scenarios, and expand the datasets using the stronger generalization capabilities of large-scale models. The fine-tuning data format is as follows:
[0087]
[0088] In the context of the digital transformation of the finance and taxation industry, in order to balance the model output accuracy and model output speed, the Tongyi Qianwen large model with a parameter volume of 7B is selected as the base model. It has excellent Chinese and English understanding and provides a solid foundation for applications in the finance and taxation field. The Tongyi Qianwen pre-trained model has rich general semantic knowledge and various general technical capabilities, but it does not have the ability to analyze sensitive questions and answers in the finance and taxation field. Therefore, a finance and taxation data set consisting of finance and taxation books, accounting standards, finance and taxation knowledge data, accounting vouchers, financial statements, etc. is used for further pre-training. By fine-tuning the large model, it can better understand and process these professional knowledge, improve its recognition efficiency of finance and taxation sensitive issues, and ensure taxation Q&A compliance.
[0089] The present invention designs a Chinese Pinyin sensitive word Trie tree according to the characteristics of Chinese Pinyin, combines the matching of a banned word library, reduces the difficulty of building a banned word library, and improves the detection efficiency of banned words.
[0090] The present invention uses a large finance and taxation model trained based on a data set in the finance and taxation field, which can effectively discover malicious finance-oriented questions in user questions and answers, and accurately identify financially illegal questions and answers.
[0091] The present invention proposes a method for identifying sensitive information in financial and taxation question and answer based on a large model, which has the following three beneficial effects:
[0092] The sensitive information recognition rate of the financial and taxation big model has been increased to 97.7%. The present invention can significantly improve the security of the big model and prevent the big model from being used for malicious acts such as identity theft, fraud and social engineering attacks, providing strong support for the digital transformation of the financial and taxation field.
[0093] Automated detection reduces the need for manual review, thereby reducing labor costs.
[0094] Identifying and effectively managing sensitive data can filter out bad information, illegal content and harmful speech, purify the network environment, and maintain the healthy and orderly development of cyberspace. It helps enterprises and platforms meet legal and compliance requirements and avoid legal risks. It also enhances user trust in the financial and taxation big model.
[0095] Figure 4 This is a structural diagram of a financial and taxation question and answer sensitive information identification system based on a large model according to a preferred embodiment of the present invention.
[0096] like Figure 4 As shown, the present invention provides a financial and taxation question and answer sensitive information identification system based on a large model, the system comprising:
[0097] Initial unit 401, used to obtain financial and taxation question data and process the obtained financial and taxation question data;
[0098] The first detection unit 402 is used to detect key prohibited words in the processed fiscal and tax question data through the established Chinese Pinyin sensitive word Trie tree;
[0099] Preferably, the initial unit 401 is used to process the obtained fiscal and tax question data and is also used for:
[0100] When the fiscal and tax question data is of the ultra-long text type, add a batch_size parameter to the fiscal and tax question data, set an adaptive context window for the fiscal and tax question data, and set text segmentation points based on specified symbols.
[0101] Preferably, the initial unit 401 is used to process the obtained fiscal and tax question data and is also used for:
[0102] Judge whether there are deformed sensitive words in the fiscal and tax question data. When it is judged that there are deformed sensitive words in the fiscal and tax question data, process the deformed sensitive words, including: converting traditional Chinese to simplified Chinese, converting case, removing special symbols, and converting to Chinese Pinyin.
[0103] Preferably, the first detection unit 402 is also used for: when it is judged that the fiscal and tax question data contains key prohibited words, return the answer template corresponding to the fiscal and tax question data based on the label set for the fiscal and tax question data.
[0104] Preferably, the first detection unit 402 is also used for: establishing a Chinese Pinyin sensitive word Trie tree:
[0105] Determine the 23 Chinese Pinyin initials after removing "u", "v", and "i", and mark the 23 Chinese Pinyin initials as nodes numbered 0 to 22;
[0106] Store Chinese sensitive words in the initials of Pinyin through nodes numbered 0 to 22;
[0107] Store sensitive words starting with numbers, non-Chinese, and English characters in the established node numbered 23;
[0108] Set a mark for whether nodes numbered 0 to 23 are terminal nodes.
[0109] The second detection unit 403 is used to detect violation sensitive words through the violation semantics in the trained large fiscal and tax model when it is judged that the fiscal and tax question data does not contain key prohibited words;
[0110] The result unit 404 is used to output the fiscal and tax question data to the normal Q&A system when no violation sensitive words are detected in the output of the large fiscal and tax model.
[0111] Preferably, the second detection unit 403 is further configured to: when the fiscal and tax large model detects a violation sensitive word, return the answer template corresponding to the fiscal and tax question data based on the labels set for the fiscal and tax question data.
[0112] Preferably, the initial unit 401 further includes training the fiscal and tax large model:
[0113] Obtain financial data, pre-label the financial data, and establish a fine-tuning data set based on the financial data and the pre-labeled labels;
[0114] Train the fiscal and tax large model through the fine-tuning data set until the detection rate of the violation sensitive words in the fine-tuning data set reaches a predetermined standard.
[0115] A fiscal and tax Q&A sensitive information recognition system according to a preferred embodiment of the present invention corresponds to a fiscal and tax Q&A sensitive information recognition method according to another preferred embodiment of the present invention, and will not be elaborated here.
[0116] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention can be implemented in various computer languages, for example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.
[0117] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0118] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in the flowchart(s) Figure 1 one or more flowcharts and / or block diagrams Figure 1 specified in one or more block diagrams or a plurality of block diagrams.
[0119] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart(s) Figure 1 one or more flowcharts and / or block diagrams Figure 1 specified in one or more block diagrams or a plurality of block diagrams.
[0120] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0121] It is obvious that those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
[0122] The present invention has been described by reference to a few embodiments. However, as is well known to those skilled in the art, other embodiments equivalent to those disclosed above of the present invention equally fall within the scope of the present invention as defined by the appended patent claims.
[0123] Generally, all terms used in the claims are to be interpreted according to their ordinary meaning in the technical field, unless otherwise explicitly defined therein. All references to "a / the [device, component, etc.]" are to be construed openly as at least one instance of the device, component, etc., unless otherwise explicitly stated. The steps of any method disclosed herein need not be performed in the exact order disclosed, unless explicitly stated.
Claims
1. A method for identifying sensitive information in financial and taxation question and answer based on a large model, the method comprising: Acquire financial and taxation question data, and process the acquired financial and taxation question data; For the processed financial and taxation question data, key prohibited words are detected through the established Chinese Pinyin sensitive word Trie tree; When it is determined that the financial and taxation question data does not contain any key prohibited words, the illegal sensitive words are detected through the illegal semantics in the trained financial and taxation big model; When no illegal sensitive words are detected in the output of the financial and taxation model, the financial and taxation question data is output to the normal question and answer system.
2. The method according to claim 1, further comprising: When it is determined that the financial and taxation question data contains key prohibited words, based on the tags set for the financial and taxation question data, an answer template corresponding to the financial and taxation question data is returned.
3. The method according to claim 1, further comprising: When the financial and taxation big model detects illegal sensitive words, based on the labels set for the financial and taxation question data, an answer template corresponding to the financial and taxation question data is returned.
4. According to the method of claim 1, the processing of the acquired financial and taxation question data comprises: When the financial and taxation question data is of an ultra-long text type, a batch_size parameter is added to the financial and taxation question data, an adaptive context window of the financial and taxation question data is set, and a text segmentation point is set based on a specified symbol.
5. According to the method of claim 1, the processing of the acquired financial and taxation question data comprises: It is determined whether there are deformed sensitive words in the financial and taxation question data. When it is determined that there are deformed sensitive words in the financial and taxation question data, the deformed sensitive words are processed, including: converting traditional Chinese to simplified Chinese, converting uppercase and lowercase, removing special symbols, and converting to Chinese pinyin.
6. The method according to claim 1 further comprises establishing a Trie tree of Chinese Pinyin sensitive words: Determine the 23 Chinese phonetic initials after removing "u", "v", and "i", and mark the 23 Chinese phonetic initials as nodes 0 to 22; The Chinese sensitive words in the initial pinyin are respectively stored through the nodes 0 to 22; The node 23 will be created to store sensitive words starting with numbers, non-Chinese characters, and English characters; Set a flag for nodes 0 to 23 to indicate whether they are terminal nodes.
7. The method according to claim 1 further comprises training a large financial and taxation model: Acquire financial data, pre-label the financial data, and establish a fine-tuning data set based on the financial data and the pre-labeled labels; The finance and taxation big model is trained by using the fine-tuning data set until the finance and taxation big model reaches a predetermined standard for the detection rate of the illegal sensitive words in the fine-tuning data set.
8. A financial and taxation question-and-answer sensitive information identification system based on a large model, the system comprising: An initialization unit, used for acquiring financial and taxation question data and processing the acquired financial and taxation question data; A first detection unit is used to detect key prohibited words on the processed financial and taxation question data through an established Chinese Pinyin sensitive word Trie tree; The second detection unit is used to detect the illegal sensitive words through the illegal semantics in the trained financial and taxation big model when it is determined that the financial and taxation question data does not contain key banned words; The result unit is used to output the financial and taxation question data to a normal question-and-answer system when no illegal sensitive words are detected in the output of the financial and taxation big model.
9. According to the system of claim 8, the first detection unit is also used for: when it is determined that the financial and taxation question data contains key banned words, based on the labels set for the financial and taxation question data, returning the answer template corresponding to the financial and taxation question data.
10. According to the system of claim 8, the second detection unit is also used to: when the financial and taxation big model detects illegal sensitive words, based on the labels set for the financial and taxation question data, return the answer template corresponding to the financial and taxation question data.
Citation Information
Cited By
Medical beauty financial data analysis method based on big data
CN120876127A