Sensitive information processing method and device, electronic equipment, medium and product

Through iterative adjustment of the optimizer and generator, the target prompt words are generated, the sensitive information generation model is trained and reliability detection is carried out, which solves the problem of malicious traffic diversion among merchants in e-commerce platforms, improves the accuracy and efficiency of information generation, and reduces transaction risks.

CN120337296APending Publication Date: 2025-07-18BEIJING WODONG TIANJUN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510503190.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Merchants in e-commerce platforms use chat tools to maliciously divert traffic, which increases user transaction risks, and it is difficult for existing technology to effectively identify and handle such behavior.

Method used

Through iterative adjustments of the optimizer and generator, the target prompt words are generated, the sensitive information generation model is trained, reliability detection and sensitive word matching is performed, sensitive thesaurus is constructed, and governance information is generated to deal with malicious drainage.

Benefits of technology

It improves the accuracy and efficiency of sensitive information generation, effectively filters out unreliable information, reduces user transaction risks, and maintains the security and order of the platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337296A_ABST
    Figure CN120337296A_ABST
Patent Text Reader

Abstract

The invention provides a sensitive information processing method and device, electronic equipment and a computer readable storage medium, and relates to the technical field of computers. The sensitive information processing method comprises the steps that iteration adjustment operation of an initial cue word is carried out based on a configured optimizer and a generator, a target cue word is generated, and the target cue word is used for configuring an output format of sensitive information corresponding to the form of a communication message; training a matched sensitive information generation model based on the target cue word and the communication message of the corresponding form, so that the sensitive information generation model outputs sensitive information with a corresponding output format; performing reliability detection on the corresponding sensitive information based on a detection rule corresponding to the form; and constructing a sensitive word library based on the sensitive information passing the reliability detection, and generating governance information for the matched sensitive words. Through the technical scheme disclosed by the invention, risk behaviors can be intervened and processed in time, the safety and order of the platform are maintained, and the transaction risk of the user is further reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Chat tools in e-commerce platforms play an important role in the communication between merchants and users because their communication functions are more targeted in the e-commerce scenario. However, with the popularization of the platform and the increase in the number of users, some merchants have started to use chat tools for malicious drainage, guiding users to other platforms for transactions. Such behavior not only causes losses to the original platform but also increases the transaction risks for users.

[0003] It should be noted that the information disclosed in the above Background Art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0004] An object of the present disclosure is to provide a sensitive information processing method, a sensitive information processing device, an electronic device, a computer-readable storage medium, and a computer program product, which can at least to some extent improve the problem in related technologies that some merchants start to use chat tools for malicious drainage, thereby increasing the transaction risks for users.

[0005] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be partially learned through the practice of the present disclosure.

[0006] According to one aspect of the present disclosure, a sensitive information processing method is provided, including: performing an iterative adjustment operation on an initial prompt word based on a configured optimizer and generator to generate a target prompt word, where the target prompt word is used to configure an output format of sensitive information corresponding to the form of a communication message; training a matching sensitive information generation model based on the target prompt word and the communication message of the corresponding form, so that the sensitive information generation model outputs the sensitive information with the corresponding output format; performing a reliability detection on the corresponding sensitive information based on a detection rule corresponding to the form; constructing a sensitive word library based on the sensitive information that passes the reliability detection, so as to perform a matching operation of sensitive words based on the sensitive word library when obtaining the communication message, and generate governance information for the matched sensitive words.

[0007] In one embodiment of the present disclosure, an iterative adjustment operation of an initial prompt is performed based on a configured optimizer and generator to generate a target prompt, including: extracting a plurality of prompt training information from training samples, where the training samples are labeled with the sensitive information; inputting the plurality of prompt training information into the optimizer for merging to generate the initial prompt; inputting the initial prompt and the training samples into the generator to generate an identification result of the sensitive information; inputting a comparison result between the identification result and the corresponding label, the length of the initial prompt, and the time taken to generate the identification result into a reward function to obtain a reward result; and feeding back the reward result and an error sample obtained based on the identification result to the optimizer to cyclically train the optimizer, so that the optimizer cyclically adjusts the initial prompt until the target prompt is obtained.

[0008] In one embodiment of the present disclosure, feeding back the reward result and an error sample obtained based on the identification result to the optimizer to cyclically train the optimizer, so that the optimizer cyclically adjusts the initial prompt until the target prompt is obtained, includes: feeding back the reward result and the error sample to the optimizer, so that the optimizer analyzes the error cause based on the error sample; aiming to improve the reward result and eliminate the error cause, calculating the gradient of the relevant factors of the optimizer based on the backpropagation algorithm, starting the cyclical training based on the direction of the gradient to cyclically adjust the initial prompt, and optimizing the reward function based on the cyclical training; and when it is detected that the number of cyclic iterations reaches a preset number of iterations, and / or the reward result reaches a preset result, determining the prompt output by the optimizer as the target prompt.

[0009] In one embodiment of the present disclosure, the forms of the communication messages include single messages and conversation messages, the training samples include a single-message drainage data set and a conversation-message drainage data set, and performing an iterative adjustment operation of an initial prompt based on a configured optimizer and generator to generate a target prompt further includes: generating a first target prompt based on the single-message drainage data set, where the first target prompt is used to configure the output format of the sensitive information corresponding to the single message; and generating a second target prompt based on the conversation-message drainage data set, where the second target prompt is used to configure the output format of the sensitive information corresponding to the conversation message.

[0010] In one embodiment of the present disclosure, training a matching sensitive information generation model based on the target prompt word and the communication message in the corresponding form, so that the sensitive information generation model outputs the sensitive information in the corresponding output format, includes: training a classification model based on the communication message to obtain an identification result of whether the communication message is a drainage message based on the output of the classification model; if it is identified as the drainage message, preprocessing the drainage message to obtain preprocessing data; training a sensitive information generation model based on the target prompt word corresponding to the form of the drainage message and the preprocessing data, so that the sensitive information generation model outputs the sensitive information.

[0011] In one embodiment of the present disclosure, training a classification model based on the communication message to obtain an identification result of whether the communication message is a drainage message based on the output of the classification model, includes: the classification model includes a binary classification BERT model and a Tiny BERT model. Fine-tuning a pre-trained BERT model with labeled data having drainage annotations to obtain the binary classification BERT model, and the binary classification BERT model is used to calculate the drainage confidence of the communication message; training the Tiny BERT model as a drainage feature calculation model based on the communication message, so that the trained Tiny BERT model performs drainage feature calculation on the communication message to obtain a feature vector related to drainage; comparing the feature vector with the vector of the known drainage feature in the preset vector library to calculate the feature similarity; adopting a fusion strategy to fuse the drainage confidence and the feature similarity to obtain a comprehensive judgment index; obtaining an identification result of whether the communication message is the drainage message based on the relationship between the comprehensive judgment index and the judgment threshold.

[0012] In one embodiment of the present disclosure, preprocessing the drainage message to obtain preprocessing data, includes: if the communication message is not a single message, organizing the non-single message into a dialogue message; for the dialogue message or if the communication message is a single message, respectively performing data filtering and data desensitization operations to obtain the preprocessing data.

[0013] In one embodiment of the present disclosure, the form of the communication message is a dialogue message, and performing reliability detection on the corresponding sensitive information based on the detection rule corresponding to the form, includes: the sensitive information generation model outputs the sensitive information, and detecting whether the corresponding drainage confidence is greater than the first confidence threshold; and wherein, if the drainage confidence is greater than the first confidence threshold, it is determined that the reliability detection is passed.

[0014] In one embodiment of the present disclosure, the form of the communication message is a single message, and the reliability of the corresponding sensitive information is detected based on the detection rule corresponding to the form, including: the sensitive information generation model outputs the sensitive information, and it is detected whether the sensitive information meets the clarity condition and the context condition; and if the sensitive information can be matched with a preset sensitive word dictionary, it is further detected whether the sensitive information meets the confidence detection condition and / or the historical recognition detection condition, wherein if the sensitive information meets the clarity condition and the context condition, and meets the confidence detection condition and / or the historical recognition detection condition, it is determined that the reliability detection is passed.

[0015] In one embodiment of the present disclosure, the confidence detection condition includes that the average confidence of the classification model is greater than a second confidence threshold and the maximum confidence of the classification model is greater than a third confidence threshold; the historical recognition detection condition includes that the number of times of identifying the sensitive information is greater than a number threshold and the proportion of identifying the sensitive information is greater than a proportion threshold.

[0016] In one embodiment of the present disclosure, when the communication message is obtained, a matching operation of sensitive words is performed based on the sensitive word library, including: when the communication message is obtained, the recognition confidence of the communication message is calculated based on the classification model; if the recognition confidence is greater than a fourth confidence threshold, a matching operation of sensitive words is performed based on the sensitive word library; if the recognition confidence is less than or equal to the fourth confidence threshold, the recognition result of whether the communication message is a drainage message is output again based on the classification model.

[0017] In one embodiment of the present disclosure, the matching operation of sensitive words based on the sensitive word library includes: detecting whether the communication message has the sensitive words recorded in the sensitive word library based on a traversal operation of the prefix tree.

[0018] According to another aspect of the present disclosure, a sensitive information processing device is provided, including: an iterative adjustment module, configured to perform an iterative adjustment operation of an initial prompt word based on a configured optimizer and a generator to generate a target prompt word, where the target prompt word is used to configure an output format of sensitive information corresponding to the form of the communication message; a training module, configured to train a matching sensitive information generation model based on the target prompt word and the communication message of the corresponding form, so that the sensitive information generation model outputs the sensitive information with the corresponding output format; a detection module, configured to perform reliability detection on the corresponding sensitive information based on the detection rule corresponding to the form; a matching module, configured to construct a sensitive word library based on the sensitive information passing the reliability detection, so as to perform a matching operation of sensitive words based on the sensitive word library when the communication message is obtained, and generate governance information for the matched sensitive words.

[0019] According to another aspect of the present disclosure, there is provided an electronic device, including: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute the sensitive information processing method of any one of the above via executing the executable instructions.

[0020] According to yet another aspect of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program, which when executed by a processor, implements the sensitive information processing method of any one of the above.

[0021] The sensitive information processing solution provided by the embodiments of the present disclosure can continuously optimize the prompt words through iterative adjustment of the optimizer and the generator, making them more accurately guide the sensitive information generation model to output sensitive information that meets the requirements. This iterative optimization method makes full use of the capabilities of the model, improves the quality of the prompt words, and thus provides a solid foundation for subsequent sensitive information generation. Training based on the target prompt words and corresponding forms of communication messages enables the model to better understand the business requirements and output format requirements. Through a large amount of data training and parameter adjustment, the performance of the model has been continuously improved, enabling it to more accurately extract and generate sensitive information from communication messages, improving the generation efficiency and accuracy of sensitive information. Strict detection of the generated sensitive information can effectively filter out unreliable sensitive information, improving the quality and reliability of sensitive information. Using the sensitive information that passes the reliability detection to construct a sensitive word library and adopting an efficient matching algorithm for sensitive word matching can quickly and accurately find sensitive words in communication messages and generate governance information for the matched sensitive words, enabling timely intervention and handling of risk behaviors, maintaining the security and order of the platform, and further reducing the transaction risks of users.

[0022] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0024] Figure 1 A schematic diagram showing the structure of a sensitive information processing system in an embodiment of the present disclosure;

[0025] Figure 2 A flowchart showing a sensitive information processing method in an embodiment of the present disclosure;

[0026] Figure 3 A schematic diagram showing another structure of the sensitive information processing system in an embodiment of the present disclosure;

[0027] Figure 4 A flowchart showing another sensitive information processing method in an embodiment of the present disclosure;

[0028] Figure 5 A flowchart showing yet another sensitive information processing method in an embodiment of the present disclosure;

[0029] Figure 6 A curve graph showing a comparison of different optimization algorithms in an embodiment of the present disclosure;

[0030] Figure 7 A curve graph showing a comparison of another different optimization algorithms in an embodiment of the present disclosure;

[0031] Figure 8 A flowchart showing yet another sensitive information processing method in an embodiment of the present disclosure;

[0032] Figure 9 A flowchart showing yet another sensitive information processing method in an embodiment of the present disclosure;

[0033] Figure 10 A flowchart showing yet another sensitive information processing method in an embodiment of the present disclosure;

[0034] Figure 11 A flowchart showing yet another sensitive information processing method in an embodiment of the present disclosure;

[0035] Figure 12 A schematic diagram showing a sensitive information processing device in an embodiment of the present disclosure;

[0036] Figure 13 A schematic diagram showing an electronic device in an embodiment of the present disclosure. Detailed implementation manners

[0037] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.

[0038] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0039] As a cutting-edge technology in the field of natural language processing, the Large Language Model (LLM) has powerful text understanding and generation capabilities. In recent years, with the development of deep learning technology, the LLM has demonstrated excellent performance in tasks such as text classification, sentiment analysis, and information extraction. Therefore, it is highly feasible and practically valuable to develop a system based on the LLM that can identify and extract malicious traffic diversion by merchants to other platforms in chat software.

[0040] In addition, fully exploiting the potential of the LLM requires one or more high-quality prompts. Prompt Engineering is an emerging technology in the field of artificial intelligence that focuses on designing and optimizing input prompts to guide language models (such as GPT-4) to generate high-quality and relevant outputs. By carefully constructing and adjusting the prompts, Prompt Engineering can significantly improve the performance of the model, making it more accurate and efficient in specific tasks.

[0041] In some embodiments, although the LLM has shown great potential in many fields, since the prompt engineering of the LLM relies on human experience, this means that experts need to deeply understand the working mechanism of the model and the requirements of the e-commerce business, and design and optimize the prompts to ensure that the model can accurately generate valuable content. Although it can bring customized and relatively accurate results, there are problems of low efficiency and difficulty in scaling.

[0042] In addition, although the introduction of large models can significantly improve the detection and prevention capabilities, to achieve this goal, a complete set of processes from data collection, model training, real-time monitoring to automatic response needs to be established.

[0043] The solution provided by this application can continuously optimize the prompt through iterative adjustment of the optimizer and the generator, making it more accurate in guiding the sensitive information generation model to output sensitive information that meets the requirements. This iterative optimization method fully utilizes the capabilities of the model, improves the quality of the prompt, and thus provides a solid foundation for subsequent sensitive information generation. Training based on the target prompt and corresponding communication messages in specific forms enables the model to better understand the business requirements and output format requirements. Through a large amount of data training and parameter adjustment, the performance of the model has been continuously improved, enabling it to extract and generate sensitive information more accurately from communication messages, improving the generation efficiency and accuracy of sensitive information. Strict detection of the generated sensitive information can effectively filter out unreliable sensitive information, improving the quality and reliability of sensitive information. Using the sensitive information that passes the reliability detection to build a sensitive word library and adopting an efficient matching algorithm for sensitive word matching can quickly and accurately find sensitive words in communication messages and generate governance information for the matched sensitive words, enabling timely intervention and handling of risk behaviors and maintaining the security and order of the platform.

[0044] For ease of understanding, several terms involved in this application are first explained below.

[0045] LLM (Large Language Model): These models are built based on deep learning technology and usually have billions or even hundreds of billions of parameters. They are trained with a large amount of text data and can understand and generate natural language.

[0046] GLM-4 is the fourth-generation General Language Model. The GLM series of models are known for their bidirectional and autoregressive generation capabilities and can achieve efficient text generation and understanding in natural language processing tasks.

[0047] GPT-4o: GPT-4o is a variant in the GPT-4 series developed by OpenAI, specifically optimized for applications in specific tasks or scenarios, providing a faster response speed and lower request cost than GPT-4.

[0048] Prompt: In large language models, Prompt refers to the input text used to guide the model to generate a response. It can be a question, instruction, or context information, aiming to clearly instruct the model to generate the required output.

[0049] APO (Auto Prompt Optimization). In large language models, APO refers to the use of automated methods to generate and optimize prompts to improve the quality and performance of model outputs. This method utilizes algorithms and machine learning techniques to explore and adjust prompts to achieve the best generation results.

[0050] BERT: Bidirectional Encoder Representations from Transformers. It is a pre-trained language model for natural language processing. Based on the Transformer architecture, it aims to understand text context bidirectionally to improve language understanding ability.

[0051] Tiny BERT is a compressed version of the BERT model, aiming to reduce the model size and computational resource requirements while maintaining its performance as much as possible. It extracts knowledge from the large BERT model through knowledge distillation technology to achieve a more efficient model.

[0052] CoSENT (Cosine Sentence Embeddings using Transformer) is a model for generating sentence embeddings. Based on the Transformer architecture, it achieves more precise semantic matching by optimizing the cosine similarity of sentence representations.

[0053] Vearch is a distributed vector search engine specialized for handling large-scale vector retrieval tasks. It combines vector search and database management functions and is suitable for application scenarios that require efficient processing and retrieval of high-dimensional vector data.

[0054] Trie Tree, also known as a dictionary tree or prefix tree, is a data structure for efficiently storing and retrieving a set of strings.

[0055] UCB (Upper Confidence Bound) is a strategy for the Multi-Armed Bandit Problem, aiming to maximize rewards within a limited time.

[0056] BP (Backpropagation). Backpropagation is a key algorithm for training artificial neural networks. It adjusts the weights in the network by calculating the gradient of the loss function to minimize the prediction error.

[0057] CoT (Chain of Thought) is a reasoning method that refers to decomposing complex problems into a series of simple steps or reasoning chains by gradually unfolding the thinking process, in order to better solve problems.

[0058] Figure 1 It is a schematic structural diagram of a computer system provided by an exemplary embodiment of the present application. The system includes: a plurality of terminals 120 and a server cluster 140.

[0059] The terminal 120 can be a mobile terminal such as a mobile phone, a game console, a tablet computer, an e-book reader, smart glasses, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a smart home device, an AR (Augmented Reality) device, a VR (Virtual Reality) device, etc. Alternatively, the terminal 120 can also be a personal computer (PC), such as a laptop computer and a desktop computer, etc.

[0060] Among them, an application program for providing sensitive information processing can be installed in the terminal 120.

[0061] The terminal 120 is connected to the server cluster 140 through a communication network. Optionally, the communication network is a wired network or a wireless network.

[0062] The server cluster 140 is a server, or consists of several servers, or is a virtualization platform, or is a cloud computing service center. The server cluster 140 is used to provide background services for the application program for providing sensitive information processing. Optionally, the server cluster 140 undertakes the main computing work, and the terminal 120 undertakes the secondary computing work; or, the server cluster 140 undertakes the secondary computing work, and the terminal 120 undertakes the main computing work; or, a distributed computing architecture is adopted between the terminal 120 and the server cluster 140 for collaborative computing.

[0063] In some alternative embodiments, the server cluster 140 is used to store sensitive information processing program information.

[0064] Optionally, the clients of the application programs installed in different terminals 120 are the same, or the clients of the application programs installed on two terminals 120 are the clients of the same type of application program on different control system platforms. Based on the differences in the terminal platforms, the specific forms of the clients of the application program can also be different. For example, the client of the application program can be a mobile phone client, a PC client, or a World Wide Web (Web) client, etc.

[0065] Those skilled in the art can understand that the number of the above-mentioned terminals 120 can be more or less. For example, the above-mentioned terminal can be only one, or there can be dozens or hundreds of the above-mentioned terminals, or even more. The embodiments of the present application do not limit the number and device types of the terminals.

[0066] Optionally, the system may further include a management device ( Figure 1 not shown), which is connected to the server cluster 140 through a communication network. Optionally, the communication network is a wired network or a wireless network.

[0067] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is usually the Internet, but it can also be any network, including but not limited to any combination of a Local Area Network (LAN), a Metropolitan Area Network (MAN), a Wide Area Network (WAN), a mobile, wired or wireless network, a private network or a virtual private network). In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent the data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some of the links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.

[0068] Next, the sensitive information processing method in the present exemplary embodiment will be described in more detail with reference to the accompanying drawings and embodiments.

[0069] AsFigure 2 As shown in Figure 2 , a sensitive information processing method according to an embodiment of the present disclosure is applied to a server and includes:

[0070] Step S202: Perform an iterative adjustment operation on the initial prompt based on a configured optimizer and generator to generate a target prompt, which is used to configure the output format of sensitive information corresponding to the form of the communication message.

[0071] Among them, the sensitive information may include account information of chat software registered on other platforms, including but not limited to WeChat IDs, QQ numbers, and account numbers of other social platforms, etc.

[0072] In some embodiments, appropriate large language models can be selected as the optimizer and generator respectively, and an APO is constructed by the optimizer and generator.

[0073] In some embodiments, the initial prompt and training data are input into the generator to generate a preliminary output of sensitive information. The optimizer analyzes the reasons for errors based on the loss value and adjusts the prompt. For example, if it is found that the format of the generated sensitive information does not meet the requirements, the description of the format in the prompt may be adjusted. The above process of generating results, evaluating results, and adjusting the prompt is repeated until the preset number of iterations is reached or the performance of the prompt meets the requirements, and the target prompt is obtained.

[0074] Step S204: Train a matching sensitive information generation model based on the target prompt and the communication message of the corresponding form, so that the sensitive information generation model outputs sensitive information with the corresponding output format.

[0075] In some embodiments, the sensitive information generation model adopts GLM-4. Two different prompts are generated through APO to process single messages and conversation messages respectively. The sensitive information generation model outputs sensitive information in the corresponding format based on the target prompt and single message of the single information, and the sensitive information generation model outputs sensitive information in the corresponding format based on the target prompt and conversation message of the conversation information.

[0076] Step S206: Perform a reliability detection on the corresponding sensitive information based on the detection rule corresponding to the form.

[0077] Step S208: Construct a sensitive word library based on the sensitive information that passes the reliability detection, so that when a communication message is obtained, a matching operation of sensitive words is performed based on the sensitive word library, and governance information is generated for the matched sensitive words.

[0078] In some embodiments, through a labeled dataset, APO is used to automatically generate prompts for extracting sensitive information from single messages and conversations. Then, all local platform chat data is used as communication messages, organized into conversation form, and data filtering, data desensitization, and drainage confidence calculation are performed in sequence. Next, using the generated prompts and the filtered data, large language model generation results for single messages and conversations are obtained and recorded in a database. Further, based on the historical data in the database for the past year, a dictionary is constructed with the strings that match the sensitive information rules as keys, and the final results are obtained according to the previously introduced rules and recorded in the final database. Finally, the obtained sensitive information is added to the sensitive word library, and messages that match the sensitive information in the sensitive word library are monitored and processed in real time.

[0079] As Figure 3 shown, assuming the sensitive information is a WeChat ID, the sensitive information processing system includes a Prompt generation module, a data layer, a model layer, a logic layer, and a governance layer.

[0080] The Prompt generation module includes single-message Prompt generation based on APO and conversation-message Prompt generation.

[0081] In some embodiments, starting from the single-message drainage dataset, a single-message Prompt is generated through APO. Based on the conversation-message drainage dataset, a conversation-message Prompt is generated through APO. A Prompt is a text instruction that guides a large language model (LLM) to generate a specific output.

[0082] The data layer, as the data source, stores the original chat data and performs operations such as session organization, data filtering, and data desensitization on the stored chat data in sequence to provide standardized and usable data for subsequent processing by the model layer.

[0083] The model layer calculates the confidence of the processed data to measure the reliability of the data-related results, performs content generation through a large language model (LLM), and stores the generation results in the LLM generation result table.

[0084] The logic layer organizes relevant information into a dictionary and performs calculations based on rules, which may be used to judge and identify specific content, and stores the identified WeChat ID-related information, which is the result of the processing by the logic layer.

[0085] The governance layer manages the WeChat ID based on the results of the logic layer and governs the messages to ensure the compliance and orderly operation of the message system.

[0086] In this embodiment, through the iterative adjustment of the optimizer and the generator, the prompt can be continuously optimized to more accurately guide the sensitive information generation model to output sensitive information that meets the requirements. This iterative optimization method makes full use of the capabilities of the model, improves the quality of the prompt, and thus provides a solid foundation for subsequent sensitive information generation. Training is carried out based on the target prompt and the corresponding form of communication messages, enabling the model to better understand the business requirements and output format requirements. Through a large amount of data training and parameter adjustment, the performance of the model has been continuously improved, enabling it to more accurately extract and generate sensitive information from communication messages, improving the generation efficiency and accuracy of sensitive information. Strict detection of the generated sensitive information can effectively filter out unreliable sensitive information, improving the quality and reliability of sensitive information. Using the sensitive information that passes the reliability detection to construct a sensitive word library and adopting an efficient matching algorithm for sensitive word matching can quickly and accurately find sensitive words in communication messages and generate governance information for the matched sensitive words, enabling timely intervention and handling of risk behaviors, maintaining the security and order of the platform, and thus reducing the transaction risks of users.

[0087] As Figure 4 shown, in an embodiment of the present disclosure, an iterative adjustment operation of the initial prompt is performed based on a configured optimizer and generator to generate a target prompt, including:

[0088] Step S402, extract multiple prompt training information from the training samples, and the training samples are labeled with sensitive information.

[0089] In some embodiments, 2400 single-message drainage data sets based on single messages and 2400 dialogue-message drainage data sets based on dialogue messages can be configured. In each data set, 2000 are randomly selected as the training set and 400 as the test set. For 200 samples in the training set, 20 samples are taken as a group to obtain 10 prompt training information.

[0090] Step S404, input the multiple prompt training information into the optimizer for merging to generate an initial prompt.

[0091] In some embodiments, after merging 10 prompt training information through the optimizer, a complete initial prompt is obtained.

[0092] In some embodiments, a suitable optimizer is selected, such as large language models like GPT-4o. Necessary configurations are made for the optimizer, including setting the parameters of the model, the training environment, etc.

[0093] Step S406, input the initial prompt and the training samples into the generator to generate the recognition result of sensitive information.

[0094] In some embodiments, a suitable generator, such as a large language model like GLM-4, is selected and necessary configurations are made to the generator to ensure its normal operation.

[0095] Step S408: The comparison result between the recognition result and the corresponding label, the length of the initial prompt, and the time taken to generate the recognition result are input into the reward function to obtain a reward result.

[0096] In some embodiments, by comparing the predicted content of the samples in the current batch with the sample labels using the prompt corresponding to the current batch and all the samples in the batch, the accuracy of this prompt on the current samples can be obtained. According to the idea of the Upper Confidence Bound (UCB), by comprehensively evaluating the model accuracy, prediction time, prompt length, etc., a reward function that can measure the search depth and breadth is configured. The reward function is shown in Equation (1).

[0097]

[0098] Where err is the model accuracy, t is the inference time, N is the current round, and na is the number of times the prompt a is selected.

[0099] Step S410: The reward result and the error samples obtained based on the recognition result are fed back to the optimizer to cyclically train the optimizer, so that the optimizer cyclically adjusts the initial prompt until the target prompt is obtained.

[0100] As Figure 5 shown, the dataset is the basic data source for the entire process. The dataset is input into the optimizer (GPT-4o). GPT-4o, as the optimizer, processes the information in the dataset and converts it into a prompt, Prompt. The generated Prompt is input into the generator (GLM-4) to generate corresponding results according to the received Prompt, that is, the generated results. The generated results will be compared with the "labels", and the "loss" is calculated through the comparison. The loss is used to measure the difference between the generated results and the expected labels, reflecting the accuracy of the model's generated results. The calculated loss will be fed back to the optimizer, and the optimizer adjusts and optimizes the subsequent processing according to the loss information, forming a closed-loop iterative optimization process to gradually improve the quality and accuracy of the generated results.

[0101] In this embodiment, prompt training information is extracted from the training samples marked with sensitive information. These information contain the key features of the sensitive information, making the optimizer more targeted when generating the initial prompt words. The initial prompt words guide the generator to process the training samples and generate the sensitive information recognition results. This process realizes the preliminary mining of sensitive information. By inputting the recognition results, the comparison results with the labels, the length of the initial prompt words, and the time taken to generate the recognition results into the reward function to obtain the reward results, the performance of the current combination of prompt words and the generator can be comprehensively and quantitatively evaluated. The reward results and the error samples are fed back to the optimizer for iterative training, prompting the optimizer to continuously adjust the initial prompt words. During this iterative process, the prompt words are gradually optimized, guiding the generator more accurately and efficiently, and gradually improving the accuracy of the generator's recognition results. The final obtained target prompt words can greatly improve the accuracy of sensitive information recognition and reduce misjudgment and missed judgment.

[0102] In one embodiment of the present disclosure, the reward results and the error samples obtained based on the recognition results are fed back to the optimizer for iterative training of the optimizer, so that the optimizer cyclically adjusts the initial prompt words until the target prompt words are obtained, including:

[0103] Feed the reward results and the error samples back to the optimizer so that the optimizer analyzes the reasons for the errors based on the error samples.

[0104] In some embodiments, according to the comparison between the recognition results and the labels, the samples with recognition errors, that is, the error samples, are found, and the error samples are analyzed to find the reasons for the recognition errors, such as the prompt words are not accurate enough, the generator's understanding of certain features is insufficient, etc.

[0105] With the goal of improving the reward results and eliminating the error reasons, calculate the gradients of the relevant factors of the optimizer based on the backpropagation algorithm, start iterative training based on the direction of the gradients to cyclically adjust the initial prompt words, and optimize the reward function based on the iterative training.

[0106] In some embodiments, with the goal of improving the reward results and eliminating the error reasons, the optimizer, based on the backpropagation algorithm, starts from the output result of the reward function and calculates the gradients of the reward function with respect to the relevant factors inside the optimizer (such as model parameters, weights in the prompt word generation logic, etc.) according to the chain rule, from back to front. The gradient indicates how the value of the reward function will change if these factors are changed. For example, if the gradient of a certain factor is positive, it means that increasing the value of this factor will increase the value of the reward function.

[0107] According to the calculated gradient, use gradient descent (or other similar optimization strategies) to adjust the optimizer. During the adjustment process, the optimizer will modify the generation method of the initial prompt, thereby adjusting the initial prompt. After each adjustment, the new initial prompt and training samples are input into the generator again, and the process of repeating the generation of recognition results and calculating the reward results forms a cyclic training. In this cycle, the reward function will also be continuously optimized according to new feedback. For example, adjust the weights of various factors in the reward function so that the reward results can more accurately reflect the performance of the prompt and the generator.

[0108] In some embodiments, the hyperparameters used in the backpropagation algorithm include batch size, Dropout, BN, and model initialization. The batch size refers to the number of samples selected for evaluation each time. Dropout is used to delete duplicate content in the prompt. It is achieved by guiding the LLM to delete and organize duplicate techniques. BN is used to summarize the content of the prompt and is achieved by guiding the LLM to summarize the original prompt. Model initialization is achieved by adding manually written prompts to the population.

[0109] When it is detected that the number of loop iterations reaches the preset number of iterations and / or the reward result reaches the preset result, the prompt output by the optimizer is determined as the target prompt.

[0110] In some embodiments, the prompt generation system evaluates two key metrics. The first metric is the accuracy of the generated Prompt on the validation set, which reflects the actual performance of the Prompt when extracting accounts from other chat software. The second metric is the length of the Chain of Thought (CoT) in the prompt, which reflects the time required for the prompt to process each piece of data and thus determines the amount of data that the system can process per day. Table 1 shows the prompts with the lowest error rate on the validation set for each method and their corresponding CoT lengths.

[0111] Table 1

[0112]

[0113] Figure 6 and Figure 7 shows the changing trends of the prompt error rate and CoT length during the optimization process for different training methods.

[0114] By comprehensively considering the stability of CoT and the inference speed, we manually selected a CoT with an error rate of 6.5% and a length of 577 generated by BN + Dropout at the 100th training step. Its complete content is:

[0115] 1. **Identify potential account strings**: Find all possible account strings in the conversation.

[0116] 2. **Process the string**: Remove line breaks, HTML tags, spaces, special characters (such as "+", "[url]", "[tel]", "[mail]"), URLs, and emoticons. Keep the minus sign ("-") and underscore ("_"), as they may be part of the account of other chat software.

[0117] 3. **Check clarity**: Confirm whether the conversation explicitly mentions asking the other party to add accounts of other chat software, or uses vague expressions such as "伽", "伽", "加", "+", "v", "wx", "可添", "+", "薇", "腹制", "卫星号", "朋友圈", "备注", "后售", "客服" and so on. Use contextual prompts to verify the legitimacy of potential accounts of other chat software.

[0118] 4. **Verify the characteristics of other chat software accounts**: Make sure that the other chat software accounts meet the characteristics, including length between 6 and 20 characters, containing letters, numbers, underscores or minus signs, and not starting with a number. Remove special characters and then verify the length to ensure that the other chat software accounts do not contain multiple consecutive underscores or minus signs.

[0119] 5. **Verify contextual rationality**: After identifying potential accounts of other chat software, further verify the rationality of the context to ensure that the string is not other identifiers such as technical parameters, product models, etc. Identify specific contextual prompt words to ensure that these words indicate that the string is a product model or advertising content rather than an account of other chat software.

[0120] 6. **Exclude specific character combinations**: Exclude specific character combinations, such as consecutive letters and numbers, which are common in technical parameters and product models, but do not conform to the characteristics of accounts in other chat software. Pay special attention to excluding strings in the format of "vatti_cn" and identify these strings as public account names.

[0121] 7. **Output result**: If both clarity and context verification are passed, the account and reasoning steps of other chat software are output; otherwise, null and reasoning steps are output.

[0122] In some embodiments, in each batch of training, the samples that generate errors are fed back to the optimizer. Then, the reasons for the errors are analyzed based on the error samples, and the content of the prompt is adjusted according to these reasons for the errors. This process is repeated. When it is detected that the number of loop iterations reaches the preset number of iterations, it indicates that enough rounds of attempts and optimizations have been carried out. Or when the reward result reaches the preset result, it means that the current combination of the prompt and the generator has reached the desired performance level. When either of these two situations occurs, the prompt output by the optimizer at this time is determined as the target prompt.

[0123] In this embodiment, the reward result and the error samples are fed back to the optimizer for error cause analysis, enabling the optimizer to accurately locate the problem and specifically improve the generation strategy of the initial prompt. Based on the loop training of the backpropagation algorithm, the relevant factors of the optimizer are continuously adjusted to gradually optimize the initial prompt, thereby making the recognition result of the generator more accurate and improving the reward result. As the loop training progresses, the reward function is also optimized to more reasonably evaluate the performance of the prompt and the generator. When the preset number of iterations is reached or the reward result meets the standard, the target prompt is determined, ensuring that the finally obtained prompt can effectively guide the generator to accurately identify sensitive information.

[0124] In an embodiment of the present disclosure, the forms of communication messages include single messages and conversation messages. The training samples include a single-message drainage data set and a conversation-message drainage data set. Based on the configured optimizer and generator, iterative adjustment operations of the initial prompt are performed to generate the target prompt, further including: generating a first target prompt based on the single-message drainage data set, where the first target prompt is used to configure the output format of the sensitive information corresponding to the single message; generating a second target prompt based on the conversation-message drainage data set, where the second target prompt is used to configure the output format of the sensitive information corresponding to the conversation message.

[0125] In some embodiments, the prompt for a single message requires the model to output the reasoning steps, the result of the clarity check, the result of the context check, and the content of the sensitive information. The prompt for a conversation message only requires the output of the drained sensitive information, and the output is empty if the conditions are not met. Through these prompts, the generalization ability of the large language model and the requirements for scenario customization are taken into account.

[0126] In some embodiments, for a single message, the model is required to output inference steps, clarity check results, context check results, and prompt words for sensitive information, fully mining the detailed information of the single message. This not only enables the large language model to exert its generalization ability and analyze the single message from different dimensions, but also can accurately meet the customized requirements for the identification and analysis of sensitive information in the scenario of a single message through the output of the inference process and various check results. For a conversation message, only the sensitive information for diversion and the prompt words that output nothing if the conditions are not met are required, fully considering the characteristics of multi-round interaction and complex information in the conversation message. Such concise and targeted prompt words not only exert the generalization ability of the large language model to process diverse conversation content, but also precisely meet the customized requirements for quickly locating key sensitive information in the scenario of conversation messages.

[0127] In one embodiment of the present disclosure, a matching sensitive information generation model is trained based on a target prompt word and a communication message in a corresponding form, so that the sensitive information generation model outputs sensitive information in a corresponding output format, including: training a classification model based on the communication message to obtain an identification result of whether the communication message is diversion information based on the output of the classification model; if it is identified as diversion information, preprocessing the diversion information to obtain preprocessing data; training the sensitive information generation model based on the target prompt word corresponding to the form of the diversion information and the preprocessing data, so that the sensitive information generation model outputs sensitive information.

[0128] In some embodiments, training a classification model based on the communication message enables the model to deeply learn various features and patterns in the communication message, so as to accurately output an identification result of whether the communication message is diversion information. Training the sensitive information generation model using the target prompt word corresponding to the form of the diversion information and the preprocessing data, the target prompt word can guide the model to focus on the key sensitive content in different forms of diversion information. Combining with high-quality preprocessing data, it prompts the sensitive information generation model to accurately learn and output sensitive information.

[0129] In this embodiment, a complete and efficient processing flow is formed from identifying diversion information, preprocessing it, to generating sensitive information, which can greatly improve the processing ability of diversion information and sensitive information in communication messages, effectively ensuring the safe and compliant operation of the communication platform and reducing the risks that may be brought by diversion behavior.

[0130] As Figure 8 shown, in one embodiment of the present disclosure, training a classification model based on the communication message to obtain an identification result of whether the communication message is diversion information based on the output of the classification model includes:

[0131] Step S802: The classification model includes a binary classification BERT model and a Tiny BERT model. The pre-trained BERT model is fine-tuned using labeled data with drainage annotations to obtain the binary classification BERT model, which is used to calculate the drainage confidence of communication messages.

[0132] In some embodiments, the classification model includes a binary classification model based on pre-trained BERT and a feature generation model based on Tiny-BERT, which is used to determine whether the text is drainage text.

[0133] Step S804: Based on the communication messages, train the Tiny BERT model as a drainage feature calculation model, so that the trained Tiny BERT model calculates the drainage features of the communication messages to obtain feature vectors related to drainage.

[0134] In some embodiments, about 50,000 pieces of labeled data are used in the training process to fine-tune the model parameters, and through the CoSENT idea and the Gamma engine of the Vearch platform, efficient matching and label determination are achieved to counter black and gray production drainage samples.

[0135] Step S806: Compare the feature vectors with the vectors of known drainage features in the preset vector library to calculate the feature similarity.

[0136] Step S808: Adopt a fusion strategy to fuse the drainage confidence and the feature similarity to obtain a comprehensive judgment index.

[0137] In some embodiments, an appropriate fusion strategy (such as weighted average) is adopted to synthesize the drainage confidence output by the binary classification BERT model and the feature similarity output by the feature calculation Tiny-BERT model. According to the performance and importance of the two models in training, different weights are assigned to their outputs, and the combined calculation is used to obtain the comprehensive judgment index.

[0138] Step S810: Obtain the recognition result of whether the communication message is drainage information based on the relationship between the comprehensive judgment index and the judgment threshold.

[0139] In some embodiments, according to the comprehensive judgment index, a reasonable threshold is set. If the index exceeds the threshold, it is determined that the communication message is drainage information; if it is lower than the threshold, it is determined that it is not drainage information, and the final recognition result is output.

[0140] In this embodiment, the pre-trained BERT model is fine-tuned using annotated data with drainage annotations to obtain a binary classification BERT model, which can deeply mine the language features and semantic information in the communication message, accurately calculate the drainage confidence of the communication message, and provide an important probabilistic basis for judging whether the message is drainage information, so that the trained Tiny BERT model can effectively extract features related to drainage, generate feature vectors, provide data support for drainage judgment from the feature level, compare the feature vector with the known drainage feature vector in the preset vector library to calculate the feature similarity, further enrich the judgment information from the feature similarity dimension, and use a fusion strategy to fuse the drainage confidence and feature similarity to obtain a comprehensive judgment index. This multi-dimensional information fusion method avoids the limitations of a single judgment basis, comprehensively and accurately reflects the degree of correlation between the communication message and the drainage information, and obtains the recognition result of whether the communication message is drainage information based on the relationship between the comprehensive judgment index and the judgment threshold, thereby improving the accuracy and reliability of the recognition of whether the communication message is drainage information, and can effectively filter out drainage information.

[0141] In one embodiment of the present disclosure, the drainage information is preprocessed to obtain preprocessed data, including:

[0142] If the communication message is not a single message, organize the non-single message into a conversation message;

[0143] For the dialogue message or if the communication message is a single message, data filtering and data desensitization operations are performed respectively to obtain preprocessed data.

[0144] like Figure 9 As shown, in some embodiments, it is crucial to build an efficient data processing Pipeline, which performs three main tasks through the data layer: conversation organization, data filtering, and data desensitization. First, the scattered single conversation data is organized into continuous conversations to effectively improve the recall rate of the overall process. Data filtering reduces the amount of data that needs to be processed by the large language model through selection based on message attributes, regular expression and simple rule filtering, and filtering based on classification models.

[0145] In some embodiments, data desensitization replaces sensitive personal information and easily misidentified information through regular expressions, reducing the probability of LLM refusing to answer and reducing misidentification. All data processing steps ensure that the information is strictly filtered and desensitized to improve the efficiency and accuracy of the system.

[0146] like Figure 10As shown, the single-message drainage dataset and the session-message drainage dataset are respectively input into the APO module. The APO processes the input data, converts the single-message drainage dataset into a single-message Prompt, and converts the session-message drainage dataset into a session-message Prompt.

[0147] Both the single-message Prompt and the session-message Prompt are input into the GLM-4 model. The GLM-4 generates a single-message generation result based on the single message and the single-message Prompt, that is, the recognition result of sensitive information in the single message, and generates a session-message generation result based on the session message and the single-message Prompt, that is, the recognition result of sensitive information in the session message.

[0148] In an embodiment of the present disclosure, the form of the communication message is a dialogue message. Reliability detection of the corresponding sensitive information is performed based on the detection rules corresponding to the form, including:

[0149] The sensitive information generation model outputs sensitive information, and detects whether the corresponding drainage confidence is greater than the first confidence threshold; and wherein, if the drainage confidence is greater than the first confidence threshold, it is determined that the reliability detection is passed.

[0150] As Figure 11 shown, in some embodiments, for dialogue data, when the accuracy of the classification model is greater than 0.75 and the large language model successfully extracts sensitive information, the data is recorded.

[0151] In an embodiment of the present disclosure, the form of the communication message is a single message. Reliability detection of the corresponding sensitive information is performed based on the detection rules corresponding to the form, including:

[0152] The sensitive information generation model outputs sensitive information, and detects whether the sensitive information meets the clarity condition and the context condition; and if the sensitive information can match the preset sensitive dictionary, continue to detect whether the sensitive information meets the confidence detection condition and / or the historical recognition detection condition. Among them, if the sensitive information meets the clarity condition and the context condition, and meets the confidence detection condition and / or the historical recognition detection condition, it is determined that the reliability detection is passed.

[0153] As Figure 11 shown, in some embodiments, the single-message generation result needs to satisfy that the clarity check result is True, that is, the expression of sensitive information in the message is clear and unambiguous without ambiguity.

[0154] In some embodiments, the context check result is also True, indicating that from the context in which the message is located, the string is reasonable as sensitive information.

[0155] In one embodiment of the present disclosure, the confidence detection conditions include that the average confidence of the classification model is greater than the second confidence threshold and the maximum confidence of the classification model is greater than the third confidence threshold; the historical recognition detection conditions include that the number of times of being recognized as sensitive information is greater than the number threshold and the proportion of being recognized as sensitive information is greater than the proportion threshold.

[0156] In some embodiments, a dictionary is constructed using the historical recognition records of single messages in the past year, and the strings (i.e., keys) that meet the above basic conditions are matched with the sensitive information in the string format extracted from the conversation data in the dictionary. If a match is found, the next step of judgment is entered.

[0157] In some embodiments, it is checked whether the average confidence of the classification model for the message is greater than 0.5, indicating that the model has a relatively high overall judgment confidence for the message being drainage information. At the same time, the highest confidence of the classification model should be greater than 0.9999, indicating that the model has a very high certainty that the message is drainage information. When both of these conditions are met, Flag1 is generated.

[0158] In some embodiments, the number of times the string is recognized as sensitive information is counted, which needs to be greater than 2 times, indicating that it has been determined as sensitive information multiple times in history. The proportion of the string being recognized as sensitive information is calculated, which should be greater than 0.5, indicating that in relevant judgments, it is considered to be sensitive information in most cases. When both of these conditions are met, Flag2 is generated.

[0159] In some embodiments, when the basic conditions are met and (Flag1 is established or Flag2 is established), the string is regarded as drainage sensitive information and recorded in the database.

[0160] In this embodiment, based on the multi-dimensional detection process, the accuracy of sensitive information recognition is effectively improved, and the probability of misjudgment and missed judgment is reduced to ensure the reliability of the subsequent generated sensitive word library.

[0161] In one embodiment of the present disclosure, when a communication message is obtained, a sensitive word matching operation is performed based on the sensitive word library, including: when a communication message is obtained, the recognition confidence of the communication message is calculated based on the classification model; if the recognition confidence is greater than the fourth confidence threshold, a sensitive word matching operation is performed based on the sensitive word library; if the recognition confidence is less than or equal to the fourth confidence threshold, the recognition result of whether the communication message is drainage information is output again based on the classification model.

[0162] In some embodiments, if the calculated recognition confidence is greater than a preset fourth confidence threshold, it indicates that the model has a high confidence in the category judgment of the communication message. At this time, the communication message is matched with a pre-constructed sensitive word library, which stores a series of sensitive words related to drainage information. Whether there is a sensitive word is searched in the communication message. If a sensitive word is found, the communication message can be determined to be drainage information. If not, it can be determined that it is not drainage information.

[0163] In some embodiments, when the recognition confidence is less than or equal to the fourth confidence threshold, it indicates that the model lacks confidence in the category judgment of the communication message. At this time, the communication message is re-input into the classification model to allow the model to perform recognition again and output the recognition result of whether the communication message is drainage information.

[0164] In this embodiment, the recognition confidence of the communication message is initially calculated by the classification model. According to the comparison result between the confidence and the threshold, different processing strategies are adopted respectively. For communication messages with high confidence, the sensitive word library matching is used to further accurately determine whether it is drainage information. For messages with low confidence, the classification model is used again for recognition. This way effectively combines the feature analysis ability of the classification model and the precise matching advantage of the sensitive word library, improves the accuracy and reliability of the recognition of drainage information in communication messages, can filter and manage communication messages more effectively, ensure the normal operation and security of the communication system, and reduce the occurrence of misjudgment and missed judgment situations.

[0165] In an embodiment of the present disclosure, the matching operation of sensitive words based on the sensitive word library includes: detecting whether there is a sensitive word recorded in the sensitive word library in the communication message based on the traversal operation of the prefix tree.

[0166] In some embodiments, by updating the recognized sensitive information to the sensitive word library in real time and using an efficient matching strategy based on the prefix tree (Trie Tree), sensitive information matching at the millisecond level is achieved. At the business level, a prompt message: "Please do not conduct transactions outside this platform to avoid economic losses caused by private transactions." is added to the samples that match sensitive information. This measure improves users' vigilance, protects the interests of the platform, and safeguards the legitimate rights and interests of users.

[0167] This system matches the account numbers of other chat software in the conversation data of this platform through the account numbers of other chat software generated by APO, effectively improving the recall rate of drainage samples by about 29.3%. The examples in Table 2 are the chat data of the account numbers of other chat software newly recognized by this system and the confidence levels of the online classification model (<0.95 not recognized).

[0168] Table 2

[0169]

[0170] In this embodiment, by using a large language model in the field of retail risk control to extract drainage accounts, through data such as classification models, semantic matching models, and statistical rules, the recall rate of drainage samples is effectively improved. By using the idea of backpropagation to automatically generate the Chain-of-Thought (CoT) part of the prompt words, compared with manually written prompt words, the thought chain generated based on backpropagation has more comprehensive reasoning logic and better reasoning effects. During the calculation process of backpropagation, regularization operations such as BN and Dropout are added through the large language model. Without affecting the effect of generating prompt words, the length of the generated prompt words is effectively shortened, and the running time during reasoning using the prompt words is improved.

[0171] It should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.

[0172] Those skilled in the art of the relevant technology can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuits", "modules", or "systems" here.

[0173] The following refers to Figure 12 to describe the sensitive information processing device 1200 according to this embodiment of the present disclosure. Figure 12 The shown sensitive information processing device 1200 is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.

[0174] The sensitive information processing device 1200 is embodied in the form of a hardware module. The components of the sensitive information processing device 1200 may include, but are not limited to: an iterative adjustment module 1202, which is used to perform iterative adjustment operations on the initial prompt word based on the configured optimizer and generator to generate a target prompt word, and the target prompt word is used to configure the output format of the sensitive information corresponding to the form of the communication message; a training module 1204, which is used to train a matching sensitive information generation model based on the target prompt word and the communication message of the corresponding form, so that the sensitive information generation model outputs sensitive information with the corresponding output format; a detection module 1206, which is used to perform reliability detection on the corresponding sensitive information based on the detection rule corresponding to the form; a matching module 1208, which is used to construct a sensitive word library based on the sensitive information that passes the reliability detection, so as to perform sensitive word matching operations based on the sensitive word library when obtaining a communication message, and generate governance information for the matched sensitive word.

[0175] Next, reference is made to Figure 13 to describe the electronic device 1300 according to this embodiment of the present disclosure. Figure 13 The shown electronic device 1300 is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0176] As Figure 13 shown, the electronic device 1300 is embodied in the form of a general computing device. The components of the electronic device 1300 may include, but are not limited to: the above-mentioned at least one processing unit 1310, the above-mentioned at least one storage unit 1320, and a bus 1330 connecting different system components (including the storage unit 1320 and the processing unit 1310).

[0177] Among them, the storage unit stores program codes, and the program codes can be executed by the processing unit 1310, so that the processing unit 1310 executes the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification. For example, the processing unit 1310 can execute steps S202 and S208 as shown in Figure 2 and other steps defined in the sensitive information processing method of the present disclosure.

[0178] The storage unit 1320 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 13201 and / or a cache storage unit 13202, and may further include a read-only storage unit (ROM) 13203.

[0179] The storage unit 1320 may also include a program / utility 13204 having a set (at least one) of program modules 13205. Such program modules 13205 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0180] The bus 1330 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.

[0181] The electronic device 1300 may also communicate with one or more external devices 1360 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device, and / or may communicate with any device that enables the electronic device 1300 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be through the input / output (I / O) interface 1350. Also, the electronic device 1300 may communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 1350. As shown in the figure, the network adapter 1350 communicates with other modules of the electronic device 1300 through the bus 1330. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc. Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described here can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to cause a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0182] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium having stored thereon a program product capable of implementing the above-described method of this specification. In some possible implementation manners, various aspects of the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.

[0183] The program product for implementing the above method according to an embodiment of the present disclosure may be a portable compact disc read-only memory (CD-ROM) and include the program code, and may run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.

[0184] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0185] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries readable program code. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing. The program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, executed as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0186] It should be noted that although several modules or units of a device for action execution are mentioned in the foregoing detailed description, such a division is not mandatory. In fact, according to an embodiment of the present disclosure, the features and functions of two or more of the above-described modules or units may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by multiple modules or units. In addition, although the various steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in that specific order, or that all of the steps shown must be performed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.

[0187] From the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on the network, including several instructions to enable a computing device (such as a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure. After considering the specification and practicing the present disclosure, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.

Claims

1. A method for processing sensitive information, characterized in that, Including: Performing iterative adjustment operations on the initial prompt word based on a configured optimizer and generator to generate a target prompt word, where the target prompt word is used to configure the output format of sensitive information corresponding to the form of the communication message; Training a matching sensitive information generation model based on the target prompt word and the communication message in the corresponding form, so that the sensitive information generation model outputs the sensitive information with the corresponding output format; Performing reliability detection on the corresponding sensitive information based on the detection rule corresponding to the form; Constructing a sensitive word library based on the sensitive information that passes the reliability detection, so that when the communication message is obtained, performing a matching operation of sensitive words based on the sensitive word library and generating governance information for the matched sensitive words.

2. The sensitive information processing method according to claim 1, wherein Performing iterative adjustment operations on the initial prompt word based on a configured optimizer and generator to generate a target prompt word, including: Extracting multiple prompt training information from the training samples, where the training samples are labeled with the sensitive information; Inputting the multiple prompt training information into the optimizer for merging to generate the initial prompt word; Inputting the initial prompt word and the training samples into the generator to generate the recognition result of the sensitive information; Inputting the comparison result between the recognition result and the corresponding label, the length of the initial prompt word, and the time for generating the recognition result into the reward function to obtain a reward result; Feeding back the reward result and the error samples obtained based on the recognition result to the optimizer to cyclically train the optimizer, so that the optimizer cyclically adjusts the initial prompt word until the target prompt word is obtained.

3. The sensitive information processing method according to claim 2, wherein Feeding back the reward result and the error samples obtained based on the recognition result to the optimizer to cyclically train the optimizer, so that the optimizer cyclically adjusts the initial prompt word until the target prompt word is obtained, including: Feeding back the reward result and the error samples to the optimizer, so that the optimizer analyzes the error cause based on the error samples; Aiming to improve the reward result and eliminate the error cause, calculating the gradient of the relevant factors of the optimizer based on the backpropagation algorithm, starting the cyclic training based on the direction of the gradient to cyclically adjust the initial prompt word, and optimizing the reward function based on the cyclic training; Detecting that the number of cyclic iterations reaches the preset number of iterations, and / or the reward result reaches the preset result, and determining the prompt word output by the optimizer as the target prompt word.

4. The sensitive information processing method according to claim 2, wherein The form of the communication message includes single-message and conversation messages, the training samples include a single-message drainage data set and a conversation-message drainage data set, and performing iterative adjustment operations on the initial prompt word based on a configured optimizer and generator to generate a target prompt word further includes: Generating a first target prompt word based on the single-message drainage data set, where the first target prompt word is used to configure the output format of the sensitive information corresponding to the single-message; Generating a second target prompt word based on the conversation-message drainage data set, where the second target prompt word is used to configure the output format of the sensitive information corresponding to the conversation message.

5. The sensitive information processing method according to claim 1, wherein Training a matching sensitive information generation model based on the target prompt and the communication message in the corresponding form, so that the sensitive information generation model outputs the sensitive information in the corresponding output format, including: Training a classification model based on the communication message to obtain an identification result of whether the communication message is drainage information based on the output of the classification model; If it is identified as the drainage information, preprocess the drainage information to obtain preprocessed data; Training a sensitive information generation model based on the target prompt corresponding to the form of the drainage information and the preprocessed data, so that the sensitive information generation model outputs the sensitive information.

6. The sensitive information processing method according to claim 5, wherein Training a classification model based on the communication message to obtain an identification result of whether the communication message is drainage information based on the output of the classification model, including: The classification model includes a binary classification BERT model and a Tiny BERT model. The pre-trained BERT model is fine-tuned using labeled data with drainage annotations to obtain the binary classification BERT model, and the binary classification BERT model is used to calculate the drainage confidence of the communication message; Training the Tiny BERT model as a drainage feature calculation model based on the communication message, so that the trained Tiny BERT model performs drainage feature calculation on the communication message to obtain a feature vector related to drainage; Comparing the feature vector with the vector of the known drainage feature in the preset vector library to calculate the feature similarity; Adopting a fusion strategy to fuse the drainage confidence and the feature similarity to obtain a comprehensive judgment index; Obtaining the identification result of whether the communication message is the drainage information based on the relationship between the comprehensive judgment index and the judgment threshold.

7. The sensitive information processing method according to claim 5, wherein Preprocessing the drainage information to obtain preprocessed data, including: If the communication message is not a single message, organizing the non-single message into a conversation message; Performing data filtering and data desensitization operations on the conversation message or if the communication message is a single message respectively to obtain the preprocessed data.

8. The sensitive information processing method according to claim 6, wherein When the form of the communication message is a conversation message, performing reliability detection on the corresponding sensitive information based on the detection rules corresponding to the form, including: The sensitive information generation model outputs the sensitive information, and detecting whether the corresponding drainage confidence is greater than the first confidence threshold; and Wherein, if the drainage confidence is greater than the first confidence threshold, it is determined that the reliability detection is passed.

9. The sensitive information processing method according to claim 6, wherein When the form of the communication message is a single message, performing reliability detection on the corresponding sensitive information based on the detection rules corresponding to the form, including: The sensitive information generation model outputs the sensitive information, and detecting whether the sensitive information meets the clarity condition and the context condition; and If the sensitive information can be matched with the preset sensitive dictionary, continue to detect whether the sensitive information meets the confidence detection condition and / or the historical identification detection condition. Among them, if the sensitive information meets the clarity condition and the context condition, and meets the confidence detection condition and / or the historical recognition detection condition, it is determined that the reliability detection is passed.

10. The sensitive information processing method according to claim 9, wherein the confidence detection condition includes that the average confidence of the classification model is greater than a second confidence threshold and the maximum confidence of the classification model is greater than a third confidence threshold; the historical recognition detection condition includes that the number of times of identifying the sensitive information is greater than a number threshold and the proportion of identifying the sensitive information is greater than a proportion threshold.

11. The sensitive information processing method according to claim 5, characterized in that, When the communication message is obtained, a matching operation of sensitive words is performed based on the sensitive word library, including: When the communication message is obtained, the recognition confidence of the communication message is calculated based on the classification model; If the recognition confidence is greater than a fourth confidence threshold, a matching operation of sensitive words is performed based on the sensitive word library; If the recognition confidence is less than or equal to the fourth confidence threshold, the recognition result of whether the communication message is a drainage message is output again based on the classification model.

12. The sensitive information processing method according to claim 1, wherein The matching operation of sensitive words based on the sensitive word library includes: Detecting whether there is a sensitive word recorded in the sensitive word library in the communication message based on a traversal operation of a prefix tree.

13. A sensitive information processing device, characterized in that, Including: An iterative adjustment module, configured to perform an iterative adjustment operation of an initial prompt word based on a configured optimizer and a generator to generate a target prompt word, where the target prompt word is used to configure an output format of sensitive information corresponding to the form of the communication message; A training module, configured to train a matching sensitive information generation model based on the target prompt word and the communication message in the corresponding form, so that the sensitive information generation model outputs the sensitive information with the corresponding output format; A detection module, configured to perform a reliability detection on the corresponding sensitive information based on a detection rule corresponding to the form; A matching module, configured to construct a sensitive word library based on the sensitive information that passes the reliability detection, so as to perform a matching operation of sensitive words based on the sensitive word library when the communication message is obtained, and generate governance information for the matched sensitive words.

14. An electronic device, characterized in that, Including: A processor; and A memory, configured to store executable instructions of the processor; wherein, the processor is configured to execute the sensitive information processing method according to any one of claims 1 to 12 by executing the executable instructions.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the sensitive information processing method according to any one of claims 1 to 12.

16. A computer program product, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, it implements the sensitive information processing method according to any one of claims 1 to 12.