Sensitive information detection method and content display method in question and answer scene

By obtaining streaming output information of large language models in question-and-answer scenarios and performing sensitive word detection, the problem of low detection efficiency in the prior art is solved, real-time detection and rejection of sensitive information is realized, and user experience is improved.

CN120124747APending Publication Date: 2025-06-10HUA DATA TECH (SHANGHAI) CO LTD
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510214151.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

When the prior art detects sensitive words for output information in question-and-answer scenarios, the detection efficiency is not high, resulting in an increase in user waiting time and affecting user experience.

Method used

By obtaining the streaming output information of the large language model, it is stored in a preset buffer, and combined with the original information into pending text. When the pending text length exceeds the preset length, sensitive word detection is performed.

Benefits of technology

Real-time detection and rejection of answers to sensitive information during the Q&A process is realized, which improves detection efficiency, reduces user waiting time, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124747A_ABST
    Figure CN120124747A_ABST
Patent Text Reader

Abstract

The invention provides a sensitive information detection method and content display method in a question and answer scene, and the method comprises the steps: obtaining the current output information of a large language model, and storing the output information to a preset buffer area; wherein the output form of the output information is streaming output; taking the output information and original information stored in a preset buffer area as a current to-be-processed text; in response to the fact that the data length of the current to-be-processed text exceeds a preset length, taking the current to-be-processed text as a target detection text; wherein the preset length is determined according to a pre-constructed sensitive word library; and performing sensitive word detection on the target detection text to obtain the detection result, thereby improving the detection speed of the sensitive information, reducing the waiting time of the user and improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data processing, and particularly to a method for detecting sensitive information and a method for presenting content in a question-and-answer scenario. Background Art

[0002] With the rapid development of artificial intelligence technology, question-and-answer scenarios based on large language models play an increasingly crucial role in many fields. However, how to efficiently identify and reject answering sensitive information during the question-and-answer interaction has become a difficult problem that needs to be solved urgently. Traditional methods for detecting sensitive information usually detect the complete output content generated by the large language model after it is generated. This means that only when the sensitive information in the entire content passes the detection will the output content be presented to the user. This detection mode not only reduces the detection efficiency but also increases the waiting time of the user, thus affecting the user experience. Therefore, it is particularly important to develop a method that can detect and reject answering sensitive information in real time during the question-and-answer process. This can not only effectively protect the information security of users and enterprises but also increase the user's trust in the large language model. In addition, with the increasingly strict laws and regulations on data privacy and information security in various countries around the world, ensuring the compliance of the question-and-answer system has also become a key technical requirement.

[0003] Currently, there are various technical means available for detecting and identifying sensitive information, including: keyword recognition and filtering, semantic analysis method, and machine learning classifier detection method; among them, keyword recognition and filtering identify sensitive information in the text through string matching with a predefined sensitive word library, but it is unable to flexibly set the detection intensity for sensitive words, resulting in a low user experience; the semantic analysis method determines whether to reject an answer by calculating the similarity between the user's question and the questions in the sensitive question library. This method relies on the model for similarity calculation and has a slow response speed in a low-resource environment; the machine learning classifier detection method requires training the model. In some dedicated large model scenarios, if special content needs to be rejected, the model needs to be retrained, and the iteration efficiency is low. Summary of the Invention

[0004] The technical problem to be solved by the present disclosure is to overcome the defect of low detection efficiency when detecting sensitive words in the output information in the question-and-answer scenario in the prior art, and to provide a method for detecting sensitive information and a method for presenting content in a question-and-answer scenario.

[0005] The present disclosure solves the above technical problem through the following technical solutions:

[0006] In the first aspect of the present disclosure, a method for detecting sensitive information in a question-and-answer scenario is provided, and the detection method includes:

[0007] Obtain the output information of the current time of the large language model, and store the output information in a preset buffer; wherein, the output form of the output information is streaming output;

[0008] Use the output information and the original information stored in the preset buffer as the current text to be processed;

[0009] In response to the data length of the current text to be processed exceeding a preset length, use the current text to be processed as the target detection text; wherein, the preset length is determined according to a pre-constructed sensitive word library;

[0010] Perform sensitive word detection on the target detection text to obtain a detection result.

[0011] Optionally, the detection method further includes:

[0012] In response to the data length of the current text to be processed being less than the preset length and the current question and answer of the large language model having ended, use the current text to be processed as the target detection text;

[0013] And / or;

[0014] In response to the detection result meeting a preset condition, release the first quantity of the first information starting from the initial character in the target detection text in the preset buffer; wherein, the first quantity is the data length corresponding to the output information of the previous time of the large language model;

[0015] And / or;

[0016] In response to the detection result not meeting the preset condition, end the current question and answer of the large language model.

[0017] Optionally, the step of performing sensitive word detection on the target detection text includes:

[0018] In response to any first word to be detected in the target detection text having a similarity greater than a first preset value with a preset sensitive word in the pre-constructed sensitive word library, determine that there is a sensitive word in the target detection text; otherwise, determine that there is no sensitive word in the target detection text;

[0019] And / or;

[0020] Before the step of obtaining the output information of the current time of the large language model, it further includes:

[0021] Obtain external input information, and perform sensitive word detection on the external input information;

[0022] In response to the similarity between any second word to be detected in the external input information and a preset sensitive word in the pre-constructed sensitive word library being greater than a second preset value, it is determined that there is a sensitive word in the external input information, and a preset rejection response is returned; otherwise, the external input information is input into the large language model for question-and-answer processing.

[0023] The second aspect of the present disclosure provides a content display method in a question-and-answer scenario. The content display method includes:

[0024] In response to the absence of a sensitive word in the external input information, the external input information is input into the large language model to obtain a detection result of sensitive word detection corresponding to the target detection text;

[0025] Based on the detection result, the target content is displayed;

[0026] Wherein, the detection result is obtained based on the detection method of sensitive information in the question-and-answer scenario described in the first aspect.

[0027] Optionally, the step of displaying the target content based on the detection result includes:

[0028] In response to the detection result not meeting the preset conditions, a preset rejection response is displayed;

[0029] In response to the detection result meeting the preset conditions, the target detection text is displayed.

[0030] The third aspect of the present disclosure provides a detection system for sensitive information in a question-and-answer scenario. The detection system includes:

[0031] An acquisition module, configured to acquire the output information of the large language model for the current time and store the output information in a preset buffer; wherein, the output form of the output information is streaming output;

[0032] A determination module, configured to use the output information and the original information stored in the preset buffer as the current text to be processed;

[0033] A response module, configured to use the current text to be processed as the target detection text when the data length of the current text to be processed exceeds a preset length; wherein, the preset length is determined according to the pre-constructed sensitive word library;

[0034] A detection module, configured to perform sensitive word detection on the target detection text to obtain a detection result.

[0035] Optionally, the detection system further includes:

[0036] The first response unit is configured to use the current text to be processed as the target detection text when the data length of the current text to be processed is less than the preset length and the current Q&A of the large language model has ended.

[0037] and / or;

[0038] The second response unit is configured to release the first quantity of first information starting from the initial character in the target detection text in the preset buffer when the detection result meets the preset conditions; wherein, the first quantity is the data length corresponding to the output information of the large language model in the previous time.

[0039] and / or;

[0040] The third response unit is configured to end the current Q&A of the large language model when the detection result does not meet the preset conditions.

[0041] Optionally, the detection module is further configured to determine that there is a sensitive word in the target detection text when the similarity between any first word to be detected in the target detection text and a preset sensitive word in the pre-constructed sensitive word library is greater than a first preset value; otherwise, determine that there is no sensitive word in the target detection text.

[0042] and / or;

[0043] The detection system further includes a first detection unit, and the first detection unit is configured to obtain external input information and perform sensitive word detection on the external input information; in response to the similarity between any second word to be detected in the external input information and a preset sensitive word in the pre-constructed sensitive word library being greater than a second preset value, determine that there is a sensitive word in the external input information and return a preset rejection content; otherwise, input the external input information into the large language model for Q&A processing; otherwise, determine that there is a sensitive word in the external input information and return a preset rejection content.

[0044] A fourth aspect of the present disclosure provides a content display system in a Q&A scenario, and the content display system includes:

[0045] A response module, configured to input the external input information into a large language model to obtain a detection result of sensitive word detection corresponding to the target detection text when there is no sensitive word in the external input information.

[0046] A display module, configured to perform display of target content based on the detection result; wherein, the detection result is obtained based on the sensitive information detection system in the Q&A scenario described in the third aspect.

[0047] Optionally, the display module is specifically configured to display a preset rejection response when the detection result does not meet the preset condition; and display the target detection text when the detection result meets the preset condition.

[0048] The fifth aspect of the present disclosure provides an electronic device, including a memory, a processor, and a computer program stored on the memory and configured to run on the processor. When the processor executes the computer program, it implements the method for detecting sensitive information in the Q&A scenario as described in the first aspect, or the content display method in the Q&A scenario as described in the second aspect.

[0049] The sixth aspect of the present disclosure provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the method for detecting sensitive information in the Q&A scenario as described in the first aspect, or the content display method in the Q&A scenario as described in the second aspect.

[0050] The seventh aspect of the present disclosure provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the method for detecting sensitive information in the Q&A scenario as described in the first aspect, or the content display method in the Q&A scenario as described in the second aspect.

[0051] On the basis of conforming to the common knowledge in the art, the above optional conditions can be combined arbitrarily to obtain various preferred examples of the present disclosure.

[0052] The positive and progressive effects of the present disclosure are as follows: The output information in the form of streaming output in the Q&A scenario is stored in a preset buffer area and combined with the original information in the preset buffer area as the text to be processed. When the text to be processed exceeds the preset length, the current text to be processed can be used as the target detection text for sensitive word detection. The present application provides a lightweight method for detecting sensitive words, instantaneously detecting sensitive words for each output information in the Q&A scenario with a pre-constructed sensitive word library, aiming to maintain high response speed and detection accuracy, improve the detection speed of sensitive information with extremely low resource consumption, reduce the user waiting time, and enhance the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a flowchart of a method for detecting sensitive information in a Q&A scenario provided by Embodiment 1 of the present disclosure;

[0054] Figure 2 It is a flowchart of a content display method in a Q&A scenario provided by Embodiment 2 of the present disclosure;

[0055] Figure 3 It is a specific example diagram of a content display method in a Q&A scenario provided by Embodiment 2 of the present disclosure;

[0056] Figure 4 It is a schematic diagram of modules of a detection system for sensitive information in a question - answering scenario provided in Embodiment 3 of the present disclosure;

[0057] Figure 5 It is a schematic diagram of modules of a content display system in a question - answering scenario provided in Embodiment 4 of the present disclosure;

[0058] Figure 6 It is a schematic diagram of the structure of an electronic device provided in Embodiment 5 of the present disclosure. Specific embodiments

[0059] The present disclosure will be further described below by way of embodiments, but the present disclosure is not limited to the scope of the described embodiments.

[0060] In the embodiments of the present disclosure, prefix words such as "first" and "second" are only used to distinguish different described objects, and have no limiting effect on the position, order, priority, quantity, content, etc. of the described objects. The use of ordinal numbers and other prefix words for distinguishing described objects in the embodiments of the present disclosure does not constitute a limitation on the described objects. The statement of the described objects refers to the description in the context of the claims or embodiments, and should not constitute an unnecessary limitation due to the use of such prefix words. In addition, in the description of this embodiment, unless otherwise specified, "a plurality" means two or more.

[0061] In the embodiments of the present disclosure, the processing of collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0062] Embodiment 1

[0063] This embodiment provides a method for detecting sensitive information in an answer scenario, as Figure 1 shown, the detection method includes:

[0064] S101. Obtain the output information of the large - language model for the current time and store the output information in a preset buffer; wherein, the output form of the output information is streaming output.

[0065] In specific implementation, streaming output means that when the large - language model generates the output information of the answer, it does not output all the content at once, but outputs word by word and sentence by sentence in real time like a typewriter. When obtaining the output information that is output in real time by the large - language model, the newly output content will be immediately stored in a preset buffer.

[0066] S102. Use the output information and the original information stored in the preset buffer as the current text to be processed.

[0067] In a specific implementation, the output information generated by the large language model previously is stored in a preset buffer, and the output information of the current time of the large language model is combined with the original information stored in the preset buffer as the current text to be processed.

[0068] In a specific example, the original information stored in the preset buffer is "I have a", and the output information of the current time of the large language model is "apple". These two parts of information are combined to generate the current text to be processed "I have an apple".

[0069] S103. In response to the data length of the current text to be processed exceeding a preset length, the current text to be processed is used as the target detection text; wherein, the preset length is determined according to a pre-constructed sensitive word library.

[0070] In a specific implementation, the pre-constructed sensitive word library contains two word libraries: "single-word library" and "multi-word library"; wherein, the preset sensitive words in the "single-word library" are all words with a relatively high degree of sensitivity, and each word is independent of each other. For example, some words related to non-polite language content can be considered words with a relatively high degree of sensitivity; while the "multi-word library" refers to multiple word bags established in the word library, and each word bag contains some specific words. These specific words have a relatively low degree of sensitivity when they appear alone, but when they appear in combination with each other, their degree of sensitivity becomes serious. In a specific example, the "multi-word library" contains two word bags, A and B. Among them, word bag A contains some specific personal names, such as company executives, etc., and word bag B contains some bad words for slandering and insulting people. When the words in word bags A and B appear alone, their degrees of sensitivity are relatively low, but when word bags A and B are combined, that is, when personal names and slanderous and other bad words appear at the same time, their degree of sensitivity becomes high.

[0071] In another specific implementation, the preset length is the length of the longest word in the "single-word library". Specifically, the "single-word library" C is [**, ***, ***, ****, **], and at this time the preset length is the length of the word "****", that is, the preset length is 4.

[0072] In a specific example, the "word library" in the preset sensitive word library can be filled based on the existing sensitive word library. For the construction of the "multi-word library", in addition to directly filling it based on the existing sensitive word library, it can also be obtained for some collected refusal answers. Specifically, an entity extraction task is constructed through a large language model to extract nouns and verbs in the refusal answers respectively and place them in different word bags, and then synonyms are generated for the words in the word bags based on the large language model to expand the word bag where the word is located. For example, when the refusal answer is "How to purchase **", based on the large language model, the verb "purchase" and the noun "**" are extracted from this refusal answer respectively, and then near-synonyms are generated for the extracted words by the large language model. Finally, the constructed word bag X is [purchase, buy, purchase, …] and the word bag Y is [**, **X, *Y*, Z**, …]. The extraction rule for words in the refusal answer can be adaptively set according to needs.

[0073] In a specific example, the refusal answers can be obtained through questionnaire surveys or screened and selected from the existing refusal answer library.

[0074] S104. Perform sensitive word detection on the target detection text to obtain a detection result.

[0075] In a specific implementation, when performing sensitive word detection on the target detection text, it is necessary to compare the similarity between the target detection text and the pre-constructed sensitive word library to determine whether there are sensitive words in the current target detection text. If a word in the sensitive word library appears in the target detection text, the detection result is not passed at this time, otherwise it is passed.

[0076] In an alternative implementation manner, the detection method further includes:

[0077] In response to the data length of the current text to be processed being less than the preset length and the current Q&A of the large language model having ended, use the current text to be processed as the target detection text;

[0078] and / or;

[0079] In response to the detection result meeting the preset conditions, release the first quantity of first information starting from the initial character in the target detection text in the preset buffer; where the first quantity is the data length corresponding to the output information of the large language model in the previous time;

[0080] and / or;

[0081] In response to the detection result not meeting the preset conditions, end the current Q&A of the large language model.

[0082] In a specific implementation, taking the end of the large language model's question and answer as a judgment condition, it can be determined that the large language model will no longer generate new output information. In this case, it means that the content of the current text to be processed has been fixed and no new content will be added. At this time, if the data length of the current text to be processed is less than the preset length, then it will be directly used as the target detection text to perform subsequent detection operations. Based on this discrimination process, it is possible to ensure sensitive word detection for all output information of the large language model to ensure that the output content meets the requirements.

[0083] In a specific implementation, if the detection result indicates passing, that is, no sensitive words appear in the current target detection text and it meets the requirements, at this time, the first quantity of the first information starting from the initial character of the target detection text in the preset buffer is released. The first quantity is the data length corresponding to the previous output information of the large language model.

[0084] In a specific example, the target detection text in the preset buffer is "The giant panda is unique to China", and the previous output information of the large language model is "unique to", then it can be determined that the first quantity is 3. If the detection result shows passing after sensitive word detection on the target detection text "The giant panda is unique to China", at this time, 3 characters starting from the initial character of the target detection text will be released. Then the released first information is "The giant panda", and at this time, the target detection text in the preset buffer after release becomes "is unique to China". In another specific example, the target detection text in the preset buffer is "Lotus is native to tropical and temperate regions of Asia", and the previous output information of the large language model is "and temperate regions", then the first quantity is 5. If the detection result shows passing after sensitive word detection on the target detection text "Lotus is native to tropical and temperate regions of Asia", at this time, 5 characters starting from the initial character will be released. Then the released first information is "Lotus is native to", and after release, the target detection text in the preset buffer becomes "tropical and temperate regions of Asia".

[0085] In a specific implementation, if the sensitive word detection result of the target detection text shows not passing, that is, it indicates that there are words in the output information of the large language model that do not meet the requirements, the large language model will immediately end the current question and answer.

[0086] In an optional implementation manner, the step of performing sensitive word detection on the target detection text includes:

[0087] In response to the similarity between any first word to be detected in the target detection text and a preset sensitive word in the pre-constructed sensitive word library being greater than a first preset value, it is determined that there is a sensitive word in the target detection text; otherwise, it is determined that there is no sensitive word in the target detection text;

[0088] and / or;

[0089] Before the step of obtaining the output information of the current time of the large language model, the following steps are further included:

[0090] Obtain external input information and perform sensitive word detection on the external input information;

[0091] In response to the similarity between any second word to be detected in the external input information and the preset sensitive words in the pre-constructed sensitive word library being greater than the second preset value, it is determined that there are sensitive words in the external input information, and a preset rejection response is returned; otherwise, the external input information is input into the large language model for question and answer processing.

[0092] In specific implementation, the first word to be detected refers to any word extracted from the target detection text. The first word to be detected is traversed and compared with all the sensitive words in the sensitive word library, and the similarity value between it and each sensitive word is calculated. During the comparison process, if there is a similarity value greater than the first preset value, it is considered that there are sensitive words in the target detection text; otherwise, it is considered that there are no sensitive words in the target detection text. Among them, the first preset value can be adaptively adjusted according to actual needs, so as to flexibly control the strictness of sensitive word detection. Since the large language model outputs information in real time, the target detection text will be continuously updated, and the frequency of sensitive word detection is relatively high. Therefore, when performing sensitive word detection on the target detection text, only the sensitive words in the "word library" of the sensitive word library need to be compared for similarity, so as to improve the detection efficiency and enhance the user experience.

[0093] In specific implementation, before obtaining the output information of the current time of the large language model, sensitive word detection is also pre-performed on the question information externally input into the large language model. Only when it is confirmed that there are no sensitive words in the external input information, the obtained external input information will be input into the large language model for question and answer processing. Specifically, the second word to be detected refers to any word extracted from the external input information. The second word to be detected is traversed and compared with all the sensitive words in the sensitive word library, and the similarity value between it and each sensitive word is calculated. During the comparison process, if there is a similarity value greater than the second preset value, it is considered that there are sensitive words in the external input information; otherwise, it is considered that there are no sensitive words in the external input information. Among them, the second preset value can be adaptively adjusted according to actual needs. Since the sensitive word detection performed on the external input information belongs to the pre-detection step of the large language model, the delay brought to the overall response is a one-time delay. Therefore, when performing sensitive word detection on the external input information, the similarity calculation can be performed with the sensitive words in both the "word library" and the "multi-word library" of the pre-constructed sensitive word library, so as to increase the accuracy of sensitive word detection for the external input information.

[0094] In a specific example, when comparing the similarity with sensitive words in a pre-built sensitive word library, string matching can be used, or the trie algorithm or the KMP algorithm (Knuth-Morris-Pratt algorithm) can be used to efficiently obtain the detection result.

[0095] In a specific example, if there is a sensitive word in the external input information, the rejection content "There is an inappropriate word in the current input information. Please re-enter." is directly returned.

[0096] In another specific example, a user feedback mechanism can be set up to determine the strictness of the current sensitive word detection, and accordingly adjust the first preset value and the second preset value in a timely manner to enhance the user experience.

[0097] This application realizes precise sensitive word detection for the input information of users and the output information of the answers of the large language model in the Q&A scenario by constructing sensitive word libraries with different characteristics. In the actual usage scenario, users can flexibly customize the rejection content according to specific needs to meet the personalized needs of different scenarios. Different from the method that relies on the model itself to detect sensitive words, the sensitive word detection method of this application runs independently of the model. Even in low-resource scenarios, it can ensure a quick response and guarantee the user experience. In addition, this application also supports a user feedback mechanism, which can continuously optimize the detection strategy and effect based on the feedback information provided by users during use, thereby continuously improving the accuracy and adaptability of the system.

[0098] Embodiment 2

[0099] This embodiment provides a content display method in a Q&A scenario. The content display method is implemented based on the sensitive information detection method in the Q&A scenario described in Embodiment 1, as Figure 2 shown, the content display method includes:

[0100] S201. In response to the absence of a sensitive word in the external input information, input the external input information into a large language model to obtain the detection result of the sensitive word detection corresponding to the target detection text;

[0101] In a specific implementation, when there is no sensitive word in the external input information, it is input into the large language model to obtain the detection result of the sensitive word detection corresponding to the target detection text.

[0102] S202. Display the target content based on the detection result; wherein, the detection result is obtained based on the sensitive information detection method in the Q&A scenario described in Embodiment 1.

[0103] In an alternative embodiment, the step of presenting the target content based on the detection result includes:

[0104] In response to the detection result not meeting the preset condition, presenting a preset rejection content;

[0105] In response to the detection result meeting the preset condition, presenting the target detection text.

[0106] In a specific implementation, if the detection result is not passed, that is, there are sensitive words in the target detection text, a preset rejection content will be directly presented to the user, and the content that has been presented before will also be refreshed and overwritten with the preset rejection content. Specifically, when the detection result is not passed, the rejection content presented at this time is "The current system is unstable, please re-enter". If the detection result shows passed, that is, there are no sensitive words in the target detection text, the target detection text will be presented to the user at this time.

[0107] In a specific example, as Figure 3 shown, after the user inputs information, the input information will first be checked for sensitive words. Only the information that passes the detection will be input into the large language model. If the sensitive word detection result of the user input information is not passed, a preset rejection content will be directly presented to the user. The user input information input into the large language model will start the question and answer processing of the large language model. The output information of the large language model will also be checked for sensitive words. If the detection passes, the output information of the large language model will be presented to the user. If the output information detection result is not passed, the preset rejection content will be presented to the user. When performing sensitive word detection on the user input information, it is necessary to perform similarity comparison based on both the "word library" and "multi-word library" in the preset sensitive word library. For the output information of the large language model, because real-time detection needs to be performed during the output of the large language model, the latency of each detection needs to be as short as possible. To improve the detection rate, only similarity comparison with the "word library" in the sensitive word library is required. At the same time, a user feedback mechanism can also be set up to update the rejection question library and sensitive word library in real time based on the user's feedback results.

[0108] Embodiment 3

[0109] Corresponding to the embodiment of the method for detecting sensitive information in the question and answer scenario in the foregoing Embodiment 1, the present disclosure also provides an embodiment of a system for detecting sensitive information in the question and answer scenario, as Figure 4 shown, the detection system includes:

[0110] An acquisition module 401, configured to acquire the output information of the large language model for the current time and store the output information in a preset buffer; wherein, the output form of the output information is stream output;

[0111] A determination module 402, configured to use the output information and the original information stored in the preset buffer as the current text to be processed;

[0112] A response module 403, configured to use the current text to be processed as the target detection text when the data length of the current text to be processed exceeds a preset length; wherein, the preset length is determined according to a pre-constructed sensitive word library;

[0113] A detection module 404, configured to perform sensitive word detection on the target detection text to obtain a detection result.

[0114] In an optional implementation manner, the detection system further includes:

[0115] A first response unit, configured to use the current text to be processed as the target detection text when the data length of the current text to be processed is less than the preset length and the current question-and-answer of the large language model has ended;

[0116] And / or;

[0117] A second response unit, configured to release the first quantity of first information starting from the initial character in the target detection text in the preset buffer when the detection result meets a preset condition; wherein, the first quantity is the data length corresponding to the previous output information of the large language model;

[0118] And / or;

[0119] A third response unit, configured to end the current question-and-answer of the large language model when the detection result does not meet a preset condition.

[0120] In an optional implementation manner, the detection module is further configured to determine that there is a sensitive word in the target detection text when the similarity between any first word to be detected in the target detection text and a preset sensitive word in the pre-constructed sensitive word library is greater than a first preset value; otherwise, determine that there is no sensitive word in the target detection text;

[0121] And / or;

[0122] The detection system further includes a first detection unit, and the first detection unit is configured to obtain external input information and perform sensitive word detection on the external input information; when the similarity between any second word to be detected in the external input information and a preset sensitive word in the pre-constructed sensitive word library is greater than a second preset value, determine that there is a sensitive word in the external input information, and return a preset rejection content; otherwise, input the external input information into the large language model for question-and-answer processing.

[0123] For the system embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution.

[0124] Embodiment 4

[0125] Corresponding to the method embodiment of content display in the Q&A scenario in the foregoing Embodiment 2, the present disclosure also provides a content display system in the Q&A scenario. The content display system in the Q&A scenario is implemented based on the detection system of sensitive information in the Q&A scenario described in Embodiment 3, as Figure 5 shown. The content display system includes:

[0126] A response module 501, configured to input the external input information into a large language model to obtain a detection result of sensitive word detection corresponding to the target detection text when there is no sensitive word in the external input information;

[0127] A display module 502, configured to display the target content based on the detection result; wherein, the detection result is obtained based on the detection system of sensitive information in the Q&A scenario described in the third aspect.

[0128] In an optional implementation manner, the display module is specifically configured to display a preset rejection content when the detection result does not meet a preset condition; and display the target detection text when the detection result meets the preset condition.

[0129] For the system embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial descriptions of the method embodiments. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution.

[0130] Embodiment 3

[0131] Figure 6The figure shows a schematic structural diagram of an electronic device according to an exemplary embodiment of the present disclosure. The electronic device includes a memory, a processor, and a computer program stored in the memory and configured to run on the processor. When the processor executes the computer program, it implements the method for detecting sensitive information in the Q&A scenario of Embodiment 1, or the content display method in the Q&A scenario of Embodiment 2. Figure 6 The illustrated electronic device 90 is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0132] As Figure 6 shown, the electronic device 90 may be presented in the form of a general-purpose computing device, for example, it may be a server device. The components of the electronic device 90 may include, but are not limited to: at least one of the above-mentioned processors 91, at least one of the above-mentioned memories 92, and a bus 93 connecting different system components (including the memory 92 and the processor 91).

[0133] The bus 93 includes a data bus, an address bus, and a control bus.

[0134] The memory 92 may include volatile memory, such as a random access memory (RAM) 921 and / or a cache memory 922, and may further include a read-only memory (ROM) 923.

[0135] The memory 92 may further include a program tool 925 (or utility) having a set of (at least one) program modules 924. Such program modules 924 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples.

[0136] The processor 91 executes various functional applications and data processing by running the computer program stored in the memory 92, such as the method for detecting sensitive information in the Q&A scenario of Embodiment 1 of the present disclosure, or the content display method in the Q&A scenario of Embodiment 2.

[0137] The electronic device 90 can also communicate with one or more external devices 94 (such as a keyboard, a pointing device, etc.). Such communication can be carried out through the input / output (I / O) interface 95. Moreover, the electronic device 90 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 96. As shown in the figure, the network adapter 96 communicates with other modules of the electronic device 90 through the bus 93. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 90, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (redundant array of independent disks) systems, magnetic tape drives, and data backup storage systems, etc.

[0138] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-described units / modules can be embodied in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided and embodied by multiple units / modules.

[0139] Embodiment 4

[0140] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the method for detecting sensitive information in the Q&A scenario of Embodiment 1 above, or the method for presenting content in the Q&A scenario of Embodiment 2.

[0141] Among them, the more specific computer-readable storage medium that can be adopted can include but not limited to: portable disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories, optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0142] Embodiment 5

[0143] The embodiments of the present disclosure also provide a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method for detecting sensitive information in the Q&A scenario of Embodiment 1 above, or the method for presenting content in the Q&A scenario of Embodiment 2.

[0144] Among them, the program code for executing the computer program product of the present disclosure can be written in any combination of one or more programming languages, and the program code can be executed entirely on the user device, partially on the user device, executed as an independent software package, partially on the user device and partially on a remote device, or entirely on a remote device.

[0145] Although the specific embodiments of the present disclosure have been described above, those skilled in the art should understand that this is only an example, and the protection scope of the present disclosure is defined by the appended claims. Without departing from the principles and essence of the present disclosure, those skilled in the art can make various changes or modifications to these embodiments, but these changes and modifications all fall within the protection scope of the present disclosure.

Claims

1. A method for detecting sensitive information in a question-and-answer scenario, characterized in that: The detection method comprises: Obtaining the current output information of the large language model, and storing the output information in a preset buffer; wherein the output information is in a streaming output form; The output information and the original information stored in the preset buffer are used as the current text to be processed; In response to the data length of the current text to be processed exceeding a preset length, the current text to be processed is used as a target detection text; wherein the preset length is determined according to a pre-constructed sensitive word library; Sensitive word detection is performed on the target detection text to obtain a detection result.

2. The detection method according to claim 1, characterized in that The detection method further comprises: In response to the data length of the current text to be processed being less than the preset length and the current question and answer of the large language model having ended, using the current text to be processed as the target detection text; and / or; In response to the detection result satisfying a preset condition, releasing a first amount of first information starting from an initial character in the target detection text in the preset buffer; wherein the first amount is the data length corresponding to the previous output information of the large language model; and / or; In response to the detection result not satisfying a preset condition, the current question and answer of the large language model is terminated.

3. The detection method according to claim 1 or 2, characterized in that: The step of performing sensitive word detection on the target detection text comprises: In response to the existence of any first to-be-detected word in the target detection text and the similarity between the preset sensitive word in the pre-built sensitive word library being greater than a first preset value, it is determined that the target detection text contains a sensitive word; otherwise, it is determined that the target detection text does not contain a sensitive word; and / or; Before the step of obtaining the current output information of the large language model, the following step is also included: Acquire external input information, and perform sensitive word detection on the external input information; In response to the existence of any second word to be detected in the external input information and the similarity between the second word to be detected and the preset sensitive word in the pre-built sensitive word library being greater than a second preset value, it is determined that there are sensitive words in the external input information, and the preset rejection content is returned; otherwise, the external input information is input into the large language model for question and answer processing.

4. A content display method in a question-and-answer scenario, characterized in that: The content display method comprises: In response to the absence of sensitive words in the external input information, the external input information is input into the large language model to obtain a detection result of sensitive word detection corresponding to the target detection text; Displaying target content based on the detection result; The detection result is obtained based on the method for detecting sensitive information in a question-and-answer scenario according to any one of claims 1 to 3.

5. The content display method according to claim 4, characterized in that: The step of displaying the target content based on the detection result includes: In response to the detection result not satisfying the preset condition, displaying preset refusal content; In response to the detection result satisfying the preset condition, the target detection text is displayed.

6. A system for detecting sensitive information in a question-and-answer scenario, characterized in that: The detection system comprises: An acquisition module, used for acquiring the current output information of the large language model, and storing the output information in a preset buffer; wherein the output information is in a streaming output form; A determination module, used for taking the output information and the original information stored in the preset buffer as the current text to be processed; A response module, configured to use the current text to be processed as a target detection text when the data length of the current text to be processed exceeds a preset length; wherein the preset length is determined according to a pre-built sensitive word library; The detection module is used to perform sensitive word detection on the target detection text to obtain a detection result.

7. A content display system in a question-and-answer scenario, characterized in that: The content display system comprises: A response module, for inputting the external input information into the large language model to obtain a detection result of sensitive word detection corresponding to the target detection text when there is no sensitive word in the external input information; A display module is used to display the target content based on the detection result; wherein the detection result is obtained based on the detection system for sensitive information in the question-and-answer scenario described in claim 6.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and used to run on the processor, characterized in that: When the processor executes the computer program, it implements the method for detecting sensitive information in the question-and-answer scenario described in any one of claims 1 to 3, or the method for displaying content in the question-and-answer scenario described in claim 4 or 5.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the method for detecting sensitive information in a question-and-answer scenario described in any one of claims 1 to 3, or the method for displaying content in a question-and-answer scenario described in claim 4 or 5.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the method for detecting sensitive information in a question-and-answer scenario as described in any one of claims 1 to 3, or the method for displaying content in a question-and-answer scenario as described in claim 4 or 5.

Citation Information

Cited By

  • Large model interaction information security filtering method and device

    CN121009895A

  • Large-scale interactive information security filtering method and device

    CN121009895B

  • Model-based information association method and device, electronic equipment and storage medium

    CN121351801A

  • Model-based information correlation method and device, electronic equipment and storage medium

    CN121351801B