Compliance risk avoidance methods, systems, devices, and media for generative artificial intelligence

By calculating the compliance potential value of the input text and constructing a three-level text database, the content compliance problem of generative artificial intelligence in compliant expression is solved, the security and reliability of the generated content are realized, and its application in scenarios with high compliance requirements is expanded.

CN122132522APending Publication Date: 2026-06-02NORTH CHINA ELECTRIC POWER UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NORTH CHINA ELECTRIC POWER UNIV
Filing Date
2026-01-12
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Generative artificial intelligence has content compliance deviations and security risks in the scope of compliant expression, making it difficult to output authoritative and reliable answers.

Method used

By calculating the compliance potential value of the input text, using a normalization algorithm to classify the levels, and constructing a three-level text database, the question-and-answer results are generated by calling different levels of the database, thus ensuring content compliance.

Benefits of technology

It enables precise quantification and early identification of compliance risks, ensuring the security and reliability of generated content and expanding the application boundaries of generative AI in scenarios with high compliance requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122132522A_ABST
    Figure CN122132522A_ABST
Patent Text Reader

Abstract

This invention provides a method, system, device, and medium for compliance risk avoidance in generative artificial intelligence. The method includes: acquiring user input text; preprocessing the input text and calculating its compliance potential value, wherein the compliance potential value is positively correlated with the compliance relevance, discourse empowerment, and group influence of each character in the input text; based on the calculated compliance potential value, using a normalization algorithm to calculate an empirical threshold for the input text, and dividing the input text into three levels according to the empirical threshold; constructing a three-level text database; generating a final question-and-answer result by calling the three-level text database according to the level of the empirical threshold corresponding to the input text; and displaying the final question-and-answer result to the user. This addresses how to ensure that the generated content does not have content compliance issues while providing users with more authoritative, comprehensive, and reliable answers when applying generative artificial intelligence in areas involving compliance expression related to content compliance.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of generative artificial intelligence technology, in particular to a compliance risk avoidance method, system, device and medium for generative artificial intelligence. BACKGROUND

[0002] The widespread application of generative artificial intelligence has become an inevitable trend, but so far different generative artificial intelligence platforms have not answered the content and problems related to compliance-related knowledge system and other compliance expression categories well, there are simplified answers, ambiguous answers or even refusal to answer, the core problem is that the network space including large models is a security risk area, and the compliance risk is the main risk type, once some compliance expression fields of words and sentences are involved, there is a possibility of content compliance deviation.

[0003] The "content generation illusion" of generative artificial intelligence is the main source of compliance risk, and the underlying corpus of generative artificial intelligence is diverse, many contents have not been strictly audited, and there are also compliance and security risks, and the content generation based on anticipation also has potential compliance and security risks.

[0004] Therefore, it is urgent to propose a compliance risk avoidance method for generative artificial intelligence to solve the technical problem of how to ensure that the generated content does not have content compliance problems and output more authoritative and reliable answers to users when applying generative artificial intelligence in the field of compliance expression category related to content compliance. SUMMARY

[0005] In order to overcome the problems in the related art, the present disclosure provides a compliance risk avoidance method, system, device and medium for generative artificial intelligence to solve the technical problem of how to ensure that the generated content does not have content compliance problems and output more authoritative and reliable answers to users when applying generative artificial intelligence in the field of compliance expression category related to content compliance in the related art.

[0006] One or more embodiments of the present specification provide a compliance risk avoidance method for generative artificial intelligence, comprising the following steps: Obtain the input text of the user, calculate the compliance potential energy value of the input text after preprocessing, the compliance potential energy value is positively correlated with the compliance correlation degree, discourse empowerment degree and group influence degree of each character in the input text; Based on the calculated compliance potential energy value, use the normalization algorithm to calculate the experience threshold value of the input text, and divide the input text into three levels according to the experience threshold value; Construct a three-level text database, and call the three-level text database to generate a final question and answer result according to the level of the experience threshold value corresponding to the input text; The final Q&A results will be displayed to the user.

[0007] Preferably, calculating the compliance potential value of the input text specifically includes the following steps: Documents from authoritative media and publishers in the nationally designated compliance field that have undergone compliance review are collected to establish a compliance and security text database. High-frequency words are mined from the compliance input text in the compliance and security text database to generate a five-level knowledge graph, with each level of words corresponding to different compliance relevance. The source of words is automatically matched through the corpus. A complete match search is performed in the files at each level of the corresponding corpus. If a word is found, the level to which the file belongs is the value of the discourse empowerment. LDA was used to perform topic word cluster analysis on the collected daily hot news. The topic weight of each word was used as the base value of the group influence of each word. The base value of each word in the topic word cluster was calculated and normalized as the group influence of each word. Calculate the compliance potential value of each word in the input text, and sum and average the compliance potential values ​​of all words to obtain the compliance potential value of the input text.

[0008] Preferably, the construction of the three-level text database specifically includes the following steps: Based on the sources of the corpus, experts reviewed publicly available documents from authoritative media in the nationally designated compliance field, normative textbooks in specific fields, and literature data from experts in authoritative publishers in the nationally designated compliance field, and established the authoritative corpus as a first-level text database. Authoritative and publicly available corpora, as determined by experts, serve as the secondary text database; The general corpus of AI large models is used as a three-level text database.

[0009] Preferably, the step of generating the final question-and-answer result by calling the three-level text database based on the experience threshold level corresponding to the input text specifically includes the following steps: The first question and answer result is intelligently generated by calling the corresponding level of the text database based on the experience threshold level corresponding to the input text. Based on the first question and answer result, a first question is generated, and the second question and answer result is intelligently generated by calling the next level text database. Based on the second question and answer result, a second question is generated, and the final question and answer result is generated by calling the next level of text database.

[0010] This specification provides one or more embodiments of a compliance risk avoidance system for generative artificial intelligence, including a calculation module, a level classification module, a generation module, and a display module; The calculation module is used to acquire the user's input text, preprocess it, and calculate the compliance potential value of the input text. The compliance potential value is positively correlated with the compliance relevance, discourse empowerment, and group influence of each character in the input text. The level division module is used to calculate the empirical threshold of the input text based on the calculated compliance potential value using a normalization algorithm, and divide the input text into three levels according to the empirical threshold; The generation module is used to construct a three-level text database and generate the final question-and-answer result by calling the three-level text database according to the level of the experience threshold corresponding to the input text. The display module is used to show the final question and answer results to the user.

[0011] Preferably, the calculation module further includes a compliance relevance unit, a discourse empowerment unit, a group influence unit, and a compliance potential value calculation unit; The compliance relevance unit is used to collect literature data from authoritative media and authoritative publishers in the nationally designated compliance field that have undergone compliance review, establish a compliance and security text database, mine high-frequency words from the compliance input text in the compliance and security text database, and generate a five-level knowledge graph, with each level of words corresponding to different compliance relevance. The discourse empowerment unit is used to automatically match word sources through the corpus and perform a complete match search in files at each level of the corresponding corpus. If a word is found, the level to which the file belongs is the value of the discourse empowerment. The group influence unit is used to perform topic word group analysis on the collected daily hot news using LDA. The topic weight of each word is used as the base value of the group influence of each word. The base value of each word in the topic word group is calculated and normalized to serve as the group influence of each word. The compliance potential value calculation unit is used to calculate the compliance potential value of each word in the input text, and to sum and average the compliance potential values ​​of all words to obtain the compliance potential value of the input text.

[0012] Preferably, the generation module further includes a text database hierarchy establishment unit, specifically configured as follows: Based on the sources of the corpus, experts reviewed publicly available documents from authoritative media in the nationally designated compliance field, normative textbooks in specific fields, and literature data from experts in authoritative publishers in the nationally designated compliance field, and established the authoritative corpus as a first-level text database. Authoritative and publicly available corpora, as determined by experts, serve as the secondary text database; The general corpus of AI large models is used as a three-level text database.

[0013] Preferably, the generation module further includes a first generation unit, a second generation unit, and a third generation unit; The first generation unit is used to intelligently generate a first question-and-answer result by calling a text database of the corresponding level based on the level of the experience threshold corresponding to the input text. The second generation unit is used to generate a first question based on the first question-and-answer result, and to intelligently generate a second question-and-answer result by calling the next-level text database; The third generation unit is used to generate a second question based on the second question-and-answer result, and then call the next level text database to intelligently generate the final question-and-answer result.

[0014] This specification provides one or more embodiments of a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the compliance risk avoidance method for generative artificial intelligence as described above.

[0015] This specification provides one or more embodiments of a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the compliance risk avoidance method for generative artificial intelligence described above.

[0016] This disclosure provides a generative artificial intelligence-based method, system, device, and medium for compliance risk avoidance. Its advantages lie in its ability to preprocess user-input text and calculate its compliance potential value by combining the text's compliance relevance, discourse empowerment, and group influence. This achieves precise quantification and early identification of compliance risks, providing a measurable standard for previously difficult-to-define compliance risks. It can identify core risk-related information in the early stages of text processing, providing precise guidance for subsequent risk management. Furthermore, it utilizes a normalization algorithm to convert the compliance potential value into an empirical threshold and divides it into three levels, achieving differentiated classification and standards for compliance risks. The evaluation process utilizes a normalization algorithm to eliminate the dimensional differences in compliance potential values ​​across different texts, making empirical thresholds comparable and ensuring the scientific rigor of the classification. A three-tiered text database is constructed and retrieved and question-and-answer results are generated according to the risk level. This achieves risk adaptation and security control of the generated content. The three-tiered text database corresponds to compliance requirements at different risk levels. The database used for high-risk texts undergoes rigorous review to effectively avoid sensitive content, while the databases for medium and low risk levels retain more information diversity while ensuring security. The final question-and-answer results are then presented to the user, achieving efficient delivery of secure content and optimized user experience. The displayed results, having undergone risk screening and compliance processing in the preceding steps, prevent users from encountering content with compliance risks, ensuring safe use. Through a technical approach of micro-quantification, dynamic grading, and layered protection, a compliance risk protection system with precision, adaptability, and practicality has been established for generative AI, effectively expanding the application boundaries of generative AI in scenarios with high compliance requirements. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in one or more embodiments of this specification or in the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 A flowchart illustrating a compliance risk avoidance method for generative artificial intelligence provided in one or more embodiments of this specification; Figure 2 A flowchart illustrating the specific implementation of the compliance risk avoidance method for generative artificial intelligence provided in one or more embodiments of this specification; Figure 3 A schematic diagram of the structure of a compliance risk avoidance system for generative artificial intelligence provided in one or more embodiments of this specification; Figure 4 This is a schematic diagram of the structure of a computer device provided for one or more embodiments of this specification. Detailed Implementation

[0019] To enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this invention.

[0020] The present invention will now be described in detail with reference to specific embodiments and accompanying drawings.

[0021] Method Implementation Examples According to embodiments of the present invention, a method for avoiding compliance risks in generative artificial intelligence is provided, such as... Figure 1 The diagram shown is a flowchart illustrating the compliance risk avoidance method for generative artificial intelligence provided in this embodiment. The compliance risk avoidance method for generative artificial intelligence according to this embodiment includes the following steps: S110. Obtain the user's input text, perform preprocessing using natural language processing, then perform semantic analysis and feature processing, and calculate the compliance potential value (PPV) of the input text. The compliance potential value (PPV) is positively correlated with the discourse empowerment, compliance relevance, and group influence of each character in the input text.

[0022] S120. Based on the calculated compliance potential value PPV, the empirical threshold of the input text is calculated using a normalization algorithm, and the input text is divided into three levels according to the empirical threshold.

[0023] Specifically, Level 1: The experience threshold is higher than 70%, the security domain is Level 1, and the security level is high security.

[0024] Level 2: The experience threshold is between 20% and 70%, the safety domain is Level 2, and the safety level is medium.

[0025] Level 3: The experience threshold is below 20%, the security domain is level 3, and the security level is low security.

[0026] S130. Construct a three-level text database, and generate the final question-and-answer result by calling the three-level text database according to the level of the experience threshold corresponding to the input text.

[0027] S140. Generate intelligent answers from the final question-and-answer results and display them to the user in a layered manner.

[0028] The method provided in this embodiment calculates the compliance potential value (PPV) of the input text by preprocessing the user's input text and combining the compliance relevance, discourse empowerment, and group influence of the text. This achieves accurate quantification and early identification of compliance risks, providing a measurable standard for previously difficult-to-define compliance risks. It can lock in core risk-related information in the early stages of text processing, providing precise guidance for subsequent risk management. The method uses a normalization algorithm to convert the compliance potential value PPV into an empirical threshold and divides it into three levels, achieving differentiated classification and standardized assessment of compliance risks. The normalization algorithm eliminates the dimensional differences in the compliance potential value PPV of different texts, making the empirical thresholds comparable and ensuring the scientific nature of the level division. A three-level text database is constructed and retrieved and question-and-answer results are generated according to the level, achieving risk adaptation and security control of the generated content. The three-level text database corresponds to the compliance requirements of different risk levels. The database called for high-risk level texts undergoes strict review to effectively avoid sensitive content, while the databases for medium and low risk levels retain more information diversity while ensuring security. The retrieved and generated results are displayed to the user, achieving efficient delivery of secure content and optimized user experience. The displayed results have undergone risk screening and compliance processing in previous steps, preventing users from encountering content with compliance risks and ensuring safe use. Through a technical approach of micro-quantification, dynamic grading, and layered protection, a compliance risk protection system that is accurate, adaptable, and practical has been built for generative AI, effectively expanding the application boundaries of generative AI in scenarios with high compliance requirements.

[0029] In one embodiment, calculating the compliance potential value (PPV) of the input text specifically includes the following steps: Literature data from authoritative media outlets and publishers in the nationally designated compliance field, after compliance review, is collected to establish a compliance and security text database. High-frequency word mining is performed on the compliance-compliant input text in this database to generate a five-level knowledge graph. Each level of words corresponds to a different compliance relevance (1, 0.8, 0.6, 0.4, 0.2). Words in the input text are retrieved from the five-level knowledge graph lexicon and assigned different values ​​according to their level; words not found in the knowledge graph lexicon are assigned a value of 0.

[0030] Discourse empowerment level (levels 1-5), with values ​​of 1, 2, 3, 4, and 5 corresponding to the level. Assignment rules: Word sources are automatically matched to the corpus, and the level is mapped according to the discourse empowerment level table corresponding to the corpus source. A complete match search is performed on files at each level of the corresponding corpus in descending order of quality. If a word is found, the level of the file is assigned the discourse empowerment level value.

[0031] We collect trending news from daily trending topics on major platforms such as Baidu, Weibo, Douyin, and Xiaohongshu over the past month. We use LDA to perform topic cluster analysis on the collected daily trending news. We use the topic weight (Di1) * (word frequency weight in the topic) of each word as the base value of the group influence of each word. We calculate the base value of each word in the topic cluster, and normalize it from high to low using 1-0 to obtain the group influence of each word.

[0032] Calculate the compliance potential value (PPV) of each word in the input text, and sum and average the compliance potential values ​​(PPV) of all words to obtain the compliance potential value (PPV) of the input text.

[0033] The compliance potential value (PPV) is positively correlated with the discourse empowerment, compliance relevance, and group influence of each character in the input text. The formula for calculating the compliance potential value (PPV) is as follows: ; C i Representing compliance relevance, with values ​​(1, 0.8, 0.6, 0.4, 0.2, 0), a content compliance knowledge base / text database is established by collecting literature data from authoritative media and publishers in the nationally designated compliance field that have undergone compliance review. For the input text in the content compliance knowledge base / text database, a publicly available knowledge graph generation tool is used to generate a five-level knowledge graph. Keywords at each level correspond to different compliance relevance values ​​(1, 0.8, 0.6, 0.4, 0.2). Words in the input text are retrieved from the five-level knowledge graph terminology, and different values ​​are assigned according to their level. Words not retrieved from the knowledge graph terminology are assigned a value of 0. Dynamic correction: Weights are adjusted based on contextual sensitivity analysis; negative sentences are adjusted by *0.02, and interrogative sentences are adjusted by *0.05. D i This represents the group influence, with a value of (0-1, two decimal places). The main tool used is the topic clustering (LDA) model. The specific idea is to collect trending news from the daily trending topics rankings of mainstream platforms such as Baidu, Weibo, Douyin, and Xiaohongshu over the past month, perform topic cluster analysis using LDA, calculate the base value of the group influence of each word based on the topic weight (Di1) * (word frequency weight in the topic), and after calculating the base values ​​of all words in the topic thesaurus, normalize them from high to low using a 1-0 approach.

[0034] Among them, the topic clustering (LDA) model is a latent Dirichlet distribution model, which extracts latent topics from the document set through unsupervised learning. Each topic is represented by a probability distribution of a set of related words.

[0035] Document-topic distribution (θ): Represents the weight of each topic in a single document, satisfying: (d=1,2,...,D); in, It is the topic weight of topic k in document d.

[0036] Topic-word distribution (φ): Represents the probability distribution of words in a single topic, satisfying: (k=1,2,...,K); in, It represents the word frequency weight of word w in topic k.

[0037] Baseline Group Influence (BIV): A comprehensive influence index of words in the global corpus, determined by topic weight and word frequency weight.

[0038] Normalized Group Impact (Norm_BIV): A standardized value that linearly maps BIV to the [0,1] interval for cross-platform comparison.

[0039] Formula for calculating the base value of group influence (Di): BIV(w) = ; Where D: Total number of documents (number of trending news items), K: Total number of topics (LDA preset parameters). The weight of topic k in document d. The weight of word w in topic k.

[0040] Calculation example: If a word appears in 3 documents: Doc1: =0.6, =0.3→0.6×0.3=0.18 Doc2: =0.4, =0.5→0.4×0.5=0.20 Doc3: =0.8, =0.1→0.8×0.1=0.08BIV=(0.18+0.20+0.08) / 3=0.153; Normalized group influence: Normalize the base value of group influence by taking the minimum value: $\min(\text{BIV})$ represents the minimum BIV value of all words.

[0041] Normalize the base value of the group influence by taking the maximum value: $\max(\text{BIV})$ represents the maximum BIV value of all words.

[0042] Normalization effect: Assuming the highest BIV = 0.82 and the lowest BIV = 0.02, and a certain word's BIV = 0.45 → Norm_BIV = (0.45 - 0.02) / (0.82 - 0.02) = 0.5375 Di = 0.54, rounded up to two decimal places.

[0043] S i The discourse empowerment level is represented by a scale of 1-5, with values ​​of 1, 2, 3, 4, and 5 corresponding to the level. The assignment rule is as follows: Lexical sources are automatically matched using a corpus, and the levels are mapped according to Table 1.

[0044] By performing exact matches on files within each corresponding level, if the word is found in the corresponding file, then the level to which the file belongs is the discourse empowerment level.

[0045] Table 1. Examples of Standards for Grading the Level of Discourse Empowerment

[0046] The method provided in this embodiment establishes a compliance and security text database by collecting literature data from authoritative media and publishers in the nationally designated compliance field that have undergone compliance review. This provides an authoritative and reliable data source for subsequent calculations, ensuring the compliance and accuracy of the compliance potential value (PPV) calculation from the source and avoiding risk misjudgment due to inappropriate reference data. High-frequency word mining is performed on the input text for content compliance, and a high-frequency word library for content compliance is generated by combining it with a content compliance knowledge graph. This accurately extracts core words that are representative of content compliance, providing a clear basis for calculating the compliance potential value (PPV) of individual words and enhancing the pertinence and professionalism of the quantitative analysis. By calculating the compliance potential value (PPV) of each word in the text and summing and averaging them, the overall compliance potential value (PPV) of the text is obtained. This quantitative approach, from micro-vocabulary to macro-text, fully considers the detailed composition of the text and comprehensively reflects the overall content compliance attributes of the text. This makes the compliance potential value (PPV) more convincing and provides a solid quantitative foundation for subsequent text level classification and compliance risk avoidance, effectively improving the accuracy and reliability of generative artificial intelligence in compliance risk control.

[0047] In one embodiment, constructing the three-level text database specifically includes the following steps: Collect more reliable corpora.

[0048] Based on the sources of the corpus, experts reviewed authoritative documents published by authoritative media in nationally designated compliance fields, normative textbooks in specific fields, and literature data from experts in nationally designated authoritative publishers in compliance fields, and established a first-level text database by collecting authoritative corpus data.

[0049] The secondary text database consists of publicly available corpora (publicly available journals and monographs related to the field of compliant expression) that have been judged and recognized by experts.

[0050] The general corpus of AI large models is used as a three-level text database.

[0051] The method provided in this embodiment constructs a first-level text database based on expert-reviewed literature data from authoritative media and publishers in nationally designated compliant fields. This ensures the highest authority and content compliance of core resources, providing the most reliable content support for the retrieval of high-risk texts and strengthening the risk defense line from the top. The second-level text database incorporates authoritative and publicly available corpora judged by experts, expanding the content scope while ensuring compliance. This maintains the rigor of authoritative resources and provides richer information sources for medium-risk texts, achieving a balance between security and practicality. The third-level text database uses a general corpus of AI large-scale models, fully utilizing the breadth and diversity of general corpora to meet the information needs of low-risk texts. At the same time, hierarchical isolation avoids interference from potential risks in general corpora on the generation of high-security content. This three-level architecture design enables texts of different risk levels to be accurately matched with resources of corresponding security levels, ensuring the controllability of compliance risks while maximizing the preservation of the information value of the content, providing a solid resource foundation for the safe and efficient application of generative artificial intelligence.

[0052] In one embodiment, the step of calling the three-level text database to generate the final question-and-answer result based on the level of the experience threshold corresponding to the input text specifically includes the following steps: Based on the pre-set experience threshold for the three-level classification of compliance potential value (PPV), the user input text is classified into three risk levels: "Level 1", "Level 2", and "Level 3".

[0053] The first question and answer result is generated intelligently by calling the corresponding text database based on the experience threshold level of the input text.

[0054] The first question is generated based on the first question and answer result, and the second question and answer result is generated intelligently by calling the next level text database.

[0055] Based on the second question and answer result, a second question is generated, and the final question and answer result is generated by calling the next level of text database.

[0056] The following specific implementation case further illustrates this point: For questions with a text security level of Level 1, it is necessary to first call the Level 1 (authoritative) corpus. Based on the Level 1 corpus, the output text result is used as the authoritative standard answer. Then, based on the output authoritative and labeled answer, a new question is generated, which calls the Level 2 (reliable) corpus. Based on the Level 2 (reliable) corpus, the output text result is used as a professional reference answer. Finally, the Level 3 (network) large model corpus is called. The result generated by the large model's own generation program after thinking and searching is used as the intelligently generated answer.

[0057] For questions with a text security level of two, it is necessary to first call the level two (reliable) corpus. Based on the level two (reliable) corpus, the output text result is used as a professional reference answer. Then, a new question is generated based on the output professional reference answer, which calls the level three (network) large model corpus. The large model's own generation program thinks and retrieves the result, which is used as the intelligently generated answer.

[0058] For questions with a text security level of three, the level three (network) large model corpus can be directly accessed. The large model's own generation program will then process and retrieve the results, which will be used as the intelligently generated answer.

[0059] like Figure 2 The diagram shown is a flowchart illustrating the specific implementation of the compliance risk avoidance method for generative artificial intelligence provided in this embodiment.

[0060] The method provided in this embodiment retrieves the first question-and-answer result by calling the corresponding database based on the empirical threshold corresponding to the input text. This ensures that the initial content output strictly matches the risk level, safeguarding the bottom line of security from the source and avoiding compliance issues caused by high-risk text directly associating with low-security-level resources. Based on the first question-and-answer result, a question is generated and the next-level database is called for retrieval. Within the security framework, targeted questions expand the information dimension, preserving the hierarchical constraints of risk prevention and control while compensating for the limitations of single-level resources through progressive retrieval, thus improving the richness and relevance of the content. The second question-and-answer result generates a question and calls the next-level database to obtain the final question-and-answer result, forming a retrieval chain of "security fallback - in-depth supplementation - comprehensive integration." This allows high-risk text to obtain necessary information support on a strict compliance basis, while medium- and low-risk text can integrate a wider range of resources through multi-level retrieval. This progressive retrieval design not only ensures accurate matching between different levels of text and corresponding databases but also breaks down the information barriers of hierarchical isolation through cross-level supplementation guided by questions. Under the premise of controllable risk, it maximizes the release of data value and provides an efficient retrieval path for generative AI to output content that is both safe and practical.

[0061] System Implementation Examples According to embodiments of the present invention, a compliance risk avoidance system for generative artificial intelligence is provided, such as... Figure 3The diagram shown is a structural schematic of the compliance risk avoidance system for generative artificial intelligence provided in this embodiment. The compliance risk avoidance system for generative artificial intelligence according to this embodiment includes a calculation module 31, a level classification module 32, a generation module 33, and a display module 34.

[0062] The calculation module 31 is used to acquire the user's input text, preprocess it, and calculate the compliance potential value (PPV) of the input text. The compliance potential value (PPV) is positively correlated with the compliance relevance, discourse empowerment, and group influence of each character in the input text.

[0063] The grading module 32 is used to calculate the empirical threshold of the input text based on the calculated compliance potential value PPV using a normalization algorithm, and to divide the input text into three grades according to the empirical threshold.

[0064] The generation module 33 is used to construct a three-level text database and generate the final question-and-answer result by calling the three-level text database according to the level of the experience threshold corresponding to the input text.

[0065] The answer display module 34 is used to display the final question and answer results to the user.

[0066] The system provided in this embodiment preprocesses the user-input text through the calculation module 31, and calculates the compliance potential value (PPV) of the input text by combining the compliance relevance, discourse empowerment, and group influence of the text. This achieves accurate quantification and early identification of compliance risks, making previously difficult-to-define compliance risks measurable. It can lock in the core information related to risks in the early stages of text processing, providing precise guidance for subsequent risk management. The level classification module 32 uses a normalization algorithm to convert the compliance potential value (PPV) into an empirical threshold and divides it into three levels, achieving differentiated classification and standardized assessment of compliance risks. The normalization algorithm eliminates differences in... The difference in the dimensions of the text compliance potential value (PPV) makes the empirical thresholds comparable, ensuring the scientific nature of the classification. The generation module 33 constructs a three-level text database and retrieves and generates question-and-answer results according to the level, achieving risk adaptation and security control of the generated content. The three-level text databases correspond to the compliance requirements of different risk levels. The database used for high-risk texts undergoes strict review to effectively avoid sensitive content, while the databases for medium and low risk levels retain more information diversity while ensuring security. The answer display module 34 presents the final question-and-answer results to the user, achieving efficient delivery of secure content and optimized user experience. The displayed results undergo risk screening and compliance processing in the preceding steps, preventing users from encountering content with compliance risks and ensuring safe use. Through a technical approach of micro-quantification, dynamic grading, and layered protection, a compliance risk protection system with accuracy, adaptability, and practicality has been built for generative AI, effectively expanding the application boundaries of generative AI in scenarios with high compliance requirements.

[0067] In one embodiment, the calculation module 31 further includes a compliance relevance unit, a discourse empowerment unit, a group influence unit, and a compliance potential value PPV calculation unit.

[0068] The compliance relevance unit is used to collect literature data from authoritative media and publishers in the nationally designated compliance field that have undergone compliance review, establish a compliance and security text database, mine high-frequency words from the compliance input text in the compliance and security text database, and generate a five-level knowledge graph, with each level of words corresponding to different compliance relevance.

[0069] The discourse empowerment unit is used to automatically match word sources through the corpus and perform a complete match search in files at each level of the corresponding corpus. If a word is found, the level to which the file belongs is the value of the discourse empowerment.

[0070] The group influence unit is used to perform topic word group analysis on the collected daily hot news using LDA. The topic weight of each word is used as the base value of the group influence of each word. The base value of each word in the topic word group is calculated and normalized to serve as the group influence of each word.

[0071] The compliance potential value (PPV) calculation unit is used to calculate the compliance potential value (PPV) of each word in the input text, and to sum and average the compliance potential values ​​(PPV) of all words to obtain the compliance potential value (PPV) of the input text.

[0072] The system provided in this embodiment establishes a compliance and security text database by collecting literature data from authoritative media and publishers in the nationally designated compliance field that have undergone compliance review. This provides an authoritative and reliable data source for subsequent calculations, ensuring the compliance and accuracy of the compliance potential value (PPV) calculation from the source and avoiding risk misjudgments caused by inappropriate reference data. High-frequency word mining is performed on the input text for content compliance, and a high-frequency word library for content compliance is generated by combining it with a content compliance knowledge graph. This accurately extracts core words that are representative of content compliance, providing a clear basis for calculating the compliance potential value (PPV) of individual words and enhancing the pertinence and professionalism of quantitative analysis. By calculating the compliance potential value (PPV) of each word in the text and summing and averaging them, the overall compliance potential value (PPV) of the text is obtained. This quantitative approach, from micro-vocabulary to macro-text, fully considers the detailed composition of the text and comprehensively reflects the overall content compliance attributes of the text. This makes the compliance potential value (PPV) more convincing and provides a solid quantitative foundation for subsequent text level classification and compliance risk avoidance, effectively improving the accuracy and reliability of generative artificial intelligence in compliance risk control.

[0073] In one embodiment, the generation module 33 further includes a text database hierarchy establishment unit, specifically configured as follows: Based on the sources of the corpus, experts reviewed publicly available documents from authoritative media in the nationally designated compliance field, normative textbooks in specific fields, and literature data from experts in authoritative publishers in the nationally designated compliance field, and established the authoritative corpus as a first-level text database. Authoritative and publicly available corpora, as determined by experts, serve as the secondary text database; The general corpus of AI large models is used as a three-level text database.

[0074] The system provided in this embodiment employs a three-tiered architecture. The first-tier text database is constructed from expert-reviewed literature data from authoritative media and publishers in nationally designated compliant fields, ensuring the highest authority and content compliance of core resources. This provides the most reliable content support for retrieving high-risk texts, thus strengthening the risk defense line from the top down. The second-tier text database incorporates authoritative and publicly available corpora approved by experts, expanding the content scope while ensuring compliance. This maintains the rigor of authoritative resources and provides richer information sources for medium-risk texts, achieving a balance between security and practicality. The third-tier text database utilizes the general corpus of AI large-scale models, fully leveraging its breadth and diversity to meet the information needs of low-risk texts. Simultaneously, hierarchical isolation prevents potential risks in the general corpus from interfering with the generation of high-security content. This three-tiered architecture design allows texts of different risk levels to be accurately matched with resources corresponding to their security levels, ensuring controllable compliance risks while maximizing the preservation of informational value, providing a solid resource foundation for the safe and efficient application of generative artificial intelligence.

[0075] In one embodiment, the generation module 33 further includes a first generation unit, a second generation unit, and a third generation unit.

[0076] The first generation unit is used to intelligently generate the first question-and-answer result by calling the corresponding level of the text database based on the experience threshold level corresponding to the input text.

[0077] The second generation unit is used to generate a first question based on the first question-and-answer result, and to intelligently generate a second question-and-answer result by calling the next-level text database.

[0078] The third generation unit is used to generate a second question based on the second question-and-answer result, and then call the next level text database to intelligently generate the final question-and-answer result.

[0079] The system provided in this embodiment generates the first question-and-answer result by calling the corresponding database according to the highest level of text, ensuring that the initial content output strictly matches the risk level, safeguarding the bottom line of security from the source, and avoiding compliance issues caused by high-risk text directly associating with low-security-level resources. Based on the first question-and-answer result, a question is generated and the next-level database is called for retrieval. Within the security framework, targeted questions are used to expand the information dimension, which not only retains the hierarchical constraints of risk prevention and control, but also makes up for the limitations of single-level resources through progressive retrieval, improving the richness and relevance of the content. The second question-and-answer result generates a question and calls the next-level database to obtain the final question-and-answer result, forming a retrieval chain of "security fallback - in-depth supplementation - comprehensive integration". This allows high-risk text to obtain necessary information support on the basis of strict compliance, while medium and low-risk text can integrate a wider range of resources through multi-level retrieval. This progressive retrieval design not only ensures the accurate matching of different levels of text with corresponding databases, but also breaks down the information barriers of hierarchical isolation through cross-level supplementation guided by questions. Under the premise of controllable risk, it maximizes the release of data value and provides an efficient retrieval path for generative AI to output content that is both safe and practical.

[0080] The embodiments of the present invention are system embodiments corresponding to the above method embodiments. The specific operations of each module processing step can be understood by referring to the description of the method embodiments, and will not be repeated here.

[0081] like Figure 4 As shown, the present invention also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the compliance risk avoidance method for generative artificial intelligence in the above embodiments, or the computer program, when executed by a processor, implements the compliance risk avoidance method for generative artificial intelligence in the above embodiments.

[0082] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0083] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for apparatus or system embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. The apparatus and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and the contents not described in detail in the specification of the present invention are known to those skilled in the art.

Claims

1. A method for mitigating compliance risks in generative artificial intelligence, characterized in that, Includes the following steps: The user's input text is obtained, preprocessed, and then the compliance potential value of the input text is calculated. The compliance potential value is positively correlated with the compliance relevance, discourse empowerment, and group influence of each character in the input text. Based on the calculated compliance potential value, the empirical threshold of the input text is calculated using a normalization algorithm, and the input text is divided into three levels according to the empirical threshold. A three-level text database is constructed, and the final question-and-answer result is generated by calling the three-level text database according to the level of the experience threshold corresponding to the input text. The final Q&A results will be displayed to the user.

2. The compliance risk avoidance method for generative artificial intelligence as described in claim 1, characterized in that, The calculation of the compliance potential value of the input text specifically includes the following steps: Documents from authoritative media and publishers in the nationally designated compliance field that have undergone compliance review are collected to establish a compliance and security text database. High-frequency words are mined from the compliance input text in the compliance and security text database to generate a five-level knowledge graph, with each level of words corresponding to different compliance relevance. The source of words is automatically matched through the corpus. A complete match search is performed in the files at each level of the corresponding corpus. If a word is found, the level to which the file belongs is the value of the discourse empowerment. LDA was used to perform topic word cluster analysis on the collected daily hot news. The topic weight of each word was used as the base value of the group influence of each word. The base value of each word in the topic word cluster was calculated and normalized as the group influence of each word. Calculate the compliance potential value of each word in the input text, and sum and average the compliance potential values ​​of all words to obtain the compliance potential value of the input text.

3. The compliance risk avoidance method for generative artificial intelligence as described in claim 1, characterized in that, The construction of the three-level text database specifically includes the following steps: Based on the sources of the corpus, experts reviewed publicly available documents from authoritative media in the nationally designated compliance field, normative textbooks in specific fields, and literature data from experts in authoritative publishers in the nationally designated compliance field, and established the authoritative corpus as a first-level text database. Authoritative and publicly available corpora, as determined by experts, serve as the secondary text database; The general corpus of AI large models is used as a three-level text database.

4. The compliance risk avoidance method for generative artificial intelligence as described in claim 1, characterized in that, The step of generating the final question-and-answer result by calling the three-level text database based on the level of the experience threshold corresponding to the input text specifically includes the following steps: The first question and answer result is intelligently generated by calling the corresponding level of the text database based on the experience threshold level corresponding to the input text. Based on the first question and answer result, a first question is generated, and the second question and answer result is intelligently generated by calling the next level text database. Based on the second question and answer result, a second question is generated, and the final question and answer result is generated by calling the next level of text database.

5. A compliance risk avoidance system for generative artificial intelligence, characterized in that, It includes a calculation module, a level classification module, a generation module, and a display module; The calculation module is used to acquire the user's input text, preprocess it, and calculate the compliance potential value of the input text. The compliance potential value is positively correlated with the compliance relevance, discourse empowerment, and group influence of each character in the input text. The level division module is used to calculate the empirical threshold of the input text based on the calculated compliance potential value using a normalization algorithm, and divide the input text into three levels according to the empirical threshold; The generation module is used to construct a three-level text database and generate the final question-and-answer result by calling the three-level text database according to the level of the experience threshold corresponding to the input text. The display module is used to show the final question and answer results to the user.

6. The compliance risk avoidance system for generative artificial intelligence as described in claim 5, characterized in that, The calculation module also includes a compliance relevance unit, a discourse empowerment unit, a group influence unit, and a compliance potential value calculation unit; The compliance relevance unit is used to collect literature data from authoritative media and authoritative publishers in the nationally designated compliance field that have undergone compliance review, establish a compliance and security text database, mine high-frequency words from the compliance input text in the compliance and security text database, and generate a five-level knowledge graph, with each level of words corresponding to different compliance relevance. The discourse empowerment unit is used to automatically match word sources through the corpus and perform a complete match search in files at each level of the corresponding corpus. If a word is found, the level to which the file belongs is the value of the discourse empowerment. The group influence unit is used to perform topic word group analysis on the collected daily hot news using LDA. The topic weight of each word is used as the base value of the group influence of each word. The base value of each word in the topic word group is calculated and normalized to serve as the group influence of each word. The compliance potential value calculation unit is used to calculate the compliance potential value of each word in the input text, and to sum and average the compliance potential values ​​of all words to obtain the compliance potential value of the input text.

7. The compliance risk avoidance system for generative artificial intelligence as described in claim 5, characterized in that, The generation module also includes a text database hierarchy establishment unit, specifically configured as follows: Based on the sources of the corpus, experts reviewed publicly available documents from authoritative media in the nationally designated compliance field, normative textbooks in specific fields, and literature data from experts in authoritative publishers in the nationally designated compliance field, and established the authoritative corpus as a first-level text database. Authoritative and publicly available corpora, as determined by experts, serve as the secondary text database; The general corpus of AI large models is used as a three-level text database.

8. The compliance risk avoidance system for generative artificial intelligence as described in claim 5, characterized in that, The generation module further includes a first generation unit, a second generation unit, and a third generation unit; The first generation unit is used to intelligently generate a first question-and-answer result by calling a text database of the corresponding level based on the level of the experience threshold corresponding to the input text. The second generation unit is used to generate a first question based on the first question-and-answer result, and to intelligently generate a second question-and-answer result by calling the next-level text database; The third generation unit is used to generate a second question based on the second question-and-answer result, and then call the next level text database to intelligently generate the final question-and-answer result.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the compliance risk avoidance method for generative artificial intelligence as described in any one of claims 1 to 4.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the compliance risk avoidance method for generative artificial intelligence as described in any one of claims 1 to 4.