Data processing method, system and device, electronic equipment, storage medium and program product

By collaborating between large and small models, unstructured operational rule texts are transformed into question-and-answer pairs, solving the problem of integrating operational rule texts in content security systems. This enables efficient and flexible risk detection and strategy adaptation, reduces labor costs, and improves system adaptability.

CN121996756APending Publication Date: 2026-05-08ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
Filing Date
2025-12-31
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing content security systems struggle to effectively integrate unstructured operational rules and texts, requiring companies to invest significant manpower in manual interpretation and conversion into actionable strategies. Furthermore, existing solutions suffer from poor flexibility, weak adaptability, and high costs, making them inadequate for adapting to dynamic and emerging risks.

Method used

Using a large-model-led and small-model-assisted approach, unstructured operational rule texts are transformed into question-answer pairs. The first preset model performs semantic analysis to generate questions and answers, and the second preset model is used to rewrite and expand the questions to generate a question-answer knowledge base for risk detection and strategy adaptation in content security systems.

Benefits of technology

It has enabled the efficient integration and practical application of unstructured operational rule texts in the content security system, reducing labor costs, improving the system's flexibility and adaptability, and enabling rapid response to rule changes and identification of dynamic risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121996756A_ABST
    Figure CN121996756A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method, system and device, electronic equipment, a storage medium and a program product. According to the scheme provided by the embodiment, a first preset model is utilized to perform semantic analysis on operation rule information provided by a first user, and a first question and an answer corresponding to the first question are generated. The operation rule information comprises internal management specifications formulated by an operation subject to which the first user belongs for operation activities and personnel behaviors. Furthermore, the first question is rewritten by using a second preset model, and a plurality of derivative questions which are semantically equivalent to the first question but are different from the first question in expression are generated. And the first question, the plurality of derivative questions and the corresponding answers form a question-answer pair, and the question-answer pair is stored in a question-answer knowledge base. The question and answer knowledge base can be used for subsequently providing a matched answer as a pickup basis when it is determined that a disposal strategy needing to be executed for a second question is pickup according to a risk detection result of the second question related to an operation subject.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of computer technology, and in particular to a data processing method, system, apparatus, electronic device, storage medium, and program product. Background Technology

[0002] With the rapid evolution of the internet content ecosystem and increasingly stringent regulatory environments, enterprises are placing higher demands on the capabilities of their content security systems. These systems not only need to possess high-precision identification capabilities for common risks but also the ability to flexibly integrate relevant unstructured operational rules (such as industry standards and internal policy manuals) to achieve collaborative content security governance through "general protection" and "personalized rules." However, because these unstructured rule texts lack a unified logical structure, they are difficult to use as a basis for decision-making within content security systems.

[0003] Therefore, there is a need to provide a technical approach that can effectively integrate unstructured operational rule texts into content security systems in order to meet the risk management needs of enterprises for "personalized compliance". Summary of the Invention

[0004] This specification provides several embodiments of a data processing method, system, apparatus, electronic device, storage medium, and program product. Among them, In a first embodiment, this specification provides a data processing method. The method includes: Obtain operational rule information provided by the first user; the operational rule information includes internal management regulations formulated by the operational entity to which the first user belongs for operational activities and personnel behavior; Using a first preset model, semantic analysis is performed on the operational rule information to generate a first question and the corresponding answer to the first question. Using a second preset model, the first problem is rewritten to generate multiple derivative problems; the derivative problems are similar problems that are semantically equivalent to the first problem but have different expressions. The first question, the plurality of derivative questions, and the answer are combined into a question-answer pair and stored in the question-answer knowledge base; The question-and-answer knowledge base is used to subsequently detect risks based on the results of a second question related to the operating entity.

[0005] In a second embodiment, this specification provides a data processing system. The system includes: The client is used to respond to input operations, obtain the operation rule information input by the first user, and send the operation rule information to the server; the operation rule text contains the internal management specifications formulated by the operation entity to which the first user belongs for operation activities and personnel behavior; On the server side, a first preset model and a second preset model are deployed. The first preset model is used to perform semantic analysis on the operational rule information to generate the first question and the answer corresponding to the first question. The second preset model is used to rewrite the first question to generate multiple derivative questions. The derivative questions are similar questions that are semantically equivalent to the first question but have different expressions. The first question, the multiple derivative questions, and the answer are combined into a question-answer pair and stored in a question-answer knowledge base. The question-answer knowledge base is used to provide a matching answer as a basis for subsequent responses when the risk detection result of the second question related to the operational entity determines that the handling strategy for the second question is to provide a proxy answer.

[0006] In a third embodiment, this specification provides a data processing apparatus. The apparatus includes: The acquisition module is used to acquire the operation rule information provided by the first user; the operation rule information includes the internal management norms formulated by the operation entity to which the first user belongs for operation activities and personnel behavior; The generation module is used to perform semantic analysis on the operation rule information using a first preset model to generate a first question and an answer corresponding to the first question; and to rewrite the first question using a second preset model to generate multiple derivative questions; the derivative questions are similar questions that are semantically equivalent to the first question but have different expressions. The storage module is used to store the first question, the plurality of derivative questions, and the answer into a question-and-answer knowledge base; wherein, the question-and-answer knowledge base is used to provide a matching answer as a basis for answering when, based on the risk detection results of the second question related to the operating entity, it is determined that the handling strategy for the second question is to answer on behalf of the operator.

[0007] In a fourth embodiment, this specification provides an electronic device including a memory and a processor, wherein the memory stores executable program instructions, and when the processor executes the program instructions, it implements the method provided in the first embodiment.

[0008] In a fifth embodiment, this specification provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed in a computer, it causes the computer to perform the method provided in the first embodiment.

[0009] In a sixth embodiment, this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the method provided in the first embodiment.

[0010] The solutions provided in the above embodiments of this specification utilize a first preset model to perform semantic analysis on the operational rule information provided by the first user, thereby generating a first question and a corresponding answer. The operational rule information includes internal management regulations formulated by the operational entity to which the first user belongs, concerning operational activities and personnel behavior. Further, a second preset model is used to rewrite the first question, generating multiple derivative questions that are semantically equivalent to the first question but express different ideas. The first question, multiple derivative questions, and corresponding answers are then combined into question-answer pairs and stored in a question-answer knowledge base. This knowledge base can be used to provide matching answers as a basis for subsequent responses when, based on the risk detection results of a second question related to the operational entity, the required handling strategy for the second question is determined to be a proxy answer. Therefore, this solution, through the collaborative approach of two models, effectively converts unstructured operational rule information (such as industry regulations and internal company manuals) into question-answer pairs, enabling the practical application of operational rule information in subsequent content security systems. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the various embodiments disclosed in this specification, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are merely examples of the various embodiments disclosed in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort. In the drawings: Figure 1 The technical process diagrams provided in this specification are for exemplary embodiments and illustrate the technical processes upon which the implementations of the methods are based. Figure 2 This is a schematic diagram of the structure of a data processing system provided for exemplary embodiments in this specification; Figure 3 A flowchart illustrating a data processing method provided as an exemplary embodiment of this specification; Figure 4 A schematic diagram of the structure of a data processing apparatus provided for exemplary embodiments in this specification; Figure 5 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this specification. Detailed Implementation

[0012] In the current environment of simultaneous digitalization and stringent regulation, industries such as finance, government affairs, and internet platforms are continuously upgrading their requirements for the accuracy and timeliness of content security audits. On the one hand, enterprises need to efficiently identify common risks such as politically sensitive, violent, and pornographic content to ensure basic compliance. On the other hand, as regulatory details become more refined and industry self-regulation requirements increase, enterprises' demand for customized content security capabilities has also grown significantly. To achieve customized content security capabilities, enterprises need to integrate relevant industry standards, internal policy manuals, and other unstructured rule texts into general content security systems, while retaining the specificity of their services during the integration process. However, because these unstructured rule texts are usually in natural language form and lack a unified logical structure and machine-executable semantics, enterprises often need to invest significant manpower to manually interpret and extract these unstructured rule texts into executable security policy rules before they can be directly used in content security systems.

[0013] Existing solutions in content security systems primarily employ, but are not limited to, the following methods for blocking inappropriate content: Option 1: Pure Rule Engine Strategy Configuration. This option relies on manually written regular expressions or keyword blacklists / whitelists to directly configure blocking rules, thereby intercepting risky content. However, this approach lacks flexibility; for example, it cannot handle implicitly risky words or semantically variant risky words (such as "leader" referring to a sensitive person). Furthermore, maintenance costs are relatively high; rule expansion can lead to conflicts, requiring frequent manual updates to the regular expressions or keyword blacklists / whitelists. Additionally, it suffers from incomplete coverage, only matching explicit violations (violations directly appearing in the text) and failing to uncover risks arising from contextual combinations.

[0014] Option 2: Single-Model End-to-End Recognition. The core of this option is to use a single Natural Language Processing (NLP) model to directly output risk labels and interception decisions. However, this approach suffers from weak generalization ability and poor adaptability to some non-standard rules (such as internal company manuals). Furthermore, its interpretability is also poor, failing to provide FAQ-level interpretable results, making it difficult for customers to verify. The FAQ consists of pre-prepared questions frequently asked by users and their concise answers. Further, it suffers from low iteration efficiency; for example, due to the high cost of model fine-tuning, it cannot quickly respond to rule changes.

[0015] Option 3: Manually Annotated Risk Database. This option relies entirely on manually annotating risk terms (FAQs) before importing them into a rule engine. However, this approach suffers from high labor costs, such as long annotation cycles and difficulty in covering multilingual and multi-scenario needs. Furthermore, it is highly subjective, with different annotators having inconsistent understandings of the same rule, potentially leading to different annotation results for the same term. Additionally, it often suffers from cold-start issues, such as the need to build rules from scratch for new domains, which are not automated.

[0016] Option 4: Static Terminology Matching. The core of this option is to use a pre-built, general risk terminology database (such as politically sensitive or terrorist terms) for matching. However, this approach suffers from poor domain transferability; for example, financial and medical terms require redundant risk terminology databases. Furthermore, the false positive rate is relatively high, with broad terms (such as "surgery") easily misidentified, and there is a lack of contextual filtering. Moreover, it lacks dynamism and cannot adapt to emerging risks (such as variations in online slang).

[0017] To address the aforementioned issues, the embodiments in this specification provide a solution. The basic idea of ​​this solution is to adopt a combination of large-model-led and small-model-assisted approach to transform unstructured rule text provided by the enterprise into question-answer pairs, thereby integrating them into the content security system.

[0018] Large models refer to deep learning models with a huge number of parameters, a large amount of training data, and strong generalization and generation capabilities. For example, large models can refer to large language models (LLMs).

[0019] Small models are a concept relative to "large models." They typically refer to machine learning or deep learning models with fewer parameters, lower computational resource requirements, and suitability for deployment on edge devices or locally.

[0020] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments in this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0021] It should be noted that, for ease of description, the accompanying drawings only show the parts related to the relevant technical solutions. Unless otherwise specified, the embodiments and features described in this specification can be combined with each other. Furthermore, the terms "first," "second," and "third" used in the embodiments of this specification are for informational purposes only and do not constitute any limitation. Moreover, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, product, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, product, or apparatus. Without further limitations, the presence of other identical or equivalent elements in the process, method, product, or apparatus that includes the stated elements is not excluded. Furthermore, in this specification, unless explicitly stated otherwise, "receiving and transmitting data" does not necessarily mean direct receiving and transmitting; it can be indirect receiving and transmitting. For example, when A receives data sent by B, it can be understood as A directly receiving the data sent by B, or it can be understood as A indirectly receiving the data sent by B through other entities such as C. Similarly, when B sends data to A, it can be understood as B sending the data directly to A, or it can be understood as B indirectly sending the data to A through other entities such as C. Here, C can be one entity, or it can be two or more entities.

[0022] Furthermore, it should be noted that specific terms are used to describe embodiments of this specification. For example, "an embodiment," "one embodiment," and / or "some embodiments" refer to a particular feature, structure, or characteristic related to at least one embodiment of this specification. Therefore, it should be emphasized and noted that "an embodiment," "one embodiment," or "an alternative embodiment" mentioned twice or more in different locations in this specification do not necessarily refer to the same embodiment. Moreover, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples, without contradiction. Although one or more embodiments of this specification provide method steps as described in the embodiments or flowcharts, it is understood that the order of steps listed in the embodiments or flowcharts is merely one possible order of execution among many steps, and does not represent the only possible order. Therefore, when the claims involve method steps, adjustments to the order of such steps, or parallel execution between steps, are also within the scope of protection of the claims.

[0023] Furthermore, it should be noted that if this manual involves user data, the user data obtained is authorized by the user and does not involve user privacy.

[0024] The embodiments provided in this specification will be described below with reference to the accompanying drawings.

[0025] The technical solutions provided in the various method embodiments described below are all based on Figure 1 The technical process shown is implemented, and Figure 1 The technical process shown can be one of the core processing flows of a content security system. For example... Figure 1 As shown, the technical process mainly includes the following steps: 1. Input data The data to be entered here includes operational rules information and risk labeling system.

[0026] Operational rule information can be, for example, custom rule text provided by the first user and related to the operating entity to which it belongs (such as banks or other enterprises). This custom rule text is usually unstructured rule text in natural language form, such as bank regulations or internal corporate manuals. Furthermore, the first user can be an enterprise-level user with content security management needs. The operating entity corresponding to an enterprise-level user refers to the organizational entity that actually provides services and assumes responsibility for content compliance, including but not limited to: financial institutions (such as banks, securities companies, and insurance companies), government agencies, internet platform companies, etc.

[0027] For example, a bank may want to deploy a content security system on its relevant service platforms (such as mobile banking apps), and requires that the system not only have general risk identification capabilities (such as detecting illegal content in public domains such as politically sensitive, terrorist, pornographic, and rumor-related content), but also specific compliance content security governance capabilities tailored to the financial industry and the bank's own regulations. In this case, the bank can submit relevant operational rules information (such as industry regulations and internal bank manuals) to integrate the system. Figure 1 The technical process shown is used to process the information so that its operational rules can be implemented and applied in the content security system.

[0028] A risk labeling system is a set of predefined risk classification standards used to accurately label the risk type of content. In its specific design, a risk labeling system typically adopts a multi-level hierarchical structure, such as a hierarchical design of "first-level risk category -> second-level risk category".

[0029] For example, the hierarchical structure of the risk labeling system is shown in Table 1 below: Table 1 The purpose of this risk labeling system is to provide a basis for determining the risk classification label to which text (such as a question entered by a second user) belongs.

[0030] It should be added here that when specifying the above risk labeling system, the first user can participate to jointly define the labels related to the operator to which they belong, such as "finance - internal compliance".

[0031] 2. Data Processing 21) Convert the rules into FAQs (Frequently Asked Questions). Here, the first pre-defined model is used to understand the key knowledge points (also known as core semantic elements) of operational rule information (such as rule text), such as concepts, terminology, and causal relationships. Then, based on the understood key knowledge points, question-and-answer pairs (containing questions and answers) of explicit and implicit knowledge points are generated. Explicit knowledge points are content directly stated in the rule text, while implicit knowledge points are boundaries or extended meanings that need to be deduced through reasoning. The reasonableness and sensitivity of the question-and-answer pairs are then verified. Reasonableness verification includes, but is not limited to, consistency checks (such as checking whether the answer faithfully reflects the original rule text, logical verification, etc.). Sensitivity verification includes, but is not limited to, sensitive word checks (such as checking whether the answer unintentionally reveals internal rule details, policy compliance checks). In practice, for example, a manual or specially pre-trained verification model can be used to verify the reasonableness and sensitivity of the question-and-answer pairs.

[0032] The aforementioned first preset model is a large model, such as a Large Language Model (LLM). Specifically, the first preset model is obtained, for example, through targeted fine-tuning training of a Large Language Model (LLM), possessing the ability to deeply analyze operational rule information and automatically generate high-quality question-answer pairs. During fine-tuning training, the training data can contain a large number of operational rule samples from different domains and corresponding question-answer pair samples. Furthermore, after the operational rule samples are input into the Large Language Model, the Large Language Model generates corresponding question-answer pairs. The loss between the generated question-answer and the corresponding question-answer pair samples is then calculated. Based on the calculated loss value, the parameters of the Large Language Model (such as temperature, penalty coefficient, etc.) are optimized. The parameter-optimized Large Language Model is the first preset model.

[0033] 22) FAQ Expansion Here, the question generated in step 21) is used as a seed question. The second preset model is used to semantically rewrite the seed question, expand the scenario, and simulate different user roles (i.e., simulate the questioning methods of different roles) to generate diverse derivative questions. The answer to the derivative question is the same as the answer to the seed question.

[0034] Alternatively, the seed question can be expanded in other ways. For example, derivative questions covering more scenarios can be generated by adjusting the question complexity and sentiment of the seed question. Adjusting the question complexity can be achieved by modifying the model parameters (such as temperature parameters) of the second preset model, while the sentiment can be guided by prompts.

[0035] For a detailed description of the specific implementation of the FAQ extension here, please refer to the relevant content in other method embodiments below.

[0036] The second preset model is a small model, such as a text rewriting model based on the BERT architecture. Its input includes a seed problem and rewriting dimensions (such as semantic rewriting, scene expansion, user role transformation, adjustment of sentiment tendency and other rewriting constraints). This enables the second preset model to rewrite the seed problem based on the rewriting dimensions, ensuring that the generated derivative problems do not deviate from the original rules.

[0037] The questions (including seed questions and corresponding derivative questions) and their answers generated by 21) and 22) above will be combined into question-answer pairs and stored in the question-answer knowledge base.

[0038] 23) Risk word extraction: Extract risk words from the question input by the second user. Here, risk word extraction refers to extracting risk words from the question input by the second user. For example, this risk word extraction can be implemented as follows: first, the question input by the second user is preprocessed (e.g., denoising, standardization), and then the preprocessed question is input into a third preset model. The third preset model performs risk analysis on the preprocessed question to extract risk words. For instance, during the risk analysis of the preprocessed question, if the third preset model determines that the question poses a safety risk, it will use matching analysis to select a suitable risk classification label for the question from a predefined risk label system. Based on the risk direction indicated by the risk classification label, it will extract core risk words from the question and assess the violation. Assessing the violation means determining whether the extracted risk words constitute a substantial violation in the current question, not just a superficial violation. The reason for assessing the violation here is that many high-risk words are neutral, and whether they constitute a violation highly depends on the context. Furthermore, the model may misjudge (e.g., "surgery" is legal in medical science popularization but illegal in black market cosmetic surgery advertisements).

[0039] The third preset model can be a large model, such as one that is fine-tuned and trained from a general large language model.

[0040] Furthermore, the extracted risk words will be filtered for neutral words to retain highly targeted violation words. Neutral words are those that do not carry obvious emotional bias, stance, praise or criticism, or sensitive attributes. Highly targeted violation words are those words or phrases that directly and explicitly point to violations, sensitive events, or prohibited content.

[0041] Specifically, the implementation of neutral word filtering may include: In one example, the third preset model can filter neutral words during the risk word extraction process. The technical means to achieve this filtering is as follows: Neutral word filtering rules (such as including manufacturer names and company names as neutral words) are added to the risk detection prompts input to the third preset model. These rules inform the third preset model which words are neutral and do not need to be extracted as risk words. Therefore, the risk detection prompts input to the third preset model can be generated based on the aforementioned preprocessed questions and neutral word filtering rules.

[0042] In another example, if the third preset model lacks neutral word filtering capabilities, its output of risky words may contain neutral words. In this case, neutral word filtering can be achieved through traffic replay verification. Traffic replay verification is based on historical text that has passed risk content review, verifying whether the blocking of risky words causes excessive disruption. Excessive disruption refers to incorrectly judging content (such as words) in the text that should pass as illegal or risky content.

[0043] For example, a specific implementation of traffic replay verification may include the following steps: Step 31: Identify the risky words to be verified Suppose that the risk words to be verified include "freedom" and "organization," and these risk words to be verified are extracted from the questions input by the second user through the second preset model.

[0044] Step 32: Select historical normal traffic samples For example, 100,000 texts that have passed the risk content review are randomly selected from the past 7 days and used as replay samples.

[0045] Step 33: Perform traffic replay The aforementioned 100,000 samples were re-tested for risky content, and the online interception logic was simulated when the aforementioned risky words to be verified were used as the interception target.

[0046] Step 34: Statistical analysis of misjudgment results For example, the false positive results are as follows: "freedom" appeared 4800 times in 100,000 samples, and was blocked 4760 times, corresponding to a false positive rate of 99.2%; "organization" appeared 3200 times in 100,000 samples, and was blocked 3100 times, corresponding to a false positive rate of 96.9%, and so on. Based on these false positive results, we can conclude that "freedom" and "organization" are high-frequency neutral words with high false positive rates. Therefore, "freedom" and "organization" can be filtered out.

[0047] Here, traffic replay verification is used to check whether the risk words are extracted correctly, thereby filtering out neutral words and reducing reliance on manual review.

[0048] In addition, based on the results of traffic replay verification to confirm the accuracy of risk word prompts, the third preset model can be iteratively optimized to improve the accuracy of risk word extraction and reduce the extraction of neutral words.

[0049] 24) Strategy Adaptation Strategy adaptation refers to determining appropriate handling strategies for questions input by a second user, such as interception, automatic answering based on question-answer pairs in a knowledge base, or transfer to manual review.

[0050] In practice, a rule engine can be used to automatically match risk words with a pre-defined control rule base based on the risk classification tags corresponding to the risk words and the intent type and sentiment expressed by the question input by the second user, thereby determining an appropriate handling strategy for the question input by the second user. The control rule base contains multiple rule entries, each containing a risk classification tag and a control rule associated with that risk classification tag. The control rule includes a handling strategy and the triggering conditions for that strategy, and the triggering conditions are configured based on at least one of the intent type, sentiment, etc.

[0051] While risk management based on risk words and control rules can cover most violation scenarios, there are still exceptions in practice where boundaries are blurred or policies allow. For example, a general subjective evaluation of a particular person is considered high-risk and should be blocked; however, if the content involves an objective interpretation or positive commentary on the person's recent public policies, it may be considered compliant expression and should be downgraded in control, not blocked, such as allowing someone to answer on behalf of another. Such controversial use cases cannot be uniformly handled by general control rules and require the introduction of a fine-grained exception configuration mechanism to avoid over-blocking or under-blocking. Therefore, in this solution, controversial use cases will adopt a "separate configuration" and / or "targeted fine-tuning" strategy to achieve precise content risk control. "Separate configuration" refers to configuring the corresponding control rules separately. "Targeted fine-tuning" uses controversial use cases as samples to fine-tune the corresponding model (such as the third preset model).

[0052] 3. Output Results The output here includes: a question-and-answer knowledge base (FAQ knowledge base), a list of risk terms and their corresponding risk classification tags, and the appropriate strategies.

[0053] The question-answering knowledge base stores multiple question-answer pairs. A question-answer pair includes multiple questions (including seed questions and derivative questions generated based on seed questions) and answers shared by multiple questions.

[0054] The risk word list includes risk words extracted from the questions input by the second user. The corresponding risk classification labels are secondary risk labels determined by classifying the questions input by the second user. Specifically, a risk classification label consists of a primary risk label and a sub-risk label belonging to that primary risk label, such as "financial-institutional inquiry" or "political-top figure."

[0055] The adaptation strategy refers to the handling strategy (such as interception, answering on behalf of, or allowing) determined for the questions input by the second user.

[0056] The technical process mentioned above is based on server-side and client-side implementation. See also... Figure 1 As shown, the first preset model, second preset model, etc., in the technical process can be deployed on the server side. The server side can be a server, server cluster, virtual server, or cloud, etc. Unstructured operational rule information, etc., can be input through the client. The client can be, but is not limited to: smartphones, smart wearable devices, tablets, laptops, desktop computers, etc.

[0057] thus, Figure 2 A data processing system according to an embodiment of this specification is also shown, the system including a client 100 and a server 200. Wherein, Client 100 is used to respond to input operations, obtain the operation rule information input by the first user, and send the operation rule information to the server; the operation rule text contains the internal management specifications formulated by the operation entity to which the first user belongs for operation activities and personnel behavior; Server 200 is deployed with a first preset model and a second preset model. The first preset model is used to perform semantic analysis on the operational rule information to generate a first question and a corresponding answer. The second preset model is used to rewrite the first question to generate multiple derivative questions. These derivative questions are semantically equivalent to the first question but have different expressions. The first question, the multiple derivative questions, and the answer are combined into a question-and-answer pair and stored in a question-and-answer knowledge base. This knowledge base is used to provide a matching answer as a basis for subsequent responses when, based on the risk detection results of a second question related to the operational entity, the required handling strategy for the second question is determined to be proxy answering.

[0058] The specific implementation of the functions of the server 200 and client 100 will be described in detail in the following method embodiments, and will not be elaborated here.

[0059] The technical solutions provided in this specification will be described below by way of method embodiments.

[0060] Figure 3 A flowchart illustrating a data processing method according to an embodiment of this specification is shown. The execution entity of this method is the server in the aforementioned system. See also... Figure 3 As shown, the data processing method includes the following steps: 102. Obtain the operation rule information provided by the first user; the operation rule information includes the internal management regulations formulated by the operation entity to which the first user belongs for operation activities and personnel behavior; 104. Using a first preset model, perform semantic analysis on the operational rule information to generate a first question and the answer corresponding to the first question; 106. Using a second preset model, the first problem is rewritten to generate multiple derivative problems; the derivative problems are similar problems that are semantically equivalent to the first problem but have different expressions. 108. Combine the first question, the plurality of derivative questions, and the question and answer to form a question-answer pair and store it in the question-answer knowledge base; The question-and-answer knowledge base is used to provide a matching answer as a basis for answering when the risk detection result of the second question related to the operating entity determines that the handling strategy for the second question is to answer on behalf of the operator.

[0061] In the aforementioned 102, the operating entities include, but are not limited to: various enterprises (such as banks and companies), and entities that conduct business activities in the form of internet platforms, such as e-commerce platforms and content service platforms. Furthermore, the first user can provide operational rule information related to the operating entity through the client application; that is, internal management regulations formulated for operational activities and personnel behavior. For example, operational rule information could be banking industry regulations or corporate internal management manuals (such as employee code of conduct, information security manuals, etc.).

[0062] In practice, the operational rule information may be provided in formats such as Word, PDF, or audio. When the operational rule information is provided in audio format, this embodiment will first use Automatic Speech Recognition (ASR) technology to convert the audio version of the operational rule information into text before proceeding with subsequent processing steps. When the operational rule information is provided in formats such as Word or PDF, the operational rule information may be unstructured plain text provided in natural language text form, or it may be non-plain text (such as containing images and / or tables). If the operational rule information is non-plain text, such as a mixed text and image document containing text content and embedded images (such as violation flowcharts), optical character recognition (OCR) can be performed on the images to extract the embedded text, and / or visual semantic analysis can be performed on the images to obtain text descriptions of the images. These obtained texts are then incorporated into subsequent processing steps.

[0063] To integrate operational rule information into the general security system, this embodiment employs the following technical approach: through steps 104 and 106, two models (i.e., the first preset model and the second preset model) are used collaboratively to convert the operational rule information into structured question-and-answer pairs. These question-and-answer pairs are then integrated as external knowledge into the content security review process to assist in risk control. Detailed implementation of integrating question-and-answer pairs as external knowledge into the content security review process will be described below and will not be elaborated upon here.

[0064] In the aforementioned 104, the first pre-defined model can be, but is not limited to, a large model that has been fine-tuned and trained beforehand, possessing strong semantic understanding, information extraction, and logical reasoning capabilities. Specifically, a large model is characterized by its massive scale, typically consisting of millions or even billions of parameters. These numerous parameters help the large model learn complex patterns in language data. For example, a large model could be a large language model (LLM).

[0065] It should be noted here that when obtaining the first preset model through fine-tuning training, domain-adaptive fine-tuning training can be performed based on the client-specific operational rule corpus. The client-specific rule corpus includes internal management documents formulated by the operating entity, which contain terminology and expression paradigms from employee handbooks, compliance guidelines, and codes of conduct.

[0066] In a specific implementation, in one feasible technical solution, the aforementioned 104 "using the first preset model to perform semantic analysis on the operational rule information to generate a first question and an answer corresponding to the first question" may include: 1042. Input the operation rule information into the first preset model, and the first preset model shall perform the following operations: 10422. Extract key knowledge points from the operational rule information; the key knowledge points include at least one of the following: concepts, terms, causal relationships, scenarios, and behaviors involved in the operational rule information; 10424. Based on the key knowledge points mentioned above, generate the first question and the answer corresponding to the first question.

[0067] The above process generates one or more initial question-and-answer pairs based on the extracted key knowledge points. Each initial question-and-answer pair includes a first question and the corresponding answer.

[0068] To facilitate understanding of the content related to step 104 above, an example is provided below.

[0069] Suppose the operational rules information contains a rule text R as follows: "The use of promises such as 'guaranteed profit' and '100% return' is strictly prohibited in the promotion of financial products." After semantic parsing by the first preset model, the rule text R extracts the following key knowledge points: prohibition of promises (behavior), "guaranteed profit" and "100% return" (terms), and "financial product promotion" (scenario). Further, based on these extracted key knowledge points, one or more initial question-and-answer pairs containing explicit and implicit knowledge points can be generated. Explicit knowledge points refer to the content directly stated in the rule text R. For example, in a generated initial question-and-answer pair containing explicit knowledge points, the first question Q might be "Which phrases cannot be used in the promotion of financial products?", and the corresponding answer A might be "The use of promises such as 'guaranteed profit' and '100% return' is prohibited." Implicit knowledge points refer to the boundaries or extended meanings that need to be deduced through reasoning. For example, in a generated initial question-and-answer pair containing implicit knowledge points, the first question Q might be "Is an expected annualized return of 8% a violation?" The answer to the first question, Q, is: "If the 'uncertain returns' or 'past performance does not predict future results' is not simultaneously indicated, it may constitute a disguised promise of returns, which is a violation."

[0070] Furthermore, to ensure the quality of the generated initial question-and-answer pairs, step 106 can be triggered only after the initial question-and-answer pairs have passed the rationality and sensitivity checks. Rationality checks include, but are not limited to, consistency checks (such as checking whether the answers faithfully adhere to the original rule text, logical verification, etc.). Rationality checks can effectively detect whether the first preset model exhibits "illusion" or overgeneralization when generating the initial question-and-answer pairs. Sensitivity checks include, but are not limited to, sensitive word checks (such as checking whether the answers accidentally reveal internal rule details, policy compliance checks). This sensitivity check can prevent the content of the generated initial question-and-answer pairs from becoming a new source of risk.

[0071] In practice, for example, a manual or specially pre-trained verification model can be used to verify the rationality and sensitivity of question-answer pairs.

[0072] Furthermore, for the initial question-answer pairs that pass validation, the first question can be used as a seed question. Using a second pre-defined model, through prompt word engineering or model parameter adjustments, the seed question can be rewritten from multiple dimensions such as semantics, scenario, user role, and sentiment. This generates diverse question variations for the original corresponding rule text, improving question coverage. The second pre-defined model can be, but is not limited to, a small model that has been fine-tuned. Small models are characterized by fewer parameters and simpler structure. For example, the second pre-defined model could be a BERT model.

[0073] Based on this, in one feasible technical solution, step 106, "using a second preset model to rewrite the first problem and generate multiple derivative problems," may include: 1062. Based on the first question and the preset rewriting dimensions, generate question generation prompts; input the question generation prompts into the second preset model, and the second preset model rewrites the first question according to the rewriting dimensions to generate the multiple derivative questions. The rewriting dimensions include, but are not limited to, at least one of the following: semantic rewriting, scenario expansion, user role transformation, and adjustment of emotional tone. Semantic rewriting methods may include synonym substitution and sentence structure transformation (such as question / rhetorical question, hypothesis). Scenario expansion refers to binding to specific service scenarios, such as loan approval. Adjusting emotional tone includes adjusting emotions / tones, adding emotional colors such as complaint, urgency, and probing. User role transformation, for example, the user role of asking the question can be changed to bank employee, bank customer, or ordinary user.

[0074] For example, taking the rewriting of the seed question by injecting different user roles as an example, based on the user roles to be injected and the first question such as "Can I post unpublished financial data of a company on social media?", a question prompt is generated. This question prompt is then input into the second preset model. The second preset model can generate multiple derivative questions, such as the following: Q1: As a customer, is it against the rules for me to post screenshots of my bank's returns on my WeChat Moments? Question 2: Can our tellers share their quarterly performance results on Douyin? In another feasible technical solution, step 106 above, "using a second preset model to rewrite the first problem and generate multiple derivative problems," may include: 1062' Adjust the generation control parameters of the second preset model; based on the adjusted generation control parameters, drive the second preset model to rewrite the first problem and generate multiple derivative problems.

[0075] In the above, the adjusted generation control parameters could be, for example, temperature parameters. For instance, by increasing the generation control parameters, the complexity of the second preset model can be improved, such as by selecting richer vocabulary and more flexible sentence structures to rewrite the first problem, thereby increasing the linguistic coverage of the derived problems to a certain extent.

[0076] Besides the methods mentioned above, such as using prompt word engineering and adjusting model parameters to rewrite the first question, other methods can be used, such as prefix generation and question fusion. This prefix generation and question fusion method refers to generating prefix text statements such as character / scene using a second preset model, and then semantically fusing (e.g., concatenating) these prefix text statements with the first question to generate corresponding derivative questions.

[0077] For example, suppose the prompt input to the second preset model is: "Please generate a question prefix from a customer's perspective, and merge this question prefix with the first question to generate derivative questions. The first question is: #####". Then: the second preset model generates prompts based on this question. The generated prefix text may include, for example, "As an ordinary user, I want to ask..." or "If I were a complaining customer,...". By merging these prefix texts with the first question, multiple corresponding derivative questions can be generated.

[0078] By using this prefix generation and question fusion method, the generation speed of derivative questions is faster, and some semantic deviations caused by the model's free play are avoided.

[0079] In the above 108, since the derivative questions generated based on the first question can share the same answer as the first question, the first question, the corresponding multiple derivative questions, and the answer corresponding to the first question can be associated with each other and stored in the question-and-answer knowledge base in the form of question-and-answer pairs for later use when needed.

[0080] Therefore, in this embodiment, the question-answering knowledge base stores multiple question-answer pairs. Each question-answer pair contains multiple questions and answers shared by these multiple questions, with the multiple questions including a first question and derivative questions generated based on the first question.

[0081] It should be noted that this embodiment will also update the question-answer pairs in the question-answer knowledge base according to the updated operation rule information when an update is detected, so as to adapt to the change in operation rule information.

[0082] Furthermore, it should be noted that while the aforementioned steps (steps 102, 104, 106, and 108) primarily introduce how to generate corresponding question-and-answer pairs based on the operational rule information provided by the client (i.e., the first user), considering that some clients may not provide specific operational rule information and only wish to provide relevant industry rules (financial industry rules), this embodiment also provides a pre-built industry question-and-answer pair library (such as a financial industry violation question-and-answer pair library). This pre-built industry question-and-answer pair library can be formed by accumulating question-and-answer pair cases related to industry rules, allowing it to be reused directly for some clients without having to perform the aforementioned steps.

[0083] Furthermore, the method provided in this embodiment may also include the following steps: 1010. Obtain the second question input by the second user; 1012. Input the second question into the third preset model, trigger the third preset model to perform risk detection, and output the risk detection result of the second question; 1014. Based on the risk detection results, determine an appropriate handling strategy for the second problem; 1016. Based on the aforementioned handling strategy, implement corresponding control measures for the second problem.

[0084] In 1010 and 1012 above, the second question input by the user can be in plain text form or in a mixed text and image form. Furthermore, the third preset model can be obtained, for example, by fine-tuning a general-purpose large language model or a multimodal model (such as a visual language model). If the third preset model is obtained by fine-tuning a large language model, then this third preset model can only handle risk detection for text-based questions. If the third preset model is obtained by fine-tuning a multimodal model, then this third preset model can only handle risk detection for questions containing multiple modalities such as text and images. Preferably, in this embodiment, the second question is in plain text form, and the third preset model is obtained by fine-tuning a large language model.

[0085] It's worth noting that for the third-preset models required in different domains, knowledge distillation techniques can be used to generate lightweight, domain-specific third-preset models (as student models) based on a pre-trained general risk detection model. The general risk detection model possesses broad semantic understanding and risk identification capabilities, while the third-preset model retains its core risk identification capabilities while adapting to the terminology and compliance boundaries of specific domains and meeting the requirements for low latency and high concurrency deployment. For example, a general risk detection model can be trained first, capable of identifying common violations such as political, terrorist, pornographic, and advertising violations. Subsequently, for financial clients, internal policy texts and labeled samples are collected, and knowledge distillation techniques are used, with the general risk detection model serving as the teacher, to train a financial domain-specific third-preset model with fewer parameters. This third-preset model achieves risk identification accuracy close to the teacher model in financial scenarios but with faster inference speeds, making it suitable for deployment on banking platforms.

[0086] In practical implementation, the second problem can be preprocessed (e.g., denoising, standardization) before being input into the third preset model. When performing risk detection, the third preset model first classifies the second problem to determine if a risk exists. If a risk is identified, it further extracts the risk content from the second problem based on the corresponding risk classification tags, thus outputting a risk detection result containing the risk content and risk classification tags. Conversely, if no risk is identified, a risk detection result containing "no risk" is output, and the risk detection process ends. The risk classification tags can be determined by matching from a predefined risk tag system based on the semantic analysis results of the input data. Therefore, this solution uses the third preset model combined with a risk tag system to achieve risk detection for the second problem, effectively reducing the false negative rate of risk content.

[0087] Among them, if the second question is in pure text form, when it is determined that there is a risk in the second question, the risk content output is a risk word. If the second question is in a form of mixed text and images, when it is determined that there is a risk in the second question, the risk content output may include at least one of a risk word and a risk area (such as a flag) marked in the image, and may also include a cross-modal association relationship between the risk word and the risk area.

[0088] Moreover, the risk word can specifically be extracted from the second question based on the risk direction indicated by the risk classification label, which can improve the accuracy of risk word extraction. The types of risk words include but are not limited to: ordinary single risk words (such as explosives), concatenated words, and adversarial variant words. The concatenated words contain at least two semantically related words, and a connection symbol (such as "&", underscore or other non-character data interference symbols) is inserted between adjacent words. Exemplarily, concatenated words include, for example, "guns & purchase". The adversarial variant words are variant expressions with unchanged semantics formed by modifying known risk words. Specifically, the adversarial variant words can include but are not limited to at least one of the following: variant words formed by inserting interference symbols within known risk words, such as "I$IS"; variant words formed by using space or line break as separator characters, such as "t a i d u"; homophonic words generated based on known risk words, for example, the homophonic words of "explosives" can be "slag explosive", "brake explosive", "sudden explosive", etc.

[0089] To improve the accuracy of risk word extraction, this embodiment supports neutral word filtering during the risk word extraction process and / or after the risk word is output. The neutral word filtering can be achieved by adding neutral word filtering rules to the risk detection prompt words of the third preset model, and / or can also be achieved through traffic playback verification.

[0090] Thus, the method provided by this embodiment may further include the following steps: S11. During the process of the third preset model extracting risk words from the second question, the candidate risk words are screened based on the preset neutral word filtering rules to exclude the candidate risk words belonging to neutral words from the input risk detection results. And / or S12. The risk words included in the risk detection results are verified by playback based on historical texts to filter out the risk words belonging to neutral words.

[0091] For the specific implementation descriptions of the above steps S11~S12, reference can be made to the relevant content described in the foregoing other embodiments in combination with Figure 1 The relevant content described is not specifically elaborated here.

[0092] In addition to the above, other example solutions can also first use the third preset model to initially extract risk words from the second question, and then push controversial risk words with confidence levels below the preset threshold or semantic ambiguity to the manual review platform for manual review. The final risk words are then determined based on the manual review results, and these results can also be used for subsequent optimization of the third preset model.

[0093] The resulting risk terms can be stored in a database, specifically a risk term library. For example, risk terms can be stored in the risk term library based on their associated risk category tags. The risk term library is a database used to store risk terms, containing various types of risk terms accumulated over time. For instance, the risk term library can use a key-value (KV) structure to store risk terms, where the key (K) represents the risk category and the value (V) represents the risk term. Furthermore, the risk term library can also serve as external knowledge for a third pre-defined model, allowing it to be used during risk term extraction.

[0094] It should be noted that, besides updating the risk terminology database based on risk words extracted from the third preset model, other methods can also be used. For example, the risk terminology database can be updated through risk popularity monitoring. Updating the risk terminology database through risk popularity monitoring means dynamically updating it in conjunction with public opinion data. Specifically, after obtaining information on hot events from external public opinion data sources, the associated words of these hot events (such as riots in a certain area, new types of fraudulent rhetoric) can be assessed for risk. If a associated word is determined to pose a content security risk, it is considered a risk word and thus added to the risk terminology database.

[0095] In the above 1014, if the risk detection result is "no risk", then the handling strategy determined for the second issue can be the release strategy, allowing the second issue to be published normally, that is, allowing the second issue to be displayed to the public (such as being displayed on a public page).

[0096] If the risk detection results contain risk classification tags, risk words, etc., it indicates that the second question is risky. In this case, further semantic analysis will be performed on the second question to extract the intent type, sentiment tendency, etc. Based on the extracted intent type and sentiment tendency, an appropriate handling strategy (such as interception or proxy answering) will be determined for the second question. Therefore, when the second problem presents a risk, step 1014 above, "based on the risk detection results, determining an appropriate handling strategy for the second problem," may include: 10142. Extract the problem features of the second problem; the problem features include at least one of the intention type (such as objective fact acquisition, subjective evaluation, policy questioning, etc.) and sentiment tendency (such as positive, neutral, negative) of the second problem; 10144. Based on the risk classification labels and the problem characteristics, determine an appropriate handling strategy for the second problem.

[0097] In the above 10144, the handling strategy determined for the second problem can be selected from a preset control rule base. The control rule base includes multiple risk classification tags, and each risk classification tag is configured with one or more control rules. The control rules include the handling strategy and the triggering conditions for the handling strategy. The triggering conditions are defined by at least one of the following: entity object, intent type, and sentiment tendency. For example, under the risk classification tag 'Political Involvement - Number One Person', a control rule configured is: when the content contains the specified entity, and the user's intent is evaluative and the sentiment tendency is non-neutral, execute the blocking strategy; when the content is a neutral policy paraphrase, execute the release strategy.

[0098] Based on this, a specific implementation of step 10144 above may include: 101442. Based on the risk classification label and the problem characteristics, select the control rules that are suitable for the second problem from the control rule base; 101444. Based on the selected control rules, determine the handling strategy for the second problem.

[0099] For example, suppose the second question is "What is our bank's employee leave policy?". Analysis determines that the risk classification type for this second question is "financial-system inquiry," and the extracted risk words include "our bank." Further analysis reveals that the intent type of this second question is "customer fact gathering," and the sentiment is "neutral." Combined with the control work under "financial-system inquiry" in the control rule base, which states "objective inquiries involving internal systems, executable proxy answering," the handling strategy for this second question can be determined as proxy answering. However, if the second question is "Is your bank's leave policy too strict?", based on the risk classification label and question characteristics (intent type: subjective evaluation, sentiment: negative), by matching with the control rule base, the handling strategy for this second question can be determined as interception (or warning, transfer to manual review, etc.).

[0100] In the above scenario 1016, if the determined handling strategy for the second question is to allow it, then the second question will be published normally, such as being displayed on the relevant public page. If the determined handling strategy for the second question is to block it, then the control measures implemented for the second question may include: prohibiting the second question from being presented externally, not generating any response content (i.e., not providing an answer), and recording a violation log for auditing. If the determined handling strategy for the second question is to provide an answer on behalf of the question, then the control measures implemented for the second question may include: hiding the second question so that it is not presented externally, and only returning the generated answer.

[0101] When the handling strategy is to provide an answer on behalf of the user, the answer to the second question can be generated and returned to the second user in the following ways, but is not limited to: After vectorizing the second question, a similarity search is performed in the question-and-answer knowledge base to select a target question similar to the second question. Then, based on the answer to the target question, an answer to the second question is generated. Specifically, the answer to the target question can be directly used as the answer to the second question, or the answer to the target question can be used as a reference to generate a new answer as the answer to the second question. For example, suppose the second question is "How do employees apply for annual leave?", and the target question similar to the second question is selected from the question-and-answer knowledge base as "What is the annual leave application process for our bank's employees?", and the answer to the target question is "Please submit your application through the ## system". Then, the answer to the target question can be directly used as the answer to the second question, or a more conversational version of the answer can be generated, such as "You can log in to the ## system to submit your annual leave application~", as the answer to the second question.

[0102] For specific implementation details of the steps described above in this embodiment, please refer to the relevant content in other embodiments, which will not be repeated here. Furthermore, the method provided in this embodiment may also include some steps disclosed in other embodiments, which can also be referred to the relevant content in other embodiments, and will not be repeated here.

[0103] The preceding text primarily described how two collaborative models were used to integrate unstructured enterprise operational rule information into a content security system. This enabled the content security system to possess not only general risk identification capabilities but also customized content security governance capabilities for enterprises. Specifically, one model (the first pre-defined model) was responsible for handling the complex semantic understanding of enterprise operational rule information and generating initial question-and-answer pairs, while the other model (the second pre-defined model) focused on scenario-based expansion of the questions in the initial question-and-answer pairs. In addition, other technical means can be employed in other example solutions to enable the content security system to possess customized content security governance capabilities for enterprises. For instance, enterprise users can provide their own pre-built risk term lists instead of the original operational rule information, thereby enabling the rapid injection of customized risk terms into the content security system. This allows the content security system to achieve customized content security governance capabilities for enterprises based on these risk term lists. This approach, which allows enterprise users to skip model-based risk term extraction and directly provide risk term lists, is well-suited for low-resource scenarios. Furthermore, for some enterprise clients, pre-built industry-specific question-and-answer pairs (such as a financial industry violation question-and-answer pair library) can be directly integrated into the content security system. This pre-built industry Q&A library can be formed by accumulating Q&A cases related to industry rules, making it directly reusable for some enterprise clients.

[0104] Furthermore, the control rules configured for some risk classification tags when describing the control rule base mentioned above can sometimes be directly reused across different enterprise clients without reconfiguration. And / or, when facing content security control in similar fields, similar risk classification tags and / or risk terms can also be directly reused, such as the risk classification tags "finance-money laundering" and "e-commerce-cash-out" which can be directly reused.

[0105] The above text combined Figure 3 Specific embodiments of the embodiments described herein have been described. It should be noted that other embodiments are within the scope of the appended claims. Furthermore, in some cases, the actions or steps described in the specification may be performed in a different order than those shown in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0106] The apparatus embodiments corresponding to the method embodiments provided in this specification are described below.

[0107] Figure 4 A schematic diagram of the structure of an apparatus provided in an exemplary embodiment of this specification is shown. For example... Figure 4As shown, the device includes: an acquisition module 21, a generation module 22, and a storage module 23. Among them, The acquisition module 21 is used to acquire the operation rule information provided by the first user; the operation rule information includes the internal management norms formulated by the operation entity to which the first user belongs for operation activities and personnel behavior; The generation module 22 is used to perform semantic analysis on the operation rule information using a first preset model to generate a first question and an answer corresponding to the first question; and to rewrite the first question using a second preset model to generate multiple derivative questions; the derivative questions are similar questions that are semantically equivalent to the first question but have different expressions. The storage module 23 is used to store the first question, the plurality of derivative questions and the answer into a question-answering pair in the question-answering knowledge base; wherein, the question-answering knowledge base is used to provide a matching answer as a basis for answering when the risk detection result of the second question related to the operating entity determines that the handling strategy to be implemented for the second question is to answer on behalf of the operator.

[0108] In one possible implementation, the generation module 22, when used to rewrite the first problem using a second preset model to generate multiple derivative problems, can specifically be used to: generate problem generation prompts based on the first problem and preset rewriting dimensions; input the problem generation prompts into the second preset model, which then rewrites the first problem according to the rewriting dimensions to generate the multiple derivative problems; and / or adjust the generation control parameters of the second preset model; and drive the second preset model to rewrite the first problem based on the adjusted generation control parameters to generate the multiple derivative problems.

[0109] In one possible implementation, the acquisition module 21 is further configured to acquire the second question input by the second user. The device also includes an execution module and a determination module. The execution module is configured to input the second question into a third preset model, triggering the third preset model to perform risk detection and output a risk detection result for the second question. The determination module is configured to determine an appropriate handling strategy for the second question based on the risk detection result. The execution module is also configured to perform corresponding control actions on the second question based on the handling strategy; wherein, when a risk is detected in the second question, the risk detection result includes a risk classification label.

[0110] In one possible implementation, the determining module, when determining an appropriate handling strategy for the second problem based on the risk detection result, has the following functions: if the risk detection result is no risk, then the handling strategy determined for the second problem is to allow it to be displayed; if the risk detection result includes a risk classification label, then extracting the problem features of the second problem, and determining an appropriate handling strategy for the second problem based on the risk classification label and the problem features; wherein the problem features include intent type and sentiment tendency.

[0111] In one possible implementation, when the aforementioned determining module is used to determine an appropriate handling strategy for the second problem based on the risk classification label and the problem characteristics, it can specifically be used to: select a control rule suitable for the second problem from a control rule base based on the risk classification label and the problem characteristics; and determine a handling strategy for the second problem according to the selected control rule; wherein the control rule base includes multiple risk classification labels and control rules configured under the risk classification labels; the control rule includes a handling strategy and a triggering condition for the handling strategy; and the triggering condition is defined by at least one of intent type and sentiment tendency.

[0112] In one feasible implementation, when a risk is detected in the second problem, the risk detection result further includes risk words; these risk words are extracted from the second problem based on the risk direction indicated by the risk classification label. The risk words include connecting words; each connecting word contains at least two semantically related words, and a connecting symbol is inserted between adjacent words.

[0113] In addition, the aforementioned storage module is also used to store the risk words into the risk word library.

[0114] In one possible implementation, when the above-mentioned generation module is used to perform semantic analysis on the operation rule information using the first preset model to generate a first question and an answer corresponding to the first question, it is specifically used to: input the operation rule information into the first preset model, and have the first preset model perform the following operations: extract key knowledge points from the operation rule information; the key knowledge points include at least one of the following: concepts, terms, causal relationships, scenarios, and behaviors involved in the operation rule information; and generate the first question and an answer corresponding to the first question based on the key knowledge points.

[0115] It should be noted that the above-mentioned devices can implement the technical solutions described in the corresponding method embodiments. The specific implementation principles of each module or unit can be found in the relevant content of the corresponding method embodiments, and will not be elaborated further here. Furthermore, for ease of description, the above devices are described by function as various modules or units. Of course, when implementing one or more of this specification, the functions of each module or unit can be implemented in one or more software and / or hardware, or a module that implements the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0116] Furthermore, embodiments of this specification also provide an electronic device. For example... Figure 5 As shown, the electronic device 900 includes a memory 91 and a processor 92.

[0117] The aforementioned memory 91 can be implemented by at least one volatile or non-volatile storage device of any type, or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. Furthermore, the memory, wholly or partially, can be integrated with the processor. The memory can contain both removable and non-removable components.

[0118] The processor 92 described above may include one or more general-purpose processors and / or special-purpose processors.

[0119] Furthermore, memory 91 may contain a non-transitory computer-readable medium storing executable program instructions 912 (e.g., compiled or uncompiled program logic and / or machine code). Processor 92 is capable of executing the program instructions 912 stored in memory to implement any methods, processes, or functions disclosed in this specification and / or the accompanying drawings. Additionally, execution of program instructions 912 by processor 92 may result in processor using corresponding data 911.

[0120] For example, the program instructions 912 described above may include an operating system 9122 (e.g., an operating system kernel, device drivers, and / or other modules) installed on the electronic device 900, and one or more application programs 9121 (e.g., a browser, social media application, or game application). Similarly, the data 911 described above may include operating system data 9112 and application data 9111. The operating system data 9112 is primarily accessible to the operating system 9122, while the application data 9111 is primarily accessible to one or more application programs 9121. The application data 9111 may reside in a file system visible or hidden from the user of the electronic device 900.

[0121] Application 9121 can communicate with operating system 9122 through one or more application programming interfaces (APIs). These APIs facilitate application 9121 in reading and / or writing application data, transmitting or receiving information via communication components, and receiving or displaying information on the user interface. In some terms, application 9121 may be simply referred to as "app". Furthermore, application 9121 can be downloaded to the electronic device through one or more online application stores or app markets. However, application 9121 can also be installed on electronic device 400 in other ways, such as through a web browser or a physical interface on electronic device 900 (e.g., a USB port). Furthermore, such as Figure 5 As shown, the electronic device also includes other components such as a communication component 93, a display 94, a power supply component 95, an audio component 96, and a user interface 97. Figure 5 The diagram only shows some components and does not mean that the electronic device 900 includes only these components. Figure 5 The components shown. Additionally... Figure 5 The components within the dashed box are optional, not mandatory, and their specific requirements depend on the product form of the electronic device 900. The electronic device 900 in this embodiment can be a terminal device such as a desktop computer, laptop computer, smartphone, or IoT device; it can also be a server-side device such as a conventional server, cloud server, or server array; or it can be an integrated device combining terminal and server-side devices. If the electronic device 900 in this embodiment is implemented as a terminal device such as a desktop computer, laptop computer, or smartphone, it may include... Figure 5 The components within the dashed box; if the electronic device 900 in this embodiment is implemented as a server-side device such as a conventional server, cloud server, or server array, then it may not include... Figure 5 The component within the dashed box.

[0122] The aforementioned communication component 93 is configured to facilitate wired or wireless communication between the device housing the communication component and other devices. The device housing the communication component 93 can access wireless networks based on communication standards, such as 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication component 93 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. Specifically, the communication component 93 includes a communication interface that enables the electronic device 900 to communicate with other devices, access networks, and transmission networks via analog or digital modulation. For example, the communication interface may include a chipset and antenna for wireless communication with a radio access network or access point. Furthermore, the communication interface can be a wired interface such as Ethernet, Token Ring, or a USB port, or a wireless interface such as Wi-Fi, Bluetooth, Global Positioning System (GPS), or a wide-area wireless interface (e.g., WiMAX or LTE). Of course, the communication interface can also support other forms of physical layer interfaces and standard or proprietary communication protocols. The communication interface may also include multiple physical communication interfaces, such as Wi-Fi, Bluetooth, and wide-area wireless interfaces.

[0123] The aforementioned display 94 includes a screen, which may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen can be implemented as a touchscreen to receive input signals from the user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.

[0124] The power supply component 95 provides power to various components of the device in which it resides. The power supply component 95 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply component resides.

[0125] The aforementioned audio component 96 can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0126] The user interface 97 described above includes receiving user input and providing output to the user. Therefore, the user interface 97 may include input components such as a keypad, keyboard, touch-sensitive or presence-sensitive panel, computer mouse, trackball, joystick, microphone, still camera, and video camera, and output components such as a display screen (which may be combined with a touch-sensitive panel), CRT, LCD, LED, display using DLP technology, printer, and other similar devices known or developed in the future. The user interface 97 may also generate auditory output via speakers, speaker jacks, audio output ports, audio output devices, headphones, and other similar devices known or developed in the future. In some embodiments, the user interface 97 may include software, circuitry, or other forms of logic capable of transmitting and receiving data from external user input / output devices. Additionally or alternatively, the electronic device 900 may support remote access from other devices via a communication interface or another physical interface (not shown). The user interface 97 may be configured to receive user input, the position and movement of which may be indicated by indicators or cursors described herein. The user interface 97 may also be configured as a display device for rendering or displaying text fragments.

[0127] Accordingly, embodiments of this specification also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium. Furthermore, embodiments of this specification also provide a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed in a computer, it causes the computer to perform actions such as... Figure 3 The method described.

[0128] This specification also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implements... Figure 3 The method described.

[0129] Those skilled in the art will recognize that the functions described in the various embodiments disclosed in this specification in one or more of the examples above can be implemented using hardware, software, firmware, or any combination thereof. When implemented in software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium.

[0130] The specific embodiments described above further illustrate the purpose, technical solutions, and beneficial effects of the multiple embodiments disclosed in this specification. It should be understood that the above descriptions are merely specific implementations of the multiple embodiments disclosed in this specification and are not intended to limit the protection scope of the multiple embodiments disclosed in this specification. Any modifications, equivalent substitutions, improvements, etc., made based on the technical solutions of the multiple embodiments disclosed in this specification should be included within the protection scope of the multiple embodiments disclosed in this specification.

Claims

1. A data processing method, characterized in that, include: Obtain operational rule information provided by the first user; The operational rules information includes the internal management regulations formulated by the operational entity to which the first user belongs for operational activities and personnel behavior; Using a first preset model, semantic analysis is performed on the operational rule information to generate a first question and the corresponding answer to the first question. Using a second preset model, the first problem is rewritten to generate multiple derivative problems; the derivative problems are similar problems that are semantically equivalent to the first problem but have different expressions. The first question, the plurality of derivative questions, and the answer are combined into a question-answer pair and stored in the question-answer knowledge base; The question-and-answer knowledge base is used to provide a matching answer as a basis for answering when the risk detection result of the second question related to the operating entity determines that the handling strategy for the second question is to answer on behalf of the operator.

2. The method according to claim 1, characterized in that, Using the second preset model, the first problem is rewritten to generate multiple derivative problems, including: Based on the first question and preset rewriting dimensions, generate question generation prompts; input the question generation prompts into the second preset model, which then rewrites the first question according to the rewriting dimensions to generate the multiple derivative questions; and / or, Adjust the generation control parameters of the second preset model; based on the adjusted generation control parameters, drive the second preset model to rewrite the first problem and generate multiple derivative problems.

3. The method according to claim 1 or 2, characterized in that, Also includes: Obtain the second question from the second user's input; The second question is input into the third preset model, which triggers the third preset model to perform risk detection and outputs the risk detection result of the second question. Based on the risk detection results, a suitable handling strategy is determined for the second problem; Based on the aforementioned handling strategy, corresponding control measures will be implemented for the second problem; When the risk of the second problem is detected, the risk detection result includes a risk classification label.

4. The method according to claim 3, characterized in that, Based on the risk detection results, a suitable handling strategy is determined for the second problem, including: If the risk detection result is no risk, then the handling strategy determined for the second problem is to allow it to be displayed. If the risk detection result includes a risk classification label, then the problem features of the second problem are extracted, and based on the risk classification label and the problem features, an appropriate handling strategy is determined for the second problem; wherein, the problem features include intent type and sentiment tendency.

5. The method according to claim 4, characterized in that, Based on the risk classification labels and the problem characteristics, a suitable handling strategy is determined for the second problem, including: Based on the risk classification labels and the problem characteristics, select control rules that are suitable for the second problem from the control rule base; Based on the selected control rules, a handling strategy will be determined for the second problem; The control rule base includes multiple risk classification labels and control rules configured under the risk classification labels; the control rules include disposal strategies and triggering conditions for the disposal strategies; the triggering conditions are configured based on at least one of intent type and sentiment tendency.

6. The method according to claim 5, characterized in that, Also includes: When the determined handling strategy for the second question is to provide a proxy answer, a target question similar to the second question is selected from the question-and-answer knowledge base, and the answer corresponding to the target question is obtained; Based on the answer to the target question, determine the answer to the second question; And, based on the aforementioned handling strategy, corresponding control measures are implemented for the second problem, including: Return the answer to the second question to the second user and prevent the second question from being displayed.

7. The method according to claim 3, characterized in that, When a risk is detected in the second problem, the risk detection result also includes risk words; the risk words are extracted from the second problem based on the risk direction indicated by the risk classification label. The risk words include connecting words; the connecting words contain at least two semantically related words, and a connecting symbol is inserted between adjacent words; Furthermore, the method further includes: The risk words are stored in the risk word database.

8. The method according to claim 1 or 2, characterized in that, Using the first preset model, semantic analysis is performed on the operational rule information to generate a first question and the corresponding answer, including: The operational rule information is input into the first preset model, and the first preset model performs the following operations: Extract key knowledge points from the operational rule information; the key knowledge points include at least one of the following: concepts, terms, causal relationships, scenarios, and behaviors involved in the operational rule information; Based on the key knowledge points mentioned above, generate the first question and the corresponding answer to the first question.

9. A data processing system, characterized in that, include: The client is used to respond to input operations, obtain the operation rule information input by the first user, and send the operation rule information to the server; the operation rule text contains the internal management specifications formulated by the operation entity to which the first user belongs for operation activities and personnel behavior; On the server side, a first preset model and a second preset model are deployed. The first preset model is used to perform semantic analysis on the operation rule information to generate the first question and the answer corresponding to the first question. The first question is rewritten using the second preset model to generate multiple derivative questions. These derivative questions are similar to the first question but have different expressions. The first question, the multiple derivative questions, and the answer are combined into a question-and-answer pair and stored in a question-and-answer knowledge base. The question-and-answer knowledge base is used to provide a matching answer as a basis for answering when the risk detection result of the second question related to the operating entity determines that the handling strategy for the second question is to provide a proxy answer.

10. A data processing apparatus, characterized in that, include: The acquisition module is used to acquire the operational rule information provided by the first user; The operational rules information includes the internal management regulations formulated by the operational entity to which the first user belongs for operational activities and personnel behavior; The generation module is used to perform semantic analysis on the operation rule information using a first preset model to generate a first question and an answer corresponding to the first question; Furthermore, using a second preset model, the first problem is rewritten to generate multiple derivative problems; the derivative problems are similar problems that are semantically equivalent to the first problem but have different expressions. The storage module is used to store the first question, the plurality of derivative questions, and the answer into a question-and-answer knowledge base; wherein, the question-and-answer knowledge base is used to provide a matching answer as a basis for answering when, based on the risk detection results of the second question related to the operating entity, it is determined that the handling strategy for the second question is to answer on behalf of the operator.

11. An electronic device, characterized in that, The method includes a memory and a processor, wherein the memory stores executable program instructions, and the processor executes the program instructions to implement the method of any one of claims 1 to 6.

12. A computer-readable storage medium, characterized in that, The storage medium stores a computer program that, when executed in a computer, causes the computer to perform the method described in any one of claims 1 to 6.

13. A computer program product, characterized in that, The computer program product includes a computer program or instructions that, when executed by a processor, implement the method of any one of claims 1 to 6.