Risk content detection method and device based on large model, equipment and medium

By applying a large-model-based risk content detection method in user-generated content, and using natural language understanding and logical ability to identify and evaluate obscure risks, the problem of identifying obscure risk content in the prior art is solved, and more accurate risk content detection is achieved.

CN119990094APending Publication Date: 2025-05-13BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510142362.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively and accurately identify obscure and risky content in user-generated content, especially in community social products.

Method used

By using the natural language understanding ability of the first large model to identify risk elements in the text to be detected and searching related risk content in the preset knowledge base, the logical ability of the second large model is used to detect it in combination with related risk content.

Benefits of technology

It realizes accurate identification of obscure risks in the detection text and accurate assessment of complex risk content, and improves the detection accuracy of risk content in user-generated content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990094A_ABST
    Figure CN119990094A_ABST
Patent Text Reader

Abstract

The invention provides a risk content detection method and device based on a large model, equipment and a medium, relates to the technical field of artificial intelligence, in particular to the technical fields of natural language processing, deep learning and the like, and can be used for application scenes such as generative retrieval, document intelligent editing, intelligent assistants, virtual assistants, intelligent e-commerce and the like. The method comprises the following steps: identifying at least one risk element in a to-be-detected text by utilizing a first large model; retrieving related risk content corresponding to the at least one risk element in example risk content included in a preset knowledge base; and performing risk content detection based on the related risk content and the to-be-detected text by using the second large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, especially to technical fields such as natural language processing and deep learning, and can be used in application scenarios such as generative retrieval, intelligent document editing, intelligent assistants, virtual assistants, and intelligent e-commerce. It specifically relates to a risk content detection method based on a large model, a risk content detection device based on a large model, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Artificial intelligence is a discipline that studies how to use computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.). It includes both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include natural language processing technology, computer vision technology, speech recognition technology, as well as machine learning / deep learning, big data processing technology, knowledge graph technology, and other major directions.

[0003] Content risk control of community social products has always been a difficult problem in the industry. How to effectively and accurately identify risky content in user-generated content is an issue that needs to be solved urgently.

[0004] The methods described in this section are not necessarily methods that have been previously conceived or used. Unless otherwise indicated, it should not be assumed that any method described in this section is considered to be prior art simply because it is included in this section. Similarly, unless otherwise indicated, the issues mentioned in this section should not be considered to have been recognized in any prior art. Summary of the invention

[0005] The present disclosure provides a risky content detection method based on a large model, a risky content detection device based on a large model, an electronic device, a computer-readable storage medium, and a computer program product.

[0006] According to one aspect of the present disclosure, a large model-based risk content detection method is provided, comprising: using a first large model to identify at least one risk element in a text to be detected; retrieving relevant risk content corresponding to the at least one risk element from example risk content included in a preset knowledge base; and using a second large model to perform risk content detection based on the relevant risk content and the text to be detected.

[0007] According to another aspect of the present disclosure, a risk content detection device based on a large model is provided, including: an identification unit, configured to use a first large model to identify at least one risk element in a text to be detected; a retrieval unit, configured to retrieve relevant risk content corresponding to the at least one risk element in the example risk content included in a preset knowledge base; and a first detection unit, configured to use a second large model to perform risk content detection based on the relevant risk content and the text to be detected.

[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method.

[0009] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the above method.

[0010] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the above method when executed by a processor.

[0011] According to one or more embodiments of the present disclosure, the present disclosure utilizes the natural language understanding capability of the first large model to identify risk elements in the text to be detected, and retrieves relevant risk content in a preset knowledge base including example risk content, and then utilizes the logical capability of the second large model in combination with the relevant risk content to detect the text to be detected, thereby enabling accurate identification of implicit risks in the text to be detected and achieving accurate assessment of complex risk content.

[0012] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The accompanying drawings exemplarily illustrate the embodiments and constitute a part of the specification, and together with the text description of the specification, are used to explain the exemplary implementation of the embodiments. The embodiments shown are for illustrative purposes only and do not limit the scope of the claims. In all drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0014] Figure 1 A schematic diagram showing an exemplary system in which the various methods described herein may be implemented according to an embodiment of the present disclosure;

[0015] Figure 2A flowchart of a risk content detection method based on a large model according to an exemplary embodiment of the present disclosure is shown;

[0016] Figure 3 A flowchart showing a process of retrieving relevant risk content in a preset knowledge base according to an exemplary embodiment of the present disclosure;

[0017] Figure 4 A flowchart showing a process of retrieving candidate risk content in a sub-knowledge base according to an exemplary embodiment of the present disclosure;

[0018] Figure 5 A flowchart showing a process of determining relevant risk content based on candidate risk content according to an exemplary embodiment of the present disclosure is shown;

[0019] Figure 6 A schematic diagram showing search text enhancement according to an exemplary embodiment of the present disclosure is shown;

[0020] Figure 7 A flowchart showing a process of using the second largest model to detect risky content according to an exemplary embodiment of the present disclosure is shown;

[0021] Figure 8 A flowchart of a method for risky content detection based on a large model according to an exemplary embodiment of the present disclosure is shown;

[0022] Fig. 9 A schematic diagram showing an expert model according to an exemplary embodiment of the present disclosure is shown;

[0023] Fig.10 A flowchart of a method for risky content detection based on a large model according to an exemplary embodiment of the present disclosure is shown;

[0024] Fig.11 A flowchart of a process of quantitatively scoring a text to be detected based on a plurality of preset potential impact dimensions according to an exemplary embodiment of the present disclosure is shown;

[0025] Fig.12 A schematic diagram showing a quantitative scoring and risk level assessment according to an exemplary embodiment of the present disclosure is shown;

[0026] Fig.13 A structural block diagram of a risky content detection device based on a large model according to an exemplary embodiment of the present disclosure is shown; and

[0027] Fig.14 A structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0028] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, the description of well-known functions and structures is omitted in the following description.

[0029] In the present disclosure, unless otherwise specified, the use of the terms "first", "second", etc. to describe various elements is not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements, and such terms are only used to distinguish one element from another element. In some examples, the first element and the second element may refer to the same instance of the element, and in some cases, based on the description of the context, they may also refer to different instances.

[0030] The terms used in the description of various examples in this disclosure are only for the purpose of describing specific examples and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element can be one or more. In addition, the term "and / or" used in this disclosure covers any one of the listed items and all possible combinations.

[0031] In the related art, the detection effect of the risk content detection method in some embodiments is not good. The harmful content involved in various risk contents is deeply hidden, has many variants, and is rich in jargon and codewords. It is very obscure, so it is difficult to solve it point by point through traditional word list models.

[0032] To solve the above problems, the present invention utilizes the natural language understanding ability of the first large model to identify risk elements in the text to be detected, and retrieves relevant risk content in a preset knowledge base including example risk content, and then utilizes the logical ability of the second large model in combination with the relevant risk content to detect the text to be detected, so as to accurately identify the implicit risks in the text to be detected and achieve accurate assessment of complex risk content.

[0033] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0034] Figure 1 FIG. 1 is a schematic diagram of an exemplary system 100 in which various methods and apparatuses described herein may be implemented according to an embodiment of the present disclosure. Figure 1, the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 may be configured to execute one or more applications.

[0035] In an embodiment of the present disclosure, the server 120 may run one or more services or software applications that enable execution of the methods of the present disclosure.

[0036] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtualized environments and virtualized environments. In some embodiments, these services may be provided as web-based services or cloud services, such as provided to users of client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.

[0037] exist Figure 1 In the configuration shown, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or a combination thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 may in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that a variety of different system configurations are possible, which may differ from the system 100. Therefore, Figure 1 is one example of a system for implementing the various methods described herein and is not intended to be limiting.

[0038] The user can use client devices 101, 102, 103, 104, 105 and / or 106 to perform human-computer interaction. The client device can provide an interface that enables the user of the client device to interact with the client device. The client device can also output information to the user via the interface. Figure 1 Only six client devices are depicted, but one skilled in the art will appreciate that the present disclosure may support any number of client devices.

[0039] Client devices 101, 102, 103, 104, 105 and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, game systems, thin clients, various messaging devices, sensors or other sensing devices, etc. These computer devices may run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT Windows Mobile OS, iOS, Windows Phone, Android. Portable handheld devices may include cellular phones, smart phones, tablet computers, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Game systems may include various handheld game devices, Internet-enabled game devices, etc. Client devices are capable of executing various different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and may use various communication protocols.

[0040] The network 110 may be any type of network known to those skilled in the art that may support data communications using any of a variety of available protocols, including but not limited to TCP / IP, SNA, IPX, etc. By way of example only, the one or more networks 110 may be a local area network (LAN), an Ethernet-based network, a token ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.

[0041] Server 120 may include one or more general purpose computers, dedicated server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running virtual operating systems, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that may be virtualized to maintain a server's virtual storage device). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0042] The computing units in the server 120 may run one or more operating systems including any of the above operating systems and any commercially available server operating systems. The server 120 may also run any of a variety of additional server applications and / or middle-tier applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.

[0043] In some implementations, server 120 may include one or more applications to analyze and consolidate data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and / or 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and / or 106.

[0044] In some embodiments, the server 120 may be a server of a distributed system, or a server combined with a blockchain. The server 120 may also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in a cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and virtual private servers (VPS) services.

[0045] The system 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. The databases 130 may reside in various locations. For example, the database used by the server 120 may be local to the server 120, or may be remote from the server 120 and may communicate with the server 120 via a network-based or dedicated connection. The databases 130 may be of different types. In some embodiments, the databases used by the server 120 may be, for example, relational databases. One or more of these databases may store, update, and retrieve data to and from the databases in response to commands.

[0046] In some embodiments, one or more of the databases 130 may also be used by applications to store application data. The databases used by the applications may be different types of databases, such as a key-value store, an object store, or a conventional store backed by a file system.

[0047] Figure 1The system 100 may be configured and operated in various ways to enable the application of various methods and apparatuses described in the present disclosure.

[0048] According to one aspect of the present disclosure, a risky content detection method based on a large model is provided. Figure 2 FIG. 2 shows a flow chart of a method 200 for detecting risky content based on a large model according to an exemplary embodiment of the present disclosure. Figure 2 As shown, the risk content detection method 200 includes: step S201, using the first large model to identify at least one risk element in the text to be detected; step S202, retrieving relevant risk content corresponding to at least one risk element in the example risk content included in the preset knowledge base; and step S203, using the second large model to perform risk content detection based on the relevant risk content and the text to be detected.

[0049] In the present disclosure, a large model (such as the first large model, the second large model, etc.) may refer to a large language model (LLM), that is, an artificial intelligence system trained on a large data set, which aims to complete tasks related to natural language or specific forms of text, such as understanding, generating, translating, and answering questions.

[0050] Therefore, by utilizing the natural language understanding capability of the first model to identify risk elements in the text to be tested, and retrieving relevant risk content from a preset knowledge base that includes example risk content, and then utilizing the logical capability of the second model in combination with the relevant risk content to detect the text to be tested, it is possible to accurately identify implicit risks in the text to be tested and achieve accurate assessment of complex risk content.

[0051] In step S201, at least one risk element is identified in the text to be detected using the first large model.

[0052] According to some embodiments, the text to be detected (also referred to as original text in the present disclosure) may include user-generated content. User-generated content may include, for example, user posts, articles, microblogs, and other content. The original user-generated content can be divided into text content, image content, and video content. The text content can be further divided into text, URLs, and emoji expressions. The pictures are parsed according to their different formats (JPG, GIF, PNG, etc.) to extract the text in the pictures. The video content needs to be extracted separately by voice-to-text conversion and frame-to-picture conversion to extract the text, and finally the text to be detected is obtained.

[0053] User-generated content (UGC) is an important part of online communities, and its unique language, symbols and expressions often hide risk information that is not easily detected. By deeply understanding the characteristics of social information content in Internet communities and combining the natural language understanding, natural language processing, logical reasoning and other capabilities of large language model technology, this disclosure provides an efficient risk content detection method that can detect implicit information in texts more quickly and accurately, more accurately judge the compliance of content, and effectively reduce risk content in online communities.

[0054] The risk element may be, for example, an element in the text to be detected that may have risk problems. In some embodiments, the risk element may include, for example, a character name, an event summary, a tone of voice, etc. in the text to be detected. In some embodiments, a corresponding prompt text may be designed to indicate that the first large model identifies at least one risk element in the text to be detected.

[0055] The preset knowledge base may include a large number of example risk contents, which may be risk contents that have been manually confirmed or determined in other ways.

[0056] In step S202, risk content corresponding to at least one risk element is retrieved from the example risk content included in the preset knowledge base.

[0057] In some embodiments, each risk element identified in the text to be detected can be detected in a preset knowledge base to obtain example risk content corresponding to at least one risk element. These example risk contents can all be used as relevant risk content, or they can be simplified or integrated in various ways to obtain relevant risk content.

[0058] In some embodiments, the search may be performed based on the identified risk elements, or based on the text to be detected itself, which is not limited here.

[0059] In step S203, the second largest model is used to perform risk content detection based on the relevant risk content and the text to be detected.

[0060] The second large model can be the same as the first large model, or another different large model can be used. The relevant risk content can be added to the original text to form a review text, and the review text can be sent to the second large model for risk content detection. In some embodiments, corresponding prompt text can be designed to instruct the second large model to perform risk content detection based on the relevant risk content and the text to be detected. In this way, the second large model can be combined with the example risk content related to the potential risk content (risk element) in the current text to be detected to perform risk content detection, thereby improving the detection effect of implicit risk content.

[0061] The above original text and review text are analyzed and organized using the semantic understanding and decision-making capabilities of the large language model to obtain the corresponding computer memory model. The memory model is used to perform in-depth algorithm fitting to obtain dynamic element nodes. Through this node information data, the large language model is used to identify abnormal text expressions in the text in terms of subject direction and contextual atmosphere.

[0062] Through the unique identification of the sub-knowledge base, the corresponding instructions are extracted from the public instruction pool for targeted combination (instructions, formats, restrictions, examples), and specific prompt instructions are obtained. Then, by combining the prompt instructions and feature parameters, the semantic understanding and decision-making capabilities of the inherent vertical categories of the large language model are used to complete the output and storage of the memory model.

[0063] Similarly, the prompt word instruction combination can be completed for each sub-knowledge base, and finally the memory model corresponding to each element type can be stored in the computer memory network. The text feature labels under different expression scenarios can be obtained as the combination strategy conditions for subsequent risk association logic judgment.

[0064] In some embodiments, a variety of different features such as individual comment type text features and group comment text features may be obtained, and corresponding features may be selected according to the scenario for risky content detection.

[0065] According to some embodiments, multiple element types may be preset, and risk elements may be identified in the text to be detected based on these element types. In other words, at least one risk element may be identified based on multiple preset element types, and each of the at least one risk element may belong to one of the multiple element types. In this way, structured risk element identification of the text to be detected can be achieved, and it is ensured that the identified risk elements all belong to a preset classification (element type).

[0066] In addition, the preset knowledge base may include multiple sub-knowledge bases corresponding to multiple element types. Each sub-knowledge base may store a large number of example risk content related to the corresponding element type. For example, for the element type "character name", the corresponding sub-knowledge base may store example risk content about multiple different characters, and these example risk contents are similar to "network stalks". When a public figure (risk element) is identified in the text to be detected, it can be searched in the sub-knowledge base corresponding to the "character name" to obtain the "network stalks" related to the public figure, that is, the relevant risk content. Then, the second largest model can be used to perform risk content detection based on the retrieved "network stalks" and the text to be detected. This method can more effectively identify the obscure content in the text to be detected, so as to accurately determine whether there is a situation of "playing stalks".

[0067] Figure 3 FIG. 3 is a flowchart of a process 300 for retrieving relevant risk content from a preset knowledge base according to an exemplary embodiment of the present disclosure. The process 300 may be used to implement step S202 in the above method 200. Figure 3 As shown, process 300 may include: step S301, for each risk element in at least one risk element, searching in the sub-knowledge base corresponding to the element type to which the risk element belongs to obtain at least one candidate risk content corresponding to the risk element; and step S302, determining relevant risk content based on at least one candidate risk content corresponding to each of the at least one risk elements.

[0068] In step S301, a search may be performed in the corresponding sub-knowledge base based on the identified risk element (eg, the name of the identified public figure), or a search may be performed in the sub-knowledge base based on the text to be detected, which is not limited here.

[0069] In some embodiments, a semantic search method may be used to retrieve candidate risk content in each sub-knowledge base. Figure 4 FIG. 4 is a flowchart of a process 400 for retrieving candidate risk content in a sub-knowledge base according to an exemplary embodiment of the present disclosure. The process 400 can be used to implement step S301 in the above process 300. Figure 4 As shown, process 400 may include: step S401, encoding the text to be detected using an embedding model to obtain a first vector; step S402, obtaining, for each risk element of at least one risk element, a plurality of second vectors corresponding to a plurality of example risk contents, the plurality of example risk contents coming from a sub-knowledge base corresponding to an element type to which the risk element belongs, and the plurality of second vectors being obtained by encoding the plurality of example risk contents using an embedding model; and step S403, semantically matching the first vector and the plurality of second vectors to determine at least one candidate risk content corresponding to the risk element.

[0070] Therefore, through the above-mentioned embedded coding and semantic matching methods, candidate risk content related to the text to be detected can be quickly obtained.

[0071] In some embodiments, the original text (text to be detected) can first be encoded through an embedding model to form a vector, and each text data (example risk content) in the sub-knowledge base corresponding to each risk element can be encoded using the same embedding model, and then the vector of the original text can be used to perform semantic matching in the sub-knowledge base corresponding to each risk element.

[0072] In an exemplary embodiment, in step S403, the first vector and the plurality of second vectors may be semantically matched based on cosine similarity. If the vectors after text conversion are displayed in a multidimensional space, they exist in the form of coordinate points, and the greater the cosine similarity, the higher the degree of semantic matching between the texts. For each risk element, one or more example risk contents with the highest degree of semantic matching may be selected as candidate risk contents.

[0073] The multiple sub-knowledge bases corresponding to the multiple element types included in the preset knowledge base are established in the following manner: obtaining sample risk data; identifying at least one sample risk element in the sample risk data; and placing the sample risk data into the sub-knowledge base corresponding to the element type to which the at least one sample risk element belongs.

[0074] In some embodiments, the sample risk data can first be grouped by title, the data content corresponding to the title can be obtained, and finally the data can be stored in a preset perception vector wide table. Then, the "batch offline reasoning" capability of the large model can be used to perform unified reasoning and prediction on the data set samples, and then synchronize the prediction results. The prediction results are parsed, and the parsed vector results are stored and updated one by one according to the preset table structure. Finally, the logical ability and understanding ability of the large model are used to identify risk elements in the data.

[0075] It is understandable that the second vector encoded with the example risk content may be stored in the preset knowledge base or the sub-knowledge base, so that the retrieval can be completed quickly.

[0076] In step S302, at least one candidate risk content corresponding to each of the at least one risk elements may be combined, processed, screened, or any combination of the above methods to obtain an example risk text.

[0077] Figure 5 FIG. 5 is a flowchart of a process 500 for determining relevant risk content based on candidate risk content according to an exemplary embodiment of the present disclosure. The process 500 may be used to implement step S302 in the above process 300. Figure 5 As shown, process 500 may include: step S501, determining the text similarity between the text to be detected and multiple candidate risk contents, the multiple candidate risk contents include at least one candidate risk content corresponding to each of at least one risk element; step S502, sorting the multiple candidate risk contents based on the text similarity; and step S503, determining the relevant risk content based on one or more candidate risk contents with the highest text similarity.

[0078] Thus, by sorting the candidate risk contents based on the text similarity with the content to be detected, and determining the relevant risk contents based on one or more candidate risk contents with the highest text similarity, it is possible to further screen out the sample risk contents that are most relevant to the content to be detected, so as to improve the detection efficiency and accuracy of the second largest model. In addition, since there are a large number of sample risk contents in the sub-knowledge base, the semantic matching method can realize the rapid retrieval of candidate risk contents in the sub-knowledge base, and by sorting the small number of candidate risk contents corresponding to at least one risk element by text similarity, it is possible to obtain sample risk contents that are more relevant and accurate to the text to be detected, thereby improving the accuracy of downstream risk content detection.

[0079] In some embodiments, for each risk element, only one example risk content with the highest semantic matching degree may be selected, and the number of candidate risk contents finally obtained is the same as the number of identified risk elements.

[0080] Figure 6 A schematic diagram of retrieval text enhancement according to an exemplary embodiment of the present disclosure is shown. In this exemplary embodiment, four risk elements are identified from the text to be detected, and these four risk elements can be sent to the corresponding sub-knowledge bases in the preset knowledge base for retrieval. Specifically, for each risk element, the sub-knowledge base corresponding to the element type to which the risk element belongs can be determined first, and then the text to be detected and the example risk content in the sub-knowledge base are semantically matched, and the example risk content with the highest degree of match is taken from the candidate risk content corresponding to the risk element. Furthermore, the candidate risk content corresponding to each risk element can be summarized, and the text similarity between the text to be detected and these candidate risk contents can be calculated. Finally, the candidate risk content most relevant to the text to be detected, that is, the relevant risk content, can be obtained by sorting based on the text similarity.

[0081] Figure 7 FIG. 7 is a flowchart of a process 700 for detecting risky content using the second largest model according to an exemplary embodiment of the present disclosure. The process 700 may be used to implement step S203 in the above method 200. Figure 7 As shown, process 700 may include: step S701, obtaining a first prompt text, the first prompt text instructing the large model to determine whether the text to be detected has multiple preset risk features, and generate reasoning results and reasoning reasons; and step S702, inputting the first prompt text, relevant risk content and the text to be detected into the second large model to obtain a first detection result.

[0082] Therefore, through the above method, it is possible to accurately detect the risk content in the text to be detected, and obtain credible and explainable reasoning results and reasoning reasons.

[0083] In an exemplary embodiment, the plurality of risk characteristics may include, for example:

[0084] 1) Topic deviation: Comments deviate from the original topic during user interaction, intentionally guide the topic, and cause dissatisfaction or suggestive discussions among other users.

[0085] 2) Differences between comment circles: Under the same topic, the discussion content of this comment circle is obviously different from that of other comment circles in terms of topic or tone, which interrupts the coherence of other comments in the comment area.

[0086] 3) Forced association and comparison: Through hints or innuendos, guide other users to make their own guesses.

[0087] In an exemplary embodiment, the plurality of risk features may include, for example, different types of risks, such as advertising, personal attack, fraud, etc. It is understandable that the above risk features are only exemplary. In actual implementation, more or other risk features may be used, which are not limited here.

[0088] Figure 8 FIG. 8 is a flowchart of a method 800 for risk content detection based on a large model according to an exemplary embodiment of the present disclosure. Figure 8 As shown, method 800 includes: step S804, in response to determining that the first detection result indicates that the text to be detected has a target risk feature among multiple risk features, determining an expert model corresponding to the target risk feature in a preset expert model library; and step S805, using the expert model to determine whether there is a logical association between the text to be detected and the target risk feature based on the relevant risk content, the text to be detected and the first detection result, so as to obtain a second detection result. It can be understood that before executing method 800, it is necessary to pre-set the expert model corresponding to each risk feature and construct a preset expert model library. Each expert model can be used to determine whether there is a logical association between the text to be detected and the risk feature corresponding to the expert model.

[0089] It can be understood that the operations and effects of steps S801 to S803 in method 800 can refer to the above description of steps S201 to S203 in method 200, and are not described in detail here.

[0090] Therefore, when the first detection result generated by the second largest model indicates that the text to be detected has the target analysis feature, the corresponding expert model is used to judge the logical relevance, thereby further improving the reliability and accuracy of risk content detection.

[0091] According to some embodiments, the preset expert model library may include multiple expert models corresponding to multiple risk features, and each of the multiple expert models may be trained using multiple training samples for the expert model. The multiple training samples for the expert model may all have risk features corresponding to the expert model.

[0092] Therefore, through the above method, each trained expert model can have the ability to judge the logical association with the corresponding risk feature.

[0093] According to some embodiments, multiple expert models can be trained using the following operations: step SA01, obtaining a risk content map, wherein the risk content map is generated based on historical business data, the risk content map includes multiple risk features, and each risk feature includes multiple association logics; step SA02, obtaining a target training sample for a target expert model among multiple expert models, and determining the real association logic of the target training sample based on the risk content map; step SA03, using the target expert model to explicitly output an inference step on whether the target training sample has a risk feature corresponding to the target expert model; and step SA04, training the target expert model based on the inference step and the real association logic.

[0094] Therefore, by obtaining the association logic related to risk characteristics from the risk content map generated based on historical business data, and allowing the expert model to explicitly output the reasoning steps, and then using the two to train the expert model, the trained expert model's ability to push the risk association logic is enhanced, and the model's interpretability and controllability are improved.

[0095] In some embodiments, each risk feature in the risk content map may include multiple risk points, and each risk point includes multiple association logics.

[0096] In some embodiments, the reasoning step can be a single step, that is, the model can be allowed to gradually participate in the process of decomposing a complex problem into sub-problems step by step and solving them in sequence.

[0097] Fig. 9A schematic diagram of an expert model according to an exemplary embodiment of the present disclosure is shown. The training method described in the above content can also be referred to as logic chain fine-tuning training. First, relevant risk content can be retrieved based on the original text (text to be detected), and the target risk features of the original text can be detected using a large model to complete the integrated content 910. Then, the corresponding MOE expert model 920 can be used to perform an associated logic judgment 930. An exemplary MOE expert model may include an MOE expert model corresponding to the risk features "advertising", "personal attack" and "fraud". The MOE expert model 920 can be obtained by performing logic chain fine-tuning training using a risk content associated logic chain 940. Each subject (i.e., risk feature) can include multiple risk points, each risk point can include multiple associated logics, and each associated logic can have several specific examples (i.e., training samples).

[0098] In an exemplary embodiment, the risk content map includes the risk feature "advertising", which has three risk points: false propaganda, illegal promotion, and the influence of bad content. The risk point "false propaganda" further includes three association logics: false exaggeration of product effects, attracting users to buy through false discounts, and using untrue user reviews. Association logic 1: One specific example of false exaggeration of product effects (i.e., training sample) can be "Teeth whitening, effective in 1 day!". By using this specific example as a training sample and the corresponding association logic 1 as the real association logic to train the expert model corresponding to the risk feature "advertising", the expert model can learn the knowledge in the risk content map and the trained expert model can have logical reasoning ability.

[0099] In some embodiments, in step S805, a prompt text may be used to instruct the expert model to determine whether there is a logical association between the text to be detected and the target risk feature. The specific content of the prompt text may be designed according to requirements.

[0100] Fig.10 FIG. 1 is a flowchart of a method 1000 for risk content detection based on a large model according to an exemplary embodiment of the present disclosure. Fig.10 As shown, method 1000 includes: step S1006, in response to determining that the second detection result indicates that there is a logical association between the text to be detected and the target risk feature, quantitatively scoring the text to be detected based on a preset plurality of potential impact dimensions to obtain a plurality of risk scores; and step S1007, determining a third detection result based on the risk thresholds and the plurality of risk scores corresponding to each of the plurality of potential impact dimensions.

[0101] It can be understood that the operations and effects of steps S1001 to S1005 in method 1000 can refer to the above description of steps S801 to S805 in method 800, and are not described in detail here.

[0102] Therefore, through the above method, the potential harm that may be caused by the content to be detected can be further evaluated, thereby obtaining a more comprehensive and accurate risk content detection result.

[0103] In some embodiments, a large model can be used to implement quantitative scoring of the text to be detected from multiple potential impact dimensions, or other natural language processing methods can be used for quantitative scoring, which is not limited here. After obtaining the risk scores corresponding to each potential impact dimension, each risk score can be compared with the corresponding risk threshold. In response to determining that one of the risk scores exceeds the corresponding threshold, the third detection result indicates that the text to be detected is at risk.

[0104] In some embodiments, the third detection result may also be a risk level determination result. Multiple risk levels may be preset, and the risk level of the current text to be detected may be determined based on multiple risk scores and corresponding risk thresholds, so that the downstream may perform different levels of processing.

[0105] According to some embodiments, each of the multiple potential impact dimensions may include multiple evaluation fields, and each of the multiple evaluation fields may have a corresponding scoring rule and a weight value.

[0106] Fig.11 FIG. 1 is a flowchart of a process 1100 for quantitatively scoring a text to be detected based on multiple preset potential impact dimensions according to an exemplary embodiment of the present disclosure. The process 1100 can be used to implement step S1006 in the above method 1000. Fig.11 As shown, process 1100 may include: step S1101, for each potential impact dimension, determining multiple field scores of the text to be detected for multiple evaluation fields based on scoring rules corresponding to each of the multiple evaluation fields corresponding to the potential impact dimension; and step S1102, determining the risk score of the text to be detected for the potential impact dimension based on the multiple field scores of the text to be detected for the multiple evaluation fields and the weight values ​​corresponding to each of the multiple evaluation fields.

[0107] Therefore, through the above method, a more refined quantification process of risk scores is achieved, thereby further improving the granularity and accuracy of subsequent risk content detection.

[0108] In some embodiments, for the text to be detected that is determined to have a logical association, a weighted calculation evaluation can be performed on it from multiple dimensions, and finally it is determined whether it is a harmful risk. First, in each dimension, the risk text is quantitatively scored according to the rules for each field. At the same time, each field in each dimension has a corresponding weight value. The value under the field is weighted and summed with the corresponding weight, and finally the score under the dimension is obtained. The specific calculation formula under each dimension is: $C_i=A_i*W_1+B_i*W_2+...+N_i*W_j$, where $C_i$ represents the weighted score under the $i$th dimension, $A_i$ represents the field $A$ under the $i$ dimension, $W_1$ represents the weight value of the field $A$ under the $i$ dimension, $B_i$ represents the field $B$ under the $i$ dimension, $W_2$ represents the weight value of the field $B$ under the $i$ dimension, and so on until $N_i$ and $W_j$. A limit value is set for the weight value under each dimension. The weighted average is calculated based on the size of the score of each dimension relative to the limit. A threshold is set for the value, and the result is generated based on the threshold: if it is greater than the threshold, it is judged as risky content; if it is less than the threshold, it is judged as not risky content.

[0109] Fig.12 A schematic diagram of quantitative scoring and risk level assessment according to an exemplary embodiment of the present disclosure is shown. Specifically, quantitative scoring can be performed under different fields included in different dimensions to obtain corresponding field scores, and then weighted calculation can be performed based on the weights corresponding to the fields to obtain scores corresponding to each dimension, and then the scores of each dimension can be compared with the threshold to generate risk level judgment results.

[0110] In an exemplary embodiment, the multiple potential impact dimensions may include:

[0111] 1) Threat dimension, including three evaluation fields: degree of harm, degree of concern, and scope of impact;

[0112] 2) Value dimension: includes three evaluation fields: platform publishing volume, platform interaction volume, and netizen popularity.

[0113] It is understandable that the above two potential impact dimensions and corresponding fields are only exemplary. In actual implementation, more or other potential impact dimensions and / or evaluation fields may be used, which are not limited here.

[0114] In some embodiments, when the third detection result indicates that there is a risk in the text to be detected, a series of processes can be triggered to generate and optimize the model.

[0115] First, a prompt text of a model training sample can be generated by inputting a prompt word and text content, and a response text of a model training sample can be generated by outputting danger prediction (reasoning result) and enhancing generated content (reasoning reason). In this way, the combination of the prompt text and the response text forms a model training sample data, and the model training sample data is stored in the training sample database.

[0116] 2) Secondly, based on the online model effect as a condition, the batch fine-tuning model link is automatically triggered.

[0117] a. After the model is fine-tuned, the model is automatically released;

[0118] b. Verify the effect of the model through the preset model effect sample library;

[0119] c. Determine whether the model is effective based on the model verification results.

[0120] 3) Finally, the published model interface information and the corresponding model verification results are stored in the fine-tuning model library. Depending on whether the model is effective, it is determined whether to replace the online model.

[0121] In summary, the present disclosure provides a risk content detection system that does not rely solely on a single model, but rather collaborates with multiple model algorithms to make task-based decisions. In some embodiments, by parsing the risk element vector in the content, combined with retrieval capability enhancement to generate the inspected content, using a large model to prompt the text to extract text features from the content, using the large model to fine-tune the risk logic chain model, using a multi-dimensional weighted algorithm to evaluate the risk threshold, automatically summarizing risk cases for risk content and returning the model for learning iteration.

[0122] According to another aspect of the present disclosure, a risky content detection device based on a large model is provided. Fig.13 FIG. 1 shows a structural block diagram of a risk content detection device 1300 based on a large model according to an exemplary embodiment of the present disclosure. Fig.13 As shown, the risk content detection device 1300 includes: an identification unit 1310, a retrieval unit 1320, and a first detection unit 1330. The identification unit is configured to identify at least one risk element in the text to be detected using the first large model; the retrieval unit is configured to retrieve relevant risk content corresponding to the at least one risk element in the example risk content included in the preset knowledge base; and the first detection unit is configured to perform risk content detection based on the relevant risk content and the text to be detected using the second large model.

[0123] The operations of the identification unit 1310, the retrieval unit 1320, and the first detection unit 1330 may correspond to the following: Figure 2The operations of step S201, step S202, and step S203 are shown. Therefore, the details of each aspect are not repeated here.

[0124] In some embodiments, at least one risk element is obtained based on a preset plurality of element types, each of the at least one risk element belongs to one of the plurality of element types, and the preset knowledge base may include a plurality of sub-knowledge bases corresponding to the plurality of element types. The retrieval unit may include a type retrieval unit and an example determination unit. The type retrieval unit may be configured to search, for each of the at least one risk element, in the sub-knowledge base corresponding to the element type to which the risk element belongs, to obtain at least one candidate risk content corresponding to the risk element; and the example determination unit may be configured to determine an example risk text based on at least one candidate risk content corresponding to each of the at least one risk element.

[0125] In some embodiments, the retrieval unit may include a first vector encoding unit, a second vector acquisition unit, and a vector matching unit. The first vector encoding unit may be configured to encode the text to be detected using an embedding model to obtain a first vector; the second vector acquisition unit may be configured to obtain, for each risk element of the at least one risk element, a plurality of second vectors corresponding to a plurality of example risk contents, the plurality of example risk contents are from a sub-knowledge base corresponding to the element type to which the risk element belongs, and the plurality of second vectors are obtained by encoding the plurality of example risk contents using an embedding model; and the vector matching unit may be configured to semantically match the first vector with the plurality of second vectors to determine at least one candidate risk content corresponding to the risk element.

[0126] In some embodiments, the example determination unit may include a similarity determination unit, a sorting unit, and a related risk content determination unit. The similarity determination unit may be configured to determine the text similarity between the to-be-detected text and a plurality of candidate risk contents, wherein the plurality of candidate risk contents may include at least one candidate risk content corresponding to each of at least one risk element; the sorting unit may be configured to sort the plurality of candidate risk contents based on the text similarity; and the related risk content determination unit may be configured to determine the one or more candidate risk contents with the highest text similarity as related risk contents.

[0127] In some embodiments, the first detection unit may include a first prompt unit and a detection result determination unit. The first prompt unit may be configured to obtain a first prompt text, the first prompt text instructing the large model to determine whether the text to be detected has multiple preset risk features, and generate an inference result and reasoning reasons; and the detection result determination unit may be configured to input the first prompt text, the relevant risk content and the text to be detected into the second large model to obtain a first detection result.

[0128] In some embodiments, the apparatus 1300 may further include an expert acquisition unit and a second detection unit. The expert acquisition unit may be configured to determine an expert model corresponding to the target risk feature in a preset expert model library in response to determining that the first detection result indicates that the text to be detected has a target risk feature among multiple risk features; and the second detection unit may be configured to use the expert model to determine whether there is a logical association between the text to be detected and the target risk feature based on the relevant risk content, the text to be detected and the first detection result, so as to obtain a second detection result.

[0129] In some embodiments, the preset expert model library may include multiple expert models corresponding to multiple risk features, and each of the multiple expert models may be trained using multiple training samples for the expert model. The multiple training samples for the expert model may all have risk features corresponding to the expert model.

[0130] In some embodiments, multiple expert models can be trained using the following operations: obtaining a risk content map, wherein the risk content map is generated based on historical business data, the risk content map includes multiple risk features, and each risk feature includes multiple association logics; obtaining a target training sample for a target expert model among multiple expert models, and determining the real association logic of the target training sample based on the risk content map; using the target expert model to explicitly output the reasoning steps of whether the target training sample has the risk feature corresponding to the target expert model; and training the target expert model based on the reasoning steps and the real association logic.

[0131] In some embodiments, the device 1300 may further include a quantitative scoring unit and a third detection unit. The quantitative scoring unit may be configured to, in response to determining that the second detection result indicates that the text to be detected is logically associated with the target risk feature, quantitatively score the text to be detected based on a plurality of preset potential impact dimensions to obtain a plurality of risk scores; and the third detection unit may be configured to determine a third detection result based on the risk thresholds and the plurality of risk scores corresponding to each of the plurality of potential impact dimensions, the third detection result indicating the risk level of the text to be detected.

[0132] In some embodiments, each of the multiple potential impact dimensions may include multiple evaluation fields, each of the multiple evaluation fields has a corresponding scoring rule and a weight value, and the quantitative scoring unit may include a field scoring unit and a risk scoring unit. The field scoring unit may be configured to determine, for each potential impact dimension, multiple field scores of the text to be detected for the multiple evaluation fields based on the scoring rules corresponding to each of the multiple evaluation fields corresponding to the potential impact dimension; and the risk scoring unit may be configured to determine the risk score of the text to be detected for the potential impact dimension based on the multiple field scores of the text to be detected for the multiple evaluation fields and the weight values ​​corresponding to each of the multiple evaluation fields.

[0133] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved are all subject to the provisions of relevant laws and regulations and do not violate public order and good morals.

[0134] According to an embodiment of the present disclosure, an electronic device, a readable storage medium and a computer program product are also provided.

[0135] refer to Fig.14 , a block diagram of an electronic device 1400 that can be used as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0136] like Fig.14 As shown, the electronic device 1400 includes a computing unit 1401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1402 or a computer program loaded from a storage unit 1408 into a random access memory (RAM) 1403. In the RAM 1403, various programs and data required for the operation of the electronic device 1400 can also be stored. The computing unit 1401, the ROM 1402, and the RAM 1403 are connected to each other via a bus 1404. An input / output (I / O) interface 1405 is also connected to the bus 1404.

[0137] Multiple components in the electronic device 1400 are connected to the I / O interface 1405, including: an input unit 1406, an output unit 1407, a storage unit 1408, and a communication unit 1409. The input unit 1406 can be any type of device that can input information to the electronic device 1400. The input unit 1406 can receive input digital or character information, and generate key signal input related to user settings and / or function control of the electronic device, and can include but is not limited to a mouse, a keyboard, a touch screen, a trackpad, a trackball, a joystick, a microphone, and / or a remote controller. The output unit 1407 can be any type of device that can present information, and can include but is not limited to a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 1408 can include but is not limited to a disk and an optical disk. The communication unit 1409 allows the electronic device 1400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and may include but is not limited to a modem, a network card, an infrared communication device, a wireless communication transceiver and / or a chipset, such as a Bluetooth device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device and / or the like.

[0138] The computing unit 1401 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1401 performs the various methods, processes, and / or processes described above. For example, in some embodiments, these methods, processes, and / or processes may be implemented as computer software programs, which are tangibly included in machine-readable media, such as storage unit 1408. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 1400 via ROM 1402 and / or communication unit 1409. When the computer program is loaded into RAM 1403 and executed by the computing unit 1401, one or more steps of the methods, processes, and / or processes described above may be performed. Alternatively, in other embodiments, the computing unit 1401 may be configured to perform these methods, processes and / or processing in any other suitable manner (eg, by means of firmware).

[0139] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0140] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0141] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0142] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0143] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0144] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services ("Virtual Private Server", or "VPS" for short). The server may also be a server for a distributed system, or a server combined with a blockchain.

[0145] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0146] Although the embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above-mentioned methods, systems and devices are merely exemplary embodiments or examples, and the scope of the present invention is not limited by these embodiments or examples, but only by the claims after authorization and their equivalent scope. Various elements in the embodiments or examples can be omitted or replaced by their equivalent elements. In addition, each step can be performed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. It is important that with the evolution of technology, many elements described herein can be replaced by equivalent elements that appear after the present disclosure.

Claims

1. A risk content detection method based on a large model, comprising: Using the first model to identify at least one risk element in the text to be detected; Retrieving relevant risk content corresponding to the at least one risk element from the example risk content included in the preset knowledge base; as well as The second largest model is used to perform risk content detection based on the relevant risk content and the text to be detected.

2. The method according to claim 1, wherein: The at least one risk element is obtained based on the identification of the preset multiple element types, each of the at least one risk element belongs to one of the multiple element types, and the preset knowledge base includes multiple sub-knowledge bases corresponding to the multiple element types. Wherein, searching the preset knowledge base for relevant risk content corresponding to the at least one risk element includes: For each risk element of the at least one risk element, searching in a sub-knowledge base corresponding to the element type to which the risk element belongs, to obtain at least one candidate risk content corresponding to the risk element; and The relevant risk content is determined based on at least one candidate risk content corresponding to each of the at least one risk element.

3. The method according to claim 2, wherein: For each risk element of the at least one risk element, searching in a sub-knowledge base corresponding to the element type to which the risk element belongs to obtain at least one candidate risk content corresponding to the risk element includes: Encoding the text to be detected by using an embedding model to obtain a first vector; For each risk element of the at least one risk element, obtaining a plurality of second vectors corresponding to a plurality of example risk contents, the plurality of example risk contents being from a sub-knowledge base corresponding to an element type to which the risk element belongs, the plurality of second vectors being obtained by encoding the plurality of example risk contents using the embedding model; and The first vector and the plurality of second vectors are semantically matched to determine at least one candidate risk content corresponding to the risk element.

4. The method according to claim 3, wherein: Determining the relevant risk content based on at least one candidate risk content corresponding to each of the at least one risk elements includes: Determining text similarity between the to-be-detected text and a plurality of candidate risk contents, wherein the plurality of candidate risk contents include at least one candidate risk content corresponding to each of the at least one risk element; sorting the plurality of candidate risk contents based on the text similarity; and The relevant risk content is determined based on one or more candidate risk content with the highest text similarity.

5. The method according to any one of claims 1 to 4, wherein: Using the second largest model to perform risk content detection based on the relevant risk content and the text to be detected includes: Obtaining a first prompt text, wherein the first prompt text instructs the large model to determine whether the to-be-detected text has a plurality of preset risk features, and to generate an inference result and an inference reason; and The first prompt text, the relevant risk content and the text to be detected are input into the second large model to obtain a first detection result.

6. The method according to claim 5, further comprising: In response to determining that the first detection result indicates that the to-be-detected text has a target risk feature among the multiple risk features, determining an expert model corresponding to the target risk feature in a preset expert model library; as well as The expert model is used to determine whether there is a logical association between the text to be detected and the target risk feature based on the relevant risk content, the text to be detected and the first detection result, so as to obtain a second detection result.

7. The method according to claim 6, wherein: The preset expert model library includes multiple expert models corresponding to the multiple risk features, each of the multiple expert models is trained using multiple training samples for the expert model, wherein the multiple training samples for the expert model all have risk features corresponding to the expert model.

8. The method according to claim 7, wherein: The multiple expert models are trained using the following operations: Obtaining a risk content map, wherein the risk content map is generated based on historical business data, the risk content map includes the multiple risk features, and each risk feature includes multiple association logics; Acquire a target training sample for a target expert model among the multiple expert models, and determine a true association logic of the target training sample according to the risk content map; A step of inferring whether the target training sample has a risk feature corresponding to the target expert model by using the target expert model to explicitly output the risk feature; and Based on the reasoning steps and the true association logic, the target expert model is trained.

9. The method according to claim 6, further comprising: In response to determining that the second detection result indicates that the text to be detected is logically associated with the target risk feature, quantitatively scoring the text to be detected based on a plurality of preset potential impact dimensions to obtain a plurality of risk scores; as well as Based on the risk thresholds corresponding to each of the multiple potential impact dimensions and the multiple risk scores, a third detection result is determined, where the third detection result indicates the risk level of the text to be detected.

10. The method according to claim 9, wherein: Each of the multiple potential impact dimensions includes multiple evaluation fields, each of the multiple evaluation fields has a corresponding scoring rule and a weight value, In response to determining that the second detection result indicates that the text to be detected is logically associated with the target risk feature, the text to be detected is quantitatively scored based on a plurality of preset potential impact dimensions to obtain a plurality of risk scores, including: For each of the potential impact dimensions, based on scoring rules corresponding to the multiple evaluation fields corresponding to the potential impact dimension, determine multiple field scores of the to-be-detected text for the multiple evaluation fields; and Based on multiple field scores of the text to be detected for the multiple evaluation fields and weight values ​​corresponding to each of the multiple evaluation fields, a risk score of the text to be detected for the potential impact dimension is determined.

11. A risk content detection device based on a large model, comprising: An identification unit, configured to identify at least one risk element in the text to be detected using the first large model; A retrieval unit, configured to retrieve relevant risk content corresponding to the at least one risk element from the example risk content included in the preset knowledge base; as well as The first detection unit is configured to perform risk content detection based on the relevant risk content and the text to be detected by using the second large model.

12. The device according to claim 11, wherein The at least one risk element is obtained based on the identification of the preset multiple element types, each of the at least one risk element belongs to one of the multiple element types, and the preset knowledge base includes multiple sub-knowledge bases corresponding to the multiple element types. Wherein, the retrieval unit comprises: a type retrieval unit configured to search, for each risk element of the at least one risk element, in a sub-knowledge base corresponding to the element type to which the risk element belongs, to obtain at least one candidate risk content corresponding to the risk element; and The example determining unit is configured to determine the relevant risk content based on at least one candidate risk content corresponding to each of the at least one risk element.

13. The device according to claim 12, wherein: The type retrieval unit comprises: A first vector encoding unit is configured to encode the text to be detected by using an embedding model to obtain a first vector; a second vector acquisition unit configured to acquire, for each of the at least one risk element, a plurality of second vectors corresponding to a plurality of example risk contents, the plurality of example risk contents being from a sub-knowledge base corresponding to an element type to which the risk element belongs, the plurality of second vectors being obtained by encoding the plurality of example risk contents using the embedding model; and The vector matching unit is configured to perform semantic matching on the first vector and the plurality of second vectors to determine at least one candidate risk content corresponding to the risk element.

14. The device according to claim 13, wherein: The example determining unit comprises: a similarity determination unit configured to determine text similarities between the to-be-detected text and a plurality of candidate risk contents, wherein the plurality of candidate risk contents include at least one candidate risk content corresponding to each of the at least one risk element; a sorting unit configured to sort the plurality of candidate risk contents based on the text similarity; and The relevant risk content determination unit is configured to determine the relevant risk content based on one or more candidate risk content with the highest text similarity.

15. The device according to any one of claims 11 to 14, wherein: The first detection unit comprises: a first prompt unit, configured to obtain a first prompt text, wherein the first prompt text instructs the large model to determine whether the to-be-detected text has a plurality of preset risk features, and to generate an inference result and an inference reason; and The detection result determination unit is configured to input the first prompt text, the relevant risk content and the text to be detected into the second large model to obtain a first detection result.

16. The apparatus according to claim 15, further comprising: an expert acquisition unit, configured to, in response to determining that the first detection result indicates that the to-be-detected text has a target risk feature among the multiple risk features, determine an expert model corresponding to the target risk feature in a preset expert model library; as well as The second detection unit is configured to use the expert model to determine whether there is a logical association between the text to be detected and the target risk feature based on the relevant risk content, the text to be detected and the first detection result to obtain a second detection result.

17. The device according to claim 16, wherein: The preset expert model library includes multiple expert models corresponding to the multiple risk features, each of the multiple expert models is trained using multiple training samples for the expert model, wherein the multiple training samples for the expert model all have risk features corresponding to the expert model.

18. The device according to claim 17, wherein: The multiple expert models are trained using the following operations: Obtaining a risk content map, wherein the risk content map is generated based on historical business data, the risk content map includes the multiple risk features, and each risk feature includes multiple association logics; Acquire a target training sample for a target expert model among the multiple expert models, and determine a true association logic of the target training sample according to the risk content map; A step of inferring whether the target training sample has a risk feature corresponding to the target expert model by using the target expert model to explicitly output the risk feature; and Based on the reasoning steps and the true association logic, the target expert model is trained.

19. The apparatus according to claim 16, further comprising: a quantitative scoring unit, configured to, in response to determining that the second detection result indicates that the text to be detected is logically associated with the target risk feature, quantitatively score the text to be detected based on a plurality of preset potential impact dimensions to obtain a plurality of risk scores; as well as The third detection unit is configured to determine a third detection result based on the risk thresholds corresponding to each of the multiple potential impact dimensions and the multiple risk scores, wherein the third detection result indicates the risk level of the text to be detected.

20. The device according to claim 19, wherein Each of the multiple potential impact dimensions includes multiple evaluation fields, each of the multiple evaluation fields has a corresponding scoring rule and a weight value, Wherein, the quantitative scoring unit includes: A field scoring unit is configured to determine, for each of the potential impact dimensions, a plurality of field scores of the to-be-detected text for the plurality of evaluation fields based on scoring rules corresponding to the plurality of evaluation fields corresponding to the potential impact dimension; and The risk scoring unit is configured to determine the risk score of the text to be detected for the potential impact dimension based on multiple field scores of the text to be detected for the multiple evaluation fields and the weight values ​​corresponding to each of the multiple evaluation fields.

21. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; wherein The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.

22. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-10.

23. A computer program product comprising a computer program, wherein: When the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Cited By

  • Large model-based risk detection method and device, intelligent agent, equipment and medium

    CN120561810A

  • Large model business risk detection method and device, medium, equipment and program product

    CN121414133A

  • Event risk detection method and device based on large model and medium

    CN121563018A