Data processing method and device, equipment, storage medium and program product
Through multimodal feature extraction and multi-level detection mechanisms, combined with large language models and multimodal knowledge bases, the problem that traditional methods have difficulty in identifying metaphors and cross-language expression attacks is solved, achieving stronger security protection capabilities and being suitable for complex scenarios such as medical and financial.
Patent Information
- Application Number
- CN202510927803.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-03
AI Technical Summary
Traditional defense methods have difficulty identifying attacks based on metaphorical expressions or cross-language expressions in AI-generated content, have weak security protection capabilities, and are difficult to adapt to complex security protection scenarios.
Through multimodal feature extraction, multimodal knowledge base to find risk cases, context enhancement and multi-level detection mechanism, attacks expressed in metaphors or cross-language expressions are identified, and large language models are used for risk assessment and defense.
It significantly reduces the risk of the model being bypassed or induced to generate illegal content, improves security protection capabilities, and adapts to the security needs of complex scenarios such as medical and financial scenarios.
Smart Images

Figure CN120744458A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to, but is not limited to, the field of information security technology, and in particular to a data processing method, apparatus, device, storage medium, and program product. Background Art
[0002] The rapid evolution of AI-generated content (AIGC) technology has driven innovation in content creation, but it has also brought complex security challenges. As model generation capabilities improve, attack methods are becoming more diverse, covert, and intelligent, placing higher demands on the real-time and adaptable nature of security protection.
[0003] Traditional defense methods rely on matching attack keywords and cannot identify attacks expressed in metaphorical or cross-language forms. They have weak security protection capabilities and are difficult to adapt to complex security protection scenarios such as healthcare and finance. Summary of the Invention
[0004] In view of this, the present application at least provides a data processing method, a computer device, a storage medium, and a program product.
[0005] The technical solution of the embodiment of the present application is implemented as follows: In one aspect, an embodiment of the present application provides a data processing method, the data processing method comprising: Performing multimodal feature extraction on the data to be processed to obtain a first multimodal feature of the data to be processed; Searching for risk cases matching the first multimodal feature from the multimodal knowledge base; Contextually enhancing the first multimodal feature based on key features of the matched risk case to obtain a second multimodal feature; The second multimodal feature is input into the large language model for processing, and a multi-level detection mechanism is used to detect the overall risk of the data to be processed at the input layer of the large language model, the multi-layer attention layer detects the critical risk of the data to be processed, and the output layer detects the word risk of the data to be processed, and outputs the risk defense strategy for the data to be processed.
[0006] On the other hand, an embodiment of the present application provides a data processing device, the data processing device comprising: a feature extraction module configured to perform multimodal feature extraction on the data to be processed to obtain a first multimodal feature of the data to be processed; a search module configured to search for risk cases matching the first multimodal feature from a multimodal knowledge base; an enhancement module configured to perform contextual enhancement on the first multimodal feature based on key features of the matched risk case to obtain a second multimodal feature; The model processing module is configured to input the second multimodal features into the large language model for processing, and adopt a multi-level detection mechanism to detect the overall risk of the data to be processed at the input layer of the large language model, the multi-layer attention layer to detect the critical risk of the data to be processed, the output layer to detect the word risk of the data to be processed, and output a risk defense strategy for the data to be processed.
[0007] On the other hand, an embodiment of the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.
[0008] On the other hand, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which implements some or all of the steps in the above method when executed by a processor.
[0009] On the other hand, an embodiment of the present application provides a computer program, including computer-readable code. When the computer-readable code is executed in a computer device, a processor in the computer device executes some or all of the steps for implementing the above method.
[0010] On the other hand, an embodiment of the present application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, some or all of the steps in the above method are implemented.
[0011] In an embodiment of the present application, by performing multimodal feature extraction on the data to be processed, a first multimodal feature containing data features of the data to be processed at different levels can be obtained; risk cases matching the first multimodal feature are searched from the multimodal knowledge base, and the first multimodal feature is contextually enhanced based on the key features of the matching risk cases, so as to obtain a second multimodal feature containing risk semantics, which facilitates subsequent models to perform risk assessment based on more sufficient risk semantics; the second multimodal feature is input into a large language model for processing, and a multi-level detection mechanism is used to detect the overall risk of the data to be processed at the input layer of the large language model, and a multi-layer attention layer detects the critical risk of the data to be processed, which can identify attacks of metaphorical expressions or cross-language expressions, and significantly reduce the risk of the model being bypassed or induced to generate illegal content; in this way, it can effectively respond to attacks of cross-modal, metaphorical or cross-language expressions, has strong security protection capabilities, can adapt to complex security protection scenarios such as medical care and finance, and meet the current higher requirements for content security compliance in large-scale applications of AIGC. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 A schematic diagram of the implementation process of a data processing method provided in an embodiment of the present application Figure 1 ; Figure 2 A schematic diagram of the implementation process of a multimodal knowledge base in a data processing method provided in an embodiment of the present application; Figure 3 A schematic diagram of the implementation process of a data processing method provided in an embodiment of the present application Figure 2 ; Figure 4 A schematic diagram of the implementation process of a data processing method provided in an embodiment of the present application Figure 3 ; Figure 5 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application; Figure 6 A hardware entity diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0013] The technical solution of the present application is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0014] In order to enable people in this technical field to better understand the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly described below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments.
[0015] The terms "first," "second," and "third," etc., as used herein, are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions, for example, including a series of steps or elements. A method, system, product, or apparatus is not necessarily limited to those steps or elements explicitly listed, but may include other steps or elements not explicitly listed or inherent to the process, method, product, or apparatus.
[0016] The present application provides a data processing method that can be executed by a processor of a computer device. The computer device may refer to a server, a laptop, a tablet computer, a desktop computer, a smart TV, a set-top box, a mobile device (such as a mobile phone, a portable video player, a personal digital assistant, a dedicated messaging device, a portable gaming device), or other device with data processing capabilities. Figure 1 As shown, the method may include the following steps 101 to 104: Step 101: Perform multimodal feature extraction on the data to be processed to obtain a first multimodal feature of the data to be processed.
[0017] Pending data refers to newly uploaded data awaiting inspection. Pending data can be a single type of data, such as text, images, videos, or audio. It can also be a combination of multiple types (i.e., multimodal data), such as embedded text and images, audio and video, text combined with audio and video, data with zero-width characters, and audio-to-text transcriptions with images.
[0018] Multimodal feature extraction is used to extract multiple different types of data features from the data to be processed. First multimodal features are multimodal features extracted from the data to be processed, used to characterize the data characteristics of the data to be processed at different levels. Each level corresponds to a data type and multiple hierarchical representations of that data type.
[0019] In some embodiments, step 101 may be specifically implemented by performing multimodal feature extraction on the data to be processed using the fine-tuned large language model to obtain first multimodal features of the data to be processed.
[0020] In some embodiments, step 101 may be implemented by performing multimodal feature extraction on the data to be processed using multiple feature extraction models to obtain first multimodal features. Each feature extraction model is used to extract data features of a corresponding type.
[0021] Step 102: Search the multimodal knowledge base for risk cases that match the first multimodal feature.
[0022] The multimodal knowledge base covers data compliance standards and various attack methods.
[0023] In some implementations, a multimodal knowledge base may be constructed based on data compliance standards and historical risk cases. Thus, using the first multimodal feature as a matching basis, one or more risk cases whose features match the first multimodal feature are searched from the multimodal knowledge base.
[0024] Step 103: Contextually enhance the first multimodal feature based on the key features of the matched risk case to obtain a second multimodal feature.
[0025] Key features are used to characterize at least one of the core text, risk feature labels, and core multimodal features of a matched risk case. Second multimodal features are enhanced multimodal features that include risk semantics, which represent at least one of the core text, risk feature labels, and core multimodal features.
[0026] In some embodiments, step 103 may be implemented by concatenating the key features of the matched risk case to the vector sequence of the first multimodal feature to obtain the second multimodal feature. The concatenation location can be set based on business needs and is not limited to the end or beginning of the vector sequence of the first multimodal feature. It can also include vector superposition of the vector sequence of the first multimodal feature.
[0027] Step 104: Input the second multimodal features into the large language model for processing, and use a multi-level detection mechanism to detect the overall risk of the data to be processed at the input layer of the large language model, detect the critical risk of the data to be processed at the multi-layer attention layer, detect the word risk of the data to be processed at the output layer, and output a risk defense strategy for the data to be processed.
[0028] The large language model is used to perform risk analysis on the data to be processed and output corresponding risk defense strategies. The multi-level detection mechanism is used to provide security protection at the input layer, multi-layer attention layer, and output layer during the model processing process. The overall risk is used to characterize whether the data to be processed is risky as a whole. The critical risk is used to characterize whether the multimodal features of the data to be processed are abnormal and / or whether there is semantic deviation; the critical risk may include whether the attention distribution of the multimodal features is abnormally focused and / or whether the contextual semantics deviate from the preset semantic track. The word risk is used to characterize whether the data to be processed contains controlled vocabulary. Controlled vocabulary refers to abnormal vocabulary that is not allowed to be published or vocabulary with abnormal risks.
[0029] In some embodiments, the specific implementation method of step 104 can be: at the input layer of the large language model, based on the risk cases in the multimodal knowledge base, the overall risk of the data to be processed is detected; at the multi-layer attention layer of the large language model, based on the weight distribution of the multi-layer attention layer or the semantic expression of the second multimodal feature, the critical risk of the data to be processed is detected; at the output layer of the large language model, the word risk of the data to be processed is detected, and finally the risk defense strategy of the data to be processed is output.
[0030] In an embodiment of the present application, by performing multimodal feature extraction on the data to be processed, a first multimodal feature containing data features of the data to be processed at different levels can be obtained; risk cases matching the first multimodal feature are searched from the multimodal knowledge base, and the first multimodal feature is contextually enhanced based on the key features of the matching risk cases, so as to obtain a second multimodal feature containing risk semantics, which facilitates the subsequent model to perform risk assessment based on more sufficient risk semantics; the second multimodal feature is input into the large language model for processing, and a multi-level detection mechanism is adopted to detect the overall risk of the data to be processed at the input layer of the large language model, and the multi-layer attention layer detects the critical risk of the data to be processed, which can identify attacks of metaphorical expressions or cross-language expressions and significantly reduce the risk of the model being bypassed or induced to generate illegal content; in this way, it can effectively respond to attacks of cross-modal, metaphorical or cross-language expressions, has strong security protection capabilities, can adapt to complex security protection scenarios such as medical care and finance, and meet the current higher requirements for content security compliance in large-scale applications of AIGC.
[0031] The present application provides a data processing method, such as Figure 2 As shown, the method can construct a multimodal knowledge base through the following steps 201 to 204: Step 201: Determine multi-level risk feature labels based on artificial intelligence service security requirements and industry security review standards.
[0032] AI service security requirements are used to characterize and assess security risks. For example, these requirements might refer to the "Basic Security Requirements for Generative AI Services." This document, through scenario-based analysis, categorizes various security risks related to corpus and generated content, and establishes specific detection methods and quantitative indicators for generative AI service security. Industry security review standards refer to the security review standards for various industries.
[0033] In some implementations, step 201 may be implemented by extracting different types of risk features from AI service security requirements and industry security review standards, and categorizing these risk features to generate multi-level risk feature labels. These different risk features may include: personal information leakage, unauthorized access, data tampering, etc.
[0034] Step 202: Determine a multimodal attack method based on historical risk cases and public texts of risk cases.
[0035] Historical risk cases refer to security risk cases that have already occurred. Publicly available risk case texts refer to publicly available texts containing risk cases, including but not limited to journals, papers, webpages, patents, etc.
[0036] Multimodal attacks refer to attacks that exploit the combination or characteristics of multiple modal data, such as text, images, audio, and video, to bypass security detection mechanisms, mislead model outputs, or obtain unauthorized data. In some implementations, multimodal attacks include, but are not limited to, cross-modal adversarial attacks, modal confusion attacks, cross-modal acquisition of unauthorized data, multimodal covert information transmission, multimodal data synonym substitution, adversarial prompt engineering (APE), and multimodal prompt attacks.
[0037] In some implementations, step 202 may be implemented by extracting multimodal attack methods from historical risk cases and public documents of risk cases. The extraction methods include, but are not limited to, manual extraction, neural network extraction, and intelligent model extraction.
[0038] Step 203: Construct a pattern index corresponding to the multimodal attack mode based on the multimodal data corresponding to the multimodal attack mode and the mapping relationship between the multimodal data.
[0039] The mapping relationship between multimodal data refers to the semantic, structural, or functional association, correspondence, or conversion relationship between different modal data. Pattern indexing is used to characterize the propagation path of multimodal attack methods and the mapping between multimodal data.
[0040] In some embodiments, step 203 may be specifically implemented by using data types such as text, image, or audio in the multimodal data as nodes and mapping relationships between the multimodal data as connection relationships between the nodes to construct a pattern index corresponding to the multimodal attack method.
[0041] In some implementations, synonym and cross-language replacements can also be performed on nodes in the pattern index to quickly lock onto semantic evolution later.
[0042] For example, for the 'collaborative illegal expression of text and images' model, a semantic mapping relationship is established between the ambiguous expressions in the text and the symbolic visual content in the image. For example, in the 'cross-language mixed spelling control word attack' model, variant nodes such as daoy ng (split spelling) and pinyin daoyon are classified and referred back to the risk feature label 'appropriation of the fruits of others' labor', thereby achieving structured recognition of spelling evasion attacks.
[0043] Step 204: construct the multimodal knowledge base based on the multi-level risk feature labels, the multimodal attack methods and the pattern index.
[0044] In some embodiments, the specific implementation method of step 204 can be: setting a basic node layer based on multi-level risk feature labels; setting a risk association layer based on pattern indexes and corresponding derivative content; setting a dynamic update layer based on risk expressions corresponding to multimodal attack methods; constructing a three-level graph structure based on the basic node layer, risk association layer and dynamic update layer to obtain a multimodal knowledge base.
[0045] In some implementations, step 204 may also be implemented through the following steps 2041 to 2044: Step 2041: Set a basic node layer based on the multi-level risk feature labels and the multimodal data.
[0046] The basic node layer contains core features and labels such as words and phrases.
[0047] In some implementations, step 2041 may be specifically implemented by using multi-level risk feature labels, keywords, and key phrases in multimodal data as the content of the basic node layer.
[0048] Step 2042: Set a risk association layer based on the pattern index, the derived index semantically similar to the pattern index, the multimodal attack method, and the derived attack method semantically similar to the multimodal attack method.
[0049] The risk association layer is used to record different types of high-risk nodes, their contextual definitions, multimodal attack methods, and their derivative attack vectors. Contextual definitions include: expansion and sub-description of high-risk nodes, synonymous substitution (including within-language and across-language substitutions) for high-risk nodes, and fuzzy representation.
[0050] In some implementations, step 2042 may be specifically implemented by using high-risk nodes in the pattern index, upper and lower concepts of high-risk nodes, multimodal attack methods, and derived attack means as the content of the risk association layer.
[0051] Step 2043: Set a dynamic update layer based on the risk expression corresponding to the multimodal attack mode.
[0052] Risk expressions are used to reflect the attack characteristics or risk characteristics of multimodal attack methods.
[0053] In some embodiments, step 2043 may be implemented by using the risk expression of the multimodal attack mode as content of the dynamic update layer, and updating the new risk expression to the dynamic update layer in real time when a new risk expression appears.
[0054] Step 2044: construct a graph structure based on the basic node layer, the risk association layer, and the dynamic update layer to obtain the multimodal knowledge base.
[0055] In some implementations, step 2044 may be specifically implemented by constructing a three-level graph structure based on a basic node layer, a risk association layer, and a dynamic update layer to obtain a multimodal knowledge base.
[0056] In some implementations, Neo4j can be used to manage the graph structure, and then supplemented with FAISS to perform large-scale vector similarity searches.
[0057] In some implementations, when new attack methods and risk cases appear, the new attack methods and risk cases are updated to the multimodal knowledge base in real time.
[0058] In the embodiment of the present application, a basic node layer is set based on multi-level risk feature labels and multimodal data, so that the basic node layer can include core features and risk feature labels; a risk association layer is set based on a pattern index, a derivative index with semantically similar semantics to the pattern index, a multimodal attack method, and a derivative attack method with semantically similar semantics to the multimodal attack method, so that the risk association layer can cover the pattern index and its derivative content, the multimodal attack method and its derivative attack method, so as to quickly lock the semantic evolution later; a dynamic update layer is set based on the risk expression corresponding to the multimodal attack method, so that the multimodal attack method and its corresponding features and labels can be quickly locked according to the dynamic update layer. A three-level graph structure is constructed based on the basic node layer, the risk association layer, and the dynamic update layer to obtain a multimodal knowledge base. The hierarchical management can improve the retrieval rate of the multimodal knowledge base, making it easy to quickly find the key features of risk cases that match the data to be processed.
[0059] The present application embodiment provides a data processing method, which can be executed by a processor of a computer device. Figure 3 As shown, the method may include the following steps 301 to 305: Step 301: Perform multimodal feature extraction on the data to be processed to obtain a first multimodal feature of the data to be processed.
[0060] In some embodiments, the specific implementation method of step 301 can be: inputting the multimodal data in the data to be processed into the corresponding feature extraction model respectively to obtain the data features corresponding to each modal data; and fusing the data features corresponding to each of the multimodal data grids to obtain the first multimodal features.
[0061] Each modal data (i.e., each type of data) corresponds to a feature extraction model, which is used to extract the data features of the corresponding modal data.
[0062] In some embodiments, for text data, a fine-tuned large language model can be used to perform feature extraction on the text data to obtain text features; for audio data, Waveform to Vector 2.0 (Wav2Vec2) can be used to perform feature extraction on the audio data to obtain audio features; for video data, contrastive language-image pre-training (CLIP) can be used to perform feature extraction on the video data to obtain video features; the data features corresponding to the multimodal data, namely text features, video features, and audio features, are spliced or aligned to obtain first multimodal features.
[0063] Step 302: Based on the similarity between the multimodal features of a plurality of reference risk cases in the multimodal knowledge base and the first multimodal features, determine risk cases to be screened from the plurality of reference risk cases.
[0064] Reference risk cases refer to risk cases in the multimodal knowledge base.
[0065] In some embodiments, similarity can be determined based on the mixed cosine-Manhattan distance between the multimodal features of multiple reference risk cases in the multimodal knowledge base and the first multimodal feature. The cases are then sorted from highest to lowest similarity, and the top K historical risk cases are selected as the risk cases to be screened. Alternatively, other similarity metrics can be used to determine the similarity between the multimodal features of the reference risk cases and the first multimodal feature.
[0066] In practice, the hyperparameter α can be set to 0.7 to find a suitable balance between semantic similarity and character differences, and the Top-K can be within the range of 10~20.
[0067] Step 303: Cluster the risk cases to be screened, remove risk cases whose semantic overlap reaches a first threshold, and obtain the matched risk cases.
[0068] In some implementations, a clustering algorithm can be used to cluster the risk cases to be screened, removing risk cases with semantic overlap reaching a first threshold to obtain matching risk cases. The first threshold can be pre-set and can be specifically set based on actual business needs to ensure that risk cases with high semantic overlap are removed while representative risk cases are retained.
[0069] Step 304: Contextually enhance the first multimodal feature based on the key features of the matched risk case to obtain a second multimodal feature.
[0070] Step 305: Input the second multimodal features into the large language model for processing, and use a multi-level detection mechanism to detect the overall risk of the data to be processed at the input layer of the large language model, detect the critical risk of the data to be processed at the multi-layer attention layer, detect the word risk of the data to be processed at the output layer, and output a risk defense strategy for the data to be processed.
[0071] In some embodiments, the step 305 of "using a multi-level detection mechanism at the input layer of the large language model to detect the overall risk of the data to be processed, multiple attention layers to detect the critical risk of the data to be processed, and an output layer to detect the word risk of the data to be processed, and outputting a risk defense strategy for the data to be processed" can be performed through the following steps 3051 to 3053: Step 3051: At the input layer of the large language model, detect whether the similarity between the second multimodal feature and the reference risk cases in the multimodal knowledge base reaches a second threshold, and output a first-level risk defense strategy if the second threshold is reached; the first-level risk defense strategy indicates that manual intervention is required.
[0072] The second threshold value may be set to 0.85. The first-level risk defense strategy indicates that the data to be processed is a high-risk item.
[0073] A high-speed interception mechanism is set up at the input layer, and the Similarity Hash Algorithm (SIMHash) is used to quickly detect the similarity between the input and historical risk cases in the multimodal knowledge base. If the similarity exceeds 0.85, the generation of the first-level risk defense strategy is directly rejected and reviewed by the auditor.
[0074] Step 3052: In the multi-layer attention layer of the large language model, detect whether the attention distribution of the second multimodal feature is abnormally focused, or whether the contextual semantics of the second multimodal feature deviates from the preset semantic track, and output a secondary risk defense strategy when the attention distribution is abnormally focused or the contextual semantics deviates from the preset semantic track; the secondary risk defense strategy represents truncation of the data to be processed.
[0075] In some embodiments, the specific implementation method of "detecting whether the attention distribution of the second multimodal feature is abnormally focused" can be: determine the weight mean of all attention heads of each attention layer, and if there is an abnormal attention layer in which the distribution of the weight mean reaches a preset first distribution threshold, determine that the attention distribution of the second multimodal feature is abnormally focused. The distribution of the weight mean reaches the preset distribution threshold means that the distribution of the weight mean shows an obvious "spike", for example: position weight>0.5, other weights<0.1. The first distribution threshold is used to judge the weight mean of all attention heads of each attention layer.
[0076] In some embodiments, "detecting whether the attention distribution of the second multimodal feature is abnormally focused" may be specifically implemented by: determining the weight distribution of each attention head in each attention layer; identifying abnormal attention heads in each attention layer whose weight distribution reaches a preset second distribution threshold; and determining abnormal focus when the number of abnormal attention heads reaches the threshold. The second distribution threshold is used to determine the weight distribution of each attention head.
[0077] In some embodiments, the specific implementation method of "detecting whether the contextual semantics of the second multimodal feature deviates from the preset semantic track" can be: mapping each modal feature in the second multimodal feature to the same semantic space to obtain the semantic expression of each modal feature; based on the similarity of the semantic expressions of different modal features, determining whether the contextual semantics of the second multimodal feature deviates from the preset semantic track. Specifically, when the similarity of the semantic expressions of different modal features does not reach the similarity threshold, it is determined that the contextual semantics of the second multimodal feature deviates from the preset semantic track. At this time, the preset semantic track refers to the semantic expression of similar modal features.
[0078] In some embodiments, the specific implementation method of "detecting whether the contextual semantics of the second multimodal feature deviates from the preset semantic track" can be: constructing a preset semantic track based on the first multimodal feature; mapping each modal feature in the second multimodal feature to the same semantic space to obtain the semantic expression of each modal feature; when the semantic expression of each modal feature does not belong to the preset semantic track, determining whether the contextual semantics of the second multimodal feature deviates from the preset semantic track. For example, if the preset semantic track is "cat" and the semantic expression of a modal feature of the second multimodal feature is "tree", it is determined to be deviated from the preset semantic track; if the preset semantic track is "cat" and the semantic expression of a modal feature of the second multimodal feature is "cat in a tree", it is determined to be not deviated from the preset semantic track.
[0079] If the input does not hit a high-risk item, the model enters the second layer of protection, which is the dynamic monitoring link based on attention distribution: when it detects abnormal focus on attention weight (such as an abnormal increase in the concentration of controlled words or high-risk words) or the conversation context deviates from the normal semantic track, the system will conduct a review and evaluation after completing this round of generation; if the risk assessment result is at a critical value, it can be cooled down (such as moderately lowering the generation temperature or truncating part of the output) and prompting manual review.
[0080] Step 3053: At the output layer of the large language model, a traceability mark is added to the output content corresponding to the second multimodal feature to detect whether there are controlled words in the output content. If there are controlled words in the output content, a third-level risk defense strategy is output; the third-level risk defense strategy represents word-level replacement or masking of the controlled words.
[0081] Traceability identification can be determined based on text, image or audio content.
[0082] During implementation, differential privacy watermarking is implemented at the output layer, and a reversible traceability mark is added to the generated text, image, or audio and video content; if residual control elements are still detected during verification of the final output, word-level replacement or masking operations can be performed to ensure that the generated results meet security requirements in terms of content compliance and traceability.
[0083] Based on the above embodiment, the data processing method provided in the embodiment of the present application further includes the following steps 306 to 309: Step 306: In a multi-round dialogue scenario, the multimodal features of the current round are merged with the multimodal features of N historical rounds before the current round to obtain a merged multimodal feature.
[0084] Step 307: Perform anomaly detection on the combined multimodal features to determine whether there are risk signs in the multi-round conversation; the risk signs include: the frequency of occurrence of sign semantics in the multi-round conversation reaches a third threshold and / or the topic of the multi-round conversation has a tendency to cross the line.
[0085] Step 308: If the risk sign exists, output a session-level risk defense strategy; Step 309: Update the multimodal knowledge base based on the multimodal attack method of the multi-round dialogue.
[0086] In a multi-round dialogue scenario, the embodiment of the present application uses multi-round context tracking to defend against the risk of attackers inducing model violations step by step. Specifically, the dialogue history is vectorized by round and saved in a short-term cache. By merging the input of the new round with the semantic information of the previous rounds, it is identified whether there are strong signs of attack. Once it is found that certain key terms appear frequently in multiple rounds of dialogue, or the dialogue topic has a clear tendency to cross the line, a higher priority security protection strategy will be triggered, and the retrieval of detailed rules for this specific field in the knowledge base can be expanded again to ensure that the monitoring and interception of subsequent potential attacks are more stringent.
[0087] The embodiment of the present application forms a continuously rolling self-learning closed loop through dynamic feedback and strategy iteration. Once a new attack sample is detected and intercepted during use, its contextual information and attack characteristics will be fed back to the system, and the variant mapping and node association of the knowledge base will be incrementally updated. With the iterative update of the knowledge base and the model, the embodiment of the present application can continuously improve the detection and response capabilities of multimodal, dynamic, and professional risks, thereby maintaining high-intensity reinforcement protection for the security of generative artificial intelligence content. This method is compatible and extensible in multiple modalities such as text, images, audio and video. Through multi-level process intervention and real-time iteration strategies, it can significantly reduce the risk of the model being bypassed or induced to generate illegal content, meeting the current higher requirements for content security compliance for large-scale applications of AIGC.
[0088] The following describes the application of the data processing method provided in the embodiments of the present application in actual scenarios.
[0089] In order to improve the content security protection capability of generative artificial intelligence, this application embodiment proposes a security control method based on dynamic knowledge enhancement (corresponding to the above-mentioned data processing method). The specific process is as follows: Figure 4 As shown, it includes the following steps 401 to 406: Step 401: Build a security knowledge enhancement model: obtain a multimodal dataset and perform domain adaptation fine-tuning on the pre-trained large language model.
[0090] Here, the security knowledge enhancement model refers to a trained large language model used to obtain risk defense strategies.
[0091] Step 402: Build a multimodal knowledge base: Integrate compliance standards and the latest security requirements to form a multi-level knowledge graph covering basic control semantics, cross-modal risk characteristics, and variant expression mapping.
[0092] Step 403: Perform risk feature context enhancement: extract multi-dimensional features based on user input, retrieve the top-K relevant risk cases through hybrid distance calculation, and use clustering algorithm to remove redundancy and retain core risk factors.
[0093] Here, the relevant risk cases correspond to the above-mentioned matching risk cases.
[0094] Step 404: Generate and execute risk defense strategies: Input the enhanced features into the large language model, output a dynamic risk defense strategy, and perform real-time correction on the input, generation process, and final output through a multi-level intervention mechanism.
[0095] Step 405: Perform multi-round dialogue tracking and fusion, continuously monitor the drift of intentions in the dialogue scenario, use the multimodal knowledge base for context comparison, and trigger targeted intervention and update the strategy when abnormal topic migration is detected.
[0096] Step 406: Conduct dynamic feedback and strategy iteration: Feedback attack samples and protection results obtained through real-time monitoring to the multimodal knowledge base, triggering incremental updates to the knowledge graph.
[0097] In some embodiments, when fine-tuning the model, the decoder structure is preferred and RoPE rotation position encoding and RMS Norm normalization layer are used to improve the ability to capture contextual features.
[0098] In some embodiments, in the above-mentioned knowledge graph, multi-level risk nodes can be set to distinguish scenarios of different severities, and sub-graph indexes can be established for professional terms to achieve refined matching.
[0099] In some implementations, to enhance the accuracy of risk retrieval, Top-K retrieval uses a hybrid metric that fuses improved cosine distance and Manhattan distance to achieve deduplication through clustering.
[0100] In some implementations, differential privacy watermarking technology can be introduced at the output layer to add traceability identification to the generated content to enable subsequent auditing and responsibility tracking.
[0101] In some embodiments, in a multi-round conversation tracking module, the conversation context is divided into blocks and vectorized to maintain continuous memory and risk perception, and advanced security detection is immediately triggered when cross-semantic scenario migration is discovered.
[0102] The embodiment of the present application builds an adaptive generative artificial intelligence security protection system by combining multiple rounds of context tracking and dynamic knowledge enhancement. The six steps are closely linked: first, by fine-tuning the model and building a dynamic knowledge base for security data in various fields, then enhancing contextual risk features and executing multi-level policies in actual interactions, and then using multiple rounds of dialogue tracking and real-time feedback loops to continuously revise and upgrade security measures. This solution can effectively respond to dynamic, cross-modal, and progressive attack threats, solves the pain points of traditional static vocabulary and single-round detection methods in terms of security and real-time performance, and innovatively realizes rolling policy updates and deep intervention based on knowledge graphs, providing strong protection for the safe application of AIGC in many management and control scenarios.
[0103] The specific implementation of the security control method based on dynamic knowledge enhancement proposed in the embodiment of the present application is as follows: Corresponding to step 401 above, the present embodiment first collects and constructs a multimodal dataset for the target application scenario (such as social media compliance review or medical Q&A anti-violation analysis), including but not limited to: text risk samples (such as illegal speech and regulated text detected on social media), cross-modal attack samples (such as image and text nesting, zero-width character injection, and audio transcription text corresponding to the original image), multi-round dialogue induction samples (simulating an attacker gradually guiding the model to output out-of-bounds content), and new obfuscated expressions such as homophones or cross-language mixing. Through the systematic organization and annotation of the above data, a multimodal dataset is constructed, laying the foundation for subsequent fine-tuning for specific fields and scenarios.
[0104] After acquiring the aforementioned dataset, a pre-trained large language model based on a decoder structure was selected as the base model. By combining risk data from the corresponding domain with general corpus, domain adaptation fine-tuning was implemented to enhance the model's ability to identify and understand multimodal risk characteristics and violation methods. To achieve better domain adaptation without large-scale updates to the model's original weights, this embodiment of the application employed a LoRA approach for fine-tuning. The LoRA parameters were: lora_r (the dimension of the low-rank matrix) = 8, lora_alpha (scaling factor α) = 16, and lora_dropout (overfitting parameter) = 0.05. Other training parameters included the AdamW optimizer, per_device_train_batch_size = 2, num_train_epochs = 5, and learning_rate = 2e-5. Through this LoRA fine-tuning scheme, the model can establish a more comprehensive representation capability for multimodal risk scenarios (such as text, image descriptions, zero-width characters, and cross-language obfuscation) with only small incremental parameter changes. Because LoRA training is more lightweight, it can also significantly reduce the hardware resources and training time required, enabling rapid, adaptive domain enhancement of large models. After completing the above fine-tuning, the resulting LoRA adaptation model can be used in various generative AI scenarios to implement compliance checks, risk reviews, and security outputs for text or multimodal content. If further expansion or updates to prevention and control strategies are required, simply retrain or merge the LoRA layer on the new scenario data to maintain robust identification and intervention capabilities for a wider range of attacks or violations.
[0105] Corresponding to step 402 above, after model training is complete, a dynamic multimodal knowledge base needs to be constructed to cover compliance standards and various attack vectors. To this end, based on the "Basic Security Requirements for Generative Artificial Intelligence Services" and the security review provisions of various industries, multi-level risk feature labels such as personal information leakage, unauthorized access, and data tampering are extracted. To accommodate multimodal scenarios, common cross-modal attack methods are extracted from historical attack cases and public literature and compiled into a unified pattern index. For example, for the "Image-Text Collaborative Illegal Expression" model, a semantic mapping relationship is established between ambiguous expressions in text and symbolic visual content in images. For example, for the "Cross-Language Mixed Spelling Control Word Attack" model, variant nodes such as "daoy ng" (split spelling) and "daoyon" (pinyin) are categorized and referenced to the risk feature label "appropriation of the fruits of others' labor," enabling structured recognition of spelling evasion attacks. Multiple categories of nodes, such as text, image, or audio, are then associated and merged using a mapping method.
[0106] For variant expressions, synonyms and cross-language replacement relationships must also be established to ensure that semantic evolution can be quickly locked. If the system detects new attack vocabulary or hidden methods, it will be quickly inserted and associated with the knowledge base to achieve online updates. The underlying data structure of the knowledge base adopts a three-level graph: the basic node layer encompasses words, phrases and labels, the risk association layer records different types of high-risk nodes and their upper and lower concepts or derived attack methods, and the dynamic update layer saves the new risk expressions that appear in increments. In order to achieve efficient retrieval in practical applications, the embodiment of the present application can use Neo4j to manage the graph content, and then use FAISS to complete large-scale vector similarity search.
[0107] Corresponding to step 403 above, after receiving the user input, it is necessary to extract risk features and enhance the context to fully capture potential violations. To this end, the text data can be fed into a fine-tuned large language model to obtain semantic vectors, and the pre-trained models of CLIP and wav2vec2 can be used for feature extraction for image and audio and video data respectively. The multimodal vectors are then spliced or aligned to obtain a comprehensive expression. The mixed cosine-Manhattan distance is then calculated in the dynamic security knowledge base to retrieve several closest risk cases. In this embodiment of the application, α is set to 0.7 to find a suitable balance between semantic similarity and character differences, and the Top-K can be within the range of 10 to 20. Applying a clustering algorithm to the search results helps to remove risk cases with high semantic overlap, ultimately retaining core cases that can represent the input features. These core case texts or corresponding labels are spliced into the main input sequence to form context-enhanced prompts, allowing subsequent models to make judgments or generate controllable outputs based on more comprehensive risk semantics.
[0108] Corresponding to step 404 above, after the context enhancement is completed, the enhancement results are fed into the fine-tuned large language model to generate a dynamic defense strategy, and the entire generation process is controlled by a multi-level strategy system. First, a high-speed interception mechanism is set up at the input layer: the SIMHash algorithm is used to quickly detect the similarity between the input and the known attack patterns in the knowledge base. If the similarity exceeds 0.85, the generation is directly rejected or reviewed by the auditor. If the input does not hit the high-risk item, the model enters the second layer of protection, which is the dynamic monitoring link based on attention distribution: when it detects an abnormal focus on the attention weight (such as an abnormal increase in the concentration of controlled words or high-risk words) or the conversation context deviates from the normal semantic track, the system will conduct a review and evaluation after completing this round of generation; if the risk assessment result is at a critical value, it can be cooled down (such as moderately reducing the generation temperature or truncating part of the output) and prompting manual review. Finally, differential privacy watermarking is implemented at the output layer, and a reversible traceability mark is added to the generated text, image, or audio and video content. If residual control elements are still detected during verification of the final output, word-level replacement or masking operations can be performed to ensure that the generated results meet security requirements in terms of content compliance and traceability.
[0109] Corresponding to the above-mentioned step 405, in a multi-round dialogue scenario, the embodiment of the present application uses multi-round context tracking to defend against the risk of attackers inducing model violations step by step. The specific approach is to vectorize the dialogue history by round and save it in a short-term cache, and identify whether there are strong signs of attack by merging the input of the new round with the semantic information of the previous rounds. Once it is found that certain key terms appear frequently in multiple rounds of dialogue, or the dialogue topic has a clear tendency to cross the line, a higher priority security protection strategy will be triggered, and the retrieval of detailed rules for this specific field in the knowledge base can be expanded again to ensure that the monitoring and interception of subsequent potential attacks are more stringent.
[0110] Corresponding to the above-mentioned step 406, the embodiment of the present application finally forms a continuously rolling self-learning closed loop through dynamic feedback and strategy iteration. Once a new attack sample is detected and intercepted during use, its contextual information and attack characteristics will be fed back to the system, and the variant mapping and node association of the knowledge base will be incrementally updated. With the iterative update of the knowledge base and the model, the present invention can continuously improve the detection and response capabilities of multimodal, dynamic, and professional risks, thereby maintaining high-intensity reinforcement protection for the security of generative artificial intelligence content. This method is compatible and generalizable in multiple modalities such as text, images, audio and video. Through multi-level process intervention and real-time iteration strategies, it can significantly reduce the risk of the model being bypassed or induced to generate illegal content, meeting the current higher requirements for content security compliance for large-scale applications of AIGC.
[0111] Based on the foregoing embodiments, an embodiment of the present application provides a data processing device, which includes the various units included and the various modules included in each unit, and can be implemented by a processor in a computer device; of course, it can also be implemented by a specific logic circuit; in the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP) or a field programmable gate array (FPGA), etc.
[0112] Figure 5 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, the data processing device 500 includes: a feature extraction module 510, a search module 520, an enhancement module 530 and a model processing module 540, wherein: The feature extraction module 510 is configured to perform multimodal feature extraction on the data to be processed to obtain a first multimodal feature of the data to be processed; The search module 520 is configured to search for risk cases matching the first multimodal feature from a multimodal knowledge base; The enhancement module 530 is configured to perform context enhancement on the first multimodal feature based on the key features of the matched risk case to obtain a second multimodal feature; The model processing module 540 is configured to input the second multimodal features into a large language model for processing, and adopt a multi-level detection mechanism to detect the overall risk of the data to be processed at the input layer of the large language model, detect the critical risk of the data to be processed at multiple attention layers, detect the word risk of the data to be processed at the output layer, and output a risk defense strategy for the data to be processed.
[0113] In some embodiments, the data processing method further includes: The knowledge base construction module is configured such that the multimodal knowledge base is obtained in the following manner: based on artificial intelligence service security requirements and industry security review standards, multi-level risk feature labels are determined; based on historical risk cases and public texts of risk cases, multimodal attack methods are determined; based on multimodal data corresponding to the multimodal attack method and a mapping relationship between the multimodal data, a pattern index corresponding to the multimodal attack method is constructed; and based on the multi-level risk feature labels, the multimodal attack method, and the pattern index, the multimodal knowledge base is constructed.
[0114] In some embodiments, the knowledge base construction module is further configured to: set a basic node layer based on the multi-level risk feature label and the multimodal data; set a risk association layer based on the pattern index, the derivative index semantically similar to the pattern index, the multimodal attack method, and the derivative attack method semantically similar to the multimodal attack method; set a dynamic update layer based on the risk expression corresponding to the multimodal attack method; construct a graph structure based on the basic node layer, the risk association layer and the dynamic update layer to obtain the multimodal knowledge base.
[0115] In some embodiments, the feature extraction module 510 is further configured to: input the multimodal data in the data to be processed into the corresponding feature extraction model respectively to obtain the data features corresponding to each modal data; and fuse the data features corresponding to each of the multimodal data grids to obtain the first multimodal features.
[0116] In some embodiments, the search module 520 is further configured to: determine the risk cases to be screened from the multiple reference risk cases in the multimodal knowledge base based on the similarity between the multimodal features of the multiple reference risk cases and the first multimodal features; cluster the risk cases to be screened, remove the risk cases whose semantic overlap reaches a first threshold, and obtain the matching risk cases.
[0117] In some embodiments, the model processing module 540 is further configured to: at the input layer of the large language model, detect whether the similarity between the second multimodal feature and the reference risk case in the multimodal knowledge base reaches a second threshold, and output a first-level risk defense strategy if the second threshold is reached; the first-level risk defense strategy representation requires manual intervention; at the multi-layer attention layer of the large language model, detect whether the attention distribution of the second multimodal feature is abnormally focused, or whether the contextual semantics of the second multimodal feature deviates from the preset semantic track, and output a second-level risk defense strategy if the attention distribution is abnormally focused or the contextual semantics deviates from the preset semantic track; the second-level risk defense strategy representation performs truncation processing on the data to be processed; at the output layer of the large language model, add a traceability mark to the output content corresponding to the second multimodal feature, detect whether there is controlled vocabulary in the output content, and output a third-level risk defense strategy if there is controlled vocabulary in the output content; the third-level risk defense strategy representation performs word-level replacement or masking processing on the controlled vocabulary.
[0118] In some embodiments, the model processing module 540 is further configured to: in a multi-round dialogue scenario, merge the multimodal features of the current round with the multimodal features of N historical rounds before the current round to obtain a merged multimodal feature; perform anomaly detection on the merged multimodal feature to determine whether there are risk signs in the multi-round dialogue; the risk signs include: the frequency of occurrence of the sign semantics in the multi-round dialogue reaches a third threshold and / or the topic of the multi-round dialogue has a tendency to cross the boundary; in the presence of the risk signs, output a dialogue-level risk defense strategy; and update the multimodal knowledge base based on the multimodal attack method of the multi-round dialogue.
[0119] The description of the above device embodiment is similar to the description of the above method embodiment and has similar beneficial effects as the method embodiment. In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to perform the methods described in the above method embodiments. For technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.
[0120] It should be noted that in the embodiments of the present application, if the above-mentioned data processing method is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application can be essentially embodied in the form of a software product, or the part that contributes to the relevant technology. The software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk. In this way, the embodiments of the present application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.
[0121] An embodiment of the present application provides a computer device, including a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the program, some or all of the steps in the above method are implemented.
[0122] The present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above method. The computer-readable storage medium may be transient or non-transient.
[0123] An embodiment of the present application provides a computer program, including computer-readable code. When the computer-readable code is run in a computer device, a processor in the computer device executes some or all of the steps for implementing the above method.
[0124] An embodiment of the present application provides a computer program product, comprising a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, the computer program implements some or all of the steps of the above-described method. The computer program product can be implemented in hardware, software, or a combination thereof. In some embodiments, the computer program product is embodied as a computer storage medium. In other embodiments, the computer program product is embodied as a software product, such as a software development kit (SDK).
[0125] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between the various embodiments, and their similarities or similarities can be referenced to each other. The descriptions of the above device, storage medium, computer program, and computer program product embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects as the method embodiments. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the description of the method embodiments of this application for understanding.
[0126] It should be noted that Figure 6 A schematic diagram of a hardware entity of a computer device in an embodiment of the present application is shown in FIG. Figure 6 As shown, the hardware entity of the computer device 600 includes: a processor 601, a communication interface 602 and a memory 603, wherein: Processor 601 generally controls the overall operation of computer device 600 .
[0127] The communication interface 602 enables the computer device to communicate with other terminals or servers through a network.
[0128] Memory 603 is configured to store instructions and applications executable by processor 601. It can also cache data to be processed or processed by processor 601 and various modules in computer device 600 (e.g., image data, audio data, voice communication data, and video communication data). This can be implemented using flash memory (FLASH) or random access memory (RAM). Data can be transmitted between processor 601, communication interface 602, and memory 603 via bus 604.
[0129] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned steps / processes does not mean the order of execution, and the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.
[0130] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0131] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0132] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0133] In addition, all functional units in the embodiments of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units.
[0134] Those skilled in the art will understand that all or part of the steps of the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.
[0135] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0136] The above is only an implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. A data processing method, characterized in that: The data processing method includes: Performing multimodal feature extraction on the data to be processed to obtain a first multimodal feature of the data to be processed; Searching a multimodal knowledge base for a risk case that matches the first multimodal feature; performing context enhancement on the first multimodal feature based on the key features of the matched risk case to obtain a second multimodal feature; The second multimodal features are input into a large language model for processing, and a multi-level detection mechanism is used to detect the overall risk of the data to be processed at the input layer of the large language model, the critical risk of the data to be processed is detected by multiple attention layers, and the word risk of the data to be processed is detected at the output layer, and a risk defense strategy for the data to be processed is output.
2. The data processing method according to claim 1, wherein: The multimodal knowledge base is obtained in the following way: Determine multi-level risk feature labels based on AI service security requirements and industry security review standards; Identify multimodal attack methods based on historical risk cases and public risk case texts; Constructing a pattern index corresponding to the multimodal attack mode based on the multimodal data corresponding to the multimodal attack mode and a mapping relationship between the multimodal data; The multimodal knowledge base is constructed based on the multi-level risk feature labels, the multimodal attack methods and the pattern index.
3. The data processing method according to claim 2, characterized in that: The constructing of the multimodal knowledge base based on the multi-level risk feature labels, the multimodal attack methods and the pattern index includes: Setting a basic node layer based on the multi-level risk feature labels and the multimodal data; Setting a risk association layer based on the pattern index, a derivative index semantically similar to the pattern index, the multimodal attack mode, and a derivative attack mode semantically similar to the multimodal attack mode; Setting a dynamic update layer based on the risk expression corresponding to the multimodal attack mode; A graph structure is constructed based on the basic node layer, the risk association layer, and the dynamic update layer to obtain the multimodal knowledge base.
4. The data processing method according to any one of claims 1 to 3, characterized in that: The extracting multimodal features of the data to be processed to obtain a first multimodal feature of the data to be processed includes: Inputting the multimodal data in the data to be processed into corresponding feature extraction models respectively to obtain data features corresponding to each modal data; The data features corresponding to the multimodal data grids are fused to obtain the first multimodal features.
5. The data processing method according to any one of claims 1 to 3, characterized in that: The step of searching the multimodal knowledge base for a risk case matching the first multimodal feature includes: Determining risk cases to be screened from the plurality of reference risk cases based on similarities between the multimodal features of the plurality of reference risk cases in the multimodal knowledge base and the first multimodal features; The risk cases to be screened are clustered, and risk cases whose semantic overlap reaches a first threshold are removed to obtain the matched risk cases.
6. The data processing method according to any one of claims 1 to 3, characterized in that: The multi-level detection mechanism is used to detect the overall risk of the data to be processed at the input layer of the large language model, detect the critical risk of the data to be processed at the multi-layer attention layer, detect the word risk of the data to be processed at the output layer, and output a risk defense strategy for the data to be processed, including: At the input layer of the large language model, detecting whether the similarity between the second multimodal feature and a reference risk case in the multimodal knowledge base reaches a second threshold, and outputting a first-level risk defense strategy if the second threshold is reached; the first-level risk defense strategy indicates that manual intervention is required; In the multi-layer attention layer of the large language model, detecting whether the attention distribution of the second multimodal feature is abnormally focused, or whether the contextual semantics of the second multimodal feature deviates from a preset semantic track, and outputting a secondary risk defense strategy when the attention distribution is abnormally focused or the contextual semantics deviates from the preset semantic track; the secondary risk defense strategy represents truncation of the data to be processed; At the output layer of the large language model, a traceability mark is added to the output content corresponding to the second multimodal feature to detect whether the output content contains controlled vocabulary. If the output content contains controlled vocabulary, a third-level risk defense strategy is output; the third-level risk defense strategy represents word-level replacement or masking of the controlled vocabulary.
7. The data processing method according to claim 1, wherein: The data processing method further includes: In a multi-round dialogue scenario, the multimodal features of the current round are merged with the multimodal features of N historical rounds before the current round to obtain a merged multimodal feature; Performing anomaly detection on the combined multimodal features to determine whether the multi-round conversations contain risk signs; the risk signs include: the frequency of occurrence of sign semantics in the multi-round conversations reaches a third threshold and / or the topic of the multi-round conversations has a tendency to cross a boundary; Outputting a conversation-level risk defense strategy when the risk signs exist; The multimodal knowledge base is updated based on the multimodal attack method of the multi-round dialogue.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the program, the steps of the method according to any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product comprising a non-transitory computer-readable storage medium storing a computer program, characterized in that: When the computer program is read and executed by a computer, the steps of the method according to any one of claims 1 to 7 are implemented.