Large model defense method based on multi-view image-text conversion

Through the large-mode defense method of multi-view graphic and text transformation, the problem of difficulty in identifying and defending cross-modal induced attacks in the existing technology is solved, effectively defending the graphic and text generation scenarios, reducing the risk of overprivileged output.

CN120145402AActive Publication Date: 2025-06-13UNIV OF SCI & TECH BEIJING +2

Patent Information

Application Number
CN202510619951.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-13
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

Existing artificial intelligence security technologies are difficult to effectively identify and defend against cross-modal-induced attacks, especially in the case of graphics and text generation, where model output is overright or potentially dangerous content.

Method used

A large-model defense method based on multi-view graphic and text transformation is adopted. By obtaining the text of the graphic and text dialogue prompt word, syntactic structural parameters and semantic parameters of the statement are extracted, multi-modal tensors are formed, and they are split into multiple structural vectors through semantic relationship density, combined with the image generation matrix, calculating semantic reconstruction offsets, positioning abnormal fragments, generating risk weight values, and finally identifying and aborting abnormal output.

Benefits of technology

It enhances cross-modal correlation analysis capabilities, accurately recognizes the mixed attack path of graphic and text, realizes two-way verification of semantic consistency of graphic and text, and reduces the probability of illegal content output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120145402A_ABST
    Figure CN120145402A_ABST
Patent Text Reader

Abstract

The invention discloses a large model defense method based on multi-view image-text conversion, and relates to the technical field of artificial intelligence security. The method comprises the steps of obtaining an image-text dialogue prompt word text, extracting syntactic structure parameters and semantic parameters of statements, splicing the syntactic structure parameters and the semantic parameters into a multi-modal tensor, and expanding a spliced vector into a plurality of structure vectors according to semantic relationship density to form a semantic structure vector group. According to the method, semantic density splitting and image space mapping of the multi-modal tensor are utilized, the cross-modal correlation analysis capability is enhanced, the image-text mixed attack path is accurately identified, and the entity mapping relation between the pixel gradient and the regional mask is combined, so that two-way verification of image-text semantic consistency is realized, and semantic fault vulnerabilities of single-modal detection are eliminated; a path intensity and instruction word dynamic weight fusion mechanism is adopted, the abnormal risk level of a statement structure is quantified, the adaptive limitation of a static rule to semantic variation is broken through, and the illegal content output probability is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence security, and particularly to a large model defense method based on multi-perspective graphic-text conversion. Background Art

[0002] The technical field of artificial intelligence security includes research on related technologies for identifying, defending against, and controlling various security threats faced by artificial intelligence systems during operation. The core content of this technical field includes model adversarial attack protection, malicious prompt detection, model robustness improvement, information injection prevention, and privacy protection. By introducing input content monitoring, model behavior constraints, and multi-modal information fusion, full-process security control is carried out on the model input, internal reasoning process, and output results, focusing on solving the problem that in open-ended conversations, graphic-text generation, and various task scenarios, the model outputs illegal, unauthorized, or potentially dangerous content due to receiving induced inputs. Artificial intelligence security technology is the core technical support for ensuring the trustworthy use of large models.

[0003] Traditional artificial intelligence security technologies rely on single-modal text detection and static rule matching, lacking an effective identification mechanism for cross-modal induced attacks. Traditional malicious prompt detection only analyzes the surface features of text and cannot cope with hidden inductions achieved through image embedding or syntactic structure camouflage, resulting in the model generating unauthorized instructions. In semantic consistency verification, one-way text comparison is used, ignoring the potential differences in multi-perspective graphic-text conversion, and it is difficult to discover logical misdirections caused by cross-modal semantic breaks. The static rule library and keyword blacklist are updated laggingly and cannot cover new semantic mutation attacks, including malicious inputs constructed through synonym replacement and syntactic restructuring. The ability to analyze complex statement structure paths is insufficient, relying on word frequency or simple dependency relationship judgment, and unable to quantify the associated risks of path span and instruction word distribution, resulting in misjudgments or missed judgments. Information injection prevention relies on manually labeled data to train classifiers and lacks real-time response capabilities for zero-shot attacks in open-domain generation scenarios, resulting in significant defense lag. The model behavior constraint mechanism is not deeply bound to input semantic reconstruction and cannot dynamically correct abnormal reasoning paths during the generation process, posing a hidden danger of out-of-control output results. Summary of the Invention

[0004] In order to solve the technical problems existing in the prior art, an embodiment of the present invention provides a large model defense method based on multi-perspective graphic-text conversion. The technical solution is as follows:

[0005] In order to achieve the above object, the present invention adopts the following technical solution. A large model defense method based on multi-perspective graphic-text conversion includes the following steps:

[0006] S1: Obtain the text of the graphic and text dialogue prompt, extract the syntactic structure parameters and semantic parameters of the sentence, splice them into a multi-modal tensor, and expand the spliced multi-modal tensor into multiple structure vectors according to the semantic relationship density to form a semantic structure vector group;

[0007] S2: Invoke the semantic structure vector group, analyze the semantic level value of each vector and divide it into multiple subsets according to the level interval, input the information of multiple subsets into the image generation channel respectively, perform regional positioning on the image space according to the vector offset angle and activation weight, construct a saliency feature distribution map, and form an image generation matrix;

[0008] S3: Invoke the image generation matrix, extract the image pixel density, edge gradient, and region mask data, calculate the entity mapping relationship through the intersection of pixel and semantic annotation, split and reorganize the corresponding sentence fragments to generate semantic vectors, compare the image embedding vectors to calculate the difference sequence, and obtain the semantic reconstruction offset;

[0009] S4: Invoke the semantic reconstruction offset, locate the abnormal fragments in the prompt according to the offset range, extract the position, connection nodes, and span data of the sentence structure path, calculate the path density, and generate a risk weight value in combination with the proportion of instruction phrase groups.

[0010] Optionally, the semantic structure vector group includes semantic level labels, position index parameters, and structural dependency markers. The image generation matrix is specifically a regional feature density map, channel index identifier, and generation coordinate label. The semantic reconstruction offset includes semantic segment mapping differences, sentence structure offset sections, and vector residual indicators. The risk weight value is specifically a path density factor, induced component proportion coefficient, and structural span ratio.

[0011] Optionally, obtaining the text of the graphic and text dialogue prompt in S1, extracting the syntactic structure parameters and semantic parameters of the sentence, splicing them into a multi-modal tensor, and expanding the spliced multi-modal tensor into multiple structure vectors according to the semantic relationship density to form a semantic structure vector group includes:

[0012] S101: Obtain the text of the graphic and text dialogue prompt, extract the part-of-speech type, position number, and syntactic dependency relationship of each word, establish a lexical dependency mapping index sequence according to the part-of-speech label and word order position, and generate a syntactic structure index sequence value;

[0013] S102: Invoke the syntactic structure index sequence value, extract the semantic label and semantic classification number corresponding to each word in the prompt, and perform tensor combination splicing according to the semantic label and syntactic mapping index to obtain a semantic combination tensor;

[0014] S103: Based on the semantic combination tensor, set a benchmark value for splitting the semantic relationship density according to the vector density and the recurrence frequency of semantic tags. Unfold the combined tensor into multiple structural vectors according to the semantic distribution range to generate a semantic structure vector group.

[0015] Optionally, in S2, call the semantic structure vector group, analyze the semantic level value of each vector and divide it into multiple subsets according to the level interval. Input the information of multiple subsets into the image generation channel respectively. According to the vector offset angle and activation weight, perform regional positioning on the image space to construct a saliency feature distribution map and form an image generation matrix, including:

[0016] S201: Call the semantic structure vector group, extract the semantic level label and corresponding structure index value of each vector, analyze the semantic level of each vector, and divide it into multiple subsets according to the level interval to generate a semantic level division cluster value;

[0017] S202: According to the semantic level division cluster value, extract the vector order and structure index number corresponding to each subset and input them into the image generation channel. Set an image region positioning index according to the offset angle value and activation weight parameter of each vector to obtain an image region positioning index value;

[0018] S203: Call the image region positioning index value, calculate the activation response intensity value of each region according to the distribution position of each group of position data in the image space, and perform partition marking on the image space according to the activation response intensity value to construct an image saliency distribution coefficient matrix and generate an image generation matrix.

[0019] Optionally, in S3, call the image generation matrix, extract the image pixel density, edge gradient and region mask data, calculate the entity mapping relationship through the intersection of pixels and semantic annotations, split and reorganize the corresponding sentence fragments to generate semantic vectors, and compare with the image embedding vector to calculate the difference sequence to obtain the semantic reconstruction offset, including:

[0020] S301: Call the image generation matrix, extract the pixel density value, edge gradient value and region mask identifier of the image, locate the semantic annotation region according to the mask identifier, perform coordinate comparison on the pixel coordinates and edge positions of the annotation region to obtain the overlapping coordinate points of the image region and semantic annotation, and generate entity mapping position parameters;

[0021] S302: Call the entity mapping position parameter, locate the corresponding sentence fragment in the original prompt word, extract the statement unit according to the word order number, convert each statement unit into a basic semantic vector, and generate a semantic structured vector sequence according to the combination to construct a structured semantic sequence;

[0022] S303: Invoke the semantic structured vector sequence, combine it with the image embedding vector, compare the semantic vectors, calculate and extract the difference sequence, and generate the semantic reconstruction offset.

[0023] Optionally, in S4, invoke the semantic reconstruction offset, locate the abnormal segment in the prompt word according to the offset range, extract the position, connection node, and span data of the statement structure path, calculate the path density, and combine the proportion of instruction phrase to generate the risk weight value, including:

[0024] S401: Invoke the semantic reconstruction offset, extract the abnormal semantic vector according to the preset semantic offset critical threshold, and extract the text segment content corresponding to the original prompt word according to the index value corresponding to the abnormal vector, and generate the abnormal segment positioning interval;

[0025] S402: Based on the abnormal segment positioning interval, extract the structure path of the statements within the interval, invoke the position numbers of each statement node in the structure path and the connection relationship between nodes, combine the node connection relationship and position data, calculate the node span value of the statement structure, and generate the structure path span numerical value;

[0026] S403: Invoke the structure path span numerical value, obtain the path density coefficient according to the number of node connections and path span of each path, and combine the proportion of instruction - type keywords in each path to generate the risk weight value.

[0027] Optionally, the specific formula for calculating the node span value of the statement structure is the following formula (1):

[0028] (1)

[0029] Where, represents the structure path span numerical value, represents the position number of the end node of the q - th connection relationship, represents the position number of the start node of the q - th connection relationship, represents the syntax connection weight of the q - th connection relationship, represents the total number of connection relationships in the structure path, represents the structural position offset of the i - th node, represents the total number of nodes within the path, represents the number of the i - th node within the path, q represents the number of the q - th connection relationship within the path.

[0030] Optionally, the method further includes:

[0031] S5: Use the risk weight value to identify the mapping nodes corresponding to the risk path. Detect the semantic distortion path through the fluctuation difference between the semantic vector and the residual information, mark it as abnormal, abort the model output, send a prompt message, and generate path intervention parameters;

[0032] The path intervention parameters include a path abort flag, an abnormal node index, and response prompt content.

[0033] Optionally, the step of using the risk weight value in S5 to identify the mapping nodes corresponding to the risk path, detecting the semantic distortion path through the fluctuation difference between the semantic vector and the residual information, marking it as abnormal, aborting the model output, sending a prompt message, and generating path intervention parameters includes:

[0034] S501: Invoke the risk weight value. According to a preset path risk threshold, identify the risk path, extract the node numbers and structural positions of the path, establish a node mapping relationship, and generate a risk path mapping node sequence;

[0035] S502: Based on the risk path mapping node sequence, extract the semantic vector data and corresponding residual information of each node. Detect the semantic distortion path by calculating the semantic deviation anomaly index of each node, and generate a semantic distortion path mark;

[0036] S503: Invoke the semantic distortion path mark, establish a path anomaly flag, abort the model output, send a response message for the path anomaly, and generate path intervention parameters.

[0037] Optionally, the specific formula for calculating the semantic deviation anomaly index of each node is the following formula (2):

[0038] (2)

[0039] Where represents the semantic deviation anomaly index of the a-th node, represents the semantic vector strength value of the a-th node, represents the residual information strength value of the a-th node, represents the deviation value between the semantic vector and the residual of the b-th node, represents the average value of the semantic vector strengths of all nodes in the risk path mapping node sequence, represents the total number of nodes in the risk path mapping node sequence, b represents the number of the b-th node participating in the calculation in the risk path mapping node sequence, and a represents the number of the single node being processed when calculating the semantic deviation anomaly index.

[0040] The beneficial effects brought by the technical solutions provided in the embodiments of the present invention at least include:

[0041] Enhance the cross-modal correlation analysis ability by splitting the semantic density of multi-modal tensors and mapping in the image space, accurately identify the text-image hybrid attack path, combine the entity mapping relationship between pixel gradients and regional masks, achieve two-way verification of text-image semantic consistency, eliminate the semantic fault vulnerability of single-modal detection, adopt a fusion mechanism of path density and dynamic weights of instruction words to quantify the abnormal risk level of sentence structures, break through the adaptability limitation of static rules to semantic variations, and reduce the probability of outputting illegal content. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0043] Figure 1 is a flowchart of a large model defense method based on multi-perspective text-image conversion provided by an embodiment of the present invention;

[0044] Figure 2 is a block diagram of a large model defense device based on multi-perspective text-image conversion provided by an embodiment of the present invention;

[0045] Figure 3 is a schematic structural diagram of a large model defense device based on multi-perspective text-image conversion provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0046] The following describes the technical solutions in the present invention with reference to the drawings.

[0047] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of the word "example" aims to present concepts in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.

[0048] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, their intended meanings are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, their intended meanings are the same.

[0049] In the embodiments of the present invention, sometimes subscripts such as W1 It may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning to be expressed is the same.

[0050] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0051] Please refer to Figure 1 , the present invention provides a technical solution, a large model defense method based on multi-view text-image conversion, including the following steps:

[0052] S1: Obtain the text of the text-image dialogue prompt, extract the syntactic structure parameters and semantic parameters of the sentence, splice them into a multi-modal tensor, and expand the spliced multi-modal tensor into multiple structure vectors according to the semantic relationship density to form a semantic structure vector group.

[0053] Optionally, the semantic structure vector group includes semantic level labels, position index parameters and structural dependency markers. The image generation matrix is specifically a regional feature density map, a channel index identifier and a generated coordinate label. The semantic reconstruction offset includes semantic segment mapping differences, sentence structure offset sections and vector residual indicators. The risk weight value is specifically a path density factor, an induced component proportion coefficient and a structural span ratio. The path intervention parameter includes a path termination marker, an abnormal node index and a response prompt content.

[0054] Optionally, the specific operation steps of S1 include S101 - S103:

[0055] S101: Obtain the text of the text-image dialogue prompt, extract the part-of-speech type, position number and syntactic dependency relationship of each word, establish a lexical dependency mapping index sequence according to the part-of-speech label and word order position, and generate a syntactic structure index sequence value.

[0056] In a feasible implementation manner, after obtaining the text of the graphic and text dialogue prompt, each word in the text is identified and extracted for its part-of-speech type, position number in the text, and the corresponding syntactic dependency relationship. First, each word in the prompt text is classified according to the language rules. The part-of-speech numbers 1 to 5 are assigned to the five categories of nouns, verbs, adjectives, adverbs, and prepositions respectively. The example text "Generate a landscape picture of blue sky and white clouds" is processed in sequence. Among them, "Generate" is identified as a verb, and the part-of-speech number is 2. "A" is a quantifier phrase and can be classified as an adverb category, numbered 4. "Of blue sky and white clouds" is treated as a modifying attributive as a whole, classified as an adjective, numbered 3. "Landscape picture" is a noun, numbered 1. After completing the part-of-speech numbering, a position number is assigned to each word in the text order, numbered continuously from 1 to n. Among them, "Generate" is numbered 1, "A" is numbered 2, "Of blue sky and white clouds" is numbered 3, and "Landscape picture" is numbered 4. Subsequently, based on the syntactic dependency relationship, the subject-predicate, verb-object, attributive-middle, etc. dependency connections between words are analyzed in sequence. It is clear that "Generate" is the core predicate, dominating "Landscape picture", "Of blue sky and white clouds" modifies "Landscape picture", and "A" modifies "Landscape picture" quantitatively. According to the extracted part-of-speech number and text position number, the lexical index value of each word is calculated. The following calculation formula (0-1) is specifically adopted:

[0057] (0-1)

[0058] Among them, is the index value of the i-th word, is the part-of-speech number (noun 1, verb 2, adjective 3, adverb 4, preposition 5), is the sequential number of the word in the text.

[0059] It is set that the part-of-speech number of "Generate" is 2 and the position number is 1. The calculation is as follows:

[0060] ;

[0061] The part-of-speech number of "Landscape picture" is 1 and the position number is 4. The calculation is as follows:

[0062] ;

[0063] The remaining words are calculated in sequence, and the extraction of the index values of all words is gradually completed. Arranged according to the word order, the syntactic structure index sequence value is finally summarized and formed.

[0064] S102: Call the syntactic structure index sequence value, extract the semantic label and semantic classification number corresponding to each word in the prompt, and perform tensor combination splicing according to the semantic label and syntactic mapping index to obtain the semantic combination tensor.

[0065] In a feasible implementation, after calling the syntactic structure index sequence value, the semantic tags and corresponding semantic classification numbers of each word in the prompt word are extracted one by one. Based on the four major categories of static objects, dynamic behaviors, spatial positions, and emotional colors, they are respectively assigned values from 1 to 4. Specifically, in the example text, "generate" corresponds to the dynamic behavior category with a number of 2, "scenery picture" belongs to the static object category with a number of 1, "blue sky and white clouds" is an adjectival modifier and belongs to the static object category with a number of 1, and "a" belongs to the quantitative modifier and does not involve a specific semantic category, with a default number of 3. After extraction, for each word, based on the previously obtained syntactic structure index sequence value and the currently extracted semantic classification number, a tensor concatenation operation is performed. According to the set rule, the index value is multiplied by 10 and then added to the semantic category number to ensure the combination of the index and the semantics. The calculation formula is as follows in formula (0-2):

[0066] (0-2)

[0067] Among them, is the semantic combination value of the i-th word, is the index value of this word, is the semantic category number of this word (static object 1, dynamic behavior 2, spatial position 3, emotional color 4).

[0068] Set the index value of "generate" to 2001 and the semantic number to 2. Substituting into the formula gives:

[0069] ;

[0070] Set the index value of "scenery picture" to 1004 and the semantic number to 1. Substituting into the formula gives:

[0071] ;

[0072] Set the index of "blue sky and white clouds" to 3003 and the semantic number 1, and get:

[0073] ;

[0074] Set the index of "‘a’" to 4002 and the semantic number 3, and get:

[0075] ;

[0076] And so on. After completing the concatenation of all words, a complete semantic combination tensor is formed, providing a data basis for subsequent analysis.

[0077] S103: According to the semantic combination tensor, according to the vector density and the frequency of repeated occurrence of the semantic tags, a semantic relationship density splitting benchmark value is set, and according to the semantic distribution range, the combination tensor is expanded into multiple structure vectors to generate a semantic structure vector group.

[0078] In a feasible implementation, according to the generated semantic combination tensor, each combination value is traversed in turn to evaluate the semantic density in the vector. A window sliding method is adopted to set 5 consecutive words for each evaluation, and the number of occurrences of the same semantic category number is counted to determine the semantic aggregation feature of the current segment. The semantic density calculation formula is as follows (0-3):

[0079] (0-3)

[0080] Where D is the semantic density, is the cumulative number of occurrences of the same semantic category number in the window, is the sliding window size, fixed at 5.

[0081] In the setting window "20041, 15031, 20041, 15031, 20041", assume that the number of occurrences of semantic number 1 (static object) is 4, substitute into the formula:

[0082] ;

[0083] If the calculated density D is greater than or equal to the preset benchmark value of 0.6, the window is judged to be a high-density area. The benchmark value comes from the static analysis. If the concentration of semantic categories in the window exceeds 60%, it is easy to form semantic focus and meet the standard of splitting into independent structural vector segments. After completing a round of window calculation, the high-density blocks are split separately, their start and end positions are marked, and the semantic categories they belong to are recorded. The whole sentence processing is completed in turn, and finally summarized and generated into a semantic structure vector group.

[0084] S2: Call the semantic structure vector group, analyze the semantic level value of each vector and divide it into multiple subsets according to the level range, input the information of multiple subsets into the image generation channel respectively, locate the region in the image space according to the vector offset angle and activation weight, construct a significant feature distribution map, and form an image generation matrix.

[0085] Optionally, the specific operation steps of S2 include S201-S203:

[0086] S201: calling the semantic structure vector group, extracting the semantic level label and the corresponding structure index value of each vector, analyzing the semantic level of each vector, and dividing it into multiple subsets according to the level interval, and generating a semantic level division cluster value.

[0087] In a feasible implementation, after calling the semantic structure vector group, the semantic level label and structure index value corresponding to each vector are extracted one by one. The semantic level label is directly obtained by extracting the semantic category number corresponding to the vector in the previously generated semantic structure vector group. The structure index value is the position index number of the semantic vector in the original prompt text. Subsequently, a numerical size analysis is performed on the semantic level value of each vector, and the semantic level value is set within a range of 1 to 10, where the value 1 represents the lowest semantic intensity and the value 10 represents the highest semantic intensity. Taking the example of the graphic dialogue prompt "Generate a landscape picture of blue sky and white clouds" as an example, assume that the semantic category of "Generate" is a dynamic behavior with a level value of 8, "landscape picture" is a static object with a level value of 7, and "of blue sky and white clouds" is a descriptive modifier with a level value of 6. After obtaining these level values, the level numerical values of all vectors are further calculated using the following formula (0-4):

[0088] (0-4)

[0089] Among them, is the average level numerical value of the semantic level values of all vectors, is the semantic level value of the th vector, n represents the total number of vectors. Assume that there are 4 vectors in the current vector group, and the level values are 8, 6, 7, and 5 respectively. Then the average level calculation is as follows:

[0090] ;

[0091] According to the average level value calculated in this way, the level interval division reference value is set, specifically: it is determined that the semantic level is in the high-level interval if it is greater than or equal to the average level value of 6.5, and in the low-level interval if it is lower than 6.5. Judging according to this rule, "Generate" with a set semantic level of 8 is classified into the high-level subset, while "of blue sky and white clouds" with a semantic level of 6 and "a" with a semantic level of 5 are classified into the low-level subset. Finally, the subsets to which all vectors belong are statistically obtained, and the semantic level division cluster value is generated.

[0092] S202: According to the semantic level division cluster value, extract the vector order and structure index number corresponding to each subset, and input them into the image generation channel. According to the offset angle value and activation weight parameter of each vector, set the image area positioning index, and obtain the image area positioning index value.

[0093] In a feasible implementation, cluster values are divided according to semantic levels, and the order information and structural index numbers of the vectors included in each subset are extracted one by one. The vector order information is the order of the vectors in the original prompt text, and the structural index number is the lexical position index value obtained in the previous step. It is set in the above-mentioned graphic and text dialogue prompt setting. Assume that the high-level subset includes "generate" with a position index of 2001 and "landscape picture" with a position index of 1004. After extracting these two indexes, they are sorted in ascending order according to their positions in the original text. 2001 of "generate" is in the front, and 1004 of "landscape picture" is in the back, obtaining an order sequence of [1004, 2001]. Subsequently, this sequence is input into the image generation channel. For each semantic vector corresponding to the structural index, the offset angle value and activation weight parameter of the vector are called. The offset angle value represents the spatial position angle of the vector relative to the center of the image space, and the activation weight parameter represents the weighted value of the semantic intensity. The specific image region positioning index is calculated using the following formula (0-5):

[0094] (0-5)

[0095] Among them, is the image region positioning index value, is the offset angle value of the vector, is the activation weight parameter of the vector. Taking "generate" as an example, assume that the offset angle value of the vector is 45 degrees and the activation weight parameter is 1.2. Substituting into the formula for calculation:

[0096] ;

[0097] Similarly, assume that the offset angle value of "landscape picture" is 30 degrees and the activation weight parameter is 1.0, then the calculation is:

[0098] ;

[0099] After calculating the positioning index values of each vector in turn, the image region positioning index value is obtained.

[0100] S203: Call the image region positioning index value, calculate the activation response intensity value of each region according to the distribution position of each group of position data in the image space, and partition and mark the image space according to the activation response intensity value to construct an image saliency distribution coefficient matrix and generate an image generation matrix.

[0101] In a feasible implementation, after calling the image region location index value, the image space position data corresponding to each group of indexes is extracted. According to the image space coordinate system, the position corresponding to the index value is mapped into the two-dimensional image space, where the image center coordinates are set as the origin (0,0). Each index value represents the coordinate angle of the corresponding image region. With the radius R fixed as the radius of the image region range (assumed to be 100 pixels), the conversion of the index value into the image space coordinates is calculated using the following formula (0-6):

[0102] (0-6)

[0103] where (x,y) represents the image space coordinates and R represents the fixed radius, is the positioning index angle value obtained from the previous calculation. It is assumed that the region positioning index value for "generation" is 54, and substituting it into the formula to calculate the coordinates:

[0104] ;

[0105] Subsequently, the activation response intensity value is calculated for each region. The intensity value depends on the distance relationship between the index angle and the image center. The closer the distance, the higher the activation intensity, and vice versa. It is assumed that the activation response intensity value is set using a linear proportional relationship, and the specific calculation is as follows in formula (0-7):

[0106] (0-7)

[0107] where, is the activation response intensity value, d represents the Euclidean distance between the region coordinates and the image center (0,0). For example, the Euclidean distance of the "generation" region coordinates (58.78, 80.90) is calculated as:

[0108] ;

[0109] Substituting into the formula:

[0110] ;

[0111] Using this method, the response intensity value of each region is calculated step by step. Regions with an intensity greater than or equal to 0.6 are classified as high-response regions, and those less than 0.6 are classified as low-response regions. After completing the partition marking, finally, an image saliency distribution coefficient matrix is constructed to generate an image generation matrix.

[0112] S3: Call the image generation matrix, extract the image pixel density, edge gradient, and region mask data. Calculate the entity mapping relationship through the intersection of pixel and semantic annotation, split and recombine the corresponding sentence fragments to generate semantic vectors, and compare with the image embedding vector to calculate the difference sequence to obtain the semantic reconstruction offset.

[0113] Optionally, the specific operation steps of S3 include S301 - S303:

[0114] S301: Invoke the image generation matrix, extract the pixel density value, edge gradient value, and region mask identifier of the image. Locate the semantic annotation region according to the mask identifier, perform coordinate comparison on the pixel coordinates and edge positions of the annotation region, obtain the overlapping coordinate points of the image region and semantic annotation, and generate the entity mapping position parameters.

[0115] In a feasible implementation, after invoking the image generation matrix, first extract the pixel density value, edge gradient value, and region mask identifier of each region in the matrix. The pixel density value refers to the number of valid pixels in the unit image region, the edge gradient value is the amplitude of pixel gray - level change in the region, and the region mask identifier represents the existence state of each region in the semantic annotation map. Subsequently, extract the corresponding semantic annotation region according to the region mask identifier, and extract the pixel coordinates and edge position coordinates in each region. For all the extracted coordinate data, perform a point - by - point comparison operation. During the comparison process, use the Euclidean distance threshold method to screen the overlapping points. Set the distance threshold to 3 pixels, and the judgment condition is that when the distance between the pixel coordinates in the image region and the edge coordinates of the semantic annotation region is less than or equal to 3 pixels, it is determined as a valid overlapping point. The distance calculation formula is as follows (0 - 8):

[0116] (0 - 8)

[0117] where d is the distance between two coordinate points, is the pixel point coordinate of the image region, is the semantic annotation edge point coordinate. Using example values, the pixel point coordinate is (120, 80), and the annotation edge point is (122, 82). The calculation is as follows:

[0118] ;

[0119] The calculation result is less than 3, so it is determined as a valid overlapping point. Process all pixels and annotation edge points in turn, count all the overlapping coordinate points that meet the conditions, and extract their region indexes and specific coordinates, and finally form the complete entity mapping position parameters.

[0120] S302: Invoke the entity mapping position parameters, locate the corresponding sentence fragments in the original prompt, extract the statement units according to the word order numbers, convert each statement unit into a basic semantic vector, and construct a structured semantic sequence according to the combination to generate a semantic structured vector sequence.

[0121] In a feasible implementation, after invoking the entity mapping position parameter, extract the region indexes associated with all overlapping coordinate points and locate the corresponding sentence fragments in the original prompt. The recognition of sentence fragments is completed through the index mapping relation table. Match the region indexes in the entity mapping position parameter with the word order numbers in the original prompt one by one, and extract each corresponding statement unit. After extraction, arrange the word order numbers in sequence to form a continuous sequence of statement units. Suppose the region indexes corresponding to the entity mapping are 5, 8, and 10, which respectively correspond to the 1st, 2nd, and 4th statement units in the original prompt. Subsequently, extract these statement units in sequence: "generate", "a", "scenery picture". For each statement unit, convert it into a basic semantic vector based on its part of speech and dependency relationship. The vector dimension is set to three-dimensional, representing action intensity, object entity weight, and spatial attribute weight respectively, and the numerical range of all is 0 - 10. Suppose "generate" is used as a verb, with an action intensity of 9, an object entity weight of 0, and a spatial attribute weight of 0; "scenery picture" is used as a noun, with an action intensity of 0, an object entity weight of 8, and a spatial attribute weight of 2. Subsequently, splice all the basic semantic vectors in sequence to combine and construct a structured semantic sequence, and finally obtain a semantic structured vector sequence.

[0122] S303: Invoke the semantic structured vector sequence, combine it with the image embedding vector, compare the semantic vectors, calculate and extract the difference sequence, and generate a semantic reconstruction offset.

[0123] In a feasible implementation, after invoking the semantic structured vector sequence, extract each semantic vector in sequence and compare it with the image embedding vector. The dimension of the image embedding vector is the same as that of the semantic vector, including three items: action intensity, object entity weight, and spatial attribute weight. Calculate the difference between the two item by item, using the absolute difference method. The difference calculation formula is as follows in formula (0 - 9):

[0124] (0 - 9)

[0125] Among them, is the difference of the corresponding item, is the numerical value of the corresponding item of the semantic vector, is the numerical value of the corresponding item of the image embedding vector. Suppose the semantic vector of "generate" is [9, 0, 0], and the image embedding vector is [7, 1, 0], then the difference of the first dimension is:

[0126] ;

[0127] The second dimension is:

[0128] ;

[0129] The third dimension is:

[0130] ;

[0131] Calculate the differences between all semantic vectors and the corresponding image embedding vectors in sequence, summarize to obtain a complete difference sequence, and finally generate a semantic reconstruction offset.

[0132] S4: Invoke the semantic reconstruction offset. According to the offset range, locate the abnormal segments in the prompt words, extract the positions, connection nodes, and span data of the statement structure path, calculate the path density, and combine with the proportion of instruction phrase to generate a risk weight value.

[0133] Optionally, the specific operation steps of S4 include S401 - S403:

[0134] S401: Invoke the semantic reconstruction offset. According to a preset semantic offset critical threshold, extract abnormal semantic vectors, and according to the index value corresponding to the abnormal vector, extract the text segment content corresponding to the original prompt word to generate an abnormal segment positioning interval.

[0135] In a feasible implementation manner, after invoking the semantic reconstruction offset, extract the difference values of each dimension of each semantic vector one by one, and compare the difference values with the preset semantic offset critical threshold. The specific setting method of the critical threshold is to calculate the average value of the overall values of the difference sequence and add a deviation coefficient. The formula is expressed as the following formula (0 - 10):

[0136] (0 - 10)

[0137] Among them, represents the semantic offset critical threshold, represents the th value in the difference sequence, represents the total number of differences in the difference sequence, k is the deviation coefficient, and its value is 20% of the average difference to ensure the stability of the threshold. For example, if the values of the difference sequence are [1.2, 1.5, 2.3, 2.0], calculate the average difference as

[0138] , then the deviation coefficient , substitute into the formula to get the threshold as:

[0139] ;

[0140] Subsequently, each vector difference is compared with this threshold one by one to determine whether the difference exceeds the threshold of 2.10. If the difference exceeds the threshold, the corresponding semantic vector is determined to be an abnormal semantic vector. Suppose the semantic vector with a difference of 2.3 significantly exceeds the threshold, then it is determined to be an abnormal vector, and the index value corresponding to the abnormal vector is extracted. Assume the index of the abnormal semantic vector is "2001", and the text fragment content corresponding to this index in the original prompt is called. If the text corresponding to the index "2001" is "generate", then locate the position of this text in the original prompt, and determine its starting number and ending number positions to obtain the complete abnormal fragment positioning interval.

[0141] S402: Based on the abnormal fragment positioning interval, extract the structural path of the statements within the interval, call the position numbers of each statement node in the structural path and the connection relationships between the nodes, and combine the node connection relationships and position data to calculate the node span value of the statement structure and generate the structural path span numerical value.

[0142] In a feasible implementation manner, the specific formula for calculating the node span value of the statement structure is the following formula (1):

[0143] (1)

[0144] Where, represents the structural path span numerical value, represents the position number of the end node of the q-th connection relationship, represents the position number of the start node of the q-th connection relationship, represents the syntactic connection weight of the q-th connection relationship, represents the total number of connection relationships in the structural path, represents the structural position offset of the i-th node, represents the total number of nodes in the path, represents the number of the i-th node in the path, and q represents the number of the q-th connection relationship in the path.

[0145] Detailed explanation of formula (1) and the derivation process of formula calculation:

[0146] The formula is used to calculate the structural path span numerical value, and the obtained result is used to quantify the span characteristics of the sentence structural path in the abnormal fragment, providing a basis for subsequent density and risk weight calculations;

[0147] Parameter meanings and setting values:

[0148] is the position number of the end node of the q-th connection relationship, set to 14, 19, 25;

[0149] The starting node position number of the q-th connection relationship is set to 8, 14, 19;

[0150] The syntactic connection weight of the q-th connection relationship. The subject-predicate relationship weight is set to 2.0, the verb-object relationship is set to 1.8, and the attributive-middle relationship is set to 1.5;

[0151] The total number of connection relationships in the structural path is set to 3;

[0152] is the structural position offset of the

[0153] -th node. The node position monitoring values are set to 8, 14, 19, 25, and the offsets are 0, 6, 11, 17 in sequence;

[0154] Substitute the parameters into the formula for calculation:

[0155] ;

[0156] ;

[0157] ;

[0158] ;

[0159] ;

[0160] ;

[0161] ;

[0162] ;

[0163] The result 8.12 indicates that the structural path span value is relatively high, indicating that the internal node connection relationship of the abnormal segment has a large span and the structural distribution is discrete. The result is an important basic parameter for subsequent risk density calculation, reflecting the overall expansibility and complexity of the current structural path.

[0164] S403: Invoke the structural path span value, obtain the path density coefficient according to the node connection number and path span of each path, and generate a risk weight value in combination with the proportion of instruction type keywords in each path.

[0165] In a feasible implementation, after calling the structural path span value, the number of node connections and the node span value in each path are extracted. For the extraction of the number of node connections, it is the total number of node connection relationships contained in each structural path. Subsequently, the density coefficient of each path is calculated one by one. The density coefficient is the ratio between the number of node connections and the path span value, and the formula is expressed as the following formula (1-1):

[0166] (1-1)

[0167] Among them, is the path density coefficient, is the number of path node connections, is the node span value. Continuing with the above example, the number of node connections is 3 (node 1 4, node 3 4, node 2 4), and the total node span value is (3 + 1 + 2) = 6. The density coefficient is calculated as:

[0168] ;

[0169] Subsequently, the number of instruction - type keyword quantities included in each path is called. Instruction - type keywords specifically include words with obvious indicating actions such as "generate", "produce", "display", etc. Assuming that the current path contains a total of 1 instruction - type keyword ("generate") and the total number of path nodes is 4, the proportion of instruction - type keywords is calculated as the following formula (1-2):

[0170] (1-2)

[0171] Among them, is the proportion of instruction - type keywords, is the number of instruction - type keywords in the path, is the total number of path nodes. Substituting the values for calculation:

[0172] ;

[0173] Finally, the path density coefficient and the proportion of instruction - type keywords are called, and the risk weight value is calculated one by one. The risk weight value is obtained through the weighted sum of the density coefficient and the proportion of instructions. The formula is expressed as the following formula (1-3):

[0174] (1-3)

[0175] Among them, is the risk weight value, is the path density coefficient, is the proportion of instruction keywords, and the weight parameter and respectively represent the path density weight and the instruction ratio weight, which are default set to 0.6 and 0.4 respectively. Substitute the values:

[0176] ;

[0177] Gradually complete all path risk weight calculations and finally generate risk weight values.

[0178] S5: Use the risk weight value to identify the mapping nodes corresponding to the risk paths. Through the fluctuation difference between the semantic vector and the residual information, detect the semantic distortion paths and mark them as abnormal, abort the model output and send a prompt message, and generate path intervention parameters.

[0179] Optionally, the specific operation steps of S5 include S501 - S503:

[0180] S501: Call the risk weight value. According to the preset path risk threshold, identify the risk paths and extract the node numbers and structural positions of the paths, establish node mapping relationships, and generate a sequence of risk path mapping nodes.

[0181] In a feasible implementation manner, after calling the risk weight value, first extract the risk weight values of each path and compare them numerically with the preset path risk threshold. The path risk threshold is set to the average value of each path risk weight plus the fluctuation adjustment coefficient. The calculation formula is the following formula (1 - 4):

[0182] (1 - 4)

[0183] Among them, is the path risk threshold, is the risk weight value of the a - th path, is the total number of paths, is the fluctuation adjustment coefficient, set to 15% of the average value. Suppose there are three paths with risk weights of 0.4, 0.6, and 0.7 respectively. Calculate the average value as , then the fluctuation adjustment coefficient , and the final threshold is calculated as:

[0184] ;

[0185] Subsequently, the risk weights of each path are compared with the threshold one by one. The paths with risk weights greater than the threshold are identified as risk paths. For example, the weight of the third path is 0.7, which is higher than the threshold of 0.6517, so it is determined as a risk path. The node numbers and structural positions of this path are extracted. The node number is the unique number of all participating nodes in the path, and the structural position is the specific serial number position of the node in the original prompt text. The node numbers are set as [5, 8, 10], corresponding to the three places of "generate", "blue sky and white clouds", and "scenery picture" in the original text. The corresponding relationship between the nodes and the text is established in order to form a node mapping relationship table, and finally, the risk path mapping node sequence is integrated.

[0186] S502: Based on the risk path mapping node sequence, extract the semantic vector data and corresponding residual information of each node, and detect the semantic distortion path by calculating the semantic deviation anomaly index of each one, and generate a semantic distortion path mark.

[0187] In a feasible implementation manner, the specific formula for calculating the semantic deviation anomaly index of each one is the following formula (2):

[0188] (2)

[0189] Calculate the semantic deviation anomaly index;

[0190] Among them, represents the semantic deviation anomaly index of the a-th node, represents the semantic vector strength value of the a-th node, represents the residual information strength value of the a-th node, represents the deviation value between the semantic vector and the residual of the b-th node, represents the average value of the semantic vector strengths of all nodes in the risk path mapping node sequence, represents the total number of nodes in the risk path mapping node sequence, b represents the serial number of the b-th participating node in the risk path mapping node sequence, and a represents the serial number of the single node being processed when calculating the semantic deviation anomaly index.

[0191] Detailed explanation of formula (2) and formula calculation derivation process:

[0192] The formula is used to calculate the semantic deviation anomaly index, and the result is used to detect the semantic distortion nodes in the risk path;

[0193] Parameter meaning and setting value:

[0194] is the semantic vector strength value of the a-th node, and it is set as , reflecting the comprehensive semantic strength of the node;

[0195] The residual information intensity value for the a-th node, set to 2.7, reflects the semantic deviation amplitude during the generation process;

[0196] It is the deviation value between the semantic vector and the residual of the b-th node in the risk path mapping node sequence. The values of nodes b1, b2, b3, b4, and b5 are set to 1.8, 2.5, 1.2, 3.1, and 2.0 respectively;

[0197] It is the total number of nodes in the risk path mapping node sequence, set to 5;

[0198] It is the average value of the semantic vector intensities of all nodes in the risk path mapping node sequence, set to 5.2, 4.8, 6.1, 5.7, and 4.9 respectively, ;

[0199] Substitute the parameters into the formula for calculation: ; ; ; ;

[0200] The result 2.05 indicates that the semantic deviation anomaly index has exceeded the safety threshold of 1.5, indicating that the semantic deviation degree corresponding to the a-th node is serious. The numerical result directly drives the confirmation of the distortion path and the subsequent abort output strategy.

[0201] S503: Invoke the semantic distortion path marker, establish a path anomaly identifier and abort the model output, send a response message for the path anomaly, and generate path intervention parameters.

[0202] In a feasible implementation, after invoking the semantic distortion path marker, the abnormal processing process is sequentially executed for all marked paths. First, a path anomaly identifier is established for the detected distortion path. The format of the anomaly identifier is the path number plus the "anomaly" label. For example, if the path is Path3, it is marked as "Path3 anomaly". Subsequently, all nodes within the path are located, and the subsequent large model output is forcibly aborted. The current path is removed from the generation process. A response message is generated for the abnormal path. The content of the response message includes the path number, the involved nodes, and the reason description text. For example, "Path3, involving nodes generation, blue sky and white clouds, with serious semantic deviation, determined as abnormal" is generated. At the same time, the path intervention parameters are calculated. The path intervention parameters are generated based on the number of nodes and the average anomaly index. The calculation formula is as follows in Equation (3):

[0203] (3)

[0204] Among them, is the path intervention parameter, is the total number of nodes in the abnormal path, is the average abnormal index of nodes. Assuming the number of nodes is 3 and the abnormal indices are 0.667, 1, and 0.4 respectively, the average abnormal index is:

[0205] ;

[0206] The final intervention parameter is:

[0207] ;

[0208] The numerical value of the path intervention parameter is recorded in the system log, and all processing procedures are completed to generate the path intervention parameter.

[0209] In the embodiment of the present invention, by using the semantic density splitting and image space mapping of the multi-modal tensor, the cross-modal correlation analysis ability is enhanced, the text-image mixed attack path is accurately identified, and the entity mapping relationship between the pixel gradient and the region mask is combined to realize the two-way verification of text-image semantic consistency, eliminate the semantic fault vulnerability of single-modal detection, adopt the path density and instruction word dynamic weight fusion mechanism to quantify the abnormal risk level of the sentence structure, break through the adaptability limitation of static rules to semantic variation, and reduce the probability of output of illegal content.

[0210] Figure 2 is a block diagram of a large model defense device based on multi-perspective text-image conversion provided by an embodiment of the present invention. This device is used for a large model defense method based on multi-perspective text-image conversion. Referring to Figure 2 , this device includes an extraction and splicing unit 210, an analysis and positioning unit 220, a splitting and comparison unit 230, a generation unit 240, and a detection unit 250. Among them:

[0211] The extraction and splicing unit 210 is used to obtain the text of the text-image dialogue prompt, extract the syntactic structure parameters and semantic parameters of the sentence, splice them into a multi-modal tensor, and expand the spliced multi-modal tensor into multiple structure vectors according to the semantic relationship density to form a semantic structure vector group;

[0212] The analysis and positioning unit 220 is used to call the semantic structure vector group, analyze the semantic grade value of each vector and divide it into multiple subsets according to the grade interval, input the information of multiple subsets into the image generation channel respectively, and perform regional positioning on the image space according to the vector offset angle and activation weight to construct a saliency feature distribution map and form an image generation matrix;

[0213] The splitting and comparison unit 230 is used to call the image generation matrix, extract image pixel density, edge gradient, and region mask data, calculate the entity mapping relationship through the intersection of pixel and semantic annotation, split and recombine the corresponding sentence fragments to generate semantic vectors, compare the image embedding vectors to calculate the difference sequence, and obtain the semantic reconstruction offset;

[0214] The generation unit 240 is used to call the semantic reconstruction offset, locate the abnormal fragments in the prompt words according to the offset range, extract the position, connection nodes, and span data of the statement structure path, calculate the path density, and generate a risk weight value in combination with the proportion of instruction phrases;

[0215] The detection unit 250 is used to utilize the risk weight value to identify the mapping nodes corresponding to the risk paths, detect the semantic distortion paths through the fluctuation difference between the semantic vectors and the residual information, mark them as abnormal, abort the model output and send a prompt message, and generate path intervention parameters.

[0216] In the embodiments of the present invention, by using the semantic density splitting of multi-modal tensors and the image space mapping, the cross-modal correlation analysis ability is enhanced, the text-image hybrid attack paths are accurately identified, the entity mapping relationship between pixel gradients and region masks is combined to achieve two-way verification of text-image semantic consistency, eliminate the semantic fault holes in single-modal detection, adopt a mechanism that fuses path density and dynamic weights of instruction words to quantify the abnormal risk level of the statement structure, break through the adaptability limitations of static rules to semantic variations, and reduce the probability of outputting illegal content.

[0217] Figure 3 It is a schematic structural diagram of a large model defense device based on multi-perspective text-image conversion provided by the embodiments of the present invention. As Figure 3 shown, the large model defense device based on multi-perspective text-image conversion may include the above-mentioned Figure 2 large model defense device shown based on multi-perspective text-image conversion. Optionally, the large model defense device 310 based on multi-perspective text-image conversion may include a first processor 2001.

[0218] Optionally, the large model defense device 310 based on multi-perspective text-image conversion may further include a memory 2002 and a transceiver 2003.

[0219] Among them, the first processor 2001, the memory 2002, and the transceiver 2003 may be connected through a communication bus, for example.

[0220] Next, in combination with Figure 3 each component of the large model defense device 310 based on multi-perspective text-image conversion will be specifically introduced:

[0221] Among them, the first processor 2001 is the control center of the large model defense device 310 based on multi-perspective graphic-text conversion, which can be a single processor or a collective term for multiple processing elements. For example, the first processor 2001 is one or more central processing units (CPUs), or can be an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention, such as: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).

[0222] Optionally, the first processor 2001 can execute various functions of the large model defense device 310 based on multi-perspective graphic-text conversion by running or executing software programs stored in the memory 2002 and calling data stored in the memory 2002.

[0223] In a specific implementation, as an embodiment, the first processor 2001 can include one or more CPUs, such as Figure 3 the CPU0 and CPU1 shown in

[0224] In a specific implementation, as an embodiment, the large model defense device 310 based on multi-perspective graphic-text conversion can also include multiple processors, such as Figure 3 the first processor 2001 and the second processor 2004 shown in. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).

[0225] Among them, the memory 2002 is used to store software programs for executing the solution of the present invention and is controlled by the first processor 2001 for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated here.

[0226] Optionally, the memory 2002 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2002 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 3 not shown) of the multi-perspective graphic-text conversion based large model defense device 310. The embodiments of the present invention do not make specific limitations in this regard.

[0227] The transceiver 2003 is used to communicate with a network device or communicate with a terminal device.

[0228] Optionally, the transceiver 2003 may include a receiver and a transmitter ( Figure 3 not shown separately). Among them, the receiver is used to implement the receiving function, and the transmitter is used to implement the sending function.

[0229] Optionally, the transceiver 2003 may be integrated with the first processor 2001 or may exist independently and be coupled to the first processor 2001 through an interface circuit ( Figure 3 not shown) of the multi-perspective graphic-text conversion based large model defense device 310. The embodiments of the present invention do not make specific limitations in this regard.

[0230] It should be noted that Figure 3 the structure of the multi-perspective graphic-text conversion based large model defense device 310 shown does not constitute a limitation on the router. The actual knowledge structure recognition device may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0231] In addition, the technical effects of the multi-perspective graphic-text conversion based large model defense device 310 may refer to the technical effects of the multi-perspective graphic-text conversion based large model defense method described in the above method embodiments, and will not be elaborated here.

[0232] It should be understood that the first processor 2001 in the embodiments of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0233] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM) or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM) and direct rambus RAM (DR RAM).

[0234] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, or magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0235] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood by referring to the context before and after.

[0236] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0237] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0238] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0239] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the devices, apparatuses, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated herein.

[0240] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be electrical, mechanical, or other forms.

[0241] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0242] In addition, the functional units in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0243] When the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0244] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A large model defense method based on multi-view image-text conversion, characterized in that: The method comprises: S1: Obtain the prompt word text of the picture-text dialogue, extract the syntactic structure parameters and semantic parameters of the sentence, and splice them into a multimodal tensor. Expand the spliced ​​multimodal tensor into multiple structure vectors according to the semantic relationship density to form a semantic structure vector group; S2: calling the semantic structure vector group, analyzing the semantic level value of each vector and dividing it into multiple subsets according to the level interval, inputting the information of the multiple subsets into the image generation channel respectively, locating the region of the image space according to the vector offset angle and the activation weight, constructing a significant feature distribution map, and forming an image generation matrix; S3: calling the image generation matrix, extracting image pixel density, edge gradient and region mask data, calculating entity mapping relationship through the intersection of pixels and semantic annotations, splitting and reorganizing corresponding sentence fragments to generate semantic vectors, comparing image embedding vectors to calculate difference sequences, and obtaining semantic reconstruction offsets; S4: Call the semantic reconstruction offset, locate the abnormal fragment in the prompt word according to the offset range, extract the position, connection nodes and span data of the sentence structure path, calculate the path density, and generate a risk weight value in combination with the instruction phrase ratio.

2. The large model defense method based on multi-view image-text conversion according to claim 1 is characterized in that: The semantic structure vector group includes semantic level labels, position index parameters and structural dependency markers; the image generation matrix is ​​specifically a regional feature density map, a channel index identifier and a generated coordinate label; the semantic reconstruction offset includes a semantic fragment mapping difference, a sentence structure offset segment and a vector residual indicator; the risk weight value is specifically a path density factor, an induced component ratio coefficient and a structural span ratio.

3. The large model defense method based on multi-view image-text conversion according to claim 1 is characterized in that: S1 obtains the prompt word text of the picture-text dialogue, extracts the syntactic structure parameters and semantic parameters of the sentence, and splices them into a multimodal tensor. According to the semantic relationship density, the spliced ​​multimodal tensor is expanded into multiple structure vectors to form a semantic structure vector group, including: S101: Obtain the prompt word text of the graphic dialogue, extract the part-of-speech type, position number and grammatical dependency relationship of each word, establish a word dependency mapping index sequence according to the part-of-speech tag and word order position, and generate a syntactic structure index sequence value; S102: calling the syntactic structure index sequence value, extracting the semantic label and semantic classification number corresponding to each word in the prompt word, and performing tensor combination and splicing according to the semantic label and the syntactic mapping index to obtain a semantic combination tensor; S103: According to the semantic combination tensor, according to the vector density and the frequency of repeated occurrence of semantic tags, a semantic relationship density splitting benchmark value is set, and according to the semantic distribution range, the combination tensor is expanded into multiple structure vectors to generate a semantic structure vector group.

4. The large model defense method based on multi-view image-text conversion according to claim 1 is characterized in that: S2 calls the semantic structure vector group, analyzes the semantic level value of each vector and divides it into multiple subsets according to the level interval, inputs the information of the multiple subsets into the image generation channel respectively, locates the region of the image space according to the vector offset angle and the activation weight, constructs a significant feature distribution map, and forms an image generation matrix, including: S201: calling the semantic structure vector group, extracting the semantic level label and the corresponding structure index value of each vector, analyzing the semantic level of each vector, and dividing it into multiple subsets according to the level interval, and generating a semantic level division cluster value; S202: Divide the cluster values ​​according to the semantic level, extract the vector sequence and structure index number corresponding to each subset, and input them into the image generation channel, set the image area positioning index according to the offset angle value and activation weight parameter of each vector, and obtain the image area positioning index value; S203: calling the image region positioning index value, calculating the activation response intensity value of each region according to the distribution position of each group of position data in the image space, and marking the image space according to the activation response intensity value, constructing the image significance distribution coefficient matrix, and generating the image generation matrix.

5. The large model defense method based on multi-view image-text conversion according to claim 1 is characterized in that: S3 calls the image generation matrix, extracts the image pixel density, edge gradient and region mask data, calculates the entity mapping relationship through the intersection of pixels and semantic annotations, splits and reorganizes the corresponding sentence fragments to generate semantic vectors, compares the image embedding vectors to calculate the difference sequence, and obtains the semantic reconstruction offset, including: S301: calling the image generation matrix, extracting the pixel density value, edge gradient value and region mask identifier of the image, locating the semantic annotation region according to the mask identifier, performing coordinate comparison between the pixel coordinates and the edge position of the annotated region, obtaining the coincident coordinate points of the image region and the semantic annotation, and generating entity mapping position parameters; S302: calling the entity mapping position parameter, locating the corresponding sentence fragment in the original prompt word, extracting the sentence unit according to the word order number, converting each sentence unit into a basic semantic vector, and constructing a structured semantic sequence according to the combination to generate a semantic structured vector sequence; S303: calling the semantic structured vector sequence, combining with the image embedding vector, comparing the semantic vectors, calculating and extracting the difference sequence, and generating a semantic reconstruction offset.

6. The large model defense method based on multi-view image-text conversion according to claim 1 is characterized in that: S4 calls the semantic reconstruction offset, locates the abnormal fragment in the prompt word according to the offset range, extracts the position, connection nodes and span data of the sentence structure path, calculates the path density, and generates the risk weight value in combination with the instruction phrase proportion, including: S401: calling the semantic reconstruction offset, extracting an abnormal semantic vector according to a preset semantic offset critical threshold, and extracting the text segment content corresponding to the original prompt word according to the index value corresponding to the abnormal vector, and generating an abnormal segment positioning interval; S402: Based on the abnormal fragment positioning interval, extract the structural path of the statement in the interval, call the position number of each statement node in the structural path and the connection relationship between the nodes, combine the node connection relationship and the position data, calculate the node span value of the statement structure, and generate the structural path span value; S403: Call the structural path span value, obtain the path density coefficient according to the number of node connections and the path span of each path, and generate a risk weight value in combination with the proportion of instruction keywords in each path.

7. The large model defense method based on multi-view image-text conversion according to claim 6 is characterized in that: The specific formula for calculating the node span value of the statement structure is as follows (1): (1) in, Represents the structure path span value, Represents the end node position number of the qth connection relationship, Represents the starting node position number of the qth connection relationship, represents the grammatical connection weight of the qth connection relationship, Represents the total number of connection relationships in the structure path, Representative The structural position offset of the node, Represents the total number of nodes in the path, Represents the path The node number is the node number, and q represents the number of the qth connection relationship in the path.

8. The large model defense method based on multi-view image-text conversion according to claim 1 is characterized in that: The method further comprises: S5: using the risk weight value, identifying the mapping node corresponding to the risk path, detecting the semantically distorted path and marking it as abnormal through the fluctuation difference between the semantic vector and the residual information, terminating the model output and sending a prompt message, and generating a path intervention parameter; The path intervention parameters include a path termination mark, an abnormal node index, and a response prompt content.

9. The large model defense method based on multi-view image-text conversion according to claim 8 is characterized in that: S5 uses the risk weight value to identify the mapping node corresponding to the risk path, detects the semantic distortion path and marks it as abnormal through the fluctuation difference between the semantic vector and the residual information, stops the model output and sends a prompt message, and generates the path intervention parameter, including: S501: calling the risk weight value, identifying the risk path according to the preset path risk threshold, extracting the node sequence number and structural position of the path, establishing a node mapping relationship, and generating a risk path mapping node sequence; S502: Based on the risk path mapping node sequence, extract the semantic vector data and corresponding residual information of each node, detect the semantic distortion path by calculating the semantic deviation anomaly index of each node, and generate a semantic distortion path marker; S503: calling the semantically distorted path mark, establishing a path anomaly identifier and terminating the model output, sending a response message of the path anomaly, and generating a path intervention parameter.

10. The large model defense method based on multi-view image-text conversion according to claim 9 is characterized in that: The specific formula for calculating the semantic deviation anomaly index of each node is as follows (2): (2) in, represents the semantic deviation anomaly index of the a-th node, Represents the semantic vector strength value of the a-th node, Represents the residual information strength value of the a-th node, Represents the deviation between the semantic vector of the b-th node and the residual, Represents the average value of the semantic vector strength of all nodes in the risk path mapping node sequence, represents the total number of nodes in the risk path mapping node sequence, b represents the number of the bth node involved in the calculation in the risk path mapping node sequence, and a represents the number of a single node being processed when calculating the semantic deviation anomaly index.

Citation Information

Patent Citations

  • Visual question and answer method and system based on fine-grained adapter

    CN118607526A

  • Financial content risk control method and system based on large model

    CN118965279A

  • Image-text content auditing method based on multi-modal large model

    CN119941157A

  • News event search method and system based on multi-level image-text semantic alignment model

    WO2023093574A1

Cited By

  • Automatic quality evaluation method for high-quality data set in cultural field

    CN120951988A

  • Information detection method and computer program product

    CN121029938A

  • Land parcel change detection method and system

    CN121236122A

  • Document processing method based on dynamic multi-level index and feature clustering

    CN121388070A

  • Engineering safety and quality intelligent evaluation method and system based on machine vision

    CN122311969A