Multimodal data sensitivity grading method and system based on semantic risk graph diffusion perception
By constructing a semantic risk graph diffusion perception method, the problem of insufficient cross-modal linkage perception capability is solved, and the accuracy and interpretability of multimodal data sensitivity identification are achieved, supporting compliance auditing and risk backtracking.
Patent Information
- Application Number
- CN202511176329.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-08-21
AI Technical Summary
Existing technologies lack cross-modal linkage perception capabilities, making it difficult to identify semantic coupling relationships between text and images, or video and audio. They are unable to effectively model "combination-sensitive" or "implicitly triggered" scenarios and cannot provide a transparent explanation of the judgment process, resulting in insufficient accuracy and practicality in multimodal sensitive data identification.
A semantic risk graph diffusion perception method is constructed. Through modality recognition and structured processing, semantic risk units are generated, semantic embedding representations are extracted, a semantic risk graph is constructed, and a heat diffusion mechanism is used to simulate risk propagation. Path matching is performed in combination with sensitive label graphs, and structured interpretation results are output.
It significantly improves the accuracy and robustness of sensitivity identification in complex scenarios, achieves traceability of judgment results and support for compliance audits, and solves the problems of weak cross-modal linkage capability and insufficient context modeling.
Smart Images

Figure CN120670963B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence security and information content recognition technology, and in particular relates to a multimodal data sensitivity classification method and system based on semantic risk graph diffusion perception. Background Technology
[0002] Current data sensitivity identification methods are mainly based on static rules or single-modal classifiers, such as sensitive keyword matching, image detection models, and speech recognition systems. While these methods have some effectiveness in specific scenarios, they have significant limitations in practical applications. Existing technologies lack cross-modal linkage perception capabilities, making it difficult to effectively identify the semantic coupling relationships between text and images, or video and audio. For example, they cannot associate a document photo with an ID number in the text. Furthermore, existing methods ignore structural dependencies between contexts, making it difficult to model "combined sensitivity" or "implicit triggering" scenarios. For instance, sensitive information scattered across different modalities may have low risk when appearing individually, but when combined, it can form high-risk content. In addition, traditional methods cannot output interpretable sensitivity determination paths, which is detrimental to compliance audits and risk backtracking, particularly in areas requiring strict compliance review. The root cause of these problems is that traditional methods, when dealing with the complex, modal-rich, and semantically linked sensitive data generated by AIGC (AI Collective Data), cannot effectively capture cross-modal semantic relationships, lack the ability to model the propagation of contextual risks, and cannot provide a transparent explanation of the determination process. These problems severely restrict the accuracy and practicality of multimodal sensitive data identification, and there is an urgent need for a sensitivity detection technology that has cross-modal modeling capabilities, supports context diffusion awareness, and can output hierarchical results and label path interpretation. Summary of the Invention
[0003] To address the aforementioned technical problems, this invention proposes a multimodal data sensitivity classification method and system based on semantic risk graph diffusion perception, thereby resolving the issues present in the prior art.
[0004] Firstly, to achieve the above objectives, this invention provides a multimodal data sensitivity classification method based on semantic risk graph diffusion perception, comprising the following steps:
[0005] Modality recognition and structured processing are performed on the input multimodal data. Cross-modal segment combinations are selected based on semantic coupling to construct semantic risk units.
[0006] Extract the semantic embedding representation of each modality segment in the semantic risk unit, map it to the shared semantic space, and aggregate it with weights to generate a fusion representation vector;
[0007] Using semantic risk units as nodes, a semantic risk graph is constructed based on semantic similarity, modal co-occurrence frequency, entity consistency, and contextual relationships.
[0008] Initialize the risk pressure value for the nodes of the semantic risk graph, and simulate risk propagation based on the graph structure using a heat diffusion mechanism to obtain the steady-state risk value of the nodes;
[0009] The steady-state risk values of all nodes are merged to construct a sample-level risk vector, which is then input into a scoring function to output the sensitivity level prediction result.
[0010] A sensitive label graph is introduced and jointly modeled with a semantic risk graph. Label path chains are generated through semantic path matching, and structured interpretation results are output.
[0011] Optionally, the process of constructing the semantic risk unit includes:
[0012] Align multimodal data along a unified time axis to generate a triplet representation that includes modality type, timestamp, and embedded features;
[0013] Calculate the semantic similarity, temporal error, and contextual consistency index of any two modal segments;
[0014] Based on a preset threshold, fragment pairs that satisfy the semantic coupling condition are selected to form a candidate combination set;
[0015] The risk relevance score of candidate combinations is calculated based on semantic fusion strength, entity consistency, and sensitive label targeting. Combinations with scores exceeding the threshold are selected to construct semantic risk units.
[0016] Optionally, the process of generating the fused representation vector includes:
[0017] Invoke a pre-trained model adapted to the modality type to extract semantic feature representations of each modality segment;
[0018] Features are mapped to a unified semantic space using modality-specific projection matrices;
[0019] A weighted fusion strategy is adopted for the multimodal embedding representations within the same semantic risk unit to generate cross-modal semantic fusion vectors.
[0020] Optionally, the process of constructing the semantic risk graph includes:
[0021] Semantic risk units are mapped to nodes in a graph, and their fused representation vectors are used as node attributes.
[0022] The edge weights between nodes are calculated based on semantic similarity, core-referential entity overlap, temporal relationship, sensitive label consistency, and modal co-occurrence frequency.
[0023] Generate an adjacency matrix with multiple relational representations to form a semantic risk graph.
[0024] Optionally, the process of the heat diffusion mechanism includes:
[0025] Initialize node risk pressure values based on label confidence, modality coverage integrity, semantic ambiguity, and entity recognition confidence.
[0026] The pressure value is iteratively updated based on the relationship between adjacent nodes until the average change is lower than the convergence threshold.
[0027] Output the set of steady-state risk pressure values for each node.
[0028] Optionally, the process of generating the tag path chain includes:
[0029] Construct a tag graph containing predefined sensitive tag nodes and their hierarchical relationships;
[0030] Soft connections are established between the semantic risk graph and the label graph, and the edge weights are calculated based on the node embedding similarity.
[0031] Search for cross-graph paths starting from high-risk nodes and generate a path chain that includes the starting node, intermediate nodes, and target label node.
[0032] Secondly, the present invention also provides a multimodal data sensitivity classification system based on semantic risk graph diffusion perception, for implementing a multimodal data sensitivity classification method based on semantic risk graph diffusion perception, the system comprising:
[0033] The multimodal preprocessing module is used to perform modality recognition and structured processing on the input multimodal data, and to screen cross-modal fragment combinations based on semantic coupling to construct semantic risk units;
[0034] The feature fusion module is used to extract the semantic embedding representation of each modality segment in the semantic risk unit, map it to the shared semantic space, and weighted aggregate it to generate a fusion representation vector;
[0035] The semantic graph construction module is used to construct a semantic risk graph based on semantic similarity, modal co-occurrence frequency, entity consistency, and contextual relationships, with semantic risk units as nodes.
[0036] The risk diffusion module is used to initialize the risk pressure value of the nodes in the semantic risk graph. Based on the graph structure, a heat diffusion mechanism is used to simulate risk propagation and obtain the steady-state risk value of the nodes.
[0037] The risk assessment module is used to construct a sample-level risk vector by fusing the steady-state risk values of all nodes, and outputs the sensitivity level prediction results by inputting the scoring function.
[0038] The interpretable output module is used to introduce a sensitive label graph and jointly model it with a semantic risk graph. It generates a label path chain through semantic path matching and outputs a structured interpretation result.
[0039] Thirdly, the present invention also provides a computer terminal device, comprising:
[0040] One or more processors;
[0041] A memory, coupled to the processor, for storing one or more programs;
[0042] When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the multimodal data sensitivity classification method based on semantic risk graph diffusion perception in the first aspect described above.
[0043] Fourthly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, it implements the steps of the multimodal data sensitivity classification method based on semantic risk graph diffusion perception in the first aspect described above.
[0044] Fifthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the multimodal data sensitivity classification method based on semantic risk graph diffusion perception in the first aspect described above.
[0045] Compared with the prior art, the present invention has the following advantages and technical effects:
[0046] This invention provides a multimodal data sensitivity classification method and system based on semantic risk graph diffusion perception. This invention achieves accurate association of cross-modal sensitive information by constructing semantic risk units, explicitly models the collaborative relationships between multimodal segments using semantic risk graphs, and dynamically simulates the propagation process of risk in context using a heat diffusion mechanism, significantly improving the accuracy and robustness of sensitivity identification in complex scenarios. Through joint path matching of sensitive label graphs and semantic risk graphs, a structured interpretation chain is generated, achieving traceability of the judgment results and support for compliance auditing. This invention effectively solves the problems of weak cross-modal linkage capability, insufficient context modeling, and poor interpretability in traditional technologies. Attached Figure Description
[0047] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:
[0048] Figure 1 This is an overall flowchart of the multimodal data sensitivity classification method and system based on semantic risk graph diffusion perception, according to an embodiment of the present invention.
[0049] Figure 2 This is a framework diagram of multimodal semantic feature extraction and SRU construction according to an embodiment of the present invention;
[0050] Figure 3 This is a framework diagram of the semantic risk graph construction and diffusion process in an embodiment of the present invention;
[0051] Figure 4 This is a diagram illustrating the interpretability framework of the tag path in an embodiment of the present invention.
[0052] Figure 5 This is a framework diagram of the multimodal sensitivity detection system according to an embodiment of the present invention. Detailed Implementation
[0053] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0054] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0055] Example 1
[0056] like Figure 1 As shown, this embodiment provides a multimodal data sensitivity classification method based on semantic risk graph diffusion perception, including:
[0057] Modality recognition and structured processing are performed on the input multimodal data. Cross-modal segment combinations are selected based on semantic coupling to construct semantic risk units.
[0058] Extract the semantic embedding representation of each modality segment in the semantic risk unit, map it to the shared semantic space, and aggregate it with weights to generate a fusion representation vector;
[0059] Using semantic risk units as nodes, a semantic risk graph is constructed based on semantic similarity, modal co-occurrence frequency, entity consistency, and contextual relationships.
[0060] Initialize the risk pressure value for the nodes of the semantic risk graph, and simulate risk propagation based on the graph structure using a heat diffusion mechanism to obtain the steady-state risk value of the nodes;
[0061] The steady-state risk values of all nodes are merged to construct a sample-level risk vector, which is then input into a scoring function to output the sensitivity level prediction result.
[0062] A sensitive label graph is introduced and jointly modeled with a semantic risk graph. Label path chains are generated through semantic path matching, and structured interpretation results are output.
[0063] Specifically, the steps include:
[0064] S1: Perform modality recognition and structured processing on the input multimodal data (including text, images, audio and video), filter cross-modal fragment combinations based on semantic coupling, and construct a semantic risk unit (SRU) as the smallest sensitive context unit;
[0065] S2: Call the pre-trained model of modality adaptation to extract the semantic embedding representation of each modality segment in SRU, and unify it to the shared semantic space through linear mapping. Then, use a weighted aggregation strategy to generate the fusion representation vector of SRU.
[0066] S3: Using all SRUs as nodes, construct graph structure edge weights based on semantic similarity, modal co-occurrence frequency, entity consistency and contextual relationship to form a semantic risk graph (SRG) with multiple relational expressions, so as to explicitly characterize the potential semantic risk collaborative structure;
[0067] S4: Initialize the risk pressure value for each SRU node, and model the risk propagation process in the graph using a heat diffusion mechanism based on the semantic graph structure, and finally obtain the steady-state risk expression value of each node;
[0068] S5: Weighted fusion of steady-state pressure values of all SRU nodes within the same original sample to construct a sample-level risk vector, which is then input into the trained sensitivity level scoring function to output discrete or continuous risk level prediction results.
[0069] S6: Introduces a sensitive label graph and jointly models it with a semantic risk graph. It generates label path chains through cross-graph semantic path matching, realizes the structured interpretation output of the sensitivity level judgment results, and supports path-level tracing and causal link analysis.
[0070] As one implementation method in this embodiment, the process of constructing the semantic risk unit includes:
[0071] Align multimodal data along a unified time axis to generate a triplet representation that includes modality type, timestamp, and embedded features;
[0072] Calculate the semantic similarity, temporal error, and contextual consistency index of any two modal segments;
[0073] Based on a preset threshold, fragment pairs that satisfy the semantic coupling condition are selected to form a candidate combination set;
[0074] The risk relevance score of candidate combinations is calculated based on semantic fusion strength, entity consistency, and sensitive label targeting. Combinations with scores exceeding the threshold are selected to construct semantic risk units.
[0075] For details, please refer to the appendix. Figure 2As shown, this embodiment provides an implementation flow for multimodal semantic feature extraction and semantic risk unit (SRU) construction. After feature extraction and alignment, the data from each modality enters the SRU construction module to form the structured input nodes required for subsequent semantic risk graph modeling. This process includes the following steps:
[0076] S1.1: Modality recognition and structuring processing are performed on the input raw multimodal data, parsing heterogeneous modal data such as text, images, audio, and video into a standardized set of segments. Each segment generates a semantic embedding vector through an adapted modality feature extraction model, and is uniformly aligned based on the time axis to construct a triple:
[0077] ;
[0078] in Modal type, For the timestamp of the fragment, This is the modality embedding feature vector. The generation model and method of the embedding vector will be explained in detail in step S2.
[0079] S1.2: For any two modal segments and Based on embedding vectors respectively timestamp Calculate semantic similarity with contextual reference information. Time error Consistency with context This is used to assess the semantic coupling strength between segments. Semantic similarity measures the degree of similarity between two modal segments in the semantic space, and is calculated using the cosine similarity of the embedding vectors, as shown in the following formula:
[0080] ;
[0081] in These are the semantic embedding vectors of the segments. A larger value indicates that the semantics are closer.
[0082] Timing alignment error refers to the degree of offset between two modal segments on the time axis, reflecting whether they occur within approximately the same time period. Its calculation formula is as follows:
[0083] ;
[0084] in , Each is a fragment , timestamp, The smaller the value, the closer the time interval between the two segments.
[0085] Contextual consistency measures whether two modal segments share similar or identical semantic context. It is measured by core-referenced entities, consistent events, and shared labels, and is defined as follows:
[0086] ;
[0087] in , Each is a fragment , The set of named entities identified in the middle The closer the value is to 1, the more consistent the entities are.
[0088] S1.3: Set a similarity threshold Time alignment error threshold and context consistency threshold When the combination of fragments satisfies the following conditions: At that time, it was considered and They exhibit strong correlation in time, semantics, and context, forming a cross-modal combination candidate pair, which is then added to the candidate set. , as a potential sensitive context structure unit.
[0089] S1.4: To further filter candidate combinations, for each combination... Calculate its portfolio risk correlation score This is used to measure the sensitivity of cross-modal segment combinations, and the specific scoring function is defined as follows:
[0090] ;
[0091] in express and The fusion tightness in the semantic vector space is achieved through attention scoring and consistency measurement after vector aggregation; express and The similarity of the included named entities is calculated based on the normalization of the degree of overlap of shared entity types or tags; This represents the probability that the combination points to a sensitive label category, generated by a trained classification model and label mapping function. These are the weighting coefficients for the three types of indicators.
[0092] S1.5: Set the scoring threshold When candidate combinations Risk score satisfy When a node is determined to have a high risk propensity, it is constructed as a semantic risk unit (SRU) and used as a basic node input in the subsequent semantic graph structure modeling process.
[0093] As one implementation method in this embodiment, the process of generating the fused representation vector includes:
[0094] Invoke a pre-trained model adapted to the modality type to extract semantic feature representations of each modality segment;
[0095] Features are mapped to a unified semantic space using modality-specific projection matrices;
[0096] A weighted fusion strategy is adopted for the multimodal embedding representations within the same semantic risk unit to generate cross-modal semantic fusion vectors.
[0097] Specifically, the above process includes:
[0098] S2.1: For each multimodal segment contained within an SRU, extract semantic feature representations by calling a pre-trained model adapted to its modality type. Let a certain modality... The original input is Its corresponding feature extractor is Then the modal feature vector is obtained. The specific implementation includes: for the text modality, the BERT model is used, with the input being segmented text fragments and the output being context semantic embedding vectors, extracting the [CLS] bits as fragment representations; for the image modality, target region detection and cropping are performed using YOLOv5 before processing, and then the cropped regions are input into the CLIP model to output visual semantic embeddings; for the audio modality, the audio fragments are segmented into frames and spectrograms are extracted, which are input into the VGGish model to output acoustic feature representations; for the video modality, continuous frames are divided into fixed-length stable action segments, which are input into the I3D (Inflated 3D ConvNet) model to extract action semantic vectors to represent temporal motion semantics.
[0099] S2.2: To achieve alignment of multimodal semantic features in the shared representation space, a linear projection matrix is designed for each modality type. This is used to map the original features to a unified semantic space. The mapping formula is as follows:
[0100] ;
[0101] in, Representing modes Middle Embedded representation of fragments in a unified semantic space For modality-specific feature extractors, For modality A specific linear mapping matrix, For modal fragment embedding in a unified semantic space.
[0102] S2.3: For multiple modal segments within the same SRU, the set of their corresponding embedding representations is denoted as:
[0103] ;
[0104] A weighted fusion strategy is used to integrate the embedding vectors in this set into a unified representation vector for SRU. This process comprehensively considers factors such as modality confidence and contextual contribution. The fusion formula is as follows:
[0105] ;
[0106] in, Indicates the first The weight coefficients of each modal segment in the fusion process are determined as follows: In the training method, a learnable multilayer perceptron is introduced during the training phase to calculate a score for each modal segment, which is then obtained through softmax normalization. In the non-training method, fixed weights are set based on prior knowledge.
[0107] Final fusion As the sole node representation of SRU, it possesses unified cross-modal semantic representation capabilities and is input into subsequent graph construction and risk reasoning modules.
[0108] As one implementation method in this embodiment, the process of constructing the semantic risk graph includes:
[0109] Semantic risk units are mapped to nodes in a graph, and their fused representation vectors are used as node attributes.
[0110] Calculate the edge weights between nodes based on semantic similarity, core-referential entity overlap, temporal relationship, sensitive label consistency, and modal co-occurrence frequency;
[0111] Generate an adjacency matrix with multiple relational representations to form a semantic risk graph.
[0112] Specifically, the above process includes:
[0113] See attached document Figure 3 This embodiment provides a mechanism for constructing a semantic risk graph and processing risk pressure diffusion and propagation. Using Semantic Risk Units (SRUs) as nodes, a multi-relationship fusion graph structure is constructed, and a graph neural network modeling method based on heat diffusion is introduced to simulate the propagation and evolution of potential risks in the semantic context, ultimately completing node-level risk representation and sample-level sensitivity level classification. Specific steps include:
[0114] S3.1: Each semantic risk unit Mapped to a node in the graph Its features are represented as node attribute vectors. The set of vertices that constitute the graph of all SRUs. The corresponding node attribute matrix is .
[0115] S3.2: Construct the set of edges in the graph Analyze any two nodes and The multidimensional coupling relationship between them, edge weights Constructed based on the following five categories of feature terms:
[0116] semantic similarity Cosine similarity is defined as the similarity between node feature vectors, and the calculation formula is as follows:
[0117] ;
[0118] in Let be the fused representation vector of the node. This represents the L2 norm.
[0119] Core Entity Overlap :measure and The degree of overlap between the included named entity sets is defined as follows:
[0120] ;
[0121] The higher the value, the more consistent the two SRUs are in terms of semantic orientation.
[0122] Time sequence indicators : used to indicate and The chronological order of events in the original sample is defined as:
[0123] ;
[0124] in They represent and The earliest timestamp in the sequence is used to convert the time difference into a proximity within the [0,1] interval. The closer the timestamps of the two SRUs are, the higher the proximity. The closer it is to 1.
[0125] Sensitive label consistency : Indicates whether two nodes point to the same sensitive label, defined as follows:
[0126] ;
[0127] in They are respectively and Predicted labels, This represents a set of sensitive categories.
[0128] Modal co-occurrence frequency :measure and The degree of overlap of modal types is defined as:
[0129] ;
[0130] in and All of these are sets of modal types contained in SRU; the higher the value, the more consistent the modal coverage.
[0131] The aforementioned features include five types of heterogeneous relationship features: semantic similarity, core-referential entity overlap, temporal sequence, sensitive label consistency, and modal co-occurrence frequency. Edge weights are calculated using a weighted fusion method. The fusion formula is as follows:
[0132] ;
[0133] in, The edge weight fusion weight coefficient satisfies During the training phase, the fusion coefficient is set as a learnable parameter and participates in the backpropagation update of the final prediction loss function after semantic graph risk propagation.
[0134] S3.3: Based on the above definition of edge weights, generate the adjacency matrix of the graph. , Let be the number of nodes in the graph, where Represents a node and The strength of the connections between them. The final semantic risk graph can be represented as triples. It possesses multimodal fusion, context alignment, and semantic coupling characteristics, providing a structural foundation for subsequent graph diffusion and risk classification.
[0135] As one implementation method in this embodiment, the process of the heat diffusion mechanism includes:
[0136] Initialize node risk pressure values based on label confidence, modality coverage integrity, semantic ambiguity, and entity recognition confidence.
[0137] The pressure value is iteratively updated based on the relationship between adjacent nodes until the average change is lower than the convergence threshold.
[0138] Output the set of steady-state risk pressure values for each node.
[0139] Specifically, the above process includes:
[0140] S4.1: Semantic Risk Graph Each node Initialize risk pressure value This serves as the initial state for graph diffusion propagation. The initial pressure value is obtained by weighting multiple attributes, primarily including label confidence. : Represents a node The confidence probability of a label being predicted as sensitive comes from the softmax output of the sensitivity discrimination model; modality coverage integrity. :express The richness of covered modes is defined as the ratio of the actual number of covered modes to the total number of available modes, with a value range of [0,1]; contextual semantic ambiguity. : indicates that a node represents a vector The semantic uncertainty is measured using vector entropy, variance, and sentence vector density estimation metrics. A higher value indicates greater uncertainty in the information; and the confidence level of named entity recognition. : Represents the average confidence score of the named entity recognition model output in SRU, used to measure the quality of entity information. The initial pressure calculation formula is as follows:
[0141] ;
[0142] in , which are the weighting coefficients for each attribute factor, and are empirically set based on the importance of different modalities or the distribution characteristics on the training set.
[0143] S4.2: After assigning the initial pressure, a multi-round iterative propagation is performed based on the graph's adjacency structure to simulate the transmission process of semantic risk in the context network. The diffusion mechanism is analogous to a heat conduction model, controlling the proportion of pressure diffused from each node to its neighboring nodes. The diffusion iteration formula is as follows:
[0144] ;
[0145] in Represents a node The set of adjacent nodes, Indicates from node To the node The propagation weight coefficients can be determined based on the edge weights. Normalization yields:
[0146] ;
[0147] This strategy ensures that the influence of adjacent nodes on the current node is distributed according to their relative strength during each diffusion.
[0148] S4.3: To ensure the graph diffusion process converges within a finite number of rounds, a stability criterion is established. If the average pressure change at all nodes between two consecutive propagation rounds is less than the convergence threshold... If the condition is met, the iteration terminates. The determination formula is as follows:
[0149] ;
[0150] After the convergence condition is met, output the set of steady-state risk pressure values for each node:
[0151] ;
[0152] in For nodes The final risk expression is used to support subsequent sample-level sensitivity risk aggregation and level determination.
[0153] S5.1: For each original sample, collect the steady-state risk pressure values of all SRU nodes it contains to form the overall risk expression vector of that sample, defined as follows:
[0154] ;
[0155] in This indicates the sample. For the first in the sample The steady-state pressure values of each SRU node. This represents the number of SRUs in the sample. If the number of SRUs is inconsistent across different samples, vector unification can be achieved through methods such as average padding, max pooling, and fixed-dimensional truncation.
[0156] S5.2: Use the scoring function fitted during the training phase to apply it to the sample risk vector. After linear weighting and normalization, the final sensitivity level score is calculated. The scoring function is defined as follows:
[0157] ;
[0158] in, A fixed weight vector reflects the relative importance of each SRU pressure to the final level. As bias terms, both are determined through supervised learning training; The risk score normalization function is: Sigmoid function in binary classification scenarios and Softmax function in multi-level decision scenarios. The output is the normalized probability distribution of the corresponding level.
[0159] S5.3: Continuous scoring results based on the output of the trained model Map it to discrete sensitivity level labels The mapping is based on the defined level range segments:
[0160] ;
[0161] in Labels indicating different sensitivity levels, such as "non-sensitive," "low-sensitivity," "medium-sensitivity," and "high-sensitivity"; segmentation thresholds. The sensitivity level labels are obtained through validation parameter tuning or empirical settings. The final output sensitivity level labels not only serve as the system's judgment results but also as a core basis for multi-module collaboration. Specifically, they include: displaying the risk compliance interface to support result visualization and policy / rule alignment; triggering automated review strategies to achieve dynamic interception and labeling of sensitive content; serving as input to the semantic path interpretation module to support causal tracing and structured explanation of sensitivity level decisions; and providing risk warnings and preliminary classification suggestions for manual review processes, improving review efficiency and accuracy.
[0162] As one implementation method in this embodiment, the process of generating the tag path chain includes:
[0163] Construct a tag graph containing predefined sensitive tag nodes and their hierarchical relationships;
[0164] Soft connections are established between the semantic risk graph and the label graph, and the edge weights are calculated based on the node embedding similarity.
[0165] Search for cross-graph paths starting from high-risk nodes and generate a path chain that includes the starting node, intermediate nodes, and target label node.
[0166] Specifically, the above process includes:
[0167] See attached document Figure 4 This embodiment provides a sensitivity determination interpretability mechanism based on semantic path matching, aiming to provide a structured label path chain for the sensitivity level output of each sample, and support path-level causal analysis and source tracing. The specific process is as follows:
[0168] S6.1: Constructing a sensitive label graph:
[0169] ;
[0170] in This represents the set of label nodes in the sensitive label graph, where each node corresponds to a sensitive category, such as "document information"; the set of edges. Represents the hierarchical relationship (superordinate / inferior), attribute inclusion relationship, or semantic approximation connection between tags; node attribute matrix This represents the embedding features of tag nodes, generating tag semantic vectors through a pre-trained language model. This tag graph can be constructed with reference to national standards, industry specifications, or sensitive word ontology, possessing a clear semantic classification and inheritance structure, providing a foundation for subsequent tag matching.
[0171] S6.2: To achieve semantic mapping between SRU nodes and label nodes, cross-graph "soft connection edges" are introduced between the semantic risk graph and the sensitive label graph, and a joint graph structure is constructed:
[0172] ;
[0173] in For SRU nodes in the semantic risk graph, This represents the set of tag nodes in the sensitive tag graph; It is the set of edges of the joint graph; The fusion characteristics of SRU nodes; Let be the set of soft connection edges; the weight of each soft connection edge is defined as:
[0174] ;
[0175] in This is the semantic vector of the SRU node. Embedded representation for label nodes, This is the cosine similarity function in the embedding space. A similarity threshold is set. Invalid connections with excessively low edge weights are removed to reduce redundant paths in the graph.
[0176] S6.3: Based on Joint Graph Starting from each high-risk node, a cross-graph path search is performed to find the shortest or maximum-weight path to the label node. The generated label path chain consists of the following components: starting node (high-risk SRU), sequence of intermediate connected nodes (semantic / modal transitions), sequence of connecting edge weights, and target label node (e.g., "document information class"). Each path can be formalized as:
[0177] ;
[0178] in The starting node of the path is the identified high-risk SRU node; This indicates an intermediate node in the path, which can be another SRU node or an intermediate label node; This represents the endpoint node of the path, corresponding to the sensitive label hit (such as "document information"). This structured path chain is used to describe the semantic link that triggers the risk, which helps in the interpretability modeling of the judgment process.
[0179] S6.4: For each sample, the sensitivity level prediction result is appended with the corresponding label path chain as a structured interpretation result, supporting path-level visualization and source tracing, as well as sensitivity trigger factor analysis. The output format is as follows:
[0180] Image node: [ID photo];
[0181] Text node: ["ID number is 330XXXX"];
[0182] Tag trigger: [ID information type];
[0183] Co-occurring with video nodes: [Patient's self-report video];
[0184] Sensitivity Level: [Highly Sensitive]
[0185] This path demonstrates the causal chain of the judgment result, key modal trigger points, tag docking logic, and contextual collaborative evidence, supporting path-level source tracing analysis, semantic trigger factor interpretation, manual audit verification assistance, and compliant and transparent display of results.
[0186] S6.5: To further improve path quality and interpretability ranking, a Graph Attention Network (GAT) is introduced for the joint graph. Embedding learning is performed to learn the connection strength between nodes across the graph and their association with labels. The embedding representation is as follows:
[0187] ;
[0188] Graph neural networks can adaptively learn the importance weights of connection edges, optimize path generation quality, and be used for path priority ranking, thereby improving the usability and reliability of the final output.
[0189] Based on this, this invention provides a multimodal data sensitivity grading method based on semantic risk graph diffusion perception, which systematically solves key problems of existing multimodal sensitivity identification methods, such as insufficient semantic fusion depth, weak contextual linkage modeling capability, and lack of interpretability of sensitivity judgment results. By constructing cross-modal semantic risk units, expressing fusion graph structures, introducing a risk pressure diffusion mechanism, and a sensitivity label graph alignment strategy, this invention achieves a structured and traceable multimodal sensitivity grading judgment process. Based on the advantages of the above method, this invention designs a multimodal data sensitivity grading system composed of six core modules. The multimodal data preprocessing module receives text, image, audio, and video data, performs modality recognition, time alignment, and segment standardization, and generates candidate semantic units. The feature extraction and unified representation module uses a pre-trained model to extract the semantic embeddings of each modality segment and maps them to a shared semantic space. The semantic risk graph construction module constructs a heterogeneous graph structure based on semantic similarity and entity coreference among segments. The graph diffusion inference module simulates the risk propagation process through a heat diffusion mechanism and generates node-level risk representations. The sensitivity level assessment module aggregates node risk vectors and outputs sample-level sensitivity levels through a scoring function. The interpretable output module constructs a sensitivity label graph and generates label path chains, providing structured judgment explanations and causal link tracing capabilities.
[0190] Example 2
[0191] In this embodiment, a computer terminal device is provided, including:
[0192] One or more processors;
[0193] A memory, coupled to the processor, for storing one or more programs;
[0194] When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the above-described multimodal data sensitivity classification method based on semantic risk graph diffusion perception.
[0195] In this embodiment, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the above-described multimodal data sensitivity classification method based on semantic risk graph diffusion perception.
[0196] In this embodiment, an electronic device is also provided, including a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to perform the steps of the multimodal data sensitivity classification method based on semantic risk graph diffusion perception described above.
[0197] In this embodiment, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps of the above-described multimodal data sensitivity classification method based on semantic risk graph diffusion perception.
[0198] The aforementioned program can run on a processor or be stored in memory (or a computer-readable medium). Computer-readable media includes both permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random-access memory (SRAM), dynamic random-access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0199] These computer programs may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes can be implemented by different modules for different steps.
[0200] This embodiment provides such an apparatus or system. The system, referred to as a multimodal data sensitivity grading system based on semantic risk graph diffusion awareness, includes:
[0201] The multimodal preprocessing module is used to perform modality recognition and structured processing on the input multimodal data, and to screen cross-modal fragment combinations based on semantic coupling to construct semantic risk units;
[0202] The feature fusion module is used to extract the semantic embedding representation of each modality segment in the semantic risk unit, map it to the shared semantic space, and weighted aggregate it to generate a fusion representation vector;
[0203] The semantic graph construction module is used to construct a semantic risk graph based on semantic similarity, modal co-occurrence frequency, entity consistency, and contextual relationships, with semantic risk units as nodes.
[0204] The risk diffusion module is used to initialize the risk pressure value of the nodes in the semantic risk graph. Based on the graph structure, a heat diffusion mechanism is used to simulate risk propagation and obtain the steady-state risk value of the nodes.
[0205] The risk assessment module is used to construct a sample-level risk vector by fusing the steady-state risk values of all nodes, and outputs the sensitivity level prediction results by inputting the scoring function.
[0206] The interpretable output module is used to introduce a sensitive label graph and jointly model it with a semantic risk graph. It generates a label path chain through semantic path matching and outputs a structured interpretation result.
[0207] As one implementation method in this embodiment, the multimodal preprocessing module includes:
[0208] The fragment alignment unit is used to align multimodal data along a unified time axis and generate a triple representation that includes modality type, timestamp, and embedding features.
[0209] The coupling filtering unit is used to calculate the semantic similarity, temporal error, and contextual consistency index of any two modal segments, and to filter candidate combinations based on a preset threshold.
[0210] The risk unit generation unit is used to calculate the risk relevance score based on semantic fusion strength, entity consistency, and sensitive label targeting, and to construct semantic risk units.
[0211] As one implementation method in this embodiment, the feature fusion module includes:
[0212] The modality feature extraction unit is used to call a pre-trained model adapted to the modality type to extract the semantic feature representation of each modality segment;
[0213] Spatial mapping unit, used to map features to a unified semantic space through a modality-specific projection matrix;
[0214] The weighted aggregation unit is used to apply a weighted fusion strategy to the multimodal embedding representations within the same semantic risk unit to generate cross-modal semantic fusion vectors.
[0215] As one implementation method in this embodiment, the semantic graph construction module includes:
[0216] The node mapping unit is used to map semantic risk units to nodes of a graph, and its fused representation vector is used as node attributes.
[0217] The edge weight calculation unit is used to calculate the edge weights between nodes based on semantic similarity, core-referenced entity overlap, temporal relationship, sensitive label consistency, and modal co-occurrence frequency.
[0218] The graph generation unit is used to generate an adjacency matrix with multiple relational representations, forming a semantic risk graph.
[0219] As one implementation method in this embodiment, the risk diffusion module includes:
[0220] The pressure initialization unit is used to initialize the node risk pressure value based on the tag confidence, modal coverage integrity, semantic ambiguity, and entity recognition confidence.
[0221] The iterative propagation unit is used to iteratively update the pressure value based on the relationship between adjacent nodes until the average change is lower than the convergence threshold.
[0222] The steady-state output unit is used to output the set of steady-state risk pressure values for each node.
[0223] As one implementation method in this embodiment, the interpretable output module includes:
[0224] The tag graph construction unit is used to construct a tag graph containing predefined sensitive tag nodes and their hierarchical relationships;
[0225] Cross-graph connection unit is used to establish soft connection edges between the semantic risk graph and the label graph, and the edge weight is calculated based on the node embedding similarity.
[0226] The path generation unit is used to search for cross-graph paths starting from high-risk nodes and generate path chains that include starting nodes, intermediate nodes, and target label nodes.
[0227] The system or apparatus is used to implement the functions of the methods in the above embodiments. Each module in the system or apparatus corresponds to each step in the method, as has been described in the method and will not be repeated here.
[0228] The above implementation method solves the problem of multimodal data sensitivity classification based on semantic risk graph diffusion perception in related technologies, thereby ensuring that the problems existing in the prior art are resolved.
[0229] Example 3
[0230] See attached document Figure 5 This embodiment provides a multimodal data sensitivity classification method and system based on semantic risk graph diffusion perception, including the following modules:
[0231] The multimodal data preprocessing module is used to receive text, image, audio and video data, complete modality recognition, time alignment and segment standardization processing, and generate candidate semantic units.
[0232] The feature extraction and unified representation module: uses a pre-trained model to extract the semantic embeddings of each modality segment and maps them to a shared semantic space;
[0233] The semantic risk graph construction module constructs a heterogeneous graph structure based on semantic similarity and coreference relationships between fragments.
[0234] The graph diffusion inference module simulates the risk propagation process through a thermal diffusion mechanism to generate node-level risk representations.
[0235] The sensitivity level assessment module aggregates node risk vectors and outputs sample-level sensitivity levels through a scoring function.
[0236] The interpretable output module constructs a sensitive label map and generates a label path chain, providing structured judgment and interpretation capabilities as well as causal link tracing capabilities.
[0237] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A multimodal data sensitivity classification method based on semantic risk graph diffusion perception, characterized in that, Includes the following steps: Modality recognition and structured processing are performed on the input multimodal data. Cross-modal segment combinations are selected based on semantic coupling to construct semantic risk units. Extract the semantic embedding representation of each modality segment in the semantic risk unit, map it to the shared semantic space, and aggregate it with weights to generate a fusion representation vector; Using semantic risk units as nodes and integrating representation vectors as node attributes, a semantic risk graph is constructed based on semantic similarity, modal co-occurrence frequency, entity consistency, and contextual relationships. Initialize the risk pressure value for the nodes of the semantic risk graph, and simulate risk propagation based on the graph structure using a heat diffusion mechanism to obtain the steady-state risk value of the nodes; The steady-state risk values of all nodes are merged to construct a sample-level risk vector, which is then input into a scoring function to output the sensitivity level prediction result. A sensitive label graph is introduced and jointly modeled with a semantic risk graph. Label path chains are generated through semantic path matching, and structured interpretation results are output. The process of the heat diffusion mechanism mentioned above includes: Initialize node risk pressure values based on label confidence, modality coverage integrity, semantic ambiguity, and entity recognition confidence. The pressure value is iteratively updated based on the relationship between adjacent nodes until the average change is lower than the convergence threshold. Output the set of steady-state risk pressure values for each node.
2. The method according to claim 1, characterized in that, The process of constructing semantic risk units includes: Align multimodal data along a unified timeline to generate a triplet representation that includes modality type, timestamp, and embedded features; Calculate the semantic similarity, temporal error, and contextual consistency index of any two modal segments; Based on a preset threshold, fragment pairs that satisfy the semantic coupling condition are selected to form a candidate combination set; Risk relevance scores for candidate combinations are calculated based on semantic fusion strength, entity consistency, and sensitive label targeting. Combinations with scores exceeding a threshold are selected to construct semantic risk units.
3. The method according to claim 1, characterized in that, The process of generating the fused representation vector includes: Invoke a pre-trained model adapted to the modality type to extract semantic feature representations of each modality segment; Features are mapped to a unified semantic space using modality-specific projection matrices; A weighted fusion strategy is adopted for the multimodal embedding representations within the same semantic risk unit to generate cross-modal semantic fusion vectors.
4. The method according to claim 1, characterized in that, The process of constructing the semantic risk graph includes: Semantic risk units are mapped to nodes in a graph, and their fused representation vectors are used as node attributes. Calculate the edge weights between nodes based on semantic similarity, core-referential entity overlap, temporal relationship, sensitive label consistency, and modal co-occurrence frequency; Generate an adjacency matrix with multiple relational representations to form a semantic risk graph.
5. The method according to claim 1, characterized in that, The process of generating the tag path chain includes: Construct a tag graph containing predefined sensitive tag nodes and their hierarchical relationships; Soft connections are established between the semantic risk graph and the label graph, and the edge weights are calculated based on the node embedding similarity. Search for cross-graph paths starting from high-risk nodes and generate a path chain that includes the starting node, intermediate nodes, and target label node.
6. A multimodal data sensitivity grading system based on semantic risk graph diffusion perception, characterized in that, The system includes: The multimodal preprocessing module is used to perform modality recognition and structured processing on the input multimodal data, and to screen cross-modal fragment combinations based on semantic coupling to construct semantic risk units; The feature fusion module is used to extract the semantic embedding representation of each modality segment in the semantic risk unit, map it to the shared semantic space, and weighted aggregate it to generate a fusion representation vector; The semantic graph construction module is used to construct a semantic risk graph based on semantic similarity, modal co-occurrence frequency, entity consistency, and contextual relationships, with semantic risk units as nodes and representation vectors as node attributes. The risk diffusion module is used to initialize the risk pressure value of the nodes in the semantic risk graph. Based on the graph structure, a heat diffusion mechanism is used to simulate risk propagation and obtain the steady-state risk value of the nodes. The risk assessment module is used to construct a sample-level risk vector by fusing the steady-state risk values of all nodes, and outputs the sensitivity level prediction results by inputting the scoring function. The interpretable output module is used to introduce a sensitive label graph and jointly model it with a semantic risk graph, generate a label path chain through semantic path matching, and output a structured interpretation result. The risk diffusion module includes: The pressure initialization unit is used to initialize the node risk pressure value based on the tag confidence, modal coverage integrity, semantic ambiguity, and entity recognition confidence. The iterative propagation unit is used to iteratively update the pressure value based on the relationship between adjacent nodes until the average change is lower than the convergence threshold. The steady-state output unit is used to output the set of steady-state risk pressure values for each node.
7. A computer terminal device, characterized in that, include: one or more processors; A memory, coupled to the processor, for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors perform the steps of the method as described in any one of claims 1-5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-5.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-5.
Citation Information
Patent Citations
Image-text content auditing method based on multi-modal large model
CN119941157A
Electric power service semantic recognition method and system based on knowledge graph
CN120297285A