Sensitive data identification method and device, computer device and readable storage medium
Patent Information
- Application Number
- CN202511224512.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-29
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2045-08-29
AI Technical Summary
[0004]然而,目前的独立判断加简单结合的数据识别方法,存在识别准确度不高的问题
[0052]上述敏感数据识别方法、装置、计算机设备、计算机可读存储介质和计算机程序产品,通过获取待识别数据对应的目标模态数据,其中目标模态数据包括文本模态数据、图像模态数据和音频模态数据中的至少两种模态数据,对目标模态数据进行特征提取,得到目标模态数据对应的模态内特征,模态内特征包括底层基础特征和高层语义特征,通过与目标模态数据的模态类型相匹配的敏感度量化模型,根据目标模态数据和目标模态数据对应的模态内特征,生成目标模态数据对应的目标敏感度符号,目标敏感度符号用于表征目标模态数据的敏感程度,通过跨模态关联模型,获取各目标敏感度符号对应的贡献权重,利用各贡献权重对各目标敏感度符号进行加权求和,得到待识别数据的综合敏感度得分,并根据综合敏感度得分和预先设置的分级敏感度得分阈值,生成待识别数据的敏感度等级和致敏模态溯源信息。通过对多模态数据进行组合分析,精准识别组合式、隐藏式敏感内容,有效对抗内容规避行为,显著提高了针对敏感数据的识别精确度。
Smart Images

Figure CN121121233B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data identification technology, and in particular to a sensitive data identification method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Technology
[0002] With the rapid development of internet and mobile communication technologies, information on online platforms is exploding in multimodal formats, including text, images, audio, and video. This presents unprecedented challenges to the regulation of online content. Currently, the automatic identification and filtering of sensitive online data is a core task in the field of content security.
[0003] Currently, each modality is usually treated as an isolated unit and judged independently, and then the results from each path are simply summarized.
[0004] However, current data identification methods that rely on independent judgment and simple combinations suffer from low accuracy. Summary of the Invention
[0005] Therefore, it is necessary to provide a sensitive data identification method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can improve the accuracy of identification in response to the above-mentioned technical problems.
[0006] Firstly, this application provides a method for identifying sensitive data, including:
[0007] Obtain the target modality data corresponding to the data to be identified; the target modality data includes at least two modality data among text modality data, image modality data, and audio modality data;
[0008] Feature extraction is performed on the target modality data to obtain the intra-modality features corresponding to the target modality data; the intra-modality features include low-level basic features and high-level semantic features;
[0009] Using a sensitivity quantification model that matches the modality type of the target modality data, a target sensitivity symbol is generated based on the target modality data and its corresponding intramodal features; the target sensitivity symbol is used to characterize the sensitivity of the target modality data.
[0010] By using a cross-modal association model, the contribution weights corresponding to each target sensitivity symbol are obtained. The contribution weights are then used to perform a weighted summation of each target sensitivity symbol to obtain the comprehensive sensitivity score of the data to be identified. Based on the comprehensive sensitivity score and the pre-set graded sensitivity score thresholds, the sensitivity level and sensitization modality tracing information of the data to be identified are generated.
[0011] In conjunction with the first aspect, in one embodiment, when the target modal data includes text modal data, feature extraction of the target modal data includes:
[0012] Remove punctuation characters from the text modal data to obtain the cleaned text modal data;
[0013] Obfuscation pattern detection is performed on the cleaned text modal data, and deobfuscation processing is performed on the cleaned text modal data based on the obfuscation pattern detection results to obtain the target text modal data; the obfuscation pattern detection results include irrelevant character insertion, single character splitting, and homophone replacement;
[0014] Feature extraction is performed on the target text modal data.
[0015] In conjunction with the first aspect, in one embodiment, a target sensitivity symbol corresponding to the number of target modalities is generated based on the target modal data and the intra-modal features corresponding to the target modal data, using a sensitivity quantification model that matches the modality type of the target modal data, including:
[0016] Using a first sensitivity quantification model that matches the modality type of the target text modality data, semantic embedding calculation is performed on the target text modality data to obtain the corresponding semantic deviation. Based on the intramodal features of the target text modality data and the pre-set sensitive topics, the corresponding keyword hit intensity is obtained. Based on the semantic deviation and keyword hit intensity, the target sensitivity symbol corresponding to the target text modality data is generated.
[0017] In conjunction with the first aspect, in an exemplary embodiment, when the target modal data includes image modal data, feature extraction of the target modal data includes:
[0018] Scaling or cropping the image modal data yields image modal data with normalized dimensions.
[0019] The pixel values of the size-normalized image modal data are mapped to a standardized data range to obtain the target image modal data;
[0020] Feature extraction is performed on the modal data of the target image.
[0021] In conjunction with the first aspect, in one embodiment, a target sensitivity symbol corresponding to the target modal data is generated based on the target modal data and its corresponding intra-modal features, using a sensitivity quantification model that matches the modality type of the target modal data. This includes:
[0022] By using a second sensitivity quantification model that matches the modality type of the target image modality data, embedded text is extracted from the target image modality data to obtain the corresponding image embedded text modality data. Target detection is then performed based on the modal features corresponding to the target image modality data, generating the corresponding spatial semantic adjacency matrix. Finally, based on the target sensitivity symbol corresponding to the image embedded text modality data and the spatial semantic adjacency matrix, the target sensitivity symbol corresponding to the target image modality data is generated.
[0023] In conjunction with the first aspect, in one embodiment, when the target modal data includes audio modal data, feature extraction of the target modal data includes:
[0024] The sampling rate in the audio modal data is unified to a preset sampling rate to obtain unified audio modal data;
[0025] The multi-channel audio modal data in the unified audio modal data is converted into single-channel audio modal data to obtain the target audio modal data;
[0026] Feature extraction is performed on the target audio modal data.
[0027] In conjunction with the first aspect, in an exemplary embodiment, a target sensitivity symbol corresponding to the number of target modalities is generated based on the target modal data and the intra-modal features corresponding to the target modal data, using a sensitivity quantification model that matches the modality type of the target modal data, including:
[0028] By using a third sensitivity quantification model that matches the modality type of the target audio modality data, text transformation is performed on the target audio modality data to obtain the corresponding baseline text. Based on the intramodal features corresponding to the target audio modality data, the corresponding prosodic features and sentiment distribution features are obtained. Based on the target sensitivity symbols, prosodic features, and sentiment distribution features corresponding to the baseline text, the target sensitivity symbols corresponding to the target audio modality data are generated.
[0029] In conjunction with the first aspect, in one embodiment, after generating the target sensitivity symbol corresponding to the target modality data, the method further includes:
[0030] Obtain the pre-set sensitivity threshold for the target modal data;
[0031] If the sensitivity represented by the target sensitivity symbol is greater than the sensitivity threshold, the data to be identified is determined to be sensitive data.
[0032] Secondly, this application also provides a sensitive data identification device, comprising:
[0033] The acquisition module is used to acquire the target modality data corresponding to the data to be identified; the target modality data includes at least two modality data among text modality data, image modality data, and audio modality data;
[0034] The feature extraction module is used to extract features from the target modality data to obtain the intra-modality features corresponding to the target modality data; the intra-modality features include low-level basic features and high-level semantic features;
[0035] The symbol generation module is used to generate target sensitivity symbols for the target modal data based on the target modal data and its corresponding intramodal features, using a sensitivity quantification model that matches the modality type of the target modal data. The target sensitivity symbols are used to characterize the sensitivity of the target modal data.
[0036] The identification module is used to obtain the contribution weights corresponding to each target sensitivity symbol through a cross-modal association model, use each contribution weight to perform a weighted summation of each target sensitivity symbol to obtain the comprehensive sensitivity score of the data to be identified, and generate the sensitivity level and sensitization modality tracing information of the data to be identified based on the comprehensive sensitivity score and the pre-set graded sensitivity score threshold.
[0037] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0038] Obtain the target modality data corresponding to the data to be identified; the target modality data includes at least two modality data among text modality data, image modality data, and audio modality data;
[0039] Feature extraction is performed on the target modality data to obtain the intra-modality features corresponding to the target modality data; the intra-modality features include low-level basic features and high-level semantic features;
[0040] Using a sensitivity quantification model that matches the modality type of the target modality data, a target sensitivity symbol is generated based on the target modality data and its corresponding intramodal features; the target sensitivity symbol is used to characterize the sensitivity of the target modality data.
[0041] By using a cross-modal association model, the contribution weights corresponding to each target sensitivity symbol are obtained. The contribution weights are then used to perform a weighted summation of each target sensitivity symbol to obtain the comprehensive sensitivity score of the data to be identified. Based on the comprehensive sensitivity score and the pre-set graded sensitivity score thresholds, the sensitivity level and sensitization modality tracing information of the data to be identified are generated.
[0042] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0043] Obtain the target modality data corresponding to the data to be identified; the target modality data includes at least two modality data among text modality data, image modality data, and audio modality data;
[0044] Feature extraction is performed on the target modality data to obtain the intra-modality features corresponding to the target modality data; the intra-modality features include low-level basic features and high-level semantic features;
[0045] Using a sensitivity quantification model that matches the modality type of the target modality data, a target sensitivity symbol is generated based on the target modality data and its corresponding intramodal features; the target sensitivity symbol is used to characterize the sensitivity of the target modality data.
[0046] By using a cross-modal association model, the contribution weights corresponding to each target sensitivity symbol are obtained. The contribution weights are then used to perform a weighted summation of each target sensitivity symbol to obtain the comprehensive sensitivity score of the data to be identified. Based on the comprehensive sensitivity score and the pre-set graded sensitivity score thresholds, the sensitivity level and sensitization modality tracing information of the data to be identified are generated.
[0047] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0048] Obtain the target modality data corresponding to the data to be identified; the target modality data includes at least two modality data among text modality data, image modality data, and audio modality data;
[0049] Feature extraction is performed on the target modality data to obtain the intra-modality features corresponding to the target modality data; the intra-modality features include low-level basic features and high-level semantic features;
[0050] Using a sensitivity quantification model that matches the modality type of the target modality data, a target sensitivity symbol is generated based on the target modality data and its corresponding intramodal features; the target sensitivity symbol is used to characterize the sensitivity of the target modality data.
[0051] By using a cross-modal association model, the contribution weights corresponding to each target sensitivity symbol are obtained. The contribution weights are then used to perform a weighted summation of each target sensitivity symbol to obtain the comprehensive sensitivity score of the data to be identified. Based on the comprehensive sensitivity score and the pre-set graded sensitivity score thresholds, the sensitivity level and sensitization modality tracing information of the data to be identified are generated.
[0052] The aforementioned sensitive data identification method, apparatus, computer equipment, computer-readable storage medium, and computer program product acquire target modal data corresponding to the data to be identified, wherein the target modal data includes at least two modal data selected from text modal data, image modal data, and audio modal data; extract features from the target modal data to obtain intra-modal features corresponding to the target modal data, the intra-modal features including low-level basic features and high-level semantic features; generate target sensitivity symbols corresponding to the target modal data based on the target modal data and the corresponding intra-modal features using a sensitivity quantification model that matches the modality type of the target modal data, the target sensitivity symbols are used to characterize the sensitivity of the target modal data; obtain the contribution weights corresponding to each target sensitivity symbol through a cross-modal association model; perform a weighted summation of each target sensitivity symbol using each contribution weight to obtain a comprehensive sensitivity score for the data to be identified; and generate the sensitivity level and sensitization modality tracing information of the data to be identified based on the comprehensive sensitivity score and a pre-set graded sensitivity score threshold. By combining and analyzing multimodal data, we can accurately identify combined and hidden sensitive content, effectively combat content evasion behaviors, and significantly improve the accuracy of sensitive data identification. Attached Figure Description
[0053] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0054] Figure 1 This is an application environment diagram of a sensitive data identification method in one embodiment;
[0055] Figure 2 This is a flowchart illustrating a sensitive data identification method in one embodiment;
[0056] Figure 3 This is a flowchart illustrating a sensitive data identification method in another embodiment;
[0057] Figure 4 This is a flowchart illustrating a method for calculating sensitivity signs in one embodiment;
[0058] Figure 5 This is a flowchart illustrating the calculation of the overall sensitivity score in another embodiment;
[0059] Figure 6 This is a flowchart illustrating the process of determining sensitivity levels in one embodiment;
[0060] Figure 7This is a structural block diagram of a sensitive data identification device in one embodiment;
[0061] Figure 8 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0062] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0063] The sensitive data identification method provided in this application can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104, or it can be located in the cloud or on another network server. Server 104 obtains target modal data corresponding to the data to be identified from terminal 102. The target modal data includes at least two modal data among text modal data, image modal data, and audio modal data. Features are extracted from the target modal data to obtain the intra-modal features corresponding to the target modal data. The intra-modal features include low-level basic features and high-level semantic features. Through a sensitivity quantification model that matches the modality type of the target modal data, a target sensitivity symbol corresponding to the target modal data is generated based on the target modal data and the intra-modal features corresponding to the target modal data. The target sensitivity symbol is used to characterize the sensitivity of the target modal data. Through a cross-modal association model, the contribution weight corresponding to each target sensitivity symbol is obtained. The weighted sum of each target sensitivity symbol is used to obtain the comprehensive sensitivity score of the data to be identified. Based on the comprehensive sensitivity score and the graded sensitivity score threshold set by the bromide, the sensitivity level and sensitization modality source information of the data to be identified are generated. The terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart in-vehicle systems, and projection devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted displays. Head-mounted displays can be virtual reality (VR) devices, augmented reality (AR) devices, and smart glasses. The server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0064] In one exemplary embodiment, such as Figure 2As shown, a sensitive data identification method is provided, which can be applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps S201 to S204. Wherein:
[0065] Step S201: Obtain the target modal data corresponding to the data to be identified; the target modal data includes at least two modal data among text modal data, image modal data and audio modal data.
[0066] Among them, text modal data can be understood as data represented in the form of text content, image modal data can be understood as data represented in the form of image information, and audio modal data can be understood as data represented in the form of audio stream signals.
[0067] For example, server 104 obtains the data to be identified from terminal 102 and performs data analysis on the data to be identified to obtain target modal data corresponding to the data to be identified, which includes at least two modal data among text modal data, image modal data and audio modal data.
[0068] Step S202: Extract features from the target modal data to obtain the intramodal features corresponding to the target modal data; the intramodal features include low-level basic features and high-level semantic features.
[0069] Low-level basic features can be understood as direct descriptions of the original data with minimal processing. These can include images (pixel-level features, color histograms, edge information, texture, local gradients and directions, local contrast, scale-invariant features), text (character and word-level statistical features, word frequency, character frequency, n-grams, preliminary vectorization of TF-IDF, etc.), audio (short-time features in the time / frequency domain, such as Mel-frequency cepstral coefficients (MFCC), short-time Fourier coefficients, energy, chroma, etc.), and video (static visual features of each frame, optical flow, color histograms, and local motion). Information, etc.); high-level semantic features can be understood as the abstract meaning, global structure and semantic concepts of data, which can include images (semantic vectors of object / scene categories, semantic labels, relational positions, layout and contextual information, sentiment / style, etc.), text (topics, sentiments, intentions, entity relationships, paragraph-level / sentence-level semantic vectors, semantic embedding, etc.), audio (speaker identity, emotional state, timbre, semantic content (such as semantic roles, speech features of keywords, etc.), and video (global representation of action categories, scene semantics, event-level information, character relationships and interactions, etc.).
[0070] Optionally, the server 104 first preprocesses at least two types of modal data, including text modal data, audio modal data, and image modal data, which are included in the target modal data, by referring to the corresponding data preprocessing methods, and then extracts features from the preprocessed target modal data to obtain the low-level basic features and high-level semantic features corresponding to the target modal data.
[0071] Step S203: Using a sensitivity quantification model that matches the modality type of the target modal data, a target sensitivity symbol is generated based on the target modal data and its corresponding intramodal features; the target sensitivity symbol is used to characterize the sensitivity of the target modal data.
[0072] Among them, the sensitivity measurement model can be understood as a machine learning model or a reinforcement learning model, and the target sensitivity symbol can be understood as a unique data that represents the sensitivity of the target modality data.
[0073] Optionally, server 104 matches and loads the corresponding sensitivity quantification model and related parameters from a preset sensitivity quantification model library according to the type identifier of the target modal data. It formats the target modal data and the intramodal features corresponding to the target modal data into the input tensor required by the sensitivity quantification model, performs forward propagation calculation of the sensitivity quantification model, obtains one or more sensitivity-related prediction values, and synthesizes the prediction values into a single-dimensional standardized target sensitivity symbol through preset calculation logic.
[0074] Step S204: Obtain the contribution weights corresponding to each target sensitivity symbol through the cross-modal association model, use each contribution weight to perform a weighted summation of each target sensitivity symbol to obtain the comprehensive sensitivity score of the data to be identified, and generate the sensitivity level and sensitization modality tracing information of the data to be identified based on the comprehensive sensitivity score and the pre-set graded sensitivity score threshold.
[0075] Among them, the contribution weight can be understood as the contribution information of the target modal data to the overall sensitivity of the data to be identified; the graded sensitivity score threshold can be understood as the upper and lower limit judgment thresholds corresponding to different sensitivity levels; and the sensitization modality tracing information can be understood as the sensitization source modal data that causes the data to be identified to be sensitive data, which can include text modal data, audio modal data and image modal data.
[0076] For example, server 104 concatenates all target sensitivity symbols corresponding to the generated target modality data into an initial fusion feature vector in a predetermined order, inputs the initial fusion feature vector into the cross-modal association model, calculates the contribution weight of each sensitivity symbol to the final comprehensive sensitivity through the attention mechanism in the cross-modal association model, and uses each contribution weight to perform a weighted summation of each target sensitivity symbol in the initial fusion feature vector to obtain the comprehensive sensitivity score of the data to be identified. Then, a graded sensitivity threshold table is read from the data storage system. The graded sensitivity threshold table defines multiple sensitivity levels and their corresponding lower threshold limits. The overall sensitivity score is compared with each lower threshold limit in the graded sensitivity threshold table in descending order. The sensitivity level corresponding to the first lower threshold limit that is less than or equal to the overall sensitivity score is determined as the sensitivity level of the data to be identified. The sensitivity symbols of which modality(s) contribute the most to the result are recorded to form the source information of the sensitive modality of the data to be identified. After the data to be identified is determined to be sensitive data, graded handling operations are performed according to the sensitivity level of the data to be identified: automatic marking for low-risk levels, alarm and manual review for medium-risk levels, and direct blocking or content desensitization for high-risk levels.
[0077] In the aforementioned sensitive data identification method, target modal data corresponding to the data to be identified is acquired. This target modal data includes at least two modalities: text, image, and audio. Feature extraction is performed on the target modal data to obtain intra-modal features, which include low-level basic features and high-level semantic features. A sensitivity quantification model matching the modality type of the target modal data is used to generate target sensitivity symbols based on the target modal data and its corresponding intra-modal features. These symbols characterize the sensitivity of the target modal data. A cross-modal association model is used to obtain the contribution weights of each target sensitivity symbol. These weights are then used to perform a weighted summation of each target sensitivity symbol to obtain a comprehensive sensitivity score for the data to be identified. Based on the comprehensive sensitivity score and a pre-set tiered sensitivity score threshold, the sensitivity level and sensitization modality source information of the data to be identified are generated. By combining and analyzing multimodal data, combined and hidden sensitive content can be accurately identified, effectively combating content avoidance behaviors and significantly improving the accuracy of sensitive data identification.
[0078] In one embodiment, when the target modal data includes text modal data, feature extraction is performed on the target modal data, including: removing punctuation characters from the text modal data to obtain cleaned text modal data; performing obfuscation pattern detection on the cleaned text modal data, and performing deobfuscation processing on the cleaned text modal data based on the obfuscation pattern detection results to obtain the target text modal data; the obfuscation pattern detection results include irrelevant character insertion, single-character splitting, and homophone replacement; and feature extraction is performed on the target text modal data.
[0079] Obfuscation mode can be understood as adding interference information to circumvent the detection of sensitive information. It can include inserting irrelevant characters, splitting single characters, and replacing homophones with different characters.
[0080] Optionally, when the target modal data includes text modal data, server 104 removes punctuation characters from the text modal data to obtain cleaned text modal data, performs obfuscation pattern detection on the cleaned text modal data, and performs deobfuscation processing on the cleaned text modal data based on the obfuscation pattern detection results. The deobfuscation processing may include identifying and restoring sensitive words that have been deliberately circumvented by inserting irrelevant characters, splitting single words, or using homophones to obtain target text modal data, and extracting features from the target text modal data.
[0081] Based on the aforementioned implementation methods, by performing content normalization and de-obfuscation processing on the text modal data, the data validity of the obtained target text modal data is improved, thereby improving the effectiveness of feature extraction from the target text modal data and thus enhancing the identification accuracy of sensitive data.
[0082] In one embodiment, a target sensitivity symbol corresponding to the target modality number is generated based on the target modality data and its corresponding intra-modal features using a sensitivity quantification model that matches the modality type of the target modality data. This includes: using a first sensitivity quantification model that matches the modality type of the target text modality data, performing semantic embedding calculations on the target text modality data to obtain the corresponding semantic deviation; obtaining the corresponding keyword hit intensity based on the intra-modal features of the target text modality data and a pre-set sensitive topic; and generating a target sensitivity symbol corresponding to the target text modality data based on the semantic deviation and the keyword hit intensity.
[0083] Among them, semantic deviation can be understood as the degree of semantic difference between the current text and the benchmark text, and keyword hit strength can be understood as the degree of similarity between the current feature and the set keyword.
[0084] For example, server 104 calculates the cosine distance between the sentence embedding vector of the target text modality data and the baseline security context embedding vector using a first sensitivity quantification model that matches the modality type of the target text modality data, obtains the corresponding semantic deviation, and calculates the keyword hit strength of the intramodal features of the target text modality data under the sensitive topic of the pre-device. Finally, based on the semantic deviation and keyword hit strength, it generates the target sensitivity symbol corresponding to the target text modality data.
[0085] According to the above implementation method, semantic deviation reflects the semantic deviation between the input and the baseline, and keyword hit intensity reflects the topic relevance. By deriving the target sensitivity symbol through two independent and complementary quantitative signals, the sensitivity judgment logic can be kept consistent under different domains and different text styles, thereby reducing cross-domain deviation.
[0086] In an exemplary embodiment, when the target modal data includes image modal data, feature extraction of the target modal data includes: scaling or cropping the image modal data to obtain size-normalized image modal data; mapping the pixel values of the size-normalized image modal data to a standardized data range to obtain target image modal data; and performing feature extraction on the target image modal data.
[0087] Size normalization can be understood as unifying image data of different sizes into image data of the same size.
[0088] Optionally, server 104 performs scaling or cropping operations on image modal data of different resolutions to unify them to the preset input size required by the sensitivity quantification model, thereby obtaining image modal data with unified size. Then, the pixel values of the image modal data with normalized size are linearly mapped to the standardized data range of [0,1] or [-1,1] to obtain target image modal data. Finally, feature extraction is performed on the target image modal data.
[0089] Based on the aforementioned implementation method, by performing size and pixel value normalization processing on the image modal data, the uniformity of the size and pixel value of the target image modal data is ensured, thereby accelerating the feature extraction speed of the target image modal data and thus improving the recognition speed of sensitive data identification.
[0090] In one embodiment, a target sensitivity symbol corresponding to the target modal data is generated based on the target modal data and its corresponding intra-modal features using a sensitivity quantification model that matches the modal type of the target modal data. This includes: extracting embedded text from the target image modal data using a second sensitivity quantification model that matches the modal type of the target image modal data to obtain corresponding image embedded text modal data; performing target detection based on the intra-modal features corresponding to the target image modal data to generate a corresponding spatial semantic adjacency matrix; and generating a target sensitivity symbol corresponding to the target image modal data based on the target sensitivity symbol corresponding to the image embedded text modal data and the spatial semantic adjacency matrix.
[0091] Among them, image-embedded text modal data can be understood as text data in image modal data, and spatial semantic adjacency matrix can be understood as the semantic association risk of detected target objects and their spatial proximity in the image.
[0092] For example, server 104 uses a second sensitivity quantification model that matches the modality type of target image modality data to extract embedded text from the target image modality data using optical character recognition technology, thereby obtaining image embedded text modality data corresponding to the target image modality data. It also performs target detection on the modal features corresponding to the target image modality data, constructs a corresponding spatial semantic adjacency matrix using the semantic association risk of the detected target object and its spatial proximity in the image, obtains the target sensitivity symbol corresponding to the image embedded text modality data in the aforementioned manner, and generates the target sensitivity symbol corresponding to the target image modality data based on the target sensitivity symbol corresponding to the image embedded text modality data and the spatial semantic adjacency matrix.
[0093] According to the aforementioned implementation method, the embedded text in the image is combined with the objects and positional relationships in the image, and the semantic relationships within the image are quantified through the spatial semantic adjacency matrix, thereby improving the fine-grainedness and robustness of sensitivity judgment and increasing the accuracy of constructing target sensitivity symbols corresponding to target image modal data.
[0094] In one embodiment, when the target modal data includes audio modal data, feature extraction of the target modal data includes: unifying the sampling rate in the audio modal data to a preset sampling rate to obtain unified audio modal data; converting the multi-channel audio modal data in the unified audio modal data into single-channel audio modal data to obtain target audio modal data; and extracting features from the target audio modal data.
[0095] Here, sampling rate can be understood as the number of consecutive samples of an audio signal per unit time, and channel can be understood as an independent transmission / output path used to carry independent sound information.
[0096] Optionally, server 104 unifies audio modal data with different sampling rates to a preset sampling rate through resampling technology to obtain unified audio modal data, converts multi-channel audio modal data in unified audio modal data into single-channel audio modal data to obtain target audio modal data, and performs feature extraction on target audio modal data.
[0097] Based on the above implementation method, by performing sampling rate and channel normalization processing on the audio modal data, the corresponding target audio modal data is obtained, which reduces the difficulty and speed of feature extraction for the target audio modal data, thereby improving the recognition speed of sensitive data identification.
[0098] In an exemplary embodiment, a target sensitivity symbol corresponding to the target modality number is generated based on the target modality data and its corresponding intramodal features using a sensitivity quantification model that matches the modality type of the target modality data. This includes: performing text conversion on the target audio modality data using a third sensitivity quantification model that matches the modality type of the target audio modality data to obtain the corresponding baseline text; obtaining the corresponding prosodic features and sentiment distribution features based on the intramodal features corresponding to the target audio modality data; and generating the target sensitivity symbol corresponding to the target audio modality data based on the target sensitivity symbol, prosodic features, and sentiment distribution features corresponding to the baseline text.
[0099] Among them, the reference text can be understood as the text modal data corresponding to the target audio modal data, the prosodic features can be understood as the changes in the beat, rhythm, strong and weak beats, pitch contours and other music structure-related information in the sentence over time, and the emotion distribution features can be understood as the statistical distribution features of emotion information in the data or signal, which are used to describe the distribution characteristics of the emotion state over time, samples or modalities.
[0100] For example, server 104 uses a third sensitivity quantification model that matches the modality type of the target audio modality data to convert the target audio modality data into text using automatic speech recognition technology to obtain the corresponding baseline text. Then, it calls the aforementioned method of calculating the target sensitivity symbol corresponding to the text modality data to calculate the target sensitivity symbol corresponding to the baseline text. It extracts prosodic features containing pitch, speech rate, and energy dimensions from the modality features corresponding to the target audio modality data, and inputs the modality features corresponding to the target audio modality features into the audio emotion recognition model to obtain the corresponding emotion distribution features. Finally, based on the target sensitivity symbol, prosodic features, and emotion distribution features corresponding to the baseline text, it generates the target sensitivity symbol corresponding to the target audio modality data.
[0101] According to the aforementioned implementation method, audio modal data is integrated with multimodal information such as text modality, prosodic features, and sentiment distribution to improve the granularity and robustness of the determination of target sensitivity symbols. Information from different modalities corroborates each other, reducing errors caused by single-modal noise and improving the accuracy of generating target sensitivity symbols corresponding to the target audio modal data, thereby enhancing the recognition accuracy of sensitive data.
[0102] In one embodiment, after generating the target sensitivity symbol corresponding to the target modality data, the method further includes: obtaining a sensitivity threshold preset for the target modality data; and determining the data to be identified as sensitive data if the sensitivity represented by the target sensitivity symbol is greater than the sensitivity threshold.
[0103] Optionally, after the server 104 has completed the calculation of the sensitivity symbols corresponding to the target modality data, it obtains the sensitivity threshold corresponding to each modality. If the sensitivity symbol generated by any modality exceeds its independent high-risk sensitivity threshold, the server will no longer perform the multimodal data fusion calculation, and will determine the current data to be identified as sensitive data. The target modality data corresponding to the sensitivity symbols that exceed the high-risk sensitivity threshold will be used as the source information of the sensitive modality of the data to be identified.
[0104] Based on the above implementation, a circuit breaker mechanism for the sensitivity symbols of single modal data is designed. When the sensitivity symbol of a certain modal data exceeds the threshold, the multimodal feature fusion step is no longer entered. Instead, the current data is directly identified as sensitive data and the current modal data is identified as sensitive modal source information. This reduces the identification steps of sensitive data and enhances the ability to deal with the identification of massive amounts of data.
[0105] In one exemplary embodiment, such as Figure 3 As shown, a specific implementation of a sensitive data identification method is provided (the following data are specific examples and are not limited to this case to implement the technical solution of this application), wherein:
[0106] Step S1: Obtain the original data to be identified, which contains at least two modalities, and parse and extract features from the original data. The parsing and feature extraction include: separating the original data into at least two of the following: text modality, image modality, and audio modality, and extracting the corresponding intra-modal features from each modal data. The intra-modal features include low-level basic features and high-level semantic features.
[0107] Suppose the system receives a social media post to be identified as raw data. The post contains both text and image modalities. Specifically, the text reads, "Brothers who urgently need money, contact me, I have connections, discreet delivery. [Image]"; the image is a promotional poster containing a QR code with the text "Scan the code to get on board, easily earn over 10,000 yuan a month." First, the system separates this raw data into independent text modal data and image modal data.
[0108] The parsing and feature extraction in step S1 include: performing content normalization and deobfuscation processing on the text modal data; the content normalization includes parsing and removing non-content tags. The deobfuscation processing is used to identify and restore sensitive words that have been deliberately circumvented by inserting irrelevant characters, splitting them into single characters, or using homophones.
[0109] When processing the text modality, the content normalization step parses and removes non-content markers, such as "[image]", resulting in the clean text: "Brothers who urgently need money, contact me, I have connections, confidential delivery." Next, the deobfuscation module checks for any obfuscation attempts. For example, if the original text is "Brothers who urgently need money," this module will restore it to "urgently need money." In this example, the original text has no obvious obfuscation. Subsequently, the system extracts features from the normalized text: for low-level basic features, it uses a built-in sensitive word dictionary to identify keywords such as "urgently need money" and "connections"; for high-level semantic features, it constructs a sensitive word feature vector. Assuming the system pre-defines three sensitive topic dimensions: [pornography, gambling, financial fraud], and based on the high correlation between the matched keywords and the topic "financial fraud," the system quantifies its strength as 0.8, while the other topics are quantified as 0. Therefore, the generated feature vector is... =[0,0,0.8].
[0110] For image modal data, size and pixel value normalization processing is performed; the size normalization is to unify images of different resolutions to the preset input size required by the sensitivity quantification model through scaling or cropping operations; the pixel value normalization is to linearly map the pixel values of the image to a standardized data range of [0,1] or [-1,1].
[0111] When processing the image modality, the system first scales the poster image to a preset input size required by the model, such as 224*224 pixels. Simultaneously, it linearly maps the 8-bit (0-255) pixel values to a floating-point range of [0,1] for normalization. Next, it extracts the high-level semantic features of the image and uses an object detection model to identify two key objects: Object 1 and Object 2. Then, for Object 2, it uses optical character recognition technology to extract the embedded text: "Scan the code to get on board and easily earn over 10,000 yuan a month." This text will be used for subsequent sensitivity calculations.
[0112] For audio modal data, sampling rate and channel normalization processing is performed; the sampling rate normalization is to unify the original audio signals with different sampling rates to a preset sampling rate through resampling technology; the channel normalization is to convert multi-channel audio streams into single-channel audio signals.
[0113] Step S2: Based on the intramodal features extracted in step S1, for each type of modal data, call the sensitivity quantification model that matches the modality, calculate and generate a sensitivity symbol that characterizes the sensitivity of the modal data.
[0114] This step will quantify the features extracted in S1.
[0115] When the modal data is text modal data, its corresponding sensitivity quantification model is used to calculate the text sensitivity symbol. The It is calculated using the following formula:
[0116] ,in, Represents the final text sensitivity symbol; It is a sensitive word feature vector, where each dimension corresponds to a preset sensitive topic, and the dimension value is the keyword hit strength of the text under the corresponding topic; It is a with The sensitive topic weight vectors of the same dimension are obtained by pre-training the model; · Represents the dot product operation of two vectors; This represents the Sigmoid activation function, used to normalize the risk to the (0,1) interval; It is a contextual semantic deviation scalar, which is obtained by calculating the cosine distance between the sentence embedding vector of the text and the baseline safe context embedding vector.
[0117] Regarding the text "Brothers who urgently need money, DM me, I have connections, discreet shipping.", we conducted... The calculation. Given. =[0,0,0.8]. Assume the pre-trained weight vector... =[0.2,0.3,0.9], representing the highest inherent risk weight for the "financial fraud" theme. First, calculate the dot product: · =0×0.2+0×0.3+0.8×0.9=0.72. Then, normalize using the Sigmoid function: σ(0.72)=1 / (1+e^(-0.72))≈0.672. Next, calculate the contextual semantic deviation. Since the semantics of this text deviate significantly from those of everyday security conversations, we assume the calculated cosine distance is 0.75. Ultimately, the text sensitivity notation is: =0.672×(1+0.75)=1.176.
[0118] When the modal data is image modal data, its corresponding sensitivity quantification model is used to calculate the image sensitivity sign. The It is calculated using the following formula:
[0119] ,in, Represents the final image sensitivity symbol; It is a spatial-semantic adjacency matrix, constructed by performing object detection on an image, with matrix elements... The value represents the detected object. and objects Semantic association risk and its spatial proximity in images; It is a with Risk weight matrix of the same dimension; This represents the Hadamard product, which is the element-wise multiplication of matrices. Represents the Frobenius norm, used to calculate the overall strength of the weighted risk matrix; This means that after extracting the embedded text in the image using optical character recognition technology, text sensitivity symbols are invoked. These are the preset weighting coefficients.
[0120] When the modal data is audio modal data, its corresponding sensitivity quantification model is used to calculate the audio sensitivity sign. The It is calculated using the following formula:
[0121] S a = S t _ a s r × [ 1 + t a n h ( W p ⋅ V p + W e ⋅ V e ) ] ,in, Represents the final audio sensitivity symbol; The representative uses automatic speech recognition technology to convert audio into text, and then calls upon a baseline text sensitivity symbol. It is a prosodic feature vector extracted from an audio signal, which includes pitch, speech rate, and energy dimensions; It is a with A prosodic risk weight vector of the same dimension; It is an emotion distribution vector obtained through an audio emotion recognition model; It is a with Emotional risk weight vector of the same dimension; This represents the hyperbolic tangent function, used to map the non-textual risk of audio to the (-1,1) interval, forming a risk modulation coefficient.
[0122] Step S3: Integrate all the sensitivity symbols generated in step S2. The integration process includes: constructing a cross-modal association model to characterize the synergistic enhancement and / or contradictory inhibition relationships between different sensitivity symbols, and weighting and aggregating each sensitivity symbol based on the cross-modal association model to calculate a comprehensive sensitivity score.
[0123] In this step, the system will use the modal sensitivity signs calculated in step S2. =1.176 and =2.036 is used for fusion. First, they are concatenated in a predetermined order to form an initial fusion feature vector: [1.176, 2.036]. Then, this vector is input into a cross-modal association model based on an attention mechanism. This model analyzes the values of each symbol and its source modality, dynamically calculating the contribution weights. In this example, due to image sensitivity... The value is much higher than Furthermore, the internal evidence is very conclusive, and the attention model assigns higher weights to the image modality. For example, the contribution weight vector output by the model is [0.3, 0.7], corresponding to text and image respectively. Finally, a weighted aggregation is performed to calculate the comprehensive sensitivity score: Score = 0.3 × 1.176 + 0.7 × 2.036 = 0.3528 + 1.4252 = 1.778.
[0124] Step S4: Compare the comprehensive sensitivity score with a preset graded sensitivity threshold to determine the sensitivity level of the original data, and generate an identification result containing the sensitivity level and sensitization modality tracing information for output.
[0125] In the final step, the system compares the overall score of 1.778 with a preset grading threshold table. The threshold table is defined as follows: [0, 0.5) represents low risk, [0.5, 1.5) represents medium risk, and [1.5, +∞) represents high risk. Since 1.778 is greater than or equal to the lower limit of 1.5 for high risk, the system determines the original data to be at the "high risk" level. Simultaneously, the system records the contribution of each modality, forming source information: the image modality contributes the most (70%). Finally, the system outputs a structured recognition result, such as: {"level":"high risk","score":1.778,"tracing":[{"modality":"image","contribution":0.7},{"modality":"text","contribution":0.3}]}.
[0126] In one embodiment, such as Figure 4 As shown, a specific implementation method for calculating sensitivity symbols is provided, wherein step S2 includes:
[0127] Step S201: Based on the type identifier of the modal data, match and load the corresponding sensitivity quantification model and related parameters from a preset model library.
[0128] During step S2, the system first identifies that the data to be processed is "text modality". Based on this type identifier, the system accesses a pre-set model library, which stores pre-trained models for different modalities. The system then matches and loads the "text-sensitive quantification model" and its associated parameters, including but not limited to sensitive topic weight vectors. Embedded vectors of baseline security contexts, etc.
[0129] Step S202: The intramodal features extracted from the modal data are formatted into the input tensor required by the loaded sensitivity quantization model.
[0130] Next, the system formats the features extracted for the text modality in step S1. Specifically, it formats the sensitive word feature vector. The vectors [0,0,0.8] and the sentence embedding vectors calculated by the BERT model are packaged together to construct a tensor that conforms to the input specifications of the loaded model. This tensor will be used as the input data for the model's forward propagation.
[0131] Step S203: Perform forward propagation calculation of the sensitivity quantification model to obtain one or more predicted values related to sensitivity, and synthesize the predicted values into a single-dimensional, standardized sensitivity symbol through a preset calculation logic.
[0132] The system feeds the formatted input tensor into the text sensitivity quantification model for forward propagation computation. Different branches within the model output predicted values related to sensitivity; for example, one branch outputs a risk value based on keyword matching, while another branch outputs a deviation value based on semantic understanding. Finally, based on the pre-defined calculation logic, these predicted values are combined into a single-dimensional, standardized sensitivity symbol. =1.176. Similarly, when processing image modalities, the system loads the image model, and... The matrix and OCR text features are formatted as input tensors. After forward propagation, the model outputs the norm of the weighted risk matrix and the sensitivity of the embedded text, which are then synthesized into a single input tensor. =2.036.
[0133] In one embodiment, such as Figure 5 As shown, a specific implementation method for calculating the comprehensive sensitivity score is provided, wherein step S3 includes:
[0134] Step S301: The sensitivity symbols of all generated modalities are concatenated into an initial fusion feature vector in a predetermined order.
[0135] At the start of step S3, the system collects all the sensitivity symbols generated in step S2 for all modalities. In this example, this refers to the text sensitivity symbols. =1.176 and image sensitivity symbol =2.036. The system concatenates these two values into an initial fusion feature vector, namely [1.176, 2.036], according to the predetermined order of "text-image".
[0136] Step S302: Input the initial fusion feature vector into the cross-modal association model, wherein the cross-modal association model is a network based on an attention mechanism, used to calculate the contribution weight of each sensitivity symbol to the final comprehensive sensitivity.
[0137] The initial fused feature vector [1.176, 2.036] is input into the cross-modal association model. This model internally implements a self-attention mechanism. Through this mechanism, the model can evaluate the relative importance of each element in the vector. The model calculates the query, key, and value matrix and generates an attention score matrix. This score matrix is normalized using the Softmax function to obtain the contribution weight of each sensitivity symbol to the final result. In this example, since 2.036 is much larger than 1.176, the attention mechanism determines that the information from the image modality is more decisive, thus calculating a higher contribution weight.
[0138] Step S303: The contribution weights are used to perform a weighted summation of each sensitivity symbol in the initial fusion feature vector to obtain the final comprehensive sensitivity score.
[0139] Based on the contribution weights calculated in step S302, for example, text weight is 0.3 and image weight is 0.7, the system performs a weighted summation of corresponding elements in the initial fused feature vector. The calculation process is as follows: Comprehensive sensitivity score = 0.3 × 1.176 + 0.7 × 2.036 = 1.778. This score of 1.778 is the final comprehensive judgment result that integrates all modal information.
[0140] In one exemplary embodiment, such as Figure 6 As shown, a specific implementation method for determining sensitivity level is provided, wherein step S4 includes:
[0141] Step S401: Read a graded sensitivity threshold table from the system configuration. The graded sensitivity threshold table defines multiple sensitivity levels and their corresponding lower limits for judgment thresholds.
[0142] When performing step S4, the system first reads a predefined hierarchical sensitivity threshold table from its system configuration. This table is shown below:
[0143]
[0144] Step S402: The comprehensive sensitivity score is compared sequentially with each lower limit of the judgment threshold in the graded sensitivity threshold table in descending order.
[0145] The system calculates a comprehensive sensitivity score of 1.778 and compares it with the lower limit of the judgment threshold in the threshold table in descending order (1.5, 0.5, 0.0). First comparison: Is 1.778 greater than or equal to 1.5? Yes.
[0146] Step S403: The sensitivity level corresponding to the first judgment threshold lower limit that is less than or equal to the comprehensive sensitivity score is determined as the final sensitivity level of the original data, and it is recorded which or which modal sensitivity symbols contribute the most to the result, so as to form the sensitization modality tracing information.
[0147] Since 1.778 met the condition of being greater than or equal to 1.5 in the first comparison, the system determined 1.5 as the first threshold less than or equal to the comprehensive score. Therefore, the system determined the sensitivity level "high risk" corresponding to this threshold as the final sensitivity level of the original data. At the same time, the system backtracked the contribution weights calculated in step S302 and found that the contribution weight of the image modality (0.7) was the largest. Therefore, the "image modality" was recorded as the main source of sensitivity, forming the source information of the sensitivity modality.
[0148] The execution path also includes a circuit breaker branch; the trigger condition for the circuit breaker branch is: after completing the sensitivity quantification for each modality, the sensitivity symbol generated by any modality exceeds its independent high-risk threshold; when the trigger condition is met, the execution path of the method will bypass the fusion processing used to generate the comprehensive sensitivity score and directly output the judgment result that the original data is sensitive.
[0149] In another possible path of this embodiment, it is assumed that the system sets an independent high-risk threshold for each modality, for example, a threshold of 2.0 for the image modality. At the end of step S2, the system calculates the image sensitivity. =2.036. At this point, the system check found that 2.036 exceeded the preset value of 2.0. The circuit breaker condition was met. Therefore, the entire process immediately interrupted the subsequent steps S3 and S4, directly determined that the original data was "high risk", and recorded the sensitizing source as "image modality".
[0150] Once the multimodal data is determined to be sensitive data, the process further includes performing tiered handling operations on the sensitive data based on the numerical range of the comprehensive sensitivity score or the triggering of the rapid circuit breaker mechanism. The tiered handling operations include: automatic marking for low-risk data, alarm and manual review for medium-risk data, and direct blocking or content desensitization for high-risk data.
[0151] Whether determined as "high-risk" through the regular process (overall score 1.778) or directly through the circuit breaker mechanism, the post in this example was ultimately classified as high-risk. The system then triggered actions associated with the "high-risk" level. For example, the system API would call the content management platform's interface to execute a "direct blocking" command, preventing the post from being publicly published; or it would perform "content anonymization," replacing the text with a warning message and overlaying or blurring the images to prevent the spread of harmful information.
[0152] Compared with the prior art, this application has the following technical advantages:
[0153] 1. By combining and analyzing multimodal data, it can accurately identify combined and hidden sensitive content, effectively combat content evasion behavior, and significantly improve the accuracy and intelligence level of identification.
[0154] 2. A circuit breaker mechanism for sensitivity symbols of single modal data is designed. When the sensitivity symbol of a certain modal data exceeds the threshold, the multimodal feature fusion step is no longer entered. Instead, the current data is directly identified as sensitive data and the current modal data is identified as the source information of the sensitive modality. This reduces the identification steps of sensitive data and enhances the ability to deal with the identification of massive data.
[0155] 3. Classify the sensitivity level of the data and perform corresponding data processing according to the sensitivity level. This enables the automated processing of a large amount of low-risk and clearly high-risk content, leaving only uncertain medium-risk content to expensive human resources for review, thereby maximizing the efficiency of review resources.
[0156] 4. Avoiding a blanket blocking of all suspected content, allowing for manual judgment in borderline cases reduces the impact of false positives on normal users and achieves a balance between security and user experience.
[0157] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0158] Based on the same inventive concept, this application also provides a sensitive data identification device for implementing the sensitive data identification method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more embodiments of the sensitive data identification device provided below can be found in the limitations of the sensitive data identification method described above, and will not be repeated here.
[0159] In one exemplary embodiment, such as Figure 7 As shown, a sensitive data identification device is provided, comprising: an acquisition module 701, a feature extraction module 702, a symbol generation module 703, and an identification module 704, wherein:
[0160] The acquisition module 701 is used to acquire the target modal data corresponding to the data to be identified; the target modal data includes at least two modal data selected from text modal data, image modal data, and audio modal data;
[0161] The feature extraction module 702 is used to extract features from the target modal data to obtain the intra-modal features corresponding to the target modal data; the intra-modal features include low-level basic features and high-level semantic features;
[0162] The symbol generation module 703 is used to generate target sensitivity symbols corresponding to the target modal data based on the target modal data and the intra-modal features corresponding to the target modal data, using a sensitivity quantification model that matches the modality type of the target modal data; the target sensitivity symbols are used to characterize the sensitivity of the target modal data;
[0163] The identification module 704 is used to obtain the contribution weights corresponding to each target sensitivity symbol through a cross-modal association model, use each contribution weight to perform a weighted summation of each target sensitivity symbol to obtain the comprehensive sensitivity score of the data to be identified, and generate the sensitivity level and sensitization modality tracing information of the data to be identified based on the comprehensive sensitivity score and the pre-set graded sensitivity score threshold.
[0164] In one embodiment, when the target modal data includes text modal data, the feature extraction module 702 is further configured to remove punctuation characters from the text modal data to obtain cleaned text modal data; perform obfuscation pattern detection on the cleaned text modal data, and perform deobfuscation processing on the cleaned text modal data based on the obfuscation pattern detection results to obtain target text modal data; the obfuscation pattern detection results include irrelevant character insertion, single character splitting, and homophone replacement; and perform feature extraction on the target text modal data.
[0165] In one embodiment, the symbol generation module 703 is further configured to perform semantic embedding calculations on the target text modal data using a first sensitivity quantification model that matches the modal type of the target text modal data, obtain the corresponding semantic deviation, obtain the corresponding keyword hit intensity based on the intramodal features of the target text modal data and the pre-set sensitive topics, and generate target sensitivity symbols corresponding to the target text modal data based on the semantic deviation and keyword hit intensity.
[0166] In an exemplary embodiment, when the target modal data includes image modal data, the feature extraction module 702 is further configured to scale or crop the image modal data to obtain image modal data with normalized dimensions; map the pixel values of the image modal data with normalized dimensions to a standardized data range to obtain the target image modal data; and perform feature extraction on the target image modal data.
[0167] In one embodiment, the symbol generation module 703 is further configured to extract embedded text from the target image modal data using a second sensitivity quantification model that matches the modal type of the target image modal data, obtain corresponding image embedded text modal data, perform target detection based on the modal features corresponding to the target image modal data, generate a corresponding spatial semantic adjacency matrix, and generate a target sensitivity symbol corresponding to the target image modal data based on the target sensitivity symbol corresponding to the image embedded text modal data and the spatial semantic adjacency matrix.
[0168] In one embodiment, when the target modal data includes audio modal data, the feature extraction module 702 is further configured to unify the sampling rate in the audio modal data to a preset sampling rate to obtain unified audio modal data; convert the multi-channel audio modal data in the unified audio modal data into single-channel audio modal data to obtain target audio modal data; and perform feature extraction on the target audio modal data.
[0169] In an exemplary embodiment, the symbol generation module 703 is further configured to perform text conversion on the target audio modal data using a third sensitivity quantification model that matches the modality type of the target audio modal data to obtain a corresponding baseline text, obtain corresponding prosodic features and sentiment distribution features based on the intramodal features corresponding to the target audio modal data, and generate target sensitivity symbols corresponding to the target audio modal data according to the target sensitivity symbols, prosodic features and sentiment distribution features corresponding to the baseline text.
[0170] In one embodiment, after generating the target sensitivity symbol corresponding to the target modal data, the sensitive data identification device is further configured to obtain a sensitivity threshold preset for the target modal data; if the sensitivity represented by the target sensitivity symbol is greater than the sensitivity threshold, the data to be identified is determined to be sensitive data.
[0171] Each module in the aforementioned sensitive data identification device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0172] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 8As shown, the computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs in the non-volatile storage media to run. The database stores the data to be identified, the target modality data of the data to be identified, the intra-modal features of the target modality data, the target sensitivity symbol of the target modality data, the sensitivity level of the data to be identified, and the sensitization modality traceability information. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a sensitive data identification method.
[0173] Those skilled in the art will understand that Figure 8 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0174] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the sensitive data identification method of the above embodiment.
[0175] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the sensitive data identification method of the above embodiment.
[0176] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the sensitive data identification method of the above embodiments.
[0177] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0178] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0179] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0180] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method of sensitive data identification, the method comprising: The method includes: Obtain the target modal data corresponding to the data to be identified; the target modal data includes at least two modal data selected from text modal data, image modal data, and audio modal data; Feature extraction is performed on the target modality data to obtain the intra-modality features corresponding to the target modality data; the intra-modality features include low-level basic features and high-level semantic features; Using a sensitivity quantification model that matches the modality type of the target modality data, a target sensitivity symbol is generated based on the target modality data and its corresponding intra-modality features. This target sensitivity symbol characterizes the sensitivity of the target modality data. If the target modality data includes text modality data, feature extraction is performed on the target text modality data. Using a first sensitivity quantification model that matches the modality type of the target text modality data, semantic embedding calculation is performed on the target text modality data to obtain the corresponding semantic deviation. Based on the intra-modality features corresponding to the target text modality data and pre-set sensitive topics, the corresponding keyword hit intensity is obtained. Based on the semantic deviation and the keyword hit intensity, a target sensitivity symbol is generated for the target text modality data. The semantic deviation represents the degree of semantic difference between the current text and the first reference text. When the target modal data includes image modal data, feature extraction is performed on the target image modal data; embedded text is extracted from the target image modal data using a second sensitivity quantification model that matches the modality type of the target image modal data to obtain corresponding image embedded text modal data; target detection is performed based on the intramodal features corresponding to the target image modal data to generate a corresponding spatial semantic adjacency matrix; and target sensitivity symbols corresponding to the target image modal data are generated based on the target sensitivity symbols corresponding to the image embedded text modal data and the spatial semantic adjacency matrix; the spatial semantic adjacency matrix represents the semantic association risk of the detected target object and its spatial proximity in the image; The sensitivity thresholds for each modality are obtained. If the sensitivity symbol generated by any modality exceeds its independent high-risk sensitivity threshold, the multimodal data fusion calculation is stopped, the current data to be identified is determined as sensitive data, and the target modality data corresponding to the sensitivity symbol exceeding the high-risk sensitivity threshold is used as the source information of the sensitized modality of the data to be identified; otherwise... By using a cross-modal association model and employing a self-attention mechanism, the contribution weights corresponding to each target sensitivity symbol are obtained. The contribution weights are then used to perform a weighted summation on each target sensitivity symbol to obtain a comprehensive sensitivity score for the data to be identified. Based on the comprehensive sensitivity score and a pre-set graded sensitivity score threshold, the sensitivity level and sensitization modality tracing information of the data to be identified are generated.
2. The method according to claim 1, characterized in that, When the target modal data includes text modal data, the feature extraction of the target modal data includes: Remove punctuation characters from the text modal data to obtain cleaned text modal data; Obfuscation pattern detection is performed on the cleaned text modal data, and deobfuscation processing is performed on the cleaned text modal data based on the obfuscation pattern detection results to obtain the target text modal data; the obfuscation pattern detection results include irrelevant character insertion, single character splitting, and homophone replacement; Feature extraction is performed on the target text modal data.
3. The method according to claim 1, characterized in that, When the target modal data includes image modal data, the feature extraction of the target modal data includes: The image modal data is scaled or cropped to obtain image modal data with normalized dimensions; The pixel values of the size-normalized image modal data are mapped to a standardized data range to obtain the target image modal data; Feature extraction is performed on the target image modal data.
4. The method according to claim 1, characterized in that, When the target modal data includes audio modal data, the feature extraction of the target modal data includes: The sampling rate in the audio modal data is unified to a preset sampling rate to obtain unified audio modal data; The multi-channel audio modal data in the unified audio modal data is converted into single-channel audio modal data to obtain the target audio modal data; Feature extraction is performed on the target audio modal data.
5. The method according to claim 4, characterized in that, The step of generating a target sensitivity symbol corresponding to the number of target modalities based on the target modal data and its corresponding intra-modal features using a sensitivity quantification model that matches the modality type of the target modal data includes: By using a third sensitivity quantification model that matches the modality type of the target audio modality data, text conversion is performed on the target audio modality data to obtain the corresponding second reference text. Based on the intramodal features corresponding to the target audio modality data, the corresponding prosodic features and sentiment distribution features are obtained. Based on the target sensitivity symbol corresponding to the second reference text, the prosodic features, and the sentiment distribution features, the target sensitivity symbol corresponding to the target audio modality data is generated.
6. A sensitive data identification device, characterized in that, The device includes: The acquisition module is used to acquire target modal data corresponding to the data to be identified; the target modal data includes at least two modal data selected from text modal data, image modal data, and audio modal data. The feature extraction module is used to extract features from the target modality data to obtain the intra-modality features corresponding to the target modality data; the intra-modality features include low-level basic features and high-level semantic features; A symbol generation module is used to generate target sensitivity symbols corresponding to the target modal data based on the target modal data and its corresponding intra-modal features, using a sensitivity quantification model that matches the modality type of the target modal data. The target sensitivity symbols characterize the sensitivity of the target modal data. When the target modal data includes text modal data, feature extraction is performed on the target text modal data. Semantic embedding calculation is performed on the target text modal data using a first sensitivity quantification model that matches the modality type of the target text modal data to obtain the corresponding semantic deviation. Based on the intra-modal features corresponding to the target text modal data and pre-set sensitive topics, the corresponding keyword hit intensity is obtained. Finally, the target text modality number is generated based on the semantic deviation and the keyword hit intensity. According to the corresponding target sensitivity symbol; the semantic deviation is the degree of semantic difference between the current text and the first reference text; when the target modal data includes image modal data, feature extraction is performed on the target image modal data; by using a second sensitivity quantification model that matches the modality type of the target image modal data, embedded text extraction is performed on the target image modal data to obtain the corresponding image embedded text modal data; target detection is performed based on the modal features corresponding to the target image modal data to generate the corresponding spatial semantic adjacency matrix; and according to the target sensitivity symbol corresponding to the image embedded text modal data and the spatial semantic adjacency matrix, the target sensitivity symbol corresponding to the target image modal data is generated; the spatial semantic adjacency matrix represents the semantic association risk of the detected target object and its spatial proximity in the image; The sensitivity thresholds for each modality are obtained. If the sensitivity symbol generated by any modality exceeds its independent high-risk sensitivity threshold, the multimodal data fusion calculation is stopped, the current data to be identified is determined as sensitive data, and the target modality data corresponding to the sensitivity symbol exceeding the high-risk sensitivity threshold is used as the source information of the sensitized modality of the data to be identified; otherwise... The identification module is used to obtain the contribution weights corresponding to each target sensitivity symbol through a cross-modal association model using a self-attention mechanism, and to perform a weighted summation of each target sensitivity symbol using each contribution weight to obtain the comprehensive sensitivity score of the data to be identified. Based on the comprehensive sensitivity score and a pre-set graded sensitivity score threshold, the module generates the sensitivity level and sensitization modality tracing information of the data to be identified.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-modal fusion sensitive information classification detection method
CN113033610A
Sensitive information risk assessment method and system and storage medium
CN113849760A