Instant secret-related information detection method and system based on IM software multi-modal feature fusion
By adopting multimodal feature fusion detection method in IM software, combined with the collaborative hierarchical architecture of client and server, the problems of low detection accuracy and unreasonable system architecture in the existing technology are solved, efficient and real-time confidential information detection is achieved, and the processing capability of multimodal data is enhanced.
Patent Information
- Application Number
- CN202510656113.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The existing IM software confidential information detection technology has problems such as low detection accuracy, unreasonable system architecture, and limitations in technical implementation, making it difficult to effectively prevent confidential information leakage.
The instant confidential information detection method of multi-modal feature fusion of IM software is adopted. Through the hierarchical detection architecture of the collaborative client and server, text feature extraction, multi-modal unified transformation and vectorization are carried out, and real-time risk assessment and interception are achieved through the hierarchical detection architecture of the collaborative client and server.
It improves the real-time and accuracy of detection, reduces the burden on servers, enhances the processing capability of multimodal data, realizes more accurate identification of complex and confidential scenarios, and has the ability to respond quickly and update dynamically.
Smart Images

Figure CN120180435A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information security technology, and particularly to an instant classified information detection method and system for multi-modal feature fusion of IM software. Background Art
[0002] With the wide application of instant messaging software, IM software is widely used for communication and file transfer in daily office work. While this improves work efficiency, it also brings security risks of classified information leakage. Especially in scenarios involving business secrets, sensitive data, or government documents, how to effectively prevent classified information from leaking through IM software is the main problem to be solved by the present invention.
[0003] Currently, the mainstream classified text information detection schemes mainly include the following categories: 1. The method based on keyword matching, such as patent CN117033629A. By presetting a sensitive word library, string matching or text occupancy ratio and occurrence probability of the communication content are used to identify classified information. This method is simple to implement, but prone to false alarms and unable to cope with evasion means such as synonym replacement.
[0004] 2. The method based on file feature recognition, such as patent CN117560215A. By analyzing features such as file format, file name, and file header to judge the classified nature of the file. However, this method can only identify known file types and is powerless for unknown types of classified files.
[0005] 3. The scheme based on server detection. All communication content is uploaded to the server for centralized detection and filtering. This scheme has the risk of privacy leakage and increases the computing burden on the server.
[0006] The above several mainstream schemes all have corresponding defects. For example: 1. Detection accuracy problem The scheme based on keyword matching only performs string-level comparison, unable to understand the semantic content of the text, and is prone to false alarms; the keyword library needs to be manually maintained and updated, and is easily evaded by means such as synonym replacement; the file feature recognition scheme overly relies on the surface features of the file, and these features are easily modified to evade detection.
[0007] 2. System architecture problem Excessive reliance on the server to implement detection, and the client only undertakes the functions of data collection and uploading; the content to be detected is transmitted in plain text or the original file, there is a risk of information leakage; all content needs to be uploaded to the server for centralized processing, bringing a large computing burden to the server; it is necessary to wait for the server level to finish processing, and the real-time performance is poor.
[0008] 3. Technical implementation limitations It only intercepts classified levels based on rules, lacking the ability to deeply understand the semantics of the text; it requires a high level of understanding of classified levels from rule maintenance personnel and lacks usability for customers; it has limited processing capabilities for multi-modal data (images, voices); the behavior is fixed after the detection result is obtained, lacking an early warning mechanism for classifying and pre-processing classified documents in advance, and also lacking a mechanism for timely processing of misjudged documents.
[0009] Based on the above existing technologies, it can be seen that there has been a long-term technical path dependence on "word frequency statistics + rule engine" in the field of classified detection. The main reason is that the industry generally believes that deep semantic models have inherent defects such as high deployment costs (requiring GPU servers) and large inference latencies (average response time > 500ms). Summary of the Invention
[0010] In view of the above situation, the main purpose of the present invention is to propose an instant classified information detection method and system for multi-modal feature fusion of IM software to solve the above technical problems.
[0011] The present invention proposes an instant classified information detection method for multi-modal feature fusion of IM software, and the method includes the following steps: Step 1, monitor the local data transmission situation. When a data transmission behavior occurs, extract text features from the data to obtain text features; Step 2, set a local early warning strategy on the client side, match the text features and the file meta-information of the data with the local early warning strategy, generate an interception instruction in real time according to the matching result, and perform real-time interception on the data transmission based on the interception instruction; Step 3, according to the matching result, compress and encapsulate the data with unclear content and upload it to the server, and send a deep analysis request to the server; Step 4, perform multi-modal unified conversion and vectorization processing on the text, image, and voice content included in the unclear data according to the deep analysis request to generate a text feature vector array; Step 5, perform a deep analysis on the text feature vector array, generate a risk assessment result and an early warning classification mark according to the deep analysis result, and based on the risk assessment result and the early warning classification mark, the client executes corresponding control instructions and dynamically updates the local early warning strategy.
[0012] The present invention also proposes an instant classified information detection system for multi-modal feature fusion of IM software. Among them, the system applies the instant classified information detection method for multi-modal feature fusion of IM software as described above. The system includes a client and a server. The client includes a local vectorization module, a file interception module, and a communication management module. The server includes an early warning management module, a deep analysis module, and a multi-modal processing module; A local vectorization module, which is used for: monitoring the local data transmission situation, and when there is a data transmission behavior, extracting text features from the data to obtain text features; A communication management module, which is used for: setting a local warning policy on the client side, matching the text features and the file meta-information of the data with the local warning policy, and generating an interception instruction in real time according to the matching result; A file interception module, which is used for: intercepting the data transmission in real time based on the interception instruction; According to the matching result, compress and encapsulate the data with unclear content and upload it to the server, and send a deep analysis request to the server; A multimodal processing module, which is used for: according to the deep analysis request, performing multimodal unified conversion and vectorization processing on the text, image and voice content contained in the unclear data to generate an array of text feature vectors; A deep analysis module, which is used for: performing deep analysis on the array of text feature vectors; A warning management module, which is used for: generating a risk assessment result and a warning classification mark according to the deep analysis result, and dynamically updating the local warning policy based on the risk assessment result and the warning classification mark.
[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. The advantages of a hierarchical detection architecture with cooperation between the client and the server; The local vectorization module uses the compressed TinyBERT model for text feature extraction, enabling the text vectorization process to be completed instantaneously on the user side. Combined with the real-time monitoring of the file interception module, the system can trigger analysis immediately at the file upload stage, achieving a millisecond-level preliminary risk assessment, with the advantages of real-time performance and low resource consumption.
[0014] 2. The advantages of multimodal unified processing; The server multimodal processing module integrates OCR (image-to-text) and TTS (speech-to-text), converts unstructured data (such as pictures and voices) into a unified text stream, and inputs it into the deep analysis module, enabling the system to cover all common file types in the IM scenario and significantly improving the detection scope.
[0015] 3. The advantages of a hierarchical warning rule processing architecture; The local preliminary risk assessment and the server deep analysis form a two-level processing mechanism. Files with obvious features are directly intercepted locally, and files with unknown risks are uploaded to the server for further verification, which can effectively reduce the server pressure. During the verification process, semantic-level understanding is adopted, combined with the database, to identify variant expressions and metaphorical content, improving the recognition accuracy in complex classified scenarios.
[0016] 4. The advantages of a continuous improvement strategy dynamic update feedback mechanism; The matching based on vector similarity enables the system to capture semantic-level correlations. The sensitive database continuously updates vector features to cover new types of classified modes. At the same time, the early warning management module supports administrators to import new sensitive files and adjust matching rules, forming a "detection - feedback - optimization" closed loop, with the ability to quickly respond to policies.
[0017] Additional aspects and advantages of the present invention will be given in part in the following description, will become apparent in part from the following description, or will be understood through the embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 is a flowchart of an instant classified information detection method for multi-modal feature fusion of an IM software proposed by the present invention; Figure 2 is a schematic structural diagram of an instant classified information detection system for multi-modal feature fusion of an IM software proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention.
[0020] These and other aspects of the embodiments of the present invention will be clear with reference to the following description and drawings. In these descriptions and drawings, some specific embodiments of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited by this.
[0021] The present invention deploys a TinyBERT (lightweight semantic conversion vector model) model on the client side to generate local vectorization capabilities; configures the distribution mechanism of the file interception module to separate and process text, image, and voice content to generate multi-modal parallel processing channels; designs the meta-information encapsulation format of the communication management module to integrate file meta-information features to generate a three-dimensional feature encapsulation data stream; establishes a sensitive document database on the server and a dynamic rule distribution channel to generate a closed-loop policy update mechanism to achieve a hierarchical processing architecture for client-server collaboration.
[0022] Please refer to Figure 1 , this embodiment provides an instant classified information detection method for multi-modal feature fusion of an IM software, and the method includes the following steps: Step 1, monitor the local data transmission situation. When a data transmission behavior occurs, extract text features from the data to obtain text features; As a preferred embodiment of the present invention, extracting text features from data to obtain text features specifically includes the following steps: Clean and sentence-process the received text content to generate a standardized text sequence; Perform sentence segmentation and sequence segmentation on the normalized text sequence to obtain a set of normalized text sequences after sentence segmentation; The standardized text sequence set after sentence segmentation is converted into word vectors using the embedding layer of the lightweight semantic conversion vector model; Perform lightweight multi-head attention calculation on word vectors to capture semantic associations and obtain attention output; The attention output is used to extract key semantic features using a shallow network after distillation; Perform mean pooling or maximum pooling on key semantic features to generate sentence-level feature vectors and obtain text features.
[0023] As a preferred embodiment of the present invention, performing lightweight multi-head attention calculation on word vectors to capture semantic associations and obtain attention output specifically includes the following steps: Linearly project the word vector to generate the query vector Q, key vector K and value vector V; Calculate the dot product similarity between the query vector Q and the key vector K, and then divide each dot product similarity by the square root of the dimension of the key vector V for normalization to obtain the preliminary attention weight matrix; Calculate the second norm of each query vector Q and key vector K respectively to obtain the second norm of the query vector Q and the second norm of the key vector K; Multiply the query vector Q's 2-norm by the key vector K's 2-norm to get the vector modulus; The vector modulus is transformed nonlinearly using the hyperbolic tangent function to obtain the fractal modulation factor, where two learnable parameters in the hyperbolic tangent function are used to control the steepness of the curve and the dynamically adjusted activation threshold. The preliminary attention weight matrix is fused with the fractal modulation factor by element-by-element multiplication to obtain the attention weight matrix fused with the dynamic modulation factor; The attention output is obtained by weighted summing the value vector V using the attention weight matrix that incorporates the dynamic modulation factor.
[0024] In a preferred embodiment of the present invention, the present invention performs a nonlinear transformation on the vector modulus through a hyperbolic tangent function, dynamically suppresses the attention weights of low-association word pairs, adds a modulus multiplication term on the basis of traditional dot product attention, strengthens the attention weights of high-association word pairs, and attenuates weakly associated word pairs whose modulus product is lower than a threshold, thereby enhancing the ability to capture semantic associations of long-tail distributions.
[0025] Step 2: Set a local warning policy on the client side, match the text features and the file meta-information of the data with the local warning policy, generate an interception instruction in real time according to the matching result, and perform real-time interception on the data transmission based on the interception instruction; As a preferred embodiment of the present invention, matching the text features and the file meta-information of the data with the local warning policy and generating an interception instruction in real time according to the matching result specifically include the following steps: Construct a feature matching rule according to the file meta-information, and scan and classify the external information of the data according to the feature matching rule to generate a preliminary classified-secret determination result; Calculate the similarity between the key classified-secret feature vectors saved locally and the text vector array to generate a semantic-level risk assessment result; According to the preliminary classified-secret determination result and the semantic-level risk assessment result, trigger operations such as blocking sending, isolation or warning, and generate a real-time interception instruction.
[0026] As a preferred embodiment of the present invention, calculating the similarity between the key classified-secret feature vectors saved locally and the text vector array to generate a semantic-level risk assessment result specifically include the following steps: Encode the file meta-information to generate a file meta-information feature vector; Quantify the confidence of the file meta-information feature vector based on the L1 norm to obtain a meta-information confidence index; Input the file meta-information feature vector into the fully connected layer, and then perform non-linear activation to obtain an enhanced file meta-information feature vector; Construct a dual-channel mapping module using a dual-channel structure, and set an independent fully connected layer in each channel. Input the text vector array into the dual-channel mapping module to capture lexical features and identify semantic patterns respectively to obtain a dual-channel mapping result; Perform feature interaction on the enhanced file meta-information feature vector and the dual-channel mapping result in an element-wise multiplication manner to obtain an interaction reinforcement result of the text segment; Select the segment feature with the largest L2 norm from all the interaction reinforcement results of the text segments to obtain an optimal text interaction feature vector; Take the sum of the L1 norms of all text segment vectors in the optimal text interaction feature vector as the complexity of the text content, and generate a dynamic weight according to the proportional relationship between the meta-information confidence index and the complexity of the text content; Perform weighted synthesis on the enhanced file meta-information feature vector and the optimal text interaction feature vector using the dynamic weight to obtain a fused feature vector, perform normalization processing on the fused feature vector, and map it to the risk score interval of 0-1 to obtain a dynamic fusion risk score; Match the dimension of the text vector array with the key classified feature vector. If the dimensions match, calculate the matching degree score of the classified feature; if not, give a default matching degree score of the classified feature as a supplementary calculation. Perform risk decision synthesis on the matching degree score of the classified feature and the dynamic fusion risk score. During the synthesis process, adjust the weight in the synthesis process according to the confidence level of the matching degree score of the classified feature to obtain the final synthesis score. Map the final synthesis score to a risk level and generate corresponding interception strategies according to the corresponding risk level.
[0027] As a preferred embodiment of the present invention, the specific steps for calculating the matching degree score of the classified feature are as follows: Regularize the time dimension of the text vector array to obtain an aligned text vector array. Perform a sliding window comparison between each key classified feature vector and the aligned text vector array, calculate the cosine similarity of the vectors within the window, and obtain the similarity score of each window. Count the positions where the similarity scores of each window exceed the set threshold to obtain the matching degree score of the classified feature.
[0028] In a preferred embodiment of the present invention, through the sliding window comparison mechanism of the local classified feature library, the historical sensitive data patterns can be accurately matched, solving the problem that traditional keyword matching cannot capture semantic features. And the dynamic feature fusion path can identify variant leakage means using synonym replacement and syntactic structure reorganization through the non-linear interaction between meta-information and text. Moreover, the classified feature comparison path is only activated when the text vector dimensions match, avoiding full-scale calculation of low-risk files to optimize the computational load and achieve the purpose of improving resource efficiency.
[0029] Step 3: According to the matching result, compress and encapsulate the content-ambiguous data and upload it to the server, and send a deep analysis request to the server. As a preferred embodiment of the present invention, compressing and encapsulating the content-ambiguous data and uploading it to the server and sending a deep analysis request to the server specifically include the following steps: Perform compression and encryption processing on the content-ambiguous data to generate a lightweight transmission data packet. Encapsulate the file meta-information of the content-ambiguous data to generate a structured analysis request. Upload the lightweight transmission data packet and the structured analysis request to the server.
[0030] Step 4: Perform multi-modal unified conversion and vectorization processing on the text, image, and voice content contained in the ambiguous data according to the deep analysis request to generate a text feature vector array. As a preferred embodiment of the present invention, performing multi-modal unified conversion and vectorization processing on the text, image, and voice content included in the ambiguous data according to the in-depth analysis request, and generating a text feature vector array specifically includes the following steps: For the text, image, and voice content included in the ambiguous data, identify the non-text regions in the image through a pre-trained convolutional neural network model, and generate a noise region coordinate mask map; Based on the noise region coordinate mask map, use an image inpainting algorithm to remove interference information and generate a denoised image; Adopt OCR to extract text from the denoised image, obtain the image text, and then perform context verification and alignment on the image text to obtain a high-precision text stream; Adopt TTS to convert the voice content into continuous speech text, frame the continuous speech text according to timestamps, and construct context dependencies in combination with the Transformer model to obtain a sequence of semantic segments with timestamps; Align the text, high-precision text stream, and sequence of semantic segments with timestamps in time series to obtain a time-synchronized multi-modal text sequence; Map the time-synchronized multi-modal text sequence to the same semantic space through a shared embedding layer, and adopt a cross-modal attention mechanism for fusion, outputting the fused multi-modal feature vector, and using the fused multi-modal feature vector as the text feature vector array.
[0031] In the preferred embodiment of the present invention, the present invention establishes an audio-image noise correlation through spatial perception denoising and cross-modal noise filtering, automatically enhances the image denoising intensity when detecting speech background noise, and designs a three-level time anchor system, effectively improving the accuracy of time series alignment.
[0032] Step 5: Perform in-depth analysis on the text feature vector array, generate a risk assessment result and a warning classification mark according to the in-depth analysis result, and based on the risk assessment result and the warning classification mark, the client executes corresponding control instructions and dynamically updates the local warning strategy.
[0033] As a preferred embodiment of the present invention, performing in-depth analysis on the text feature vector array and generating a risk assessment result and a warning classification mark according to the in-depth analysis result specifically includes the following steps: Calculate the cosine similarity between the text feature vector array and the feature vectors in the sensitive document database, and screen out the Top-K candidate matching items; Perform context reasoning on the Top-K candidate matching items through a large model to generate context relevance; Generate a comprehensive classified probability score based on the cosine similarity and the context relevance, and use the comprehensive classified probability score as the risk assessment result; Load the hierarchical warning rules configured by the administrator, match the thresholds of the hierarchical warning rules according to the classified probability score, and generate corresponding warning classification marks.
[0034] In a preferred embodiment of the present invention, a sensitive document database and a more accurate complete BERT model or a variant model customized for specific tasks are set up. Based on the vectorized data uploaded by the database and the client, the vector similarity is calculated, which can quickly perform feature comparison, evaluate file features and obtain preliminary evaluation results. This process can generate an analysis report for filing, and the results are returned to the client and the warning management module for execution. The database includes existing sensitive document vector features and file features, and will automatically analyze new vector features and file features of newly imported sensitive files by the administrator through a high-precision model, or customize features according to the administrator's needs after reporting the analysis report, thereby realizing incremental update of the database. This process can further develop a steganography file analyzer for analyzing difficult-to-analyze binary files, etc., and can also deploy a large model to handle customer misjudgment appeals, which are all optional.
[0035] Moreover, the advantage of in-depth analysis by the server lies in high accuracy, fast speed, no resource limitations, and "association analysis" can use more advanced technologies such as large models and semantic rearrangement models to deeply analyze and calculate semantic similarity calculation and matching of classified categories.
[0036] As a preferred embodiment of the present invention, dynamically updating the local warning policy specifically includes the following steps: According to the risk assessment result, confirm whether the data with unclear content belongs to classified data or sensitive data; If it belongs to classified data or sensitive data, trigger the policy update requirement and obtain the corresponding file meta-information features; Extract key features from the data with unclear content that belongs to classified data or sensitive data to obtain key classified text feature vectors; Integrate the file meta-information features, key classified text feature vectors, and warning classification marks into an update package, and encrypt and send it to the client to update the local warning policy.
[0037] In this step, the warning management module analyzes the results to determine the warning level. At the same time, the administrator will maintain and configure the warning rules, and issue the local warning policy of the client in real time, providing data support for system optimization through the archived management of warning records. The administrator can feedback to the overall system by inputting new matching warning rules, or import new sensitive files and feedback them to the in-depth analysis module, and distribute them to the client interception module in grades.
[0038] Please refer to Figure 2, this embodiment also provides an instant classified information detection system for IM software multi-modal feature fusion. Among them, the system applies the instant classified information detection method for IM software multi-modal feature fusion as described above. The system includes a client and a server. The client includes a local vectorization module, a file interception module, and a communication management module. The server includes an early warning management module, a deep analysis module, and a multi-modal processing module; The local vectorization module is used for: monitoring the local data transmission situation. When a data transmission behavior occurs, extracting text features from the data to obtain text features; The communication management module is used for: setting a local early warning policy on the client, matching the text features and the file meta-information of the data with the local early warning policy, and generating an interception instruction in real time according to the matching result; The file interception module is used for: performing real-time interception on the data transmission based on the interception instruction; According to the matching result, compressing and encapsulating the data with unclear content and uploading it to the server, and sending a deep analysis request to the server; The multi-modal processing module is used for: according to the deep analysis request, performing multi-modal unified conversion and vectorization processing on the text, image, and voice content included in the unclear data to generate a text feature vector array; The deep analysis module is used for: performing deep analysis on the text feature vector array; The early warning management module is used for: generating a risk assessment result and an early warning classification mark according to the deep analysis result, and dynamically updating the local early warning policy based on the risk assessment result and the early warning classification mark.
[0039] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are shown in sequence according to the indication of the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily need to be executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages does not necessarily need to be sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.
[0040] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.
[0041] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0042] The above-described embodiments merely represent several implementation manners of the present invention, and the description thereof is relatively specific and detailed, but should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
Claims
1. A method for detecting instant confidential information by fusion of multimodal features of IM software, characterized in that: The method comprises the following steps: Step 1: monitor the local data transmission situation, and when data transmission behavior occurs, extract text features from the data to obtain text features; Step 2: Set a local warning strategy on the client, match the text features and the file meta information of the data with the local warning strategy, generate an interception instruction in real time according to the matching result, and intercept the data transmission in real time based on the interception instruction; Step 3: According to the matching results, the data with unclear content is compressed and packaged and uploaded to the server, and a deep analysis request is sent to the server; Step 4: Perform multimodal unified conversion and vectorization processing on the text, image and voice content contained in the ambiguous data according to the deep analysis request to generate a text feature vector array; Step 5: Perform a deep analysis on the text feature vector array, generate risk assessment results and warning classification marks based on the deep analysis results, and based on the risk assessment results and warning classification marks, the client executes corresponding control instructions and dynamically updates the local warning strategy.
2. The instant confidential information detection method based on multimodal feature fusion of IM software according to claim 1 is characterized in that: In step 1, extracting text features from data to obtain text features specifically includes the following steps: Clean and sentence-process the received text content to generate a standardized text sequence; Perform sentence segmentation and sequence segmentation on the normalized text sequence to obtain a set of normalized text sequences after sentence segmentation; The standardized text sequence set after sentence segmentation is converted into word vectors using the embedding layer of the lightweight semantic conversion vector model; Perform lightweight multi-head attention calculation on word vectors to capture semantic associations and obtain attention output; The attention output is used to extract key semantic features using a shallow network after distillation; Perform mean pooling or maximum pooling on key semantic features to generate sentence-level feature vectors and obtain text features.
3. The instant confidential information detection method based on multimodal feature fusion of IM software according to claim 2 is characterized in that: Performing lightweight multi-head attention calculation on word vectors to capture semantic associations and obtain attention output specifically includes the following steps: Linearly project the word vector to generate the query vector Q, key vector K and value vector V; Calculate the dot product similarity between the query vector Q and the key vector K, and then divide each dot product similarity by the square root of the dimension of the key vector V for normalization to obtain the preliminary attention weight matrix; Calculate the second norm of each query vector Q and key vector K respectively to obtain the second norm of the query vector Q and the second norm of the key vector K; Multiply the query vector Q's 2-norm by the key vector K's 2-norm to get the vector modulus; The vector modulus is transformed nonlinearly using the hyperbolic tangent function to obtain the fractal modulation factor, where two learnable parameters in the hyperbolic tangent function are used to control the steepness of the curve and the dynamically adjusted activation threshold. The preliminary attention weight matrix is fused with the fractal modulation factor by element-by-element multiplication to obtain the attention weight matrix fused with the dynamic modulation factor; The attention output is obtained by weighted summing the value vector V using the attention weight matrix that incorporates the dynamic modulation factor.
4. The instant confidential information detection method based on multimodal feature fusion of IM software according to claim 3 is characterized in that: In step 2, matching the text features and the file meta information of the data with the local warning strategy, and generating an interception instruction in real time according to the matching result specifically includes the following steps: Construct feature matching rules based on file metadata, and scan and classify external data information based on the feature matching rules to generate preliminary confidentiality determination results; Generate semantic-level risk assessment results by calculating similarity between the locally saved key confidential feature vectors and the text vector array; Based on the preliminary confidentiality determination results and semantic-level risk assessment results, blocking, isolation or warning operations are triggered and real-time interception instructions are generated.
5. The instant confidential information detection method based on multimodal feature fusion of IM software according to claim 4 is characterized in that: The similarity calculation between the locally saved key confidential feature vector and the text vector array is performed to generate the semantic-level risk assessment results, which specifically includes the following steps: Encode the file meta information to generate a file meta information feature vector; The confidence index of the meta-information is obtained by quantifying the confidence of the feature vector of the file meta-information based on the L1 norm. The file meta information feature vector is input into the fully connected layer, and then nonlinear activation is performed to obtain the enhanced file meta information feature vector; A dual-channel mapping module is constructed using a dual-channel structure, and an independent fully connected layer is set in each channel. The text vector array is input into the dual-channel mapping module to capture lexical features and recognize semantic patterns, and obtain dual-channel mapping results. The enhanced file meta-information feature vector and the dual-channel mapping result are subjected to feature interaction by element-wise product to obtain the interactive enhancement result of the text fragment; Select the segment feature with the largest L2 norm from the interactive reinforcement results of all text segments to obtain the optimal text interaction feature vector; The sum of the L1 norms of all text segment vectors in the optimal text interaction feature vector is used as the complexity of the text content, and a dynamic weight is generated based on the proportional relationship between the meta-information confidence index and the complexity of the text content; The enhanced file meta-information feature vector and the optimal text interaction feature vector are weighted and synthesized using dynamic weights to obtain a fused feature vector, which is normalized and mapped to a risk score interval of 0-1 to obtain a dynamic fused risk score; The text vector array is dimensionally matched with the key confidential feature vector. If the dimensions match, the confidential feature matching score is calculated; if not, the default confidential feature matching score is given as a supplementary calculation; The confidential feature matching score and the dynamic fusion risk score are synthesized for risk decision making. During the synthesis process, the weights in the synthesis process are adjusted according to the confidence level of the confidential feature matching score to obtain the final synthetic score. The final synthetic score is mapped to the risk level, and the corresponding interception strategy is generated according to the corresponding risk level.
6. The instant confidential information detection method based on multimodal feature fusion of IM software according to claim 5 is characterized in that: In step 3, the data with unclear content is compressed and packaged and uploaded to the server, and a deep analysis request is sent to the server, which specifically includes the following steps: Compress and encrypt data with unclear content to generate lightweight transmission data packets; Encapsulate file meta information of ambiguous data and generate structured analysis requests; Upload lightweight transmission data packets and structured analysis requests to the server.
7. The instant confidential information detection method based on multimodal feature fusion of IM software according to claim 6 is characterized in that: In step 4, the text, image and voice content contained in the ambiguous data are converted and vectorized in a multi-modal unified manner according to the deep analysis request, and the generation of a text feature vector array specifically includes the following steps: The text, image and voice content contained in the unclear data are used to identify the non-text area in the image through a pre-trained convolutional neural network model to generate a noise area coordinate mask map; Based on the noise region coordinate mask map, an image restoration algorithm is used to remove interference information and generate a denoised image. Use OCR to extract text from the denoised image to obtain the image text, and then perform context verification and alignment on the image text to obtain a high-precision text stream; Use TTS to convert speech content into continuous speech text, divide the continuous speech text into frames according to timestamps, and combine the Transformer model to build context dependencies to obtain a sequence of semantic fragments with timestamps. The intent classification model and the sentiment polarity model are used to perform intent recognition and sentiment analysis on the semantic fragment sequences with timestamps, respectively, to obtain semantic fragments with scene labels and sentiment scores. According to the word frequency and scene label in the semantic fragment with scene label and sentiment score, the TF-IDF weighted calculation is performed on the semantic fragment with scene label and sentiment score to obtain the structured semantic fragment; Align the timestamps of the text, the timestamps of the high-precision text stream, and the timestamps of the structured semantic fragments to obtain a time-synchronized multimodal text sequence; The time-synchronized multimodal text sequences are mapped to the same semantic space through a shared embedding layer and fused using a cross-modal attention mechanism. The fused multimodal feature vector is output and used as the text feature vector array.
8. The instant confidential information detection method based on multimodal feature fusion of IM software according to claim 7 is characterized in that: In step 5, a deep analysis is performed on the text feature vector array, and generating a risk assessment result and a warning classification mark according to the deep analysis result specifically includes the following steps: Calculate the cosine similarity between the text feature vector array and the feature vector in the sensitive document database to filter out the Top-K candidate matches; Use a large model to perform contextual reasoning on the Top-K candidate matches and generate contextual relevance. Generate a comprehensive confidentiality probability score based on cosine similarity and context relevance, and use the comprehensive confidentiality probability score as the risk assessment result; Load the hierarchical warning rules configured by the administrator, match the threshold of the hierarchical warning rules according to the confidentiality probability score, and generate the corresponding warning classification mark.
9. The instant confidential information detection method based on multimodal feature fusion of IM software according to claim 8 is characterized in that: In step 5, dynamically updating the local early warning strategy specifically includes the following steps: Based on the risk assessment results, confirm whether the data with unclear content is classified or sensitive data; If it is confidential or sensitive data, it triggers the policy update requirement and obtains the corresponding file meta-information features; Extract key features of confidential or sensitive data with unclear content to obtain key confidential text feature vectors; The file meta-information features, key confidential text feature vectors and warning classification marks are integrated into an update package, and sent to the client in an encrypted manner to update the local warning strategy.
10. An instant confidential information detection system based on multimodal feature fusion of IM software, characterized in that: The system applies the instant confidential information detection method of IM software multimodal feature fusion as described in any one of claims 1 to 9, and the system includes a client and a server, the client includes a local vectorization module, a file interception module and a communication management module, and the server includes an early warning management module, a deep analysis module and a multimodal processing module; The local vectorization module is used to: monitor the local data transmission situation, and when data transmission occurs, extract text features from the data to obtain text features; The communication management module is used to: set a local warning strategy on the client, match the text features and the file meta information of the data with the local warning strategy, and generate interception instructions in real time according to the matching results; The file interception module is used to: intercept data transmission in real time based on the interception instruction; According to the matching results, the data with unclear content is compressed and packaged and uploaded to the server, and a deep analysis request is sent to the server; The multimodal processing module is used to: perform multimodal unified conversion and vectorization processing on the text, image and voice content contained in the ambiguous data according to the deep analysis request, and generate a text feature vector array; The deep analysis module is used to: perform deep analysis on the text feature vector array; The early warning management module is used to generate risk assessment results and early warning classification marks according to the in-depth analysis results, and dynamically update the local early warning strategy based on the risk assessment results and early warning classification marks.
Citation Information
Patent Citations
Real-time monitoring method and system based on file content received by mobile terminal
CN112422739A
Text data acquisition method and device, equipment and storage medium
CN115294575A
Medical data item asset sensitivity identification method and system, terminal and medium
CN118116609A
Sensitive text filtering method based on end-side cloud collaboration
CN118193720A
Multi-modal data detection method applied to database security
CN118585677A
Cited By
Scenarized hierarchical management and control method for education application in data space
CN121029833A
Multi-modal large model-based law enforcement case handling document intelligent auditing system and method
CN121212989A
Secret-related document classification and exchange permission management and control method based on natural language processing
CN121389190A