Instant classified information detection method and system based on multimodal feature fusion of IM software

Through client-server collaborative architecture and multimodal data processing, the TinyBERT model and in-depth analysis technology are used to solve the accuracy and real-time problem of confidential information detection of IM software, and efficient identification and real-time interception of confidential information are achieved.

CN120180435BActive Publication Date: 2025-07-29JIANGXI TONGRUI INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510656113.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-07-29
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The existing IM software confidential information detection technology has problems such as low detection accuracy, poor real-time performance, large computing burden, limited multimodal data processing capabilities and easy to be avoided.

Method used

The client-server collaborative architecture is adopted, and the local text feature extraction and preliminary risk assessment is used to use the TinyBERT model, real-time interception is performed in combination with the file interception module, and the unstructured data is converted into text streams through the multi-modal processing module, and precise verification is performed in combination with the server's in-depth analysis, and early warning strategies are updated dynamically.

Benefits of technology

A preliminary risk assessment at millisecond level has been realized, which significantly improves detection accuracy and real-time performance, reduces the burden on the server, and can identify variant expressions and metaphorical content in complex and confidential scenarios, and has the ability to respond quickly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180435B_ABST
    Figure CN120180435B_ABST
Patent Text Reader

Abstract

The present invention proposes an instant classified information detection method and system for multimodal feature fusion of IM software. By deploying the TinyBERT model on the client side, local vectorization capabilities are generated; the distribution mechanism of the profile interception module is configured to separate and process text, image, and voice content, generating multimodal parallel processing channels; the meta-information encapsulation format of the communication management module is designed to integrate file meta-information features, generating a three-dimensional feature encapsulation data stream; a server sensitive document database and a dynamic rule distribution channel are established to generate a closed-loop policy update mechanism, so as to implement a hierarchical processing architecture for client-server collaboration. The present invention can trigger analysis immediately at the file upload stage, achieving millisecond-level preliminary risk assessment, and having the advantages of real-time performance and low resource consumption. When encountering data with unknown risks, it is uploaded to the server for further verification to improve the recognition accuracy in complex classified scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information security technology, and particularly relates to an instant classified information detection method and system for multi-modal feature fusion of IM software. Background Art

[0002] With the wide application of instant messaging software, IM software is widely used for communication and file transfer in daily office work. While this improves work efficiency, it also brings security risks of classified information leakage. Especially in scenarios involving commercial secrets, sensitive data, or government documents, how to effectively prevent the leakage of classified information through IM software is the main problem to be solved by the present invention.

[0003] Currently, the mainstream classified text information detection schemes mainly include the following categories:

[0004] 1. The method based on keyword matching, such as patent CN117033629A. By presetting a sensitive word library, string matching or text occupancy ratio and occurrence probability are used to identify classified information in the communication content. This method is simple to implement, but prone to false alarms and unable to cope with evasion means such as synonym replacement.

[0005] 2. The method based on file feature recognition, such as patent CN117560215A. By analyzing features such as file format, file name, and file header to judge the classified nature of the file. However, this method can only identify known file types and is powerless for unknown types of classified files.

[0006] 3. The scheme based on server detection. All communication content is uploaded to the server for centralized detection and filtering. This scheme has the risk of privacy leakage and increases the computing burden on the server.

[0007] The above several mainstream schemes all have corresponding defects. For example:

[0008] 1. Detection accuracy problem

[0009] The scheme based on keyword matching only conducts string-level comparison, unable to understand the semantic content of the text, and prone to false alarms; the keyword library needs to be manually maintained and updated, and is easily evaded by means such as synonym replacement; the file feature recognition scheme overly relies on the surface features of the file, and these features are easily modified to evade detection.

[0010] 2. System architecture problem

[0011] Over-reliance on the server to implement detection, the client only undertakes the functions of data collection and uploading; transmitting the content to be detected in plain text or the original file has the risk of information leakage; all content needs to be uploaded to the server for centralized processing, bringing a large computing burden to the server; it is necessary to wait for the server layer to finish processing, and the real-time performance is poor.

[0012] 3. Technical implementation limitations

[0013] It only intercepts classified levels based on rules, lacking the ability to deeply understand the semantics of the text; it has high requirements for rule maintenance personnel to understand classified levels and lacks usability for customers; it has limited processing ability for multi-modal data (images, voices); the behavior is fixed after the detection result is obtained, lacking an early warning mechanism for classifying and pre-processing classified documents in advance and also lacking a mechanism for timely processing misjudged documents.

[0014] Based on the above existing technologies, it can be seen that there has been a long-term technical path dependence on "word frequency statistics + rule engine" in the field of classified detection. The main reason is that the industry generally believes that deep semantic models have inherent defects such as high deployment costs (requiring GPU servers) and large inference latencies (average response time > 500ms). Summary of the Invention

[0015] In view of the above situation, the main purpose of the present invention is to propose an instant classified information detection method and system for multi-modal feature fusion of IM software to solve the above technical problems.

[0016] The present invention proposes an instant classified information detection method for multi-modal feature fusion of IM software, and the method includes the following steps:

[0017] Step 1, monitor the local data transmission situation. When a data transmission behavior occurs, extract text features from the data to obtain text features;

[0018] Step 2, set a local early warning strategy on the client side, match the text features and the file meta-information of the data with the local early warning strategy, generate an interception instruction in real time according to the matching result, and perform real-time interception on the data transmission based on the interception instruction;

[0019] Step 3, according to the matching result, compress and encapsulate the data with unclear content and upload it to the server, and send a deep analysis request to the server;

[0020] Step 4, perform multi-modal unified conversion and vectorization processing on the text, image and voice content included in the unclear data according to the deep analysis request to generate a text feature vector array;

[0021] Step 5, perform deep analysis on the text feature vector array, generate a risk assessment result and an early warning classification mark according to the deep analysis result, and based on the risk assessment result and the early warning classification mark, the client executes corresponding control instructions and dynamically updates the local early warning strategy.

[0022] The present invention also provides an instant classified information detection system for multi-modal feature fusion of an IM software. Among them, the system applies the instant classified information detection method for multi-modal feature fusion of the IM software as described above. The system includes a client and a server. The client includes a local vectorization module, a file interception module, and a communication management module. The server includes an early warning management module, a deep analysis module, and a multi-modal processing module;

[0023] The local vectorization module is used for: monitoring the local data transmission situation. When a data transmission behavior occurs, extracting text features from the data to obtain text features;

[0024] The communication management module is used for: setting a local early warning policy on the client, matching the text features and the file meta-information of the data with the local early warning policy, and generating an interception instruction in real time according to the matching result;

[0025] The file interception module is used for: performing real-time interception on the data transmission based on the interception instruction;

[0026] According to the matching result, compressing and encapsulating the data with unclear content and uploading it to the server, and sending a deep analysis request to the server;

[0027] The multi-modal processing module is used for: according to the deep analysis request, performing multi-modal unified conversion and vectorization processing on the text, image, and voice content included in the unclear data to generate a text feature vector array;

[0028] The deep analysis module is used for: performing deep analysis on the text feature vector array;

[0029] The early warning management module is used for: generating a risk assessment result and an early warning classification mark according to the deep analysis result, and dynamically updating the local early warning policy based on the risk assessment result and the early warning classification mark.

[0030] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0031] 1. The advantage of the hierarchical detection architecture coordinated by the client and the server;

[0032] The local vectorization module uses the compressed TinyBERT model to extract text features, enabling the text vectorization process to be instantaneously completed on the user side. Combined with the real-time monitoring of the file interception module, the system can trigger analysis immediately at the file upload stage, achieving a millisecond-level preliminary risk assessment, with the advantages of real-time performance and low resource consumption.

[0033] 2. The advantage of multi-modal unified processing;

[0034] The server multi-modal processing module integrates OCR (image-to-text) and TTS (speech-to-text), converts unstructured data (such as pictures and voices) into a unified text stream, and inputs it into the in-depth analysis module, enabling the system to cover all common file types in the IM scenario and significantly expanding the detection scope.

[0035] 3. Advantages of the hierarchical early warning rule processing architecture;

[0036] The local preliminary risk assessment and the server in-depth analysis form a two-level processing mechanism. Files with obvious features are directly intercepted locally, and those with unknown risks are uploaded to the server for further verification, which can effectively relieve the server pressure. During the verification process, semantic-level understanding is adopted, combined with the database, to identify variant expressions and metaphorical content, improving the recognition accuracy in complex classified scenarios.

[0037] 4. Advantages of the dynamic update feedback mechanism of the continuous improvement strategy;

[0038] The matching based on vector similarity enables the system to capture the semantic-level relevance. The sensitive database continuously updates vector features to cover new classified modes. At the same time, the early warning management module supports the administrator to import new sensitive files and adjust the matching rules, forming a "detection - feedback - optimization" closed loop, with the ability to quickly respond to policies.

[0039] The additional aspects and advantages of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the embodiments of the present invention. Description of the Drawings

[0040] Figure 1 It is a flowchart of the instant classified information detection method for IM software multi-modal feature fusion proposed by the present invention;

[0041] Figure 2 It is a schematic structural diagram of the instant classified information detection system for IM software multi-modal feature fusion proposed by the present invention. Detailed Embodiments

[0042] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention.

[0043] Referring to the following description and drawings, these and other aspects of the embodiments of the present invention will become clear. In these descriptions and drawings, some specific embodiments of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited thereto.

[0044] In the present invention, by deploying the TinyBERT (lightweight semantic conversion vector model) model on the client side, local vectorization capabilities are generated; a distribution mechanism for the profile interception module is configured to separate and process text, image, and voice content to generate a multi-modal parallel processing channel; a meta-information encapsulation format for the communication management module is designed to integrate file meta-information features to generate a three-dimensional feature encapsulation data stream; a server sensitive document database and a dynamic rule distribution channel are established to generate a closed-loop policy update mechanism, so as to implement a hierarchical processing architecture for client-server collaboration.

[0045] Please refer to Figure 1 , this embodiment provides an instant classified information detection method for multi-modal feature fusion of IM software, and the method includes the following steps:

[0046] Step 1, monitor the local data transmission situation. When a data transmission behavior occurs, extract text features from the data to obtain text features;

[0047] As a preferred embodiment of the present invention, extracting text features from the data to obtain text features specifically includes the following steps:

[0048] Clean and split sentences of the received text content to generate a normalized text sequence;

[0049] Perform sentence splitting and sequence segmentation on the normalized text sequence to obtain a set of standardized text sequences after sentence splitting;

[0050] Convert text characters in the set of standardized text sequences after sentence splitting into word vectors using the embedding layer of the lightweight semantic conversion vector model;

[0051] Perform lightweight multi-head attention calculation on the word vectors to capture semantic associations and obtain an attention output;

[0052] Extract key semantic features from the attention output using a distilled shallow network;

[0053] Perform mean pooling or max pooling on the key semantic features to generate a sentence-level feature vector and obtain text features.

[0054] As a preferred embodiment of the present invention, performing lightweight multi-head attention calculation on the word vectors to capture semantic associations and obtain an attention output specifically includes the following steps:

[0055] Perform linear projection on the word vectors to generate a query vector Q, a key vector K, and a value vector V;

[0056] Calculate the dot product similarity between the query vector Q and the key vector K, and then divide each dot product similarity by the square root of the dimension of the key vector V for normalization processing to obtain a preliminary attention weight matrix;

[0057] Calculate the two - norm of each query vector Q and key vector K respectively to obtain the two - norm of the query vector Q and the two - norm of the key vector K;

[0058] Multiply the two - norm of the query vector Q by the two - norm of the key vector K to obtain the vector norm length;

[0059] Use the hyperbolic tangent function to perform a non - linear transformation on the vector norm length to obtain a fractal modulation factor. In the hyperbolic tangent function, two learnable parameters are respectively used to control the steepness of the curve and the activation threshold for dynamic adjustment;

[0060] Fuse the preliminary attention weight matrix and the fractal modulation factor in an element - by - element multiplication manner to obtain an attention weight matrix that incorporates the dynamic modulation factor;

[0061] Use the attention weight matrix that incorporates the dynamic modulation factor to perform a weighted sum on the value vector V to obtain the attention output.

[0062] In a preferred embodiment of the present invention, the present invention performs a non - linear transformation on the vector norm length through the hyperbolic tangent function, dynamically suppressing the attention weights of low - correlation word pairs, adding a modulus - multiplication term on the basis of the traditional dot - product attention, strengthening the attention weights of high - correlation word pairs, and at the same time attenuating the weak - correlation word pairs with a modulus - product lower than the threshold, enhancing the ability to capture semantic associations in the long - tail distribution.

[0063] Step 2: Set a local warning strategy on the client side, match the text features and the file meta - information of the data with the local warning strategy, and according to the matching result, generate an interception instruction in real - time, and perform real - time interception on the data transmission based on the interception instruction;

[0064] As a preferred embodiment of the present invention, matching the text features and the file meta - information of the data with the local warning strategy and generating an interception instruction in real - time according to the matching result specifically includes the following steps:

[0065] Construct a feature matching rule according to the file meta - information, and scan and classify the external information of the data according to the feature matching rule to generate a preliminary classified - as - secret determination result;

[0066] Calculate the similarity between the key classified - as - secret feature vectors saved locally and the text vector array to generate a semantic - level risk assessment result;

[0067] According to the preliminary classified - as - secret determination result and the semantic - level risk assessment result, trigger operations such as blocking sending, isolation, or warning, and generate a real - time interception instruction.

[0068] As a preferred embodiment of the present invention, calculating the similarity between the key classified - as - secret feature vectors saved locally and the text vector array to generate a semantic - level risk assessment result specifically includes the following steps:

[0069] Encode the file meta-information to generate a file meta-information feature vector;

[0070] Quantify the confidence of the file meta-information feature vector based on the L1 norm to obtain a meta-information confidence index;

[0071] Input the file meta-information feature vector into a fully connected layer, and then perform non-linear activation to obtain an enhanced file meta-information feature vector;

[0072] Construct a dual-channel mapping module using a dual-channel structure, and set independent fully connected layers in each channel. Input the text vector array into the dual-channel mapping module to capture lexical features and identify semantic patterns respectively, and obtain a dual-channel mapping result;

[0073] Perform feature interaction on the enhanced file meta-information feature vector and the dual-channel mapping result in an element-wise product manner to obtain an interaction reinforcement result of the text segment;

[0074] Select the segment feature with the largest L2 norm from the interaction reinforcement results of all text segments to obtain an optimal text interaction feature vector;

[0075] Use the sum of the L1 norms of all text segment vectors in the optimal text interaction feature vector as the complexity of the text content, and generate a dynamic weight according to the proportional relationship between the meta-information confidence index and the complexity of the text content;

[0076] Perform weighted synthesis on the enhanced file meta-information feature vector and the optimal text interaction feature vector using the dynamic weight, obtain a fused feature vector, perform normalization processing on the fused feature vector, map it to the risk score interval of 0-1, and obtain a dynamic fusion risk score;

[0077] Match the dimensions of the text vector array with the key classified feature vector. If the dimensions match, calculate the classified feature matching degree score; if not, give a default classified feature matching degree score as a supplementary calculation;

[0078] Perform risk decision synthesis on the classified feature matching degree score and the dynamic fusion risk score, and adjust the weight in the synthesis process according to the confidence level of the classified feature matching degree score to obtain a final synthesis score. Map the final synthesis score to a risk level and generate a corresponding interception strategy according to the corresponding risk level.

[0079] As a preferred embodiment of the present invention, the specific steps for calculating the classified feature matching degree score are as follows:

[0080] Regularize the time dimension of the text vector array to obtain an aligned text vector array;

[0081] Perform a sliding window comparison between each key confidential feature vector and the aligned text vector array, calculate the cosine similarity of the vectors within the window, and obtain the similarity score for each window.

[0082] Statistically analyze the positions where the similarity scores of each window exceed the set threshold to obtain the matching degree score of the confidential features.

[0083] In a preferred embodiment of the present invention, through the sliding window comparison mechanism of the local confidential feature library, historical sensitive data patterns can be accurately matched, solving the problem that traditional keyword matching cannot capture semantic features. And the dynamic feature fusion path can identify variant leakage means using synonym replacement and syntactic structure recombination through the non-linear interaction between meta-information and text. Moreover, the confidential feature comparison path is only activated when the text vector dimensions match, avoiding full-scale calculations for low-risk files to optimize the computational load and achieve the purpose of improving resource efficiency.

[0084] Step 3: According to the matching result, compress and encapsulate the data with unclear content and upload it to the server, and send a deep analysis request to the server.

[0085] As a preferred embodiment of the present invention, compressing and encapsulating the data with unclear content and uploading it to the server and sending a deep analysis request to the server specifically include the following steps:

[0086] Perform compression and encryption processing on the data with unclear content to generate a lightweight transmission data packet.

[0087] Encapsulate the file meta-information of the data with unclear content to generate a structured analysis request.

[0088] Upload the lightweight transmission data packet and the structured analysis request to the server.

[0089] Step 4: Perform multi-modal unified conversion and vectorization processing on the text, image, and voice content included in the unclear data according to the deep analysis request to generate a text feature vector array.

[0090] As a preferred embodiment of the present invention, performing multi-modal unified conversion and vectorization processing on the text, image, and voice content included in the unclear data according to the deep analysis request to generate a text feature vector array specifically includes the following steps:

[0091] For the text, image, and voice content included in the unclear data, identify the non-text areas in the image through a pre-trained convolutional neural network model to generate a noise area coordinate mask map.

[0092] Based on the noise area coordinate mask map, use an image inpainting algorithm to remove interference information and generate a denoised image.

[0093] Use OCR to extract text from the denoised image to obtain the image text, and then perform context verification and alignment on the image text to obtain a high-precision text stream;

[0094] Use TTS to convert the speech content into continuous speech text, frame the continuous speech text according to timestamps, and build context dependencies in combination with the Transformer model to obtain a sequence of semantic segments with timestamps;

[0095] Align the text, high-precision text stream, and sequence of semantic segments with timestamps in time series to obtain a time-synchronized multi-modal text sequence;

[0096] Map the time-synchronized multi-modal text sequence to the same semantic space through a shared embedding layer, and use a cross-modal attention mechanism for fusion to output a fused multi-modal feature vector, and use the fused multi-modal feature vector as an array of text feature vectors.

[0097] In a preferred embodiment of the present invention, the present invention establishes an audio-image noise correlation through spatial perception denoising and cross-modal noise filtering, automatically enhances the image denoising intensity when detecting speech background noise, and designs a three-level time anchor system to effectively improve the accuracy of time series alignment.

[0098] Step 5: Deeply analyze the array of text feature vectors, generate a risk assessment result and a warning classification mark according to the deep analysis result, and based on the risk assessment result and the warning classification mark, the client executes corresponding control instructions and dynamically updates the local warning strategy.

[0099] As a preferred embodiment of the present invention, deeply analyzing the array of text feature vectors and generating a risk assessment result and a warning classification mark according to the deep analysis result specifically includes the following steps:

[0100] Calculate the cosine similarity between the array of text feature vectors and the feature vectors in the sensitive document database, and screen out the Top-K candidate matching items;

[0101] Perform context reasoning on the Top-K candidate matching items through a large model to generate a context correlation degree;

[0102] Generate a comprehensive classified probability score based on the cosine similarity and the context correlation degree, and use the comprehensive classified probability score as the risk assessment result;

[0103] Load the classification warning rules configured by the administrator, match the threshold of the classification warning rules according to the classified probability score, and generate corresponding warning classification marks.

[0104] In a preferred embodiment of the present invention, a sensitive document database and a more accurate complete BERT model or a variant model specially tuned to meet specific tasks are set up. Vector similarity is calculated based on the database and vectorized data uploaded by the client, which can quickly compare features, evaluate file features and obtain preliminary evaluation results. This process can generate an analysis report for archiving, and the results are returned to the client and the early warning management module for execution. The database includes existing sensitive document vector features and file features. It will analyze new vector features and file features based on new sensitive files imported by the administrator through a high-precision model, or customize features according to the administrator's needs after submitting the analysis report, thereby achieving incremental updates to the database. This process can further develop an enclosed file analyzer for analyzing binary files that are difficult to analyze, etc. according to actual needs, and can also deploy large models to handle customer complaints of misjudgment. These are all optional.

[0105] The advantages of server-based deep analysis are high accuracy, fast speed, and no resource limitations. In addition, "association analysis" can use more advanced technologies such as large models, semantic rearrangement models, and other tools to perform in-depth analysis and calculations on semantic similarity calculations and matching of confidential categories.

[0106] As a preferred embodiment of the present invention, dynamically updating the local early warning strategy specifically includes the following steps:

[0107] Based on the risk assessment results, confirm whether the data with unclear content is classified or sensitive data;

[0108] If the data is confidential or sensitive, it triggers the policy update requirement and obtains the corresponding file metadata features;

[0109] Extract key features from confidential or sensitive data with unclear content to obtain key confidential text feature vectors;

[0110] The file metadata features, key confidential text feature vectors and warning classification marks are integrated into an update package and sent encrypted to the client to update the local warning strategy.

[0111] In this step, the Alert Management Module analyzes the results and determines the alert level. Administrators also maintain and configure alert rules and distribute local alert policies to clients in real time. Alert records are archived and managed to provide data support for system optimization. Administrators can input new matching alert rules and provide feedback to the overall system. They can also import new sensitive files and feed them into the Deep Analysis Module, which then distributes them to the client's interception module in a graded manner.

[0112] Please refer to Figure 2, this embodiment also provides an instant classified information detection system for multi-modal feature fusion of an IM software. Among them, the system applies the instant classified information detection method for multi-modal feature fusion of the IM software as described above. The system includes a client and a server. The client includes a local vectorization module, a file interception module, and a communication management module. The server includes a warning management module, a deep analysis module, and a multi-modal processing module;

[0113] The local vectorization module is used to: monitor the local data transmission situation. When a data transmission behavior occurs, extract text features from the data to obtain text features;

[0114] The communication management module is used to: set a local warning policy on the client, match the text features and the file meta-information of the data with the local warning policy, and generate an interception instruction in real time according to the matching result;

[0115] The file interception module is used to: perform real-time interception on the data transmission based on the interception instruction;

[0116] According to the matching result, compress and encapsulate the data with unclear content and upload it to the server, and send a deep analysis request to the server;

[0117] The multi-modal processing module is used to: according to the deep analysis request, perform multi-modal unified conversion and vectorization processing on the text, image, and voice content included in the unclear data, and generate an array of text feature vectors;

[0118] The deep analysis module is used to: perform in-depth analysis on the array of text feature vectors;

[0119] The warning management module is used to: generate a risk assessment result and a warning classification mark according to the deep analysis result, and dynamically update the local warning policy based on the risk assessment result and the warning classification mark.

[0120] It should be understood that although the steps in the flowcharts of the embodiments of the present invention are displayed in sequence according to the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in each embodiment may include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily need to be executed at the same moment, but can be executed at different moments. The execution order of these sub-steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0121] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), and the like.

[0122] In the description of this specification, the description referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0123] The above-described embodiments merely represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.

Claims

1. An instant classified information detection method for multimodal feature fusion of IM software, characterized in that The method includes the following steps: Step 1, monitor the local data transmission situation. When a data transmission behavior occurs, extract the text features of the data to obtain text features; Step 2, set a local warning policy on the client side, match the text features and the file meta-information of the data with the local warning policy, generate an interception instruction in real time according to the matching result, and perform real-time interception on the data transmission based on the interception instruction; Step 3, according to the matching result, compress and encapsulate the data with unclear content and upload it to the server, and send a deep analysis request to the server; Step 4, perform multi-modal unified conversion and vectorization processing on the text, image and voice content included in the unclear data according to the deep analysis request, and generate a text feature vector array; Step 5, perform deep analysis on the text feature vector array, generate a risk assessment result and a warning classification mark according to the deep analysis result, and based on the risk assessment result and the warning classification mark, the client executes corresponding control instructions and dynamically updates the local warning policy; In the said Step 4, performing multi-modal unified conversion and vectorization processing on the text, image and voice content included in the unclear data according to the deep analysis request to generate a text feature vector array specifically includes the following steps: For the text, image and voice content included in the unclear data, identify the non-text areas in the image through a pre-trained convolutional neural network model to generate a noise area coordinate mask map; Based on the noise area coordinate mask map, use an image repair algorithm to remove interference information and generate a denoised image; Adopt OCR to extract text from the denoised image, obtain the image text, and then perform context verification and alignment on the image text to obtain a high-precision text stream; Adopt TTS to convert the voice content into continuous speech text, frame the continuous speech text according to timestamps, and construct a context dependence relationship in combination with the Transformer model to obtain a sequence of semantic segments with timestamps; Perform intent recognition and sentiment analysis on the sequence of semantic segments with timestamps through an intent classification model and a sentiment polarity model respectively to obtain semantic segments with scene labels and sentiment scores; According to the word frequencies and scene labels in the semantic segments with scene labels and sentiment scores, perform TF-IDF weighted calculation on the semantic segments with scene labels and sentiment scores to obtain structured semantic segments; Align the timestamps of the text, the timestamps of the high-precision text stream and the timestamps of the structured semantic segments to obtain a time-synchronized multi-modal text sequence; Map the time-synchronized multi-modal text sequence to the same semantic space through a shared embedding layer, and adopt a cross-modal attention mechanism for fusion, output the fused multi-modal feature vector, and use the fused multi-modal feature vector as the text feature vector array; In the said Step 5, performing deep analysis on the text feature vector array and generating a risk assessment result and a warning classification mark according to the deep analysis result specifically includes the following steps: Calculate the cosine similarity between the text feature vector array and the feature vectors in the sensitive document database to screen out the Top-K candidate matching items; Perform context reasoning on the Top-K candidate matches through a large model to generate context relevance; Generate a comprehensive classified probability score based on cosine similarity and context relevance, and use the comprehensive classified probability score as the risk assessment result; Load the hierarchical warning rules configured by the administrator, match the threshold of the hierarchical warning rules according to the classified probability score, and generate the corresponding warning classification mark.

2. The instant classified information detection method for IM software multi-modal feature fusion according to claim 1, characterized in that, In the step 1, text feature extraction is performed on the data, and the specific steps for obtaining text features are as follows: Clean and split sentences of the received text content to generate a normalized text sequence; Perform sentence splitting and sequence segmentation on the normalized text sequence to obtain a set of standardized text sequences after sentence splitting; Use the embedding layer of the lightweight semantic conversion vector model for the set of standardized text sequences after sentence splitting to convert text characters into word vectors; Perform lightweight multi-head attention calculation on the word vectors to capture semantic associations and obtain an attention output; Use the distilled shallow network to extract key semantic features from the attention output; Perform mean pooling or max pooling on the key semantic features to generate sentence-level feature vectors and obtain text features.

3. The instant classified information detection method for multi-modal feature fusion of IM software according to claim 2, characterized in that, Performing lightweight multi-head attention calculation on the word vectors to capture semantic associations and obtain an attention output specifically includes the following steps: Perform linear projection on the word vectors to generate a query vector Q, a key vector K, and a value vector V; Calculate the dot product similarity between the query vector Q and the key vector K, and then divide each dot product similarity by the square root of the dimension of the key vector V for normalization processing to obtain a preliminary attention weight matrix; Calculate the L2 norm of each query vector Q and key vector K respectively to obtain the L2 norm of the query vector Q and the L2 norm of the key vector K; Multiply the L2 norm of the query vector Q and the L2 norm of the key vector K to obtain the vector norm; Perform a non-linear transformation on the vector norm using the hyperbolic tangent function to obtain a fractal modulation factor, where in the hyperbolic tangent function, two learnable parameters are respectively used to control the steepness of the curve and the activation threshold for dynamic adjustment; Fuse the preliminary attention weight matrix and the fractal modulation factor in an element-wise multiplication manner to obtain an attention weight matrix fused with the dynamic modulation factor; Use the attention weight matrix fused with the dynamic modulation factor to perform weighted summation on the value vector V to obtain an attention output.

4. The instant classified information detection method for multimodal feature fusion of IM software according to claim 3, characterized in that, In the step 2, matching the text features and the file meta-information of the data with the local warning policy, and generating an interception instruction in real time according to the matching result specifically includes the following steps: Construct a feature matching rule according to the file meta-information, and scan and classify the external information of the data according to the feature matching rule to generate a preliminary classified determination result; Calculate the similarity between the key classified feature vectors saved locally and the text vector array to generate a semantic-level risk assessment result; Trigger operations such as blocking sending, isolation, or warning according to the preliminary classified determination result and the semantic-level risk assessment result, and generate a real-time interception instruction.

5. The instant classified information detection method for IM software multi-modal feature fusion according to claim 4, characterized in that Calculating the similarity between the key classified feature vectors saved locally and the text vector array to generate a semantic-level risk assessment result specifically includes the following steps: Encode the file meta-information to generate a file meta-information feature vector; Quantify the confidence of the file meta - information feature vector based on the L1 norm to obtain the meta - information confidence index; Input the file meta - information feature vector into the fully - connected layer, and then perform non - linear activation to obtain the enhanced file meta - information feature vector; Construct a dual - channel mapping module using a dual - channel structure, and set independent fully - connected layers in each channel. Input the text vector array into the dual - channel mapping module to capture lexical features and identify semantic patterns respectively, and obtain the dual - channel mapping result; Perform feature interaction on the enhanced file meta - information feature vector and the dual - channel mapping result in an element - wise multiplication manner to obtain the interaction - enhanced result of the text segment; Select the segment feature with the largest L2 norm from the interaction - enhanced results of all text segments to obtain the optimal text interaction feature vector; Use the sum of the L1 norms of all text segment vectors in the optimal text interaction feature vector as the complexity of the text content, and generate a dynamic weight according to the proportional relationship between the meta - information confidence index and the complexity of the text content; Perform weighted synthesis on the enhanced file meta - information feature vector and the optimal text interaction feature vector using the dynamic weight to obtain the fused feature vector. Normalize the fused feature vector and map it to the risk score interval of 0 - 1 to obtain the dynamic fusion risk score; Match the dimensions of the text vector array and the key classified - secret feature vector. If the dimensions match, calculate the classified - secret feature matching degree score; if they do not match, give a default classified - secret feature matching degree score as a supplementary calculation; Perform risk - decision synthesis on the classified - secret feature matching degree score and the dynamic fusion risk score, and adjust the weight in the synthesis process according to the confidence level of the classified - secret feature matching degree score to obtain the final synthesis score. Map the final synthesis score to a risk level and generate the corresponding interception strategy according to the corresponding risk level.

6. The instant classified information detection method for multimodal feature fusion of IM software according to claim 5, wherein In step 3, compressing, encapsulating and uploading the content - ambiguous data to the server and sending a deep - analysis request to the server specifically includes the following steps: Perform compression and encryption processing on the content - ambiguous data to generate a lightweight transmission data packet; Encapsulate the file meta - information of the content - ambiguous data to generate a structured analysis request; Upload the lightweight transmission data packet and the structured analysis request to the server.

7. The instant classified information detection method for IM software multi-modal feature fusion according to claim 1, characterized in that In step 5, dynamically updating the local warning strategy specifically includes the following steps: According to the risk assessment result, confirm whether the content - ambiguous data belongs to classified - secret data or sensitive data; If it belongs to classified - secret data or sensitive data, trigger the policy update requirement and obtain the corresponding file meta - information features; Extract the key features of the content - ambiguous data belonging to classified - secret data or sensitive data to obtain the key classified - secret text feature vector; Integrate the file meta - information features, the key classified - secret text feature vector and the warning classification mark into an update package, and encrypt and send it to the client to update the local warning strategy.

8. An instant classified information detection system for multimodal feature fusion of IM software, characterized in that, The system applies the instant classified information detection method for multimodal feature fusion of IM software as described in any one of claims 1 to 7. The system includes a client and a server. The client includes a local vectorization module, a file interception module, and a communication management module. The server includes an early warning management module, a deep analysis module, and a multimodal processing module; The local vectorization module is used to: monitor the local data transmission situation, and when a data transmission behavior occurs, extract text features from the data to obtain text features; The communication management module is used to: set a local early warning policy on the client, match the text features and the file meta-information of the data with the local early warning policy, and generate an interception instruction in real time according to the matching result; The file interception module is used to: perform real-time interception of data transmission based on the interception instruction; According to the matching result, compress and encapsulate the data with unclear content and upload it to the server, and send a deep analysis request to the server; The multimodal processing module is used to: perform multimodal unified conversion and vectorization processing on the text, image, and voice content included in the unclear data according to the deep analysis request, and generate a text feature vector array; The deep analysis module is used to: perform deep analysis on the text feature vector array; The early warning management module is used to: generate a risk assessment result and an early warning classification mark according to the deep analysis result, and dynamically update the local early warning policy based on the risk assessment result and the early warning classification mark.

Citation Information

Patent Citations

  • Real-time monitoring method and system based on file content received by mobile terminal

    CN112422739A

  • Text data acquisition method and device, equipment and storage medium

    CN115294575A