Sensitive information isolation method and system and computer equipment

By collecting and analyzing multi-platform, multi-modal data in real time, dynamically updating the sensitive database and generating isolation strategies, the problems of slow sensitive information isolation updates and weak data source awareness in existing technologies are solved, achieving efficient and accurate sensitive information isolation.

CN121145263APending Publication Date: 2025-12-16QUANZHOU INST OF INFORMATION ENG
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511679573.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2025-12-16

AI Technical Summary

Technical Problem

Existing technologies for sensitive information identification and isolation are slow to update and have weak data source awareness, making it difficult to meet the modern data security field's demand for efficient and accurate isolation.

Method used

It collects and analyzes multimodal data from multiple platforms in real time, dynamically refreshes the sensitive database and generates isolation strategies, and updates the sensitive database through matching analysis and evaluation results to achieve precise isolation across time and platforms.

Benefits of technology

It enables real-time and precise isolation of sensitive information, improves the adaptability and accuracy of isolation strategies, and dynamically updates the sensitive database to adapt to changes in different data sources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121145263A_ABST
    Figure CN121145263A_ABST
Patent Text Reader

Abstract

The invention provides a sensitive information isolation method and system and computer equipment. The method comprises the steps that multi-source data from multiple data propagation platforms are collected and analyzed in real time, and multiple pieces of feature data are obtained; performing matching analysis on the plurality of feature data to obtain a matching result used for reflecting the matching degree of each feature data and the initial sensitive category in the dynamic sensitive database; according to the matching result and the data source identifier, generating an isolation strategy corresponding to each piece of source data within a preset response time limit; according to the isolation strategy, performing data isolation on the corresponding source data to obtain a plurality of isolated data; and detecting residual sensitive features and error isolation normal features of the isolation data to form an evaluation result, and reinjecting the detected residual sensitive features and error isolation normal features as incremental samples into the initial data set according to the evaluation result to update the dynamic sensitive database. According to the invention, dynamic identification, hierarchical isolation and closed-loop updating of various data sensitive information can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data security technology, and in particular to a method and system for isolating sensitive information, as well as computer equipment. Background Technology

[0002] With the widespread application of deep learning technology in identifying sensitive information in single data types such as text, images, videos, and audio, effectively isolating sensitive information has become an urgent problem to be solved. Existing technologies mainly focus on the identification and classification of sensitive information, but existing isolation methods suffer from slow updates to sensitive information databases and weak awareness of data source origins, making it difficult to meet the demands of modern data security for efficient and accurate isolation. Summary of the Invention

[0003] This application provides a sensitive information isolation method and system, as well as a computer device, which aims to integrate multimodal data such as text, images, and videos from multiple dissemination platforms in real time, dynamically refresh the sensitive information database and adjust the isolation strategy to achieve time-limited, cross-platform, and precise isolation.

[0004] In a first aspect, embodiments of this application provide a method for isolating sensitive information. The method includes: real-time acquisition and parsing of multi-source data from various data dissemination platforms to obtain multiple feature data, whereby the multi-source data includes text, images, and / or video frames, and each source data has a data source identifier, including the dissemination platform type, the number of audience members, and a dissemination timestamp; performing matching analysis on the multiple feature data to obtain a matching result reflecting the degree of matching between each feature data and an initial sensitive category in a dynamic sensitive database, whereby the dynamic sensitive database is constructed from an initial dataset, which includes initial feature data corresponding to the initial sensitive category; generating an isolation strategy corresponding to each source data within a preset response time limit based on the matching result and the data source identifier, whereby the isolation strategy is used to reduce, remove, or hide the identifiability of portions of the feature data that have been matched as sensitive categories; performing data isolation on the corresponding source data according to the isolation strategy to obtain multiple isolated data; detecting residual sensitive features and falsely isolated normal features in the isolated data to form an evaluation result, and then using the detected residual sensitive features and falsely isolated normal features as incremental samples to inject back into the initial dataset to update the dynamic sensitive database based on the evaluation result.

[0005] Secondly, embodiments of this application provide a sensitive information isolation system. The sensitive information isolation system includes a data acquisition and parsing module, a matching and analysis module, a strategy generation module, a data isolation module, and an evaluation and update module. The data acquisition and parsing module is used to acquire and parse multi-source data from various data dissemination platforms in real time to obtain multiple feature data. The multi-source data includes text, images, and / or video frames. Each source data has a data source identifier, which includes the dissemination platform type, the number of audience members, and the dissemination timestamp. The matching and analysis module is used to perform matching analysis on the multiple feature data to obtain a matching result reflecting the degree of matching between each feature data and an initial sensitive category in a dynamic sensitive database. The dynamic sensitive database is constructed from an initial dataset. The initial dataset includes initial feature data corresponding to the initial sensitive category; the strategy generation module is used to generate an isolation strategy corresponding to each source data within a preset response time limit based on the matching result and the data source identifier, the isolation strategy being used to reduce, remove or hide the identifiability of the part of the feature data that has been matched as a sensitive category; the data isolation module is used to perform data isolation on the corresponding source data according to the isolation strategy to obtain multiple isolated data; the evaluation and update module is used to detect residual sensitive features and falsely isolated normal features of the isolated data to form an evaluation result, so as to use the detected residual sensitive features and falsely isolated normal features as incremental samples to be injected back into the initial dataset to update the dynamic sensitive database according to the evaluation result.

[0006] Thirdly, embodiments of this application provide a computer device, the computer device including a memory and a processor, the memory being used to store a computer program; the processor being used to execute the computer program to implement the above-mentioned sensitive information isolation method.

[0007] The aforementioned sensitive information isolation method and system, as well as computer equipment, collect and parse text, image / video frames from multiple platforms in real time, normalize them to obtain feature data, generate isolation strategies within a preset response time limit according to the matching results and data source identifiers (platform, audience, timestamp), and then inject residual and mis-isolated features back into the sensitive database to achieve dynamic expansion of sensitive categories and time-limited accurate isolation. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0009] Figure 1This is a first flowchart of a sensitive information isolation method provided in an embodiment of this application.

[0010] Figure 2 A flowchart of step S101 provided for an embodiment of this application.

[0011] Figure 3 A flowchart of step S102 provided in the embodiments of this application.

[0012] Figure 4 A flowchart of step S103 provided in the embodiments of this application.

[0013] Figure 5 A flowchart of step S105 provided in the embodiments of this application.

[0014] Figure 6 This is a second flowchart of a sensitive information isolation method provided in an embodiment of this application.

[0015] Figure 7 The third flowchart is a method for isolating sensitive information provided in the embodiments of this application.

[0016] Figure 8 This is a structural block diagram of a sensitive information isolation system provided in an embodiment of this application.

[0017] Figure 9 This is a schematic diagram of the internal structure of a computer device that uses a sensitive information isolation method as described in an embodiment of this application.

[0018] Figure 10 This is a first schematic diagram illustrating the construction of a dynamic sensitive database, as provided in an embodiment of this application.

[0019] Figure 11 This is a second schematic diagram illustrating the construction of a dynamic sensitive database, as provided in an embodiment of this application.

[0020] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0022] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar planned objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data are interchangeable where appropriate; in other words, the described embodiments are implemented according to a sequence other than that illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, may also include other content; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0023] It should be noted that the use of terms such as "first" and "second" in this application is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" and "second" may explicitly or implicitly include one or more of that feature. Furthermore, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.

[0024] Please refer to Figure 1 This is a first flowchart of the sensitive information isolation method provided in this application embodiment. This application provides a sensitive information isolation method for real-time sensitive information isolation of multi-source data from various data dissemination platforms such as social media platforms, enterprise internal systems, mobile terminals, and IoT devices, including but not limited to text messages, image files, video frames, and audio streams. The sensitive information isolation method includes steps S101-S105.

[0025] Step S101: Collect and analyze multi-source data from multiple data transmission platforms in real time to obtain multiple feature data.

[0026] In step S101, the source data refers to data of different data types disseminated within various data dissemination platforms, such as text, images, video frames, audio, point clouds, etc. Each source data has a data source identifier. The data source identifier is used to reflect relevant parameters during the dissemination of the corresponding source data, including the dissemination platform type, the number of audience members, and the dissemination timestamp. For example, if a data source identifier records the dissemination platform type as "Weibo," the number of audience members as 3000, and the dissemination timestamp as "2025-10-20-12:00:00," it indicates that the corresponding source data was disseminated on Weibo on 2025-10-20-12:00:00, and the current number of audience members is 3000. Feature data includes text feature data and / or image feature data. Among them, text feature data corresponds to text data, and image feature data corresponds to image data and / or video frame data. In this application, each source data can be collected simultaneously in a single form such as text, images, and video frames, or it can be collected in a combined whole. For example, when multi-source data is collected in the form of product manuals, the source data can be data consisting of text containing product parameters and images representing product illustrations. In other words, the data type of feature data can be broadened accordingly as the data type of multi-source data expands.

[0027] Please refer to Figure 2 This is a flowchart of sub-step S101 provided in the embodiments of this application. Different feature data can be parsed in corresponding data forms. Specifically, text data can be parsed in the form of word vectors, and image feature data can be parsed in the form of pixel intensity values, color histograms, etc. More specifically, real-time acquisition and parsing of multi-source data from multiple data dissemination platforms to obtain multiple feature data includes steps S1011-S1014.

[0028] Step S1011: Identify each source data.

[0029] In step S1011, the data types corresponding to each source data are identified, and corresponding feature data are obtained in different ways accordingly.

[0030] Step S1012: When the source data contains text, the corresponding source data is parsed using the first parsing method to obtain the corresponding text feature data.

[0031] In step S1012, the first parsing method includes, but is not limited to, natural language processing of the corresponding text in the form of word vectors, term frequency-inverse document frequency (TF-IDF), latent Dirichlet Allocation (LDA) model, etc.

[0032] Step S1013: When the source data contains images and / or video frames, the corresponding source data is parsed using the second parsing method to obtain the corresponding image feature data.

[0033] In step S1013, the second parsing method includes, but is not limited to, processing with convolutional neural networks, Hough transform, scale-invariant feature transform (SIFT), edge detection, and texture analysis in the form of pixel intensity values ​​and color histograms.

[0034] Step S1014: Normalize the text feature data and / or image feature data corresponding to each source data to form multiple feature data.

[0035] In step S1014, due to scale differences between different data types in the source data caused by factors such as acquisition methods and encoding formats, the parsed text feature data and / or image feature data need to be normalized to adjust the different feature data to a target scale range that facilitates unified processing. Specifically, when different feature data are represented in numerical form, the values ​​corresponding to all feature data are linearly adjusted proportionally to the same preset numerical range. Taking a patent document website as the data dissemination platform and patent documents as the source data as an example, assuming the patent documents contain text descriptions and image illustrations, both text feature data and image feature data can be parsed. Since text feature data is represented in the form of word vectors and image feature data is represented in the form of pixel intensity values, the two differ significantly in value. Therefore, through normalization, the values ​​corresponding to these feature data at different scales can be linearly scaled and adjusted to the same preset numerical range to eliminate the scale difference between text feature data and image feature data.

[0036] Understandably, when the data types of the source data in multi-source data are not limited to text data, image data, and video frame data, the corresponding feature data can be parsed using appropriate parsing methods based on the data types of the source data to be isolated. For example, when the source data includes audio, Mel-Frequency Cepstral Coefficients (MFCCs) and Mel spectrograms can be used to parse the corresponding feature data from the audio data, and the parsed feature data can be normalized. In other words, as the data types of source data continue to expand, this application can parse the expanded multi-source data using the parsing methods corresponding to the multi-source data when the data types are expanded, obtain the corresponding feature data, and normalize the parsed feature data, so that the multi-source data after the data types are expanded can still perform subsequent matching analysis and isolation strategy generation in a standardized feature space.

[0037] In the above embodiments, by normalizing the parsed feature data, the compatibility of multi-source data is achieved, the robustness of subsequent matching analysis and isolation strategy generation is improved, and misjudgments and errors caused by differences in data scale are reduced.

[0038] Step S102: Perform matching analysis on multiple feature data to obtain matching results that reflect the degree of matching between each feature data and the initial sensitive category in the dynamic sensitive database.

[0039] In step S102, the dynamic sensitive database is constructed from the initial dataset. The initial dataset includes initial sensitive categories and initial feature data corresponding to the initial sensitive categories. The matching results can be used to determine whether data isolation of the feature data is necessary by reflecting the degree of matching of each feature data. In this application, the initial feature data can be a variety of general sensitive information that has been pre-sorted, solidified, and verified within various data dissemination platforms, including but not limited to personal identity information (such as ID card number, bank card number, mobile phone number, etc.), trade secrets (such as unpublished financial reports of enterprises, contract amounts, source code fragments, etc.), internal sensitive text (such as titles of classified documents, internal project codes, executive itineraries, etc.), image sensitive elements (such as portrait photos, scanned copies of documents, close-ups of license plates, etc.), unauthorized video frames (such as live broadcast footage of meetings, original surveillance videos, etc.). The initial sensitive categories include but are not limited to the categories covered by the initial dataset, as well as categories customized according to the requirements for sensitive information isolation. Understandably, for various categories, several verified initial feature data can be associated as the initial sensitive categories of the dynamic sensitive database, which can then serve as the core reference for real-time comparison in subsequent processes. This allows for quick determination of whether the collected multi-source data hits the known initial sensitive categories, thereby supporting real-time sensitive information isolation.

[0040] like Figure 10 As shown, Figure 10 This illustrates the process of constructing the dynamic sensitive database BASE_0 from the initial datasets corresponding to data source identifiers A, B, C, etc. To reflect the different sensitivity categories within each data dissemination platform, Figure 10This application only illustrates the categories of generally sensitive information, without indicating the number of audiences or dissemination timestamps. It first identifies historical sensitive data from various data dissemination platforms and then categorizes this data according to the respective platforms. For example, historical sensitive data of varying sensitivity levels, such as ID card numbers, mobile phone numbers, license plates, and original videos, can be obtained from the data dissemination platform corresponding to data source identifier A0, and the sensitivity categories of this data can be identified. Then, different categories are obtained from data source identifier A0 as initial sensitivity categories, and different examples within the same sensitivity category are further obtained from this historical sensitive data, such as the differentiated composition of ID card numbers from different regions. Different examples under other sensitivity categories can also be further obtained from this historical sensitive data to obtain initial feature data corresponding to different initial sensitivity categories and form an initial dataset.

[0041] In this embodiment, after obtaining a certain number of initial datasets under the data source identifiers, the initial datasets are written into the dynamic sensitive database BASE_0 to indicate the initial sensitive data and their corresponding sensitive categories. This allows the sensitive information database BASE_0 to directly compare the data with the existing initial feature data in subsequent processes, or to obtain the potential feature data corresponding to different sensitive categories from the initial dataset and the rules for forming the feature data through deep learning for comparison, thereby supporting real-time sensitive information isolation.

[0042] Step S103: Based on the matching results and the data source identifier, generate an isolation strategy corresponding to each source data within a preset response time limit.

[0043] In step S103, the isolation strategy is used to reduce, remove, or hide the identifiability of portions of the feature data that have been matched as sensitive categories. The isolation strategy for each source data can be implemented within a preset response time limit. This preset response time limit can be pre-set based on the data source identifier and the actual needs of real-time sensitive information isolation. For example, when the sensitive data dissemination node is a statutory holiday (such as Spring Festival, National Day, etc.), the preset response time limit can be adjusted accordingly to mitigate the impact of timeliness differences on the dissemination of sensitive information.

[0044] Step S104: According to the isolation strategy, the corresponding source data is isolated to obtain multiple isolated data.

[0045] Step S105: Detect residual sensitive features and falsely isolated normal features of the isolated data to form an evaluation result. Based on the evaluation result, the detected residual sensitive features and falsely isolated normal features are used as incremental samples to be injected back into the initial dataset to update the dynamic sensitive database.

[0046] The process of obtaining matching results, generating isolation strategies for each source data, isolating source data based on each source data to obtain isolated data, and evaluating the isolated data to obtain evaluation results, in order to update the dynamic sensitive database BASE_0, will be explained in detail below.

[0047] Please refer to Figure 6 This is a second flowchart of the sensitive information isolation method provided in this application embodiment. To ensure that subsequent isolation strategies, data isolation, and detection and evaluation all have adaptive and real-time capabilities, this embodiment obtains target isolation parameters in advance to determine the generation rules of the isolation strategy, the execution strength of data isolation, and the setting of the evaluation threshold, so that the sensitive information isolation process continuously maintains optimal performance according to the source data and its corresponding data source identifier. Specifically, the sensitive information isolation method in this application also includes steps S201-S202.

[0048] Step S201: Feed the preset training set into the preset training network to train and obtain the isolated parameter model.

[0049] In step S201, the preset training network can be an initial deep learning network. The preset training set includes multiple sample isolation data after pre-separation of data. Each sample isolation data includes the true sensitive category, the source data before isolation, the source data after isolation, and the sample feature data corresponding to the source data before isolation. The isolation parameter model is used to output the isolation parameters. The true sensitive category includes the sample sensitive category and the potential sensitive category. The sample sensitive category can be obtained through the sample feature data, and the potential sensitive category is the sensitive category (if any) corresponding to the residual sensitive features and the erroneously isolated normal features detected by evaluating the source data after isolation. That is to say, by evaluating the source data after isolation, this application can obtain different true sensitive categories based on the evaluation effect, and thus obtain different isolation parameter models under different evaluation effects, thereby improving the accuracy of the subsequent target isolation parameters.

[0050] Step S202: Fit the isolation parameters output by the isolation parameter model to the true sensitive category to obtain the target isolation parameters.

[0051] In step S202, the target isolation parameter is used to adjust the matching analysis, isolation strategy generation, evaluation detection, and / or the refresh of the dynamic sensitive database, including the dynamic update frequency, matching level threshold, isolation sub-strategy threshold, and evaluation threshold. In one embodiment, the isolation parameter model obtained in step S201 pre-outputs an initial isolation parameter according to its purpose. The initial isolation parameter and the target isolation parameter correspond one-to-one. Using the true sensitive category as the supervision label, a fitting function is used to optimize the hyperparameters of the initial isolation parameter. The fitting function can be a function that minimizes cross-entropy, mean squared error loss function, etc. The fitted isolation parameter is used to iteratively train on a preset training set until the prediction error of the isolation parameter model converges to a preset tolerance. The isolation parameter of the most recent iteration is used as the target isolation parameter for real-time reference in subsequent isolation strategies, data isolation, and evaluation stages.

[0052] In the above embodiments, isolated sample data that has already undergone data isolation is used as training samples to continuously train and fit an isolation parameter model, thereby dynamically generating target isolation parameters. This allows the system to automatically correct the sensitivity level classification, update cycle, and isolation strategy priority based on the data source and its corresponding data source identifier, significantly improving isolation accuracy and achieving adaptive optimization of the sensitive information database. The following sections will describe the application scenarios of the target isolation parameters to illustrate how to obtain matching results, generate isolation strategies, perform data isolation, and obtain evaluation results through detection and assessment, in order to update the dynamic sensitive database BASE_0.

[0053] Please refer to Figure 7 This is a third flowchart of the sensitive information isolation method provided in the embodiments of this application. The sensitive information isolation method further includes steps S301-S302.

[0054] Step S301: Set the dynamic update frequency to the current update frequency of the dynamic sensitive database.

[0055] Step S302: Dynamically expand the initial dataset according to the current update frequency, so as to update the dynamic sensitive database based on the expanded initial dataset.

[0056] In step S302, an update task for the dynamic sensitive database is triggered based on the current update frequency. Specifically, starting from time T0, the pipeline of "acquiring multi-source data - parsing to obtain feature data - deduplicating to obtain the initial dataset" is automatically started every Δt to expand upon the previously updated initial dataset. Here, Δt is determined by the current update frequency, and time T0 is the time when the update task is triggered. After triggering the update task for the dynamic sensitive database, the updated initial dataset is written to the dynamic sensitive database BASE_0 for refresh. Understandably, the updated initial dataset can be either newly added initial sensitive categories or newly added initial feature data, or both. The newly added initial sensitive categories and / or newly added initial feature data can be obtained in real-time from the pre-selected data source identifiers according to the current update frequency to update the dynamic sensitive database in real time.

[0057] Please refer to Figure 3 This is a flowchart of sub-step S102 provided in the embodiments of this application. The matching result in this application includes a matching score and a matching level. The matching score measures the degree of matching between the sensitive category corresponding to the feature data and the dynamic sensitive database BASE_0 using a numerical value. The matching level is one of multiple preset matching levels. A matching level threshold is used to determine the matching score range corresponding to each preset matching level. Steps S1021-S1023 include performing matching analysis on multiple feature data to obtain a matching result reflecting the degree of matching between each feature data and the initial sensitive category in the dynamic sensitive database.

[0058] Step S1021: Feed the initial dataset into the preset training model to learn the matching degree model.

[0059] In step S1021, the preset training model can be an initial deep learning model. The matching degree model is used to describe the correlation between feature data and matching scores. The matching level threshold can be set to the matching score range corresponding to the preset matching level according to the actual needs of matching accuracy. Specifically, the matching level threshold can be expressed as... .in, It can be used to determine the upper and lower limits of the matching score range corresponding to each preset matching level. The scores can be equal or unequal, allowing the scale of the matching score range corresponding to different preset matching levels to be adjustable. In this application, the matching degree model can learn the mapping relationship between feature data and matching scores, enabling rapid scoring of the input feature data subsequently.

[0060] For example, for an 18-digit string, the matching model learns the mapping relationship through the following analysis process: First, the 18-digit string is parsed into the following six categories of feature data: Length: scored according to compliance; First 6 digits: Regional code: compared with the Ministry of Civil Affairs' regional code database, scored according to whether it matches; 7th-14th digits: Birth date: verifies the legality of the date, scored according to legality; 15th-17th digits: Sequence code: uses parity to map gender, scored according to whether the verification logic is valid; 18th digit: Check code: verified using the GB11643 check algorithm, scored according to whether the verification is valid. Scoring is done separately; overall regularization matching of the numeric string: scoring is done separately based on whether it conforms to the 18-digit pure numeric format; then, the above six types of feature data are concatenated into an input vector and input into the matching degree model, which learns weights to output continuous matching scores within a preset numerical range; next, using the "real ID number" label as a positive sample and the "non-ID number" label as a negative sample, the cross-entropy loss is minimized, so that the matching degree model learns the mapping relationship that "when all six types of feature data are matched, the matching score approximately reaches the maximum value of the preset numerical range, and when any feature data fails, the matching score drops significantly." When performing matching analysis on the sensitive category of "ID number," its corresponding preset matching level can be assigned to high matching level, medium matching level, and low matching level respectively. At this time, the matching level threshold is used... Assign upper and lower limits to the matching score range corresponding to different preset matching levels, based on the mapping relationship learned earlier. That is, when the matching score > When the match rating falls into the high match level; when >Matching score> When the match score falls into the medium match level; when the match score is < At that time, the match score falls into a low match level.

[0061] It is important to note that the analysis process illustrated above is merely one example of the analysis process corresponding to the mapping relationship of "18-digit string - ID number," and is not a limitation on this mapping relationship. Furthermore, the "18-digit string - ID number" mapping relationship described above is only one example of a mapping relationship learned by the matching degree model, and is not a limitation on this mapping relationship. The matching degree model can learn the corresponding mapping relationships by learning the analysis processes corresponding to different sensitivity categories, thereby achieving sufficient matching analysis of different types of feature data.

[0062] Step S1022: Based on the matching degree model, perform matching analysis on multiple feature data to obtain the matching score for each feature data.

[0063] In step S1022, the matching degree model learns at least the mapping relationship between each initial feature data and its corresponding matching score in the dynamic sensitive database. For different feature data, the matching degree model can obtain the corresponding matching score according to the corresponding mapping relationship.

[0064] Step S1023: Use the matching level threshold to determine the matching score range corresponding to each preset matching level, and classify the matching score into the corresponding matching level to obtain the matching result of each source data.

[0065] Please refer to Figure 4 This is a flowchart of step S103 provided in the embodiments of this application. Based on the matching result and data source identifier, generating an isolation strategy corresponding to each source data within a preset response time limit includes steps S1031-S1033.

[0066] Step S1031: Obtain the sensitivity of each source data according to the data source identifier.

[0067] In step S1031, the sensitivity level can be obtained as follows: First, assign corresponding weights to each factor included in the data source identifier. For example, the weights include the dissemination platform weight, the audience size weight, and the dissemination time weight. The dissemination platform weight, audience size weight, and dissemination time weight correspond to the dissemination platform type, the audience size, and the dissemination timestamp, respectively. Then, the sensitivity level is calculated according to the expression below: Sensitivity = Dissemination platform weight * Dissemination platform factor + Audience size weight * Dissemination range factor + Dissemination time weight * Dissemination time factor; The propagation platform factor, propagation range factor, and propagation time factor are determined through isolation sub-strategy thresholds to differentiate the various factors included under different data source identifiers with varying degrees of importance. Once the sensitivity of a particular source data is calculated, the sensitivity of the remaining source data can also be calculated using the same method.

[0068] Step S1032: Using the isolation sub-policy threshold, assign the corresponding isolation sub-policy to each preset matching level.

[0069] In step S1032, the isolation sub-strategy can be a sensitive isolation method for different source data under different data source identifiers, including masking, desensitization, and / or encryption. For example, for text, the isolation sub-strategy includes, but is not limited to, keyword masking, summary truncation, ellipsis replacement, and watermark embedding. For images and / or video frames, the isolation sub-strategy includes, but is not limited to, region mosaic, target erasure, resolution reduction, color shifting, and frame discarding / replacement. For audio, the isolation sub-strategy includes, but is not limited to, audio replacement / overlay and frequency domain perturbation: slight frequency shifting or noise addition to sensitive frequency bands. The isolation sub-strategy in this application can include one or more isolation methods. That is, for data that needs to be isolated, data isolation can be performed using a single data isolation method or a combination of multiple data isolation methods.

[0070] Furthermore, different isolation sub-strategies correspond to different levels of isolation. Further, before executing step S1032, the sensitive information isolation method also includes the following steps: using isolation sub-strategy thresholds, multiple preset matching levels are sorted according to their isolation levels, so that corresponding isolation sub-strategies are selected sequentially when generating isolation strategies, thereby adjusting the priority of different isolation sub-strategies.

[0071] Step S1033: Select the corresponding isolation sub-strategy according to the corresponding matching level to generate the isolation strategy for each source data.

[0072] Please refer to Figure 5 The flowchart below shows step S105, which is provided in the embodiment of this application. Steps S1051-S1052 include detecting residual sensitive features and falsely isolated normal features of the isolated data to form an evaluation result, and then using the detected residual sensitive features and falsely isolated normal features as incremental samples to inject back into the initial dataset to update the dynamic sensitive database.

[0073] Step S1051: When residual sensitive features and falsely isolated normal features are detected from the isolation data, the corresponding residual sensitive features and falsely isolated normal features are compared with the evaluation threshold to form an evaluation result.

[0074] In step S1051, the evaluation results include, but are not limited to, residual rate, false isolation rate, user feedback data, isolation response time, resource utilization rate, rule adaptation coverage, and anti-interference capability. Specifically, the residual rate measures the proportion of residual sensitive features still detected after isolation; the false isolation rate measures the proportion of incorrectly isolated normal features; the isolation response time determines the time taken from identifying sensitive information through matching analysis to executing data isolation; the resource utilization rate reflects the consumption of system computing power and storage by the data isolation operation; the adaptation coverage rate provides feedback on the coverage ratio of different sensitive categories of sensitive information by the isolation sub-strategy; and the anti-interference capability measures the proportion of feature data accurately corresponding to sensitive information in noisy data (such as data containing a large amount of irrelevant content).

[0075] Step S1052: Based on the evaluation results, the isolated data that did not reach the evaluation threshold are injected back into the initial dataset as incremental samples to update the dynamic sensitive database, and the isolation parameter model is updated synchronously.

[0076] In step S1052, the sensitive categories corresponding to the isolated data that did not reach the evaluation threshold are obtained to obtain the true sensitive categories, as well as the source data before data isolation. This allows the residual sensitive features and incorrectly isolated normal features in the corresponding source data to be combined as incremental samples and synchronized to the initial dataset to update the dynamic sensitive database. Additionally, the incremental samples are also synchronized to a preset training set, and when the number of new isolated sample data reaches a certain amount, an update to the isolation parameter model is triggered to update the target isolation parameters.

[0077] like Figure 11 As shown, the source data corresponding to data source identifiers A_0, B_0, C_0, etc., are collected as multi-source data. After the sensitive information isolation method, if feature data such as pay slips, bank statements, images of classified locations, and noisy audio are detected and evaluated but not isolated, the feature data such as "pay slips, bank statements, images of classified locations, and noisy audio" that are still detected and not isolated, along with the corresponding real sensitive categories, are expanded to the initial dataset to update the dynamic sensitive database BASE_0. The feature data such as "pay slips, bank statements, images of classified locations, and noisy audio" are also added as incremental samples to the preset training set. When the number of incremental samples reaches a certain amount, the isolation parameter model is updated to update the target isolation parameters.

[0078] Please refer to Figure 8 This is a structural block diagram of the sensitive information isolation system provided in the embodiments of this application. This application also provides a sensitive information isolation system 10. The sensitive information isolation system 10 includes a data acquisition and parsing module 1, a matching and analysis module 2, a policy generation module 3, a data isolation module 4, and an evaluation and update module 5.

[0079] The acquisition and parsing module 1 is used to acquire and parse multi-source data from multiple data dissemination platforms in real time to obtain multiple feature data. The multi-source data includes text, images and / or video frames. Each source data has a data source identifier, which includes the dissemination platform type, the number of audiences, and the dissemination timestamp.

[0080] Matching analysis module 2 is used to perform matching analysis on multiple feature data to obtain matching results that reflect the degree of matching between each feature data and the initial sensitive category in the dynamic sensitive database. The dynamic sensitive database is constructed from the initial dataset, which includes the initial feature data corresponding to the initial sensitive category.

[0081] The strategy generation module 3 is used to generate an isolation strategy corresponding to each source data within a preset response time limit based on the matching results and the data source identifier. The isolation strategy is used to reduce, remove or hide the identifiability of the part of the feature data that has been matched as a sensitive category.

[0082] The data isolation module 4 is used to isolate the corresponding source data according to the isolation strategy to obtain multiple isolated data.

[0083] The evaluation update module 5 is used to detect residual sensitive features and falsely isolated normal features of the isolated data to form an evaluation result. Based on the evaluation result, the detected residual sensitive features and falsely isolated normal features are used as incremental samples to be injected back into the initial dataset to update the dynamic sensitive database.

[0084] Please refer to Figure 9 This is a schematic diagram of the internal structure of a computer device using the application of a sensitive information isolation method provided in the embodiments of this application.

[0085] like Figure 9 As shown, the computer device 100 includes a memory 901 and a processor 902. The processor 902 is used to execute computer program instructions stored in the memory 901 to implement a sensitive information isolation method.

[0086] The memory 901 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 901 can be an internal storage unit of a computer device, such as a hard disk. In other embodiments, the memory 901 can be an external storage device of a computer device, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc., configured in the computer device. Furthermore, the memory 901 can include both internal and external storage units of the computer device. The memory 901 can be used not only to store application software and various types of data installed on the computer device, such as code for sensitive information isolation methods, but also to temporarily store data that has been output or will be output.

[0087] Furthermore, the computer device 100 also includes a bus 903. The bus 903 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 9 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0088] Furthermore, the computer device 100 may also include a display component 904. The display component 904 may be an LED display, a liquid crystal display, a touch-screen liquid crystal display, or an organic light-emitting diode (OLED) touchscreen, etc. The display component 904 may also be appropriately referred to as a display device or display unit, used to display information processed in the computer device 100 and to display a visual user interface.

[0089] Furthermore, the computer device 100 may also include a communication component 905. The communication component 905 may optionally include a wired communication component and / or a wireless communication component (such as a Wi-Fi communication component, a Bluetooth communication component, etc.), which is typically used to establish a communication connection between the computer device 100 and other computer devices.

[0090] Figure 9 Only a computer device 100 with some components and a method for implementing sensitive information isolation is shown. Those skilled in the art will understand that... Figure 9 The structure shown does not constitute a limitation on the computer device 100 and may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0091] In the above embodiments, the implementation can be achieved, in whole or in part, through software, hardware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, in the form of a computer program product.

[0092] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to embodiments of the present invention is generated. The computer device may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state disk (SSD)).

[0093] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0094] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between apparatuses or units through some interfaces, and may be electrical, mechanical, or other forms.

[0095] The unit described as a separate component may or may not be physically separate. The component shown as a unit may or may not be a physical unit; that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0096] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist independently, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0097] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes: USB flash drives, portable hard disks, read-only storage media (ROM), random access storage media (RAM), magnetic disks, optical disks, and other media capable of storing program code.

[0098] In the above embodiments, text, image / video frames from multiple platforms are collected and parsed in real time. After normalization to obtain feature data, an isolation strategy is generated within a preset response time limit according to the matching results and data source identifiers (platform, audience, timestamp). Then, residual and mis-isolated features are injected back to update the sensitive database, so as to realize dynamic expansion of sensitive categories and time-limited accurate isolation.

[0099] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

[0100] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0101] The above-listed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. A method for isolating sensitive information, characterized in that, The sensitive information isolation method includes: Multiple feature data are obtained by real-time collection and analysis of multi-source data from various data dissemination platforms. The multi-source data includes text, images and / or video frames. Each source data has a data source identifier, which includes the dissemination platform type, the number of audiences, and the dissemination timestamp. The multiple feature data are matched and analyzed to obtain a matching result that reflects the degree of matching between each feature data and the initial sensitive category in the dynamic sensitive database. The dynamic sensitive database is constructed from an initial dataset, which includes initial feature data corresponding to the initial sensitive category. Based on the matching results and the data source identifier, an isolation strategy corresponding to each source data is generated within a preset response time limit. The isolation strategy is used to reduce, remove or hide the identifiability of the part of the feature data that has been matched as a sensitive category. According to the isolation strategy, data isolation is performed on the corresponding source data to obtain multiple isolated data; The residual sensitive features and falsely isolated normal features of the isolated data are detected to form an evaluation result. Based on the evaluation result, the detected residual sensitive features and falsely isolated normal features are used as incremental samples to be injected back into the initial dataset to update the dynamic sensitive database.

2. The sensitive information isolation method as described in claim 1, characterized in that, The feature data includes text feature data and / or image feature data; Real-time acquisition and analysis of multi-source data from various data dissemination platforms yields multiple feature data, including: Each of the source data is identified; When the source data contains text, the corresponding source data is parsed using the first parsing method to obtain the corresponding text feature data; When the source data contains images and / or video frames, the corresponding source data is parsed using the second parsing method to obtain the corresponding image feature data; The text feature data and / or image feature data corresponding to each source data are normalized to form the multiple feature data.

3. The sensitive information isolation method as described in claim 1, characterized in that, The sensitive information isolation method also includes: An isolation parameter model is obtained by feeding a preset training set into a preset training network. The preset training set includes multiple sample isolation data obtained after data isolation. Each sample isolation data includes the true sensitive category, the source data before isolation, the source data after isolation, and the sample feature data corresponding to the source data before isolation. The isolation parameters output by the isolation parameter model are fitted to the true sensitive categories to obtain the target isolation parameters. The target isolation parameters are used to adjust the matching analysis, isolation strategy generation, evaluation detection, and / or the refresh of the dynamic sensitive database.

4. The sensitive information isolation method as described in claim 3, characterized in that, The target isolation parameters include a dynamic update frequency; the sensitive information isolation method further includes: Set the dynamic update frequency to the current update frequency of the dynamic sensitive database; The initial dataset is dynamically expanded according to the current update frequency, so as to update the dynamic sensitive database based on the expanded initial dataset.

5. The sensitive information isolation method as described in claim 3, characterized in that, The matching result includes a matching score and a matching level, wherein the matching level is one of multiple preset matching levels; the target isolation parameter also includes a matching level threshold; matching analysis is performed on the multiple feature data to obtain a matching result reflecting the degree of matching between each feature data and the initial sensitive category in the dynamic sensitive database, including: The initial dataset is fed into a preset training model to learn a matching degree model, which is used to describe the correlation between feature data and matching scores; Based on the matching degree model, a matching score for each feature data is obtained by performing matching analysis on the multiple feature data. The matching score range corresponding to each preset matching level is determined by using the matching level threshold, and the matching score is assigned to the corresponding matching level to obtain the matching result of each source data.

6. The sensitive information isolation method as described in claim 5, characterized in that, The target isolation parameters also include isolation sub-policy thresholds; based on the matching results and the data source identifier, an isolation policy corresponding to each source data is generated within a preset response time limit, including: Based on the data source identifier, the sensitivity of each source data is obtained; Using the isolation sub-policy threshold, a corresponding isolation sub-policy is assigned to each preset matching level. The isolation sub-policy includes masking, desensitization and / or encryption. According to the corresponding matching level, select the corresponding isolation sub-strategy to generate the isolation strategy for each source data.

7. The sensitive information isolation method as described in claim 6, characterized in that, Different isolation sub-strategies correspond to different levels of isolation; The sensitive information isolation method also includes: Using the isolation sub-policy threshold, the multiple preset matching levels are sorted according to the degree of isolation, so that the corresponding isolation sub-policies are selected in sequence when generating the isolation policy.

8. The sensitive information isolation method as described in claim 3, characterized in that, The target isolation parameter also includes an evaluation threshold; detecting residual sensitive features and falsely isolated normal features of the isolated data to form an evaluation result, and then using the detected residual sensitive features and falsely isolated normal features as incremental samples to inject back into the initial dataset to update the dynamic sensitive database based on the evaluation result, including: When residual sensitive features and falsely isolated normal features are detected from the isolation data, the corresponding residual sensitive features and falsely isolated normal features are compared with the evaluation threshold to form the evaluation result; Based on the evaluation results, the isolated data that did not reach the evaluation threshold were injected back into the initial dataset as incremental samples to update the dynamic sensitive database, and the isolation parameter model was updated synchronously.

9. A sensitive information isolation system, characterized in that, The sensitive information isolation system includes: The acquisition and parsing module is used to acquire and parse multi-source data from multiple data dissemination platforms in real time to obtain multiple feature data. The multi-source data includes text, images and / or video frames. Each source data has a data source identifier, which includes the dissemination platform type, the number of audiences, and the dissemination timestamp. The matching analysis module is used to perform matching analysis on the multiple feature data to obtain a matching result that reflects the degree of matching between each feature data and the initial sensitive category in the dynamic sensitive database. The dynamic sensitive database is constructed from an initial dataset, which includes initial feature data corresponding to the initial sensitive category. The strategy generation module is used to generate an isolation strategy corresponding to each source data within a preset response time limit based on the matching result and the data source identifier. The isolation strategy is used to reduce, remove or hide the identifiability of the part of the feature data that has been matched as a sensitive category. The data isolation module is used to isolate the corresponding source data according to the isolation strategy to obtain multiple isolated data. An evaluation and update module is used to detect residual sensitive features and falsely isolated normal features of the isolated data to form an evaluation result, and to use the detected residual sensitive features and falsely isolated normal features as incremental samples to be injected back into the initial dataset to update the dynamic sensitive database based on the evaluation result.

10. A computer device, characterized in that, The computer device includes: Memory, used to store computer programs; and A processor for executing the computer program to implement the sensitive information isolation method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Computer sensitive data intelligent identification method based on artificial intelligence

    CN119312165A

  • Dynamic sensitive data outbound risk assessment method and system based on multi-source risk information

    CN120470590A

  • Dynamic data isolation method based on multi-dimensional security situation assessment

    CN120614202A

  • Multi-level dynamic isolation data processing system and method

    CN120781371A

  • Systems and methods for auto discovery of sensitive data in applications or databases using metadata via machine learning techniques

    US20220067185A1