USB electronic information anti-disclosure storage method and device

By deduplication, outlier processing and feature extraction of electronic information data in the USB flash drive, combined with the risk feature matching model, intelligent risk identification and differentiated encryption of USB flash drive data are achieved, and the problem of insufficient risk determination and encryption effects of USB flash drive data in the prior art is solved, and data security and resource utilization efficiency are improved.

CN120105458APending Publication Date: 2025-06-06HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510207029.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing USB flash drive electronic information anti-leakage system is not convenient for risk determination of its internal information data, and the encryption effect in single-fold encryption is poor, resulting in waste of resources and insufficient security.

Method used

By deduplication and outlier processing of electronic information data in USB, segmentation and numbering, extracting feature sets, building a risk feature matching model, and encrypting high-level and low-level levels according to risk levels.

Benefits of technology

It realizes intelligent risk identification of electronic information data, reduces security vulnerabilities caused by data errors, improves the protection efficiency of sensitive data, reduces the processing overhead of non-sensitive data, and meets the needs of staff.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105458A_ABST
    Figure CN120105458A_ABST
Patent Text Reader

Abstract

The invention discloses a USB electronic information anti-leakage storage method and device. The method comprises the steps that all electronic information data in a USB are obtained; duplicate removal and abnormal value processing are carried out on all electronic information data in the USB, it is ensured that data input into the encryption process is accurate and reliable, security vulnerabilities caused by data errors are reduced, encryption of different levels is carried out according to the risk levels of the data, high-risk data are encrypted at high levels, and the security vulnerabilities caused by data errors are reduced. Low-level encryption is adopted for low-risk data, sensitive data can be protected more effectively through the differentiated encryption strategy, intelligent risk identification of electronic information data is achieved through feature extraction and construction of a risk feature matching model, potential high-risk data are found and marked in time, and the safety of the electronic information data is improved. According to the method, risk classification and encryption processing are carried out on the data, so that an enterprise can better comply with related laws and regulations and industrial standards, and legal risks caused by data leakage are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic information leakage prevention storage, and in particular to a USB electronic information leakage prevention storage method and device. Background Art

[0002] Electronic information refers to the text, icons, images, audio, video and other file information records generated by computers and other electronic devices and stored in digital form. As an information resource, electronic information is a true record of an enterprise's production, technology, scientific research and business operations. It is an intangible asset that develops in sync with the enterprise and plays an important role in various aspects of enterprise management.

[0003] As a mobile storage device, U disk is deeply favored by users for its convenience and portability. It has become an indispensable information transmission tool in people's daily life and work. At the same time, it also brings risks such as losing equipment and leaking data. Therefore, people have higher and higher requirements for its security and confidentiality. However, the existing U disk electronic information leakage prevention system is not convenient for risk assessment of the information data inside it. If all data are encrypted, it will lead to excessive waste of resources and it is difficult to meet the needs of staff. In addition, most of the existing encryption technologies are single encryption forms, and the encryption effect is still poor. Summary of the invention

[0004] In order to solve the above technical problems, the present invention provides a USB electronic information anti-leakage storage method and device. The technical solution solves the problem that the existing USB disk electronic information anti-leakage system proposed in the above background technology is not convenient for risk assessment of the information data inside it. If all data are encrypted, it will lead to excessive waste of resources and it is difficult to meet the needs of staff. In addition, most of the existing encryption technologies are in the form of single encryption, and the encryption effect is still poor.

[0005] In order to achieve the above purpose, the technical solution adopted by the present invention is: In a first aspect of the present invention, a USB electronic information leakage prevention storage method is provided, comprising: Acquire all electronic information data in the USB, wherein all electronic information data includes domains, packages, transactions, transmissions, and other types, and pre-process the electronic information data, wherein the pre-processing includes data deduplication and outlier processing; Segmenting and numbering all pre-processed electronic information data, and extracting features from each segment of electronic information data to obtain a feature set of the electronic information data; Based on the risk feature word bag pre-set in the anti-leakage system, a risk feature matching model is constructed; Inputting the feature set of the electronic information data into the risk feature matching model, obtaining the risk feature matching value, and performing comparison to obtain high-risk electronic information data and low-risk electronic information data; High-risk electronic information data and low-risk electronic information data are encrypted at high levels and low levels respectively.

[0006] Preferably, the deduplication of electronic information data specifically includes the following steps: Mark different types of electronic information data in the form of coordinates, and arrange the different types of electronic information into columns in order; Compare the duplication of the first coordinate data of each column with the subsequent coordinate data. If duplicate data is found, delete the subsequent coordinate data. After the comparison is completed, the second coordinate data is compared with the subsequent coordinate data for duplication. If there is duplication, the subsequent coordinates are deleted; Repeat the above steps until all coordinate data are compared.

[0007] Preferably, the outlier processing of electronic information data specifically includes the following steps: Mark different types of electronic information data in the form of coordinates, and arrange the different types of electronic information into columns in order; Calculate the deviation value between the coordinate data and the adjacent coordinate data as k; When the deviation value k is greater than 100%, the data coordinate is an outlier and the data coordinate is replaced by the mean of the adjacent data. Otherwise, no replacement is required. The deviation formula is:

[0008] Preferably, the step of segmenting and numbering all pre-processed electronic information data, extracting features from each segment of electronic information data, and obtaining a feature set of electronic information data specifically comprises the following steps: Segment all different types of electronic information data after pre-processing; All different types of electronic information data are numbered as S1, S2, ..... Sn; Perform feature extraction on different types of electronic information data in each number segment at the same time, obtain and count each feature in the electronic information data of each number; All features are grouped together.

[0009] Preferably, the risk feature matching model is:

[0010] in, is the matching degree between the input feature set and the pre-set j-th risk level, is the total number of feature sets corresponding to the jth risk level, is the total number of features in the input feature set, is the total number of identical features between the feature data of the input feature set and the feature set corresponding to the jth risk level, is the correlation weight between the feature data of the j-th risk level and the feature set of the l-th input and the same feature in the feature set corresponding to the j-th risk level, To find the maximum value function.

[0011] Preferably, the step of inputting the feature set of the electronic information data into the risk feature matching model, obtaining the risk feature matching value, and performing comparison to obtain the high-risk electronic information data and the low-risk electronic information data specifically comprises the following steps: Inputting the feature set of the electronic information data into the risk feature matching model to obtain the risk feature matching value; Compare the risk feature matching value with the preset threshold; If the risk feature matching value is greater than or equal to the preset threshold, the data is determined to be high-risk electronic information data; If the risk feature matching value is less than the preset threshold, the data is judged to be low-risk electronic information data.

[0012] Preferably, the high-level encryption specifically includes the following steps: According to the content of the electronic information data, the electronic information data is segmented to obtain multiple segments of power communication sub-data; Each piece of electronic information data is segmented according to a fixed length to obtain a plurality of encoding vectors; Generate a dummy encoding vector for each encoding vector; According to the virtual code vector, each code vector is encrypted based on the local encryption sequence to obtain each encrypted code vector; splicing multiple encrypted coding vectors to obtain each segment of encrypted electronic information data; Each segment of encrypted electronic information data is encrypted according to the global encryption sequence to obtain double-encrypted electronic information data.

[0013] Preferably, the low-level encryption specifically includes the following steps: According to the content of the electronic information data, the electronic information data is segmented to obtain multiple segments of power communication sub-data; Multiple segments of electronic information data are shuffled in a disordered manner to obtain low-level encrypted electronic information data.

[0014] In a second aspect of the present invention, a USB electronic information leakage prevention storage system is also provided, comprising: An acquisition module, the acquisition module is used to acquire all electronic information data in the USB, the electronic information data includes domains, packages, transactions and transmissions, and pre-process the electronic information data, the pre-processing includes data deduplication and outlier processing; A preprocessing module, which is used to segment and number all the preprocessed electronic information data, and extract features from each segment of the electronic information data to obtain a feature set of the electronic information data; A model building module, wherein the model building module is used to build a risk feature matching model based on a risk feature word bag pre-set in the anti-leakage system; A calculation and comparison module, wherein the calculation and comparison module is used to input the feature set of the electronic information data into the risk feature matching model, obtain the risk feature matching value, and perform a comparison to obtain high-risk electronic information data and low-risk electronic information data; An encryption module is used to perform high-level encryption and low-level encryption on high-risk electronic information data and low-risk electronic information data respectively.

[0015] In a third aspect of the present invention, an electronic device is further provided. The electronic device comprises at least one processor; and a memory connected to the at least one processor in communication; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the method of the first aspect of the present invention.

[0016] Compared with the prior art, the present invention provides a USB electronic information leakage prevention storage method and device, which has the following beneficial effects: 1. The present invention deduplicates and processes outliers for all electronic information data in the USB, thereby ensuring that the data input into the encryption process is accurate and reliable, reducing security vulnerabilities caused by data errors, and performing different levels of encryption according to the risk level of the data. High-risk data uses high-level encryption, and low-risk data uses low-level encryption. This differentiated encryption strategy can more effectively protect sensitive data. Through feature extraction and the construction of risk feature matching models, intelligent risk identification of electronic information data is achieved, and potential high-risk data is discovered and marked in a timely manner. By performing risk classification and encryption processing on data, the present invention enables enterprises to better comply with relevant laws, regulations and industry standards, and reduce the legal risks faced due to data leakage.

[0017] 2. When encrypting electronic information of a high-risk level, the present invention divides the electronic information data according to its content to obtain multiple segments of power communication sub-data, which helps to classify and process the data according to different data types or business requirements. Each segment of the electronic information data is segmented according to a fixed length to obtain multiple coding vectors, thereby increasing the granularity of data processing and allowing subsequent operations to be more finely controlled. A virtual coding vector is generated for each coding vector, increasing the complexity of the data and making it more difficult for attackers to analyze or predict the original data. Each coding vector is encrypted according to a virtual coding vector and a local encryption sequence, providing a first level of data protection. Even if part of the data is intercepted, it is difficult for an attacker to decrypt the complete information. Each segment of the encrypted electronic information data is re-encrypted according to the global encryption sequence, thereby achieving double encryption and further improving data security. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A schematic diagram of a USB electronic information leakage prevention storage method in the present invention; Figure 2 A schematic diagram of a method for deduplicating electronic information data in the present invention; Figure 3 A schematic diagram of a method for processing outliers in electronic information data in the present invention; Figure 4 A schematic diagram of a method for obtaining a feature set of electronic information data in the present invention; Figure 5 A schematic diagram of a method for obtaining high-risk electronic information data and low-risk electronic information data in the present invention; Figure 6 It is a schematic diagram of the high-level encryption method of the present invention; Figure 7 A schematic diagram of a low-level encryption method of the present invention; Figure 8 A schematic diagram of a USB electronic information leakage prevention storage system in the present invention; Fig. 9 A block diagram of an exemplary electronic device capable of implementing embodiments of the present invention is shown; Among them, 900 is an electronic device, 901 is a computing unit, 902 is a ROM, 903 is a RAM, 904 is a bus, 905 is an I / O interface, 906 is an input unit, 907 is an output unit, 908 is a storage unit, and 909 is a communication unit. DETAILED DESCRIPTION

[0019] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are only examples, and those skilled in the art may think of other obvious variations.

[0020] Example 1 Please refer to Figure 1 As shown, a USB electronic information leakage prevention storage method and device, in a first aspect of the present invention, a USB electronic information leakage prevention storage method is provided, comprising: S101, obtaining all electronic information data in the USB, including domains, packages, transactions, transmissions, etc., and preprocessing the electronic information data, including data deduplication and outlier processing; S102, segmenting and numbering all pre-processed electronic information data, and extracting features from each segment of electronic information data to obtain a feature set of the electronic information data; S103, constructing a risk feature matching model based on the risk feature word bag pre-set in the anti-leakage system; S104, inputting the feature set of the electronic information data into the risk feature matching model, obtaining the risk feature matching value, and performing comparison to obtain high-risk electronic information data and low-risk electronic information data; S105. Perform high-level encryption and low-level encryption on high-risk electronic information data and low-risk electronic information data respectively.

[0021] It will be understood by those skilled in the art that the present invention ensures that the data input into the encryption process is accurate and reliable by deduplicating and processing outliers for all electronic information data in the USB, thereby reducing security vulnerabilities caused by data errors, and performing different levels of encryption according to the risk level of the data, with high-risk data using high-level encryption and low-risk data using low-level encryption. This differentiated encryption strategy can more effectively protect sensitive data while reducing the processing overhead of non-sensitive data. Through feature extraction and the construction of risk feature matching models, intelligent risk identification of electronic information data is achieved, and potential high-risk data is discovered and marked in a timely manner, providing an important basis for subsequent encryption processing. The extraction of feature sets simplifies data preparation before encryption and improves encryption efficiency. The present invention enables enterprises to better comply with relevant laws, regulations and industry standards and reduce legal risks due to data leakage by performing risk classification and encryption processing on data.

[0022] Please refer to Figure 2 As shown, deduplication of electronic information data specifically includes the following steps: S201, marking different types of electronic information data in the form of coordinates, and arranging the different types of electronic information into columns in order; S202, comparing the duplication of the first coordinate data of each column with the subsequent coordinate data, and if duplicate data is found, deleting the subsequent coordinate data; S203, after the comparison is completed, the second coordinate data is compared with the subsequent coordinate data for repetition, and if there is repetition, the subsequent coordinates are deleted; S204, repeat the above steps until all coordinate data are compared.

[0023] It can be understood by those skilled in the art that the present invention can effectively identify and delete duplicate data points by gradually comparing the coordinate data of each column, which is crucial for reducing data redundancy and improving data quality. After removing duplicate data, the space required for data storage can be significantly reduced, which is especially important in an environment with limited resources, and can save storage costs and improve storage efficiency. The existence of duplicate data will increase the complexity of data processing and analysis. By removing these duplicate data, the subsequent data processing steps can be simplified and the processing speed can be improved.

[0024] Please refer to Figure 3 As shown, the outlier processing of electronic information data specifically includes the following steps: S301, marking different types of electronic information data in the form of coordinates, and arranging different types of electronic information into columns in order; S302, calculating the deviation value k between the coordinate data and the adjacent coordinate data; S303, when the deviation value k is greater than 100%, the data coordinate is an outlier, and the data coordinate is replaced by the mean of the adjacent data, otherwise no replacement is required; The deviation formula is:

[0025] Please refer to Figure 4 As shown, all the pre-processed electronic information data are segmented and numbered, and feature extraction is performed on each segment of the electronic information data. The feature set of the electronic information data is obtained, which specifically includes the following steps: S401, segmenting all different types of electronic information data after preprocessing; S402, numbering all different types of electronic information data, recording as S1, S2, ..... Sn; S403, extracting features of different types of electronic information data in each number segment at the same time, obtaining and counting each feature in each numbered electronic information data; S404: All features are grouped together.

[0026] It can be understood by those skilled in the art that, in the present invention, data is organized in an orderly manner through segmentation and numbering, which is convenient for subsequent processing and analysis. Each segment has a unique identifier (number), which helps to track and manage data. Feature extraction is performed on the data of each numbered segment to identify key information or patterns in the data. Feature extraction is an important step in data mining and machine learning. It helps to convert data into a numerical form that can be used for prediction or classification. By forming a set of all features, a comprehensive feature library can be created for subsequent data analysis, model training and prediction. The integrity of the feature set is crucial to the accuracy and reliability of the model. The feature set provides a basis for advanced data analysis (such as cluster analysis, classification analysis, association rule mining, etc.), which helps to reveal hidden patterns, outliers and potential values ​​in the data. By segmenting and numbering the data, data of multiple segments can be processed in parallel, improving the efficiency of data processing. Feature extraction and counting can be completed at the same time, reducing the time cost of subsequent analysis. The specific Python code and comments are as follows: import pandas as pd from collections import defaultdict # Assume that the preprocessed electronic information data is stored in multiple DataFrames, which are given here in list form preprocessed_data = [df1, df2, ..., dfn] # df1, df2, ..., dfn are pandas DataFrame objects # Segment all different types of electronic information data after preprocessing # In this example, each DataFrame has been treated as a segment # Number all different types of electronic information data segments = {f'S{i+1}': df for i, df in enumerate(preprocessed_data)} # Perform feature extraction on different types of electronic information data in each number segment at the same time, obtain each feature in the electronic information data of each number and count it feature_counts = defaultdict(int) for segment_id, df in segments.items(): # Assume that feature extraction is based on counting values ​​of certain columns # Here we take a hypothetical column 'feature_column' as an example if 'feature_column' in df.columns: feature_values ​​= df['feature_column'].value_counts() for feature, count in feature_values.items(): # Use segment number and feature name as key and count as value feature_counts[f'{segment_id}_{feature}'] += count # Form all features into a set (this actually forms a dictionary, the key is the feature name, the value is the count) # But in order to conform to the concept of "set" (not containing duplicate elements), we can convert the feature name into a set # But since we need count information, it is more appropriate to use a dictionary form features_set = {key.split('_')[1] for key in feature_counts.keys()}# Extract only the feature name (excluding the segment number) # But note that this loses the count information. In order to retain the complete information, we should use a dictionary # It is more appropriate to use the feature_counts dictionary directly, which contains all features and their counts # features_with_counts = feature_counts # Keep features and count information # Output feature set (feature name only, without count. If count is needed, use features_with_counts) print("Feature set (feature name only) :", features_set) # Output features and count information # for feature, count in features_with_counts.items(): # print(f"Feature {feature.split('_')[1]}: count {count}") The risk feature matching model is:

[0027] in, is the matching degree between the input feature set and the pre-set j-th risk level, is the total number of feature sets corresponding to the jth risk level, is the total number of features in the input feature set, is the total number of identical features between the feature data of the input feature set and the feature set corresponding to the jth risk level, is the correlation weight between the feature data of the j-th risk level and the feature set of the l-th input and the same feature in the feature set corresponding to the j-th risk level, To find the maximum value function.

[0028] Please refer to Figure 5 As shown, the feature set of electronic information data is input into the risk feature matching model, the risk feature matching value is obtained, and a comparison is performed to obtain high-risk electronic information data and low-risk electronic information data, which specifically includes the following steps: S501, inputting the feature set of electronic information data into the risk feature matching model to obtain the risk feature matching value; S502, comparing the risk feature matching value with a preset threshold; S503: If the risk feature matching value is greater than or equal to a preset threshold, the data is determined to be high-risk electronic information data; S504: If the risk feature matching value is less than a preset threshold, the data is determined to be low-risk electronic information data.

[0029] Please refer to Figure 6 As shown, high-level encryption specifically includes the following steps: S601, dividing the electronic information data according to the content of the electronic information data to obtain multiple segments of power communication sub-data; S602, segmenting each segment of electronic information data according to a fixed length to obtain multiple encoding vectors; S603, generating a virtual coding vector for each coding vector; S604, encrypting each code vector according to the virtual code vector and based on the local encryption sequence to obtain each encrypted code vector; S605, concatenating multiple encryption coding vectors to obtain each segment of encrypted electronic information data; S606. Encrypt each segment of encrypted electronic information data according to the global encryption sequence to obtain double-encrypted electronic information data.

[0030] It will be understood by those skilled in the art that segmenting the electronic information data according to its content to obtain multiple segments of power communication sub-data helps to classify and process the data according to different data types or business requirements, segmenting each segment of the electronic information data according to a fixed length to obtain multiple coding vectors increases the granularity of data processing, allowing subsequent operations to be more finely controlled, generating a virtual coding vector for each coding vector increases the complexity of the data, and makes it more difficult for attackers to analyze or predict the original data, encrypting each coding vector according to the virtual coding vector and the local encryption sequence provides a first level of data protection, and even if part of the data is intercepted, it is difficult for an attacker to decrypt the complete information, and re-encrypting each segment of the encrypted electronic information data according to the global encryption sequence implements double encryption, further improving data security.

[0031] The specific Python code and comments are as follows: import numpy as np import hashlib from Crypto.Cipher import AES from Crypto.Util.Padding import pad, unpad from Crypto.Random import get_random_bytes # Assume that the electronic information data is a long string or binary data # In actual applications, it may be obtained from a file, database or network Electronic information data = b"Here is the content of your electronic information data..." # According to the content of electronic information data, the electronic information data is segmented to obtain multiple segments of power communication sub-data # The segmentation logic here depends on your specific needs, such as segmentation by specific delimiter, fixed size or content mode # The following is a simple example of splitting by fixed size Sub-segment size = 1024 # Assume that each sub-segment size is 1024 bytes Sub-data list = [electronic information data[i:i+sub-data segment size] for i in range(0, len(electronic information data), sub-data segment size)] #Divide each piece of electronic information data into segments of fixed length to obtain multiple encoding vectors # The fixed length here depends on your encryption algorithm and storage / transmission requirements Encoding vector size = 64 #Assume each encoding vector size is 64 bytes Encoding vector list = [[subdata[j:j+encoding vector size] for j in range(0, len(subdata), encoding vector size)] for subdata in subdata list] # S603: Generate a dummy encoding vector for each encoding vector # The generation of dummy encoding vectors can be based on some transformation of the original encoding vector or random generation # The following is a simple example using a randomly generated dummy encoding vector (should be safer in real applications) dummy_encoded_vector_list = [[[get_random_bytes(encoded_vector_size) for _ in range(len(vector_group))] for vector_group in segments] for segments in encoded_vector_list] # According to the virtual coding vector, based on the local encryption sequence, encrypt each coding vector # The local encryption sequence and encryption algorithm here depend on your security requirements # The following is an example of encryption using the AES encryption algorithm and a randomly generated key def localencrypt(encodingvector, key): cipher = AES.new(key, AES.MODE_CBC) ct_bytes = cipher.encrypt(pad(encodingvector, AES.block_size)) return ct_bytes local_key = get_random_bytes(32) # AES-256 key length EncryptedEncodedVectorList = [[[LocalEncryption(vector, localKey) for vector in vectorGroup]for vectorGroup in segment] for segment in EncodedVectorList] # Concatenate multiple encrypted coding vectors to obtain each segment of encrypted electronic information data Encrypted subdata list = [b''.join(b''.join(vector group) for vector group in segment) for segment in encrypted encoding vector list] # According to the global encryption sequence, encrypt each piece of encrypted electronic information data to obtain double encrypted electronic information data # The global encryption sequence and encryption algorithm here also depend on your security requirements # You can use the same algorithm as local encryption, but with a different key Global key = get_random_bytes(32) # Another key for AES-256 def globalencrypt(data, key): cipher = AES.new(key, AES.MODE_CBC) ct_bytes = cipher.encrypt(pad(data, AES.block_size)) return ct_bytes Double encrypted electronic information data list = [global encryption (subdata, global key) for subdata in encrypted subdata list] # Now, the double encrypted electronic information data list contains the encryption results of each piece of data # You can store this data in a file, a database, or transmit it over the network Please refer to Figure 7 As shown, low-level encryption specifically includes the following steps: S701, dividing the electronic information data according to the content of the electronic information data to obtain multiple segments of power communication sub-data; S702, shuffling multiple segments of electronic information data in a disordered manner to obtain low-level encrypted electronic information data.

[0032] In a second aspect of the present invention, a USB electronic information leakage prevention storage system is also provided, the system 800 comprising: The acquisition module 810 is used to acquire all electronic information data in the USB, including domains, packages, transactions, and transmissions, and preprocess the electronic information data, including data deduplication and outlier processing; A preprocessing module 820 is used to segment and number all the preprocessed electronic information data, and extract features from each segment of the electronic information data to obtain a feature set of the electronic information data; A model building module 830, which is used to build a risk feature matching model based on a risk feature word bag pre-set in the anti-leakage system; A calculation and comparison module 840 is used to input the feature set of the electronic information data into the risk feature matching model, obtain the risk feature matching value, and perform a comparison to obtain high-risk electronic information data and low-risk electronic information data; The encryption module 850 is used to perform high-level encryption and low-level encryption on high-risk electronic information data and low-risk electronic information data respectively.

[0033] In the third aspect of the present invention, an electronic device is also provided.

[0034] Fig. 9 A schematic block diagram of an electronic device 900 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.

[0035] The electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0036] Multiple components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0037] The computing unit 901 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as methods S100~S600. For example, in some embodiments, methods S101~S105 may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the methods S101~S105 described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to execute methods S101 - S105 in any other appropriate manner (eg, by means of firmware).

[0038] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0039] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.

[0040] In the context of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0041] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0042] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0043] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0044] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.

[0045] In summary, the present invention ensures that the data input into the encryption process is accurate and reliable by deduplicating and processing outliers for all electronic information data in the USB, reduces security vulnerabilities caused by data errors, and performs different levels of encryption according to the risk level of the data. High-risk data uses high-level encryption, and low-risk data uses low-level encryption. This differentiated encryption strategy can more effectively protect sensitive data while reducing the processing overhead of non-sensitive data. Through feature extraction and the construction of risk feature matching models, intelligent risk identification of electronic information data is achieved, and potential high-risk data is discovered and marked in time, providing an important basis for subsequent encryption processing. The extraction of feature sets simplifies data preparation before encryption and improves encryption efficiency. The present invention enables enterprises to better comply with relevant laws, regulations and industry standards and reduce legal risks due to data leakage by risk classification and encryption processing of data.

[0046] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions only describe the principles of the present invention. The present invention may be subject to various changes and improvements without departing from the spirit and scope of the present invention. These changes and improvements fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the attached claims and their equivalents.

Claims

1. A USB electronic information leakage prevention storage method, characterized in that: include: Acquire all electronic information data in the USB, wherein all electronic information data includes domains, packages, transactions, transmissions, and other types, and pre-process the electronic information data, wherein the pre-processing includes data deduplication and outlier processing; Segmenting and numbering all pre-processed electronic information data, and extracting features from each segment of electronic information data to obtain a feature set of the electronic information data; Based on the risk feature word bag pre-set in the anti-leakage system, a risk feature matching model is constructed; Inputting the feature set of the electronic information data into the risk feature matching model, obtaining the risk feature matching value, and performing comparison to obtain high-risk electronic information data and low-risk electronic information data; High-risk electronic information data and low-risk electronic information data are encrypted at high levels and low levels respectively.

2. A USB electronic information leakage prevention storage method according to claim 1, characterized in that: Deduplication of electronic information data specifically includes the following steps: Mark different types of electronic information data in the form of coordinates, and arrange the different types of electronic information into columns in order; Compare the duplication of the first coordinate data of each column with the subsequent coordinate data. If duplicate data is found, delete the subsequent coordinate data. After the comparison is completed, the second coordinate data is compared with the subsequent coordinate data for duplication. If there is duplication, the subsequent coordinates are deleted; Repeat the above steps until all coordinate data are compared.

3. A USB electronic information leakage prevention storage method according to claim 2, characterized in that: The processing of outliers on electronic information data specifically includes the following steps: Mark different types of electronic information data in the form of coordinates, and arrange the different types of electronic information into columns in order; Calculate the deviation value between the coordinate data and the adjacent coordinate data as k; When the deviation value k is greater than 100%, the data coordinate is an outlier and the data coordinate is replaced by the mean of the adjacent data. Otherwise, no replacement is required. The deviation formula is:

4. A USB electronic information leakage prevention storage method according to claim 3, characterized in that: The steps of segmenting and numbering all the pre-processed electronic information data, extracting features from each segment of the electronic information data, and obtaining a feature set of the electronic information data specifically include the following steps: Segment all different types of electronic information data after pre-processing; All different types of electronic information data are numbered as S1, S2, ..... Sn; Perform feature extraction on different types of electronic information data in each number segment at the same time, obtain and count each feature in the electronic information data of each number; All features are grouped together.

5. A USB electronic information leakage prevention storage method according to claim 4, characterized in that: The risk feature matching model is: ; in, is the matching degree between the input feature set and the pre-set j-th risk level, is the total number of feature sets corresponding to the jth risk level, is the total number of features in the input feature set, is the total number of identical features between the feature data of the input feature set and the feature set corresponding to the jth risk level, is the correlation weight between the feature data of the j-th risk level and the feature set of the l-th input and the same feature in the feature set corresponding to the j-th risk level, To find the maximum value function.

6. A USB electronic information leakage prevention storage method according to claim 5, characterized in that: The step of inputting the feature set of the electronic information data into the risk feature matching model, obtaining the risk feature matching value, and performing comparison to obtain the high-risk electronic information data and the low-risk electronic information data specifically includes the following steps: Inputting the feature set of the electronic information data into the risk feature matching model to obtain the risk feature matching value; Compare the risk feature matching value with the preset threshold; If the risk feature matching value is greater than or equal to the preset threshold, the data is determined to be high-risk electronic information data; If the risk feature matching value is less than the preset threshold, the data is judged to be low-risk electronic information data.

7. A USB electronic information leakage prevention storage method according to claim 6, characterized in that: The high-level encryption specifically includes the following steps: According to the content of the electronic information data, the electronic information data is segmented to obtain multiple segments of power communication sub-data; Each piece of electronic information data is segmented according to a fixed length to obtain a plurality of encoding vectors; Generate a dummy encoding vector for each encoding vector; According to the virtual code vector, each code vector is encrypted based on the local encryption sequence to obtain each encrypted code vector; splicing multiple encrypted coding vectors to obtain each segment of encrypted electronic information data; Each segment of encrypted electronic information data is encrypted according to the global encryption sequence to obtain double-encrypted electronic information data.

8. A USB electronic information leakage prevention storage method according to claim 7, characterized in that: The low-level encryption specifically includes the following steps: According to the content of the electronic information data, the electronic information data is segmented to obtain multiple segments of power communication sub-data; Multiple segments of electronic information data are shuffled in a disordered manner to obtain low-level encrypted electronic information data.

9. A USB electronic information anti-leakage storage system, used to implement a USB electronic information anti-leakage storage method as claimed in any one of claims 1 to 8, characterized in that: include: An acquisition module, the acquisition module is used to acquire all electronic information data in the USB, the electronic information data includes domains, packages, transactions and transmissions, and pre-process the electronic information data, the pre-processing includes data deduplication and outlier processing; A preprocessing module, which is used to segment and number all the preprocessed electronic information data, and extract features from each segment of the electronic information data to obtain a feature set of the electronic information data; A model building module, wherein the model building module is used to build a risk feature matching model based on a risk feature word bag pre-set in the anti-leakage system; A calculation and comparison module, wherein the calculation and comparison module is used to input the feature set of the electronic information data into the risk feature matching model, obtain the risk feature matching value, and perform a comparison to obtain high-risk electronic information data and low-risk electronic information data; An encryption module is used to perform high-level encryption and low-level encryption on high-risk electronic information data and low-risk electronic information data respectively.

10. An electronic device comprising at least one processor; and a memory connected in communication with the at least one processor; characterized in that: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.