A data identification method, device and storage medium based on feature vector matching
By using feature vector matching methods in a trusted execution environment, the core or importance of data is automatically identified, and the problems of low efficiency and poor adaptability of traditional data classification and grading methods are solved, and efficient and accurate data classification and recognition are achieved.
Patent Information
- Application Number
- CN202510089450.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-21
AI Technical Summary
Traditional data classification and grading methods rely on manual annotation and static rule databases, resulting in low efficiency and poor adaptability, and cannot meet the real-time classification needs in the big data environment.
Using a method based on feature vector matching, the core or importance of the data is automatically identified by obtaining and processing the description information of the data in a trusted execution environment, the Euclidean distance of the feature vector is calculated, and the weight and probability value of the data are calculated based on the matching results and data scale.
It improves the efficiency and accuracy of data processing, is highly adaptable, can quickly process large-scale data, reduce mismatch and misreport, and enhances the privacy and security of data.
Smart Images

Figure CN119513674B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a data classification and grading identification method based on feature vector matching. Background Art
[0002] With the widespread application of big data technology and the continuous accumulation of data resources, data has become an important asset for various industries, enterprises and even government departments. Especially in the fields of medical care, finance, energy, transportation, etc., the value of data is getting higher and higher. However, at the same time, data security issues are becoming increasingly prominent. How to accurately and effectively protect different types of data has become an urgent problem to be solved worldwide. In the data security management framework, data classification and grading, as a basic task, is directly related to the accuracy, efficiency and security of data protection. By reasonably classifying and grading data, it can ensure that sensitive data is strictly protected to prevent data leakage, abuse and improper access.
[0003] At present, the methods of data classification and grading mainly rely on rule engines or manual definitions. Traditional methods often rely on the experience of experts to formulate rules, and judge the sensitivity or importance of data through static rule bases or keyword-based matching. Although this method can provide certain effects in some relatively simple application scenarios, its limitations are becoming increasingly apparent as the scale of data expands and the application scenarios become more complex, as shown in the following aspects:
[0004] Traditional data classification and grading methods usually require manual labeling of large amounts of data and the establishment of a static rule base based on the labeling results. This method is highly dependent on experts and requires a lot of time and resources for data labeling and rule definition. As time goes by, the maintenance and update of the rule base becomes a cumbersome task. With dynamically changing business needs and growing data volumes, the update speed of the static rule base is far from meeting the needs of real-time data processing, resulting in a significant reduction in the timeliness of data classification and grading.
[0005] Traditional methods often use rule matching to judge data during data classification and grading. Due to the large and complex rule base, the matching process requires a huge amount of calculation when processing large amounts of data, and the algorithm is not optimized enough, resulting in low processing efficiency. In a big data environment, traditional methods cannot meet the needs of fast and real-time classification. This makes data classification and grading often become a bottleneck that affects the efficiency of data security management, and it cannot reflect data changes and the characteristics of newly added data types in real time.
[0006] The accuracy of traditional rule matching methods is limited by the completeness and accuracy of the rule base. The rule base only contains fixed keywords and templates, and cannot perform in-depth semantic analysis of data. For complex and diverse actual data, rule matching is prone to omissions (failure to identify important data) or false positives (misjudging unimportant data as important data). For example, when processing structured data, the rules may not accurately capture the implicit features in the data, or have weak recognition capabilities for unstructured data (such as text, images, etc.), resulting in unsatisfactory matching results.
[0007] As industry applications continue to deepen, the feature dimensions and scale of data are gradually increasing, and the speed of data itself is also changing faster and faster. For example, data in the medical industry continuously generates new examination items, treatment plans, and patient information, and data in the financial industry continuously adds new financial products, transaction records, etc. Traditional data classification and grading methods are difficult to flexibly adapt to changes in new data types and features, or it is difficult to update the rule base in a short period of time. The diversity and variability of data features require data classification and grading methods to be highly adaptable and flexible, but traditional methods find it difficult to do this, resulting in frequent failure of the rule base or failure to cover the latest data features.
[0008] Traditional data classification and grading methods mostly rely on manual definition and rule-driven, lacking support for deep learning and intelligent analysis of data content. With the rapid development of technologies such as artificial intelligence and machine learning, intelligent data processing and analysis have become possible. Compared with traditional methods, intelligent methods can automatically extract features from data and perform classification and grading through technologies such as deep learning and natural language processing (NLP), significantly improving processing efficiency and accuracy. However, many current data classification and grading systems still rely on static rules and manual intervention, and fail to make full use of modern intelligent technologies to improve the automation level of the system. Summary of the invention
[0009] In order to solve the above technical problems, the present application provides a data recognition method, device and storage medium based on feature vector matching. The technical solution of the present application is described below:
[0010] The first aspect of the present application provides a data recognition method based on feature vector matching, comprising:
[0011] Based on the trusted execution environment TEE of the data provider, obtain the database description information, table description information and field description information of the data to be tested;
[0012] Based on the trusted execution environment TEE of the data provider, the data to be detected is segmented to obtain a feature vector of the data to be detected;
[0013] Loading a feature vector library of predefined core data and important data into the TEE;
[0014] For each feature vector to be matched in the feature vector library, the Euclidean distance between the feature vectors and all predefined core data and important data feature vectors is calculated through matrix operations;
[0015] Determine the number of feature vectors that match the feature vector to be matched according to the calculation result;
[0016] Calculate the weight value of the data according to the feature vector matching results in the description information, table description information and field description information of the database, wherein the weight of the description information is 0.2, the weight of the table description information is 0.5, and the weight of the field description information is 0.3;
[0017] According to the matching results and the data size, the probability value of the data to be detected belonging to the core data or important data is calculated, and the probability value is obtained by weighted calculation based on the number of matches and the data size;
[0018] It is determined whether the probability value exceeds a predetermined threshold value. If so, it is determined that the data to be detected is core data or important data.
[0019] Optionally, the table description information is metadata at the database level, including the name, type, and structure of the database;
[0020] The database description information is the metadata of the data table, including the table name, field information, and index status;
[0021] The field description information is a description of each field, including field name, data type, field length, primary key or foreign key attributes;
[0022] The step of obtaining database description information, table description information and field description information of the data to be detected includes:
[0023] Determine a matching driver according to the database type, the driver being JDBC or ODBC;
[0024] Configure the database address, port, username and password;
[0025] Obtain the database description information through the system table;
[0026] Extract table_name, table_type, and table_comment fields;
[0027] Query the information_schema.COLUMNS system table to obtain the field description information.
[0028] Optionally, for each feature vector to be matched, calculating the Euclidean distance between the feature vectors of all predefined core data and important data by matrix operation includes:
[0029] Calculate using the following formula:
[0030]
[0031] Among them, v represents the feature vector to be matched, and the matrix M represents the shape of (m,n), where m is the number of predefined feature vectors and n is the vector dimension;
[0032] The Euclidean distance between v and each row in M is calculated by the above operation.
[0033] Optionally, determining the number of feature vectors matching the feature vector to be matched according to the calculation result includes:
[0034] Determine the distance threshold threshold;
[0035] Construct all the output Euclidean distances d into a distance list;
[0036] For each distance value d in the distance list, determine whether d≤threshold is satisfied;
[0037] If satisfied, it is determined that the corresponding feature vector matches the predefined feature vector;
[0038] Count the number of eigenvectors that meet the conditions.
[0039] Optionally, the weight value of the data is calculated according to the feature vector matching results in the description information, table description information and field description information of the database, wherein the weight of the description information is 0.2, the weight of the table description information is 0.5, and the weight of the field description information is 0.3, and is calculated by the following formula:
[0040] Weight value=(W_db*N_db+W_table*N_table+W_field*N_field)*S
[0041] Among them, W_db represents the weight of the database description information, N_db represents the matching number of the feature vector of the database description information, W_table represents the weight of the table description information, N_table represents the matching number of the feature vector of the table description information, W_field represents the weight of the field description information, N_field represents the matching number of the feature vector of the field description information, and S is the data scale factor, which represents the influencing factor of the data scale and is used to adjust the calculation results according to the scale of the data to be detected.
[0042] Optionally, the calculating of the probability value of the data to be detected belonging to the core data or important data according to the matching result and the data scale, wherein the probability value is obtained by weighted calculation based on the number of matches and the data scale, includes:
[0043] Obtain matching results, including the number of matches M_total and N_db, N_table, and N_field;
[0044] Get the number of records N_records and the number of fields N_fields of the data to be tested;
[0045] The S is calculated by the following formula:
[0046]
[0047] Among them, N_records-ref is the number of reference records, indicating the preset standard size, and N_fields-ref is the number of reference fields, indicating the preset standard size;
[0048] Obtaining the weight value;
[0049] The weight value is adjusted according to the M_total by the following formula:
[0050]
[0051] Among them, M_ref is the reference value of the number of matches, which represents the total number of matches of typical core data or important data;
[0052] W_adjusted represents the adjusted weight value;
[0053] The adjusted weight value is combined with the data scale factor through the following formula to calculate the final probability value:
[0054] P_final=W_adjusted*S
[0055] Where P represents the final probability value.
[0056] Optionally, the determining that the probability value exceeds a predetermined threshold, and if so, determining that the data to be detected is core data or important data includes:
[0057] Determine a preset first threshold and a second threshold, wherein the first threshold is greater than the second threshold;
[0058] If the probability value satisfies that the probability value is greater than the first threshold, determining that the data to be detected is core data;
[0059] If the probability value satisfies the condition that it is greater than the second threshold value and less than the first threshold value, it is determined that the data to be detected is important data;
[0060] If the probability value satisfies and is smaller than the second threshold, it is determined that the data to be detected is important data.
[0061] The second aspect of the present application provides a data recognition system based on feature vector matching, comprising:
[0062] An acquisition unit, used to acquire database description information, table description information and field description information of the data to be detected based on the trusted execution environment TEE of the data provider;
[0063] A word segmentation unit is used to perform word segmentation processing on the data to be detected based on the trusted execution environment TEE of the data supplier to obtain a feature vector of the data to be detected;
[0064] A loading unit, used for loading a feature vector library of predefined core data and important data into the TEE;
[0065] A distance calculation unit, used for calculating the Euclidean distance between each feature vector to be matched in the feature vector library and all predefined core data and important data feature vectors through matrix operation;
[0066] A quantity determination unit, used to determine the quantity of feature vectors matching the feature vector to be matched according to the calculation result;
[0067] A weight value calculation unit, used to calculate the weight value of the data according to the feature vector matching results in the description information, table description information and field description information of the database, wherein the weight of the description information is 0.2, the weight of the table description information is 0.5, and the weight of the field description information is 0.3;
[0068] A probability value calculation unit, used to 1 calculate the probability value of the data to be detected belonging to the core data or important data according to the matching result and the data scale, wherein the probability value is obtained by weighted calculation based on the number of matches and the data scale;
[0069] The judgment unit is used to judge whether the probability value exceeds a predetermined threshold value, and if so, determine that the data to be detected is core data or important data.
[0070] A third aspect of the present application provides a device for data recognition based on feature vector matching, the device comprising:
[0071] Processor, memory, input-output unit, and bus;
[0072] The processor is connected to the memory, the input and output unit, and the bus;
[0073] The memory stores a program, and the processor calls the program to execute the first aspect and any optional method in the first aspect.
[0074] A fourth aspect of the present application provides a computer-readable storage medium, on which a program is stored. When the program is executed on a computer, the program executes the first aspect and any optional method in the first aspect.
[0075] It can be seen from the above technical solutions that this application has the following advantages:
[0076] 1. This method avoids manual intervention and complex manual labeling through the automated feature vector matching process, thereby improving the efficiency of data processing. Traditional data classification and recognition methods usually rely on manual rules or statically defined rule bases, while this method can automatically classify and recognize based on data features, has strong adaptability, and can quickly process large-scale data.
[0077] 2. This method relies on feature vector matching, can automatically process and identify new data types, and adapt to the dynamic changes of data features. Compared with the traditional rule engine method, it can maintain high recognition accuracy and robustness in an environment where the amount of data and the type of data are constantly changing.
[0078] 3. Through feature vector matching, this method can be adjusted according to different data features. Whether it is table description information, field description information or data content itself, it can adapt to the needs of various data types by adjusting the feature vector.
[0079] 4. By calculating the Euclidean distance between the feature vector to be matched and the predefined core data and important data, the similarity between feature vectors can be measured more accurately, reducing the number of mismatches and omissions. This precise distance measurement method can effectively improve the accuracy of classification results.
[0080] 5. By performing weighted calculation on database description information, table description information and field description information, this method adds weight factors of multiple dimensions on the basis of feature vector matching, making data recognition more comprehensive and accurate.
[0081] 6. By comprehensively considering the weights of database description information, table description information, and field description information, the method can perform weighted evaluation of data from multiple dimensions to avoid excessive influence of a single factor on the data judgment results. This comprehensive calculation method can better adapt to diversified data scenarios and improve the ability to process complex data.
[0082] 7. The extraction, matching, and calculation of feature vectors are all performed in the TEE environment. The data provider's data to be tested and the predefined core data feature library will not be exposed to external systems or networks, thereby preventing data leakage and unauthorized access. The output only includes the final classification results, avoiding the exposure of detailed information such as feature vectors and calculation processes, further ensuring the privacy of data.
[0083] 8. TEE provides a hardware isolation mechanism to ensure the authenticity and integrity of the calculation and prevent malicious tampering of the algorithm or data. Using the remote proof function of TEE, the data recipient can verify that the calculation results are generated by a trusted hardware environment, enhancing the credibility of the matching results.
[0084] 9. Use the matrix parallel operations (such as SIMD and vectorized computing) supported by TEE to achieve batch processing of feature vectors and greatly improve matching efficiency. Complete local calculations on the supplier side to avoid delays and communication overhead caused by data transmission to external processing, and improve real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] In order to more clearly illustrate the technical solution in the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0086] Figure 1 A schematic diagram of an embodiment of a data recognition method based on feature vector matching provided in this application;
[0087] Figure 2 A schematic diagram of an embodiment of a flow chart for threshold value judgment in a data recognition method based on feature vector matching provided in the present application;
[0088] Figure 3 A schematic diagram of the structure of an embodiment of a data recognition system based on feature vector matching provided in this application;
[0089] Figure 4 This is a schematic diagram of the structure of an embodiment of a data recognition device based on feature vector matching provided in this application. DETAILED DESCRIPTION
[0090] It should be noted that the data recognition method based on feature vector matching provided in this application can be applied to a terminal or a system, or a server. For example, the terminal can be a smart phone or a computer, a tablet computer, a smart TV, a smart watch, a portable computer terminal, or a fixed terminal such as a desktop computer. For the convenience of explanation, this application uses the terminal as an example for explanation.
[0091] See also Figure 1 , Figure 1 A schematic flow chart of an embodiment of a data identification method based on feature vector matching provided in the present application, the data identification method based on feature vector matching includes:
[0092] S101, based on the trusted execution environment TEE of the data provider, obtain database description information, table description information and field description information of the data to be detected;
[0093] The purpose of this step is to obtain the context information of the data to be detected, including database description information, table description information, and field description information. This information helps to understand the structure and background of the data so as to perform feature vector matching on the data later.
[0094] In this embodiment, all data description information is acquired and processed within the TEE to prevent external systems or malicious programs from stealing sensitive data. The data to be detected is completely isolated from the predefined core data, and the acquired data description information is only used for feature vector extraction. Before data acquisition, TEE proves its credibility to the data provider and receiver through remote attestation to ensure that data processing is only performed in a trusted environment. When acquiring, first start TEE, load the trusted data acquisition module and related keys. Prove the credibility of the environment to the supplier through remote attestation. The supplier enters the database connection information (such as host, port, user name, password) in TEE. And use the TLS protocol to establish an encrypted connection. Perform metadata queries to obtain description information of databases, tables, and fields. Generate field sample statistics to assist in generating field description information. Unify the format of description information and store it in TEE memory.
[0095] When storing, the description information is stored in a structured form in an isolated area within the TEE, preparing for subsequent feature vector extraction.
[0096] In this embodiment, the entire data acquisition and processing process is completed in TEE, preventing external systems from accessing sensitive information and protecting the confidentiality of the supplier's data.
[0097] In this step, you can access the database metadata (such as the metadata view or information architecture provided by the database management system) to extract relevant descriptive information. Specifically:
[0098] Database description information refers to the basic information of the database itself, including the type, creation time, size, storage method, etc. of the database. Table description information refers to the structural information of the data table, including the table name, field, data type, constraint, index, relationship, etc. Field description information involves detailed information of a single field, describing the field name, type, default value, length, whether it is a primary key, foreign key relationship, etc. These metadata are extracted by accessing the database system table or using a database query language (such as SQL).
[0099] In an optional embodiment, the database description information is metadata of a data table, including table name, field information, and index status;
[0100] The field description information is a description of each field, including field name, data type, field length, primary key or foreign key attributes;
[0101] The step of obtaining database description information, table description information and field description information of the data to be detected includes:
[0102] Determine a matching driver according to the database type, the driver being JDBC or ODBC;
[0103] Configure the database address, port, username and password;
[0104] Obtain the database description information through the system table;
[0105] Extract table_name, table_type, and table_comment fields;
[0106] Query the information_schema.COLUMNS system table to obtain the field description information.
[0107] In this optional embodiment, the correct driver can be selected to connect to different types of databases. JDBC (JavaDatabaseConnectivity) is a standard database connection interface for Java applications, while ODBC (OpenDatabaseConnectivity) is a standard that provides a common access interface for different databases. For example, the MySQL database requires the com.mysql.cj.jdbc.Driver driver, and the Oracle database requires the oracle.jdbc.OracleDriver driver. For the ODBC driver, the database is connected by configuring a data source (DSN). ODBC is used for cross-platform connections between different systems.
[0108] The purpose of configuring the database address, port, user name, and password is to set the necessary connection information required to connect to the database and ensure that the connection can be established successfully.
[0109] Configure the database address (IP or domain name), port, user name, and password in the connection string. For example, for JDBC, the connection string can be as follows:
[0110] Stringurl="jdbc:mysql: / / localhost:3306 / mydatabase";Stringusername="root";Stringpassword="password";
[0111] For ODBC, you can configure the database connection parameters by setting the DSN (DataSourceName).
[0112] The purpose of obtaining database description information through system tables is to extract metadata information from the system tables of the database and obtain structural information about the database, including descriptions of tables and fields.
[0113] Different database systems usually provide information schemas or system tables for developers to query the database structure. Most relational databases (such as MySQL and PostgreSQL) follow the ANSI SQL standard and support the use of information_schema to obtain metadata information. In MySQL, information_schema is a virtual database that contains metadata information such as all tables, columns, indexes, etc. in the database. You can query the system table to obtain database description information.
[0114] The purpose of extracting the table_name, table_type, and table_comment fields is to extract table-related information from the system tables of the database, such as table name, table type (normal table, view, etc.) and table comment (optional descriptive text).
[0115] You can extract this information by querying the information_schema.tables table. Here is a code example:
[0116] SELECT table_name, table_type, table_commentFROM information_schema.tablesWHERE table_schema = 'your_database_name';
[0117] table_name: The name of the data table.
[0118] table_type: The type of table, which may be BASE TABLE (regular table) or VIEW (view).
[0119] table_comment: The comment for the table, if any.
[0120] Query the information_schema.COLUMNS system table to obtain field description information. The purpose is to obtain the description information of each field, including the field name, data type, field length, primary key or foreign key attributes, etc.
[0121] You can query the information_schema.columns system table to obtain field description information. The field information table contains a detailed description of each column in the table.
[0122] Here is a code example:
[0123] SELECT column_name, data_type, character_maximum_length, is_nullable,column_default,
[0124] column_key, extraFROM information_schema.columnsWHERE table_schema ='your_database_name'AND table_name = 'your_table_name';
[0125] The following is an explanation of the various fields in the above code example:
[0126] column_name: field name.
[0127] data_type: The data type of the field (such as VARCHAR, INT, DATE, etc.).
[0128] character_maximum_length: The maximum length of the field (applicable to character type fields).
[0129] is_nullable: whether NULL is allowed.
[0130] column_default: The default value of the field, if any.
[0131] column_key: Identifies whether the field is a primary key (PRI), foreign key (MUL), etc.
[0132] extra: Additional information about the field, such as whether it is automatically incremented (auto_increment).
[0133] Step S101 establishes a secure, trustworthy, and reliable foundation for core data identification by acquiring multi-layer description information of databases, tables, and fields in TEE. This step not only ensures the privacy of data acquisition, but also enhances the credibility and applicability of the entire identification method.
[0134] S102, based on the trusted execution environment TEE of the data provider, performing word segmentation processing on the data to be detected to obtain a feature vector of the data to be detected;
[0135] This step aims to process the data to be detected and convert it into feature vector form for subsequent feature vector matching and distance calculation.
[0136] In this embodiment, data segmentation processing and feature vector extraction are completed within TEE, ensuring that the original data and intermediate results are not exposed to the external environment. TEE ensures that the execution environment of the segmentation algorithm and feature vector extraction cannot be tampered with, and the output results are credible. The segmentation processing results can provide credible proof to the data supplier and receiver through TEE's remote proof to ensure that the processing process meets expectations.
[0137] For the text data to be detected (such as database field values or document content), a word segmentation algorithm is used to split it into individual words or keywords. For example, for Chinese text, a Chinese word segmentation algorithm such as Jieba word segmentation can be used; for English text, a common word segmentation method (such as space-based splitting) is used.
[0138] Generate feature vectors by vectorizing the segmented data. Vectorization methods include:
[0139] TF-IDF (TermFrequency-InverseDocumentFrequency): used to measure the importance of each word in a document.
[0140] Word2Vec or GloVe: These methods can map words into a vector space of fixed dimensionality, capturing the semantic relationship between words.
[0141] BERT deep learning model: Generate more semantic word vectors through contextual information.
[0142] In step S102, the data to be detected is segmented and a feature vector is generated. This process can be achieved through text processing and vectorization technology. The following provides some specific implementation methods and processes:
[0143] Data preprocessing (word segmentation):
[0144] Word segmentation is the process of breaking down text data into individual words or phrases. For Chinese and English, the strategies and tools for word segmentation are different. For Chinese, word frequency-based word segmentation algorithms (such as Jieba) are usually used, while English can be split by space, or more advanced tools can be used for more complex text parsing.
[0145] Chinese word segmentation (using jieba)
[0146] Jieba is a Chinese word segmentation library that can efficiently segment Chinese text. Jieba uses a maximum probability algorithm based on word frequency to identify the most likely word segmentation points.
[0147] Here is a code example using jieba for word segmentation:
[0148] Import jieba
[0149] #Example Chinese text
[0150] text="Data recognition method based on feature vector matching"
[0151] #Use jieba for word segmentation
[0152] words = jieba.cut(text)
[0153] # Output word segmentation results
[0154] print("".join(words))
[0155] Output example:
[0156] Data recognition method based on feature vector matching
[0157] English word segmentation (based on spaces):
[0158] For English text, the word segmentation method can be to split the sentence into words based on spaces. For example:
[0159] Sample English text:
[0160] vbnet
[0161] This is a sample text for vectorization.
[0162] text = "This is a sample text for vectorization."
[0163] words = text.split()# Split based on spaces print(words)
[0164] Output example:
[0165] ['This', 'is', 'a', 'sample', 'text', 'for', 'vectorization.']
[0166] Of course, in an optional way, English can also use more complex word segmentation tools such as NLTK or SpaCy, which can handle more complex language structures, such as stemming, stop word removal, etc.
[0167] Vectorized method:
[0168] After word segmentation, the text data needs to be converted into feature vectors. Vectorization can use a variety of methods to represent the semantic information of text data. The following are several common vectorization methods.
[0169] TF-IDF (Term Frequency-Inverse Document Frequency)
[0170] TF-IDF is a common statistical method to measure the importance of words in a text. It combines the frequency of a word in a single document (TF) and the inverse document frequency (IDF) in the entire document collection.
[0171] Here is a vectorized word segmentation code example:
[0172] from sklearn.feature_extraction.text import TfidfVectorizer
[0173] # Assuming there are multiple documents
[0174] documents = [
[0175] "Data recognition method based on feature vector matching",
[0176] "Data recognition methods include feature vector matching",
[0177] "Algorithm for matching feature vectors to data" ]
[0179] # Initialize TF-IDF vectorizer
[0180] vectorizer = TfidfVectorizer()
[0181] # Vectorize the document
[0182] X = vectorizer.fit_transform(documents)
[0183] # Output TF-IDF matrix
[0184] print(X.toarray())
[0185] Output example: [[0.0.0.58808018 0.0.] [0.0.58808018 0.0.0.] [0.58808018 0.0.0.0.]]
[0189] In this step, word segmentation and feature vectorization of the text data to be detected are important steps in text analysis. Through appropriate word segmentation algorithms and vectorization methods, the original text can be converted into feature vectors that are easy for computer processing, and then subsequent feature matching and similarity calculations can be performed.
[0190] The choice of specific methods (such as TF-IDF, Word2Vec, BERT, etc.) depends on the type of data and the application scenario. For example, BERT can better handle contextual semantics and is suitable for complex natural language processing tasks.
[0191] S103, loading a feature vector library of predefined core data and important data into the TEE;
[0192] In this step, the feature vector library is pre-defined by the system and contains feature vectors of core data and important data. These feature vectors are generated based on domain expertise or historical data analysis. Feature vectors can be stored in a structured format (such as matrix form) to facilitate subsequent efficient operations. During the storage phase, the feature vector library can be encrypted using advanced encryption algorithms (such as AES-256) to ensure the confidentiality of the feature library. Start a trusted execution environment (such as Intel SGX or ARM TrustZone) and load the code modules and trusted key management modules required for execution. Verify the authenticity of the TEE through the remote attestation function to ensure that the feature vector library is only loaded into a trusted environment.
[0193] The predefined feature vector library is transferred from external storage to TEE. During the process, the data is encrypted using a transport layer security protocol (such as TLS) to prevent the data from being stolen or tampered with during transmission. The data provider verifies the public key of the TEE to ensure that the loading process can only be executed by a legitimate TEE instance.
[0194] After the feature vector library is transmitted to TEE, the trusted key management module inside TEE uses the private key to decrypt it and restore it to operable plaintext data.
[0195] The decrypted feature vector library is mapped to a dedicated memory area inside the TEE. This memory area is isolated from the outside to prevent external programs from accessing it.
[0196] You can use a hash algorithm (such as SHA-256) to verify the feature vector library loaded into the TEE and compare it with the pre-stored hash value to ensure that the transmission and decryption process has not been tampered with. Verify whether the loaded feature vector library version is consistent with the latest version to prevent loading old versions or invalid data.
[0197] The feature vector library loaded into TEE is stored in TEE's exclusive memory, which is completely invisible to external operating systems and applications. Access to the feature vector library is controlled by trusted code inside TEE, and any external request must be verified through a secure API before access.
[0198] If the core data or important data feature vector library needs to be updated, the existing feature library can be safely replaced at runtime through a similar encrypted transmission and loading mechanism. The system maintains the version number and update log of the feature vector library to ensure that the data is up to date at all times.
[0199] After the feature vector library is successfully loaded, the system enters the calculation phase, and the core data matching algorithm can directly call the feature vector library inside the TEE to perform batch matrix operations. The matching results and intermediate data are processed and retained inside the TEE and are not exposed to the outside.
[0200] In this embodiment, the loading and storage of data are completely isolated from the outside to avoid the risk of feature library leakage. This step provides a reliable foundation for the data matching algorithm in the TEE environment, ensuring that the algorithm runs in a safe and trusted environment to avoid the risk of data leakage or tampering.
[0201] S103, for each feature vector to be matched in the feature vector library, calculating the Euclidean distance between each feature vector and all predefined core data and important data feature vectors through matrix operation;
[0202] In this step, the similarity between the feature vector to be matched and the feature vectors of predefined core data and important data is calculated to determine the similarity between the data to be detected and the existing data.
[0203] First, the feature vectors of the data to be detected and the feature vectors of all predefined core data and important data are stored in a matrix. Each row represents the feature vector of a data item.
[0204] For each feature vector to be matched, the Euclidean distance between it and all predefined core data and important data feature vectors is calculated.
[0205] Specifically, it can be calculated by the following formula:
[0206]
[0207] Among them, v represents the feature vector to be matched, and the matrix M represents the shape of (m,n), where m is the number of predefined feature vectors and n is the vector dimension;
[0208] The Euclidean distance between v and each row in M is calculated by the above operation.
[0209] The following is a further description of this step:
[0210] First, the feature vector v to be matched is represented as a 1×n vector.
[0211] The matrix M is taken as a matrix consisting of m predefined eigenvectors, where each row is an eigenvector and the shape of the matrix is m×n.
[0212] In order to calculate the Euclidean distance, matrix operations can be used to achieve efficient calculation. Assume that the dimensions of the feature vector v to be matched and the matrix M meet the calculation requirements, that is, v is a 1×n vector and the matrix M is an m×n matrix.
[0213] First, calculate the difference between v and M. You can use matrix broadcasting to expand v into a matrix with the same shape as M. Then calculate the sum of the squares of the differences and finally take the square root to get the Euclidean distance.
[0214] Assume the following data:
[0215] v=[v1,v2,...,vn] is the feature vector to be matched (shape is 1×n).
[0216] is an m×n matrix, where each row is a predefined eigenvector.
[0217] First, the difference between the feature vector to be matched and each row of the feature vector in the matrix is calculated. In order to make the shape of the matrix compatible, the feature vector v to be matched can be expanded into a matrix of shape m×n, which is the same as the matrix M. Then they are subtracted element by element.
[0218] Here is a code example:
[0219] importnumpyasnp
[0220] #Assume that the feature vector v and matrix M to be matched
[0221] v=np.array([v1,v2,v3,...,vn]) #1xn vector
[0222] M=np.array([[m11,m12,...,m1n], #mxn matrix
[0223] [m21,m22,...,m2n], ...
[0224] [mm1,mm2,...,mmn]])
[0225] #Calculate the difference between the feature vector v to be matched and each row of the matrix M
[0226] diff=Mv#mxn
[0227] #Calculate Euclidean distance (by row)
[0228] euclidean_distances=np.sqrt(np.sum(diff**2,axis=1))
[0229] The Euclidean distance is the square root of the sum of the squares of the feature vector differences. Use diff**2 to get the square difference of each dimension, and then use np.sum(diff**2,axis=1) to find the sum of the square differences of all dimensions for each row (that is, between each predefined feature vector and the feature vector to be matched). Finally, use np.sqrt() to calculate the Euclidean distance of each row.
[0230] In an alternative approach, if the number of feature vectors to be calculated is large, the Euclidean distance calculation may become very time-consuming. The following optimization method can be used:
[0231] Use NumPy vectorized calculations to improve computational efficiency.
[0232] Using the cdist() function in the SciPy library to directly calculate the distance between two sets of vectors can speed up the calculation of large-scale data.
[0233] In this step, the Euclidean distance between the feature vector to be matched and all predefined core data and important data feature vectors can be effectively calculated. The Euclidean distance can measure the similarity of data and is used for subsequent feature vector matching and data classification.
[0234] S104, determining the number of feature vectors that match the feature vector to be matched according to the calculation result;
[0235] The calculated Euclidean distance is used to determine the number of feature vectors similar to the feature vector to be matched for subsequent processing.
[0236] According to the calculated Euclidean distance value, select several predefined feature vectors with the smallest distance. You can set a threshold to select those predefined feature vectors whose Euclidean distance to the feature vector of the data to be detected is less than the threshold, or select several feature vectors with the closest distance (for example, the first 5 or the first 10).
[0237] These matching feature vectors represent the data types or categories that are most relevant to the data to be detected.
[0238] See also Figure 2 , specifically, it can be:
[0239] S1041, determining a distance threshold value;
[0240] S1042, constructing all the output Euclidean distances d into a distance list;
[0241] S1043. For each distance value d in the distance list, determine whether d≤threshold is satisfied. If so, execute step S1044; if not, execute step S1046.
[0242] S1044, determining that the corresponding feature vector matches a predefined feature vector;
[0243] S1045. Count the number of eigenvectors that meet the conditions.
[0244] S1046. Compare the next distance value d.
[0245] S105, calculating the weight value of the data according to the feature vector matching results in the description information, table description information and field description information of the database, wherein the weight of the description information is 0.2, the weight of the table description information is 0.5, and the weight of the field description information is 0.3;
[0246] S106. Calculate the probability value of the data to be detected belonging to the core data or important data according to the matching result and the data size, wherein the probability value is obtained by weighted calculation based on the number of matches and the data size;
[0247] In this step, the weight value of each data to be detected is calculated based on the weighting of the feature vector matching result and the data description information, thereby reflecting the importance of the data.
[0248] According to the different importance of database description information, table description information and field description information, different weight values are assigned to them respectively. In this embodiment:
[0249] The weight of database description information is 0.2.
[0250] The weight of table description information is 0.5.
[0251] The weight of field description information is 0.3.
[0252] According to the matching results, the matching degree of each data is multiplied by its corresponding weight to obtain the weight value of each data item. Taking into account the matching degree of all relevant information, the following formula can be used:
[0253] Weight value=(W_db*N_db+W_table*N_table+W_field*N_field)*S;
[0254] Among them, W_db represents the weight of the database description information, N_db represents the matching number of the feature vector of the database description information, W_table represents the weight of the table description information, N_table represents the matching number of the feature vector of the table description information, W_field represents the weight of the field description information, N_field represents the matching number of the feature vector of the field description information, and S is the data scale factor, which represents the influencing factor of the data scale and is used to adjust the calculation results according to the scale of the data to be detected.
[0255] The data scale factor S is a factor that reflects the impact of data scale on the result. It can be calculated as follows:
[0256]
[0257] Among them, N_records-ref is the number of reference records, indicating the preset standard scale, and N_fields-ref is the number of reference fields, indicating the preset standard scale.
[0258] Furthermore, the present application provides a more specific method for calculating the probability value, including:
[0259] Obtain matching results, including the number of matches M_total and N_db, N_table, and N_field;
[0260] Get the number of records N_records and the number of fields N_fields of the data to be tested;
[0261] The S is calculated by the following formula:
[0262]
[0263] Among them, N_records-ref is the number of reference records, indicating the preset standard size, and N_fields-ref is the number of reference fields, indicating the preset standard size;
[0264] Obtaining the weight value;
[0265] The weight value is adjusted according to the M_total by the following formula:
[0266]
[0267] Among them, M_ref is the reference value of the number of matches, which represents the total number of matches of typical core data or important data;
[0268] W_adjusted represents the adjusted weight value;
[0269] The adjusted weight value is combined with the data scale factor through the following formula to calculate the final probability value:
[0270] P_final=W_adjusted*S
[0271] Where P represents the final probability value.
[0272] Here is a specific example:
[0273] Assume the following information:
[0274] N_db=3
[0275] N_table=4
[0276] N_field=2
[0277] S=1.5
[0278] The weights are: W_db=0.2, W_table=0.5, W_field=0.3;
[0279] The following is a code example of a specific calculation:
[0280] #Weight and number of matches
[0281] W_db=0.2
[0282] W_table=0.5
[0283] W_field=0.3
[0284] N_db=3#Number of database description information matches
[0285] N_table=4#Number of table description information matches
[0286] N_field=2#Number of field description information matches
[0287] S=1.5#Data scale factor
[0288] #Calculate the total weight value
[0289] weight_value=(W_db*N_db+W_table*N_table+W_field*N_field)*S
[0290] print("The calculated weight value is:",weight_value)
[0291] The calculated weight value is: 4.8
[0292] Through this implementation, the total weight value of the data to be detected can be calculated based on the matching results of the database description information, table description information and field description information, combined with their respective weights. The data scale factor SSS can be used to adjust the weight calculation results to adapt to data of different scales. This method will help evaluate the importance of the data to be detected, so as to further process the classification, screening or priority sorting of the data.
[0293] S107: Determine whether the probability value exceeds a predetermined threshold value. If so, determine whether the data to be detected is core data or important data.
[0294] Based on the calculated probability value, determine whether the data to be detected is core data or important data.
[0295] A threshold can be set, and if the calculated probability value exceeds the threshold, the data is judged to be core data or important data. Otherwise, the data is considered not to be core data or important data.
[0296] The threshold can be adjusted according to the actual application scenario to control the strictness of data classification. For example, the threshold can be adjusted according to actual needs to optimize the system's ability to identify core data or important data.
[0297] In an optional embodiment, the judgment can be made in the following manner:
[0298] Determine a preset first threshold and a second threshold, wherein the first threshold is greater than the second threshold;
[0299] If the probability value satisfies that the probability value is greater than the first threshold, determining that the data to be detected is core data;
[0300] If the probability value satisfies being greater than the second threshold and less than the first threshold, determine that the data to be detected is important data;
[0301] If the probability value satisfies being less than the second threshold, determine that the data to be detected is important data.
[0302] In this alternative embodiment, first determine the thresholds:
[0303] First threshold (Threshold1): Used to determine whether the data to be detected is core data. This is a relatively high value, indicating a very high degree of data matching.
[0304] Second threshold (Threshold2): Used to determine whether the data to be detected is important data. This is a relatively low value, indicating a relatively low degree of data matching, but still having a certain value.
[0305] If the probability value is greater than the first threshold (P > Threshold1P), then determine the data to be detected as "core data".
[0306] If the probability value is greater than the second threshold and less than or equal to the first threshold (Threshold2 < P ≤ Threshold1), then determine the data to be detected as "important data".
[0307] If the probability value is less than the second threshold (P ≤ Threshold2), then determine the data to be detected as "important data".
[0308] The following is an example for illustration:
[0309] Suppose the thresholds can be determined according to business requirements or data analysis. For example, the thresholds can be set by methods such as statistical analysis and historical data.
[0310] Set the first threshold (Threshold1) to 0.8.
[0311] Set the second threshold (Threshold2) to 0.5.
[0312] Suppose the probability value P of the data to be detected has been calculated, obtained through weighted calculation in the previous steps.
[0313] Through conditional judgment statements, the data to be detected can be classified.
[0314] The following provides a code example:
[0315] # Assume the probability value we have calculated
[0316] P = 0.75 # Probability value of the data to be detected
[0317] #Threshold setting
[0318] Threshold1=0.8#The first threshold is used to determine the core data
[0319] Threshold2=0.5#The second threshold is used to determine important data
[0320] #Judge the type of data to be detected
[0321] ifP>Threshold1:
[0322] data_type="core data" #If the probability value is greater than the first threshold, the data is core data
[0323] elifThreshold2 <P<=Threshold1:
[0324] data_type="Important data" #If the probability value is between the second threshold and the first threshold, the data is important data
[0325] else:
[0326] data_type="normal data" #If the probability value is less than or equal to the second threshold, the data is normal data
[0327] # Output results
[0328] print(f"The type of data to be detected is: {data_type}")
[0329] ifP>Threshold1:
[0330] If the probability value PPP is greater than the first threshold, it means that the data to be detected has a very high matching degree and is therefore regarded as “core data”.
[0331] elifThreshold2 <P<=Threshold1::
[0332] If the probability value PPP is between the second threshold and the first threshold, it means that the matching degree of the data to be detected is high, but it does not meet the standard of core data, and is therefore regarded as "important data".
[0333] else:
[0334] If the probability value is less than or equal to the second threshold, it means that the matching degree of the data to be detected is low, and therefore it is classified as "ordinary data" or unimportant data.
[0335] Assume P=0.75 and the threshold is set as follows:
[0336] Threshold1=0.8
[0337] Threshold2=0.5
[0338] In this example, P falls between the second threshold and the first threshold (i.e., 0.5<0.75≤0.8), so it will be classified as "significant data".
[0339] Through the method provided in this embodiment, it is possible to determine whether the data to be detected belongs to "core data" or "important data" according to the set threshold. This classification method can be flexibly adjusted according to the matching degree of the data and different threshold settings, thereby helping to accurately identify and classify the data. This implementation method has good adaptability, and the threshold can be adjusted according to actual needs to adapt to different application scenarios.
[0340] The above describes in detail the embodiments of the method provided in the present application. The following describes the device, system and storage medium provided in the present application:
[0341] See also Figure 2 The present application provides an embodiment of a data recognition system based on feature vector matching, the embodiment comprising:
[0342] The acquisition unit 301 is used to acquire database description information, table description information and field description information of the data to be detected based on the trusted execution environment TEE of the data provider;
[0343] The word segmentation unit 302 is used to perform word segmentation processing on the data to be detected based on the trusted execution environment TEE of the data supplier to obtain a feature vector of the data to be detected;
[0344] A loading unit 303, used for loading a feature vector library of predefined core data and important data into the TEE;
[0345] A distance calculation unit 304 is used to calculate the Euclidean distance between each feature vector to be matched and all predefined core data and important data feature vectors through matrix operations;
[0346] A quantity determination unit 305 is used to determine the quantity of feature vectors matching the feature vector to be matched according to the calculation result;
[0347] A weight value calculation unit 306, used to calculate the weight value of the data according to the feature vector matching results in the description information, table description information and field description information of the database, wherein the weight of the description information is 0.2, the weight of the table description information is 0.5, and the weight of the field description information is 0.3;
[0348] The probability value calculation unit 307 is used to calculate the probability value of the data to be detected belonging to the core data or the important data according to the matching result and the data scale, and the probability value is obtained by weighted calculation based on the number of matches and the data scale;
[0349] The judging unit 308 is configured to judge whether the probability value exceeds a predetermined threshold value, and if so, determine that the data to be detected is core data or important data.
[0350] Optionally, the table description information is metadata at the database level, including the name, type, and structure of the database;
[0351] The database description information is the metadata of the data table, including the table name, field information, and index status;
[0352] The field description information is a description of each field, including field name, data type, field length, primary key or foreign key attributes;
[0353] The acquisition unit 301 is specifically used for:
[0354] Determine a matching driver according to the database type, the driver being JDBC or ODBC;
[0355] Configure the database address, port, username and password;
[0356] Obtain the database description information through the system table;
[0357] Extract table_name, table_type, and table_comment fields;
[0358] Query the information_schema.COLUMNS system table to obtain the field description information.
[0359] Optionally, the distance calculation unit 304 is specifically used for:
[0360] Calculate using the following formula:
[0361]
[0362] Among them, v represents the feature vector to be matched, and the matrix M represents the shape of (m,n), where m is the number of predefined feature vectors and n is the vector dimension;
[0363] The Euclidean distance between v and each row in M is calculated by the above operation.
[0364] Optionally, the distance calculation unit 304 is specifically used for:
[0365] Determine the distance threshold threshold;
[0366] Construct all the output Euclidean distances d into a distance list;
[0367] For each distance value d in the distance list, determine whether d≤threshold is satisfied;
[0368] If satisfied, it is determined that the corresponding feature vector matches the predefined feature vector;
[0369] Count the number of eigenvectors that meet the conditions.
[0370] Optionally, the probability value calculation unit 307 is specifically used to:
[0371] Weight value=(W_db*N_db+W_table*N_table+W_field*N_field)*S
[0372] Among them, W_db represents the weight of the database description information, N_db represents the matching number of the feature vector of the database description information, W_table represents the weight of the table description information, N_table represents the matching number of the feature vector of the table description information, W_field represents the weight of the field description information, N_field represents the matching number of the feature vector of the field description information, and S is the data scale factor, which represents the influencing factor of the data scale and is used to adjust the calculation results according to the scale of the data to be detected.
[0373] Optionally, the probability value calculation unit 307 is specifically used to:
[0374] Obtain matching results, including the number of matches M_total and N_db, N_table, and N_field;
[0375] Get the number of records N_records and the number of fields N_fields of the data to be tested;
[0376] The S is calculated by the following formula:
[0377]
[0378] Among them, N_records-ref is the number of reference records, indicating the preset standard size, and N_fields-ref is the number of reference fields, indicating the preset standard size;
[0379] Obtaining the weight value;
[0380] The weight value is adjusted according to the M_total by the following formula:
[0381]
[0382] Among them, M_ref is the reference value of the number of matches, which represents the total number of matches of typical core data or important data;
[0383] W_adjusted represents the adjusted weight value;
[0384] The adjusted weight value is combined with the data scale factor through the following formula to calculate the final probability value:
[0385] P_final=W_adjusted*S
[0386] Where P represents the final probability value.
[0387] The determination unit 308 is specifically used for:
[0388] Determine a preset first threshold and a second threshold, wherein the first threshold is greater than the second threshold;
[0389] If the probability value satisfies that the probability value is greater than the first threshold, determining that the data to be detected is core data;
[0390] If the probability value satisfies the condition that it is greater than the second threshold value and less than the first threshold value, it is determined that the data to be detected is important data;
[0391] If the probability value satisfies and is smaller than the second threshold, it is determined that the data to be detected is important data.
[0392] See also Figure 4 , the present application also provides a data recognition device based on feature vector matching, comprising:
[0393] Processor 401, memory 402, input and output unit 403, bus 404;
[0394] The processor 401 is connected to the memory 402, the input and output unit 403 and the bus 404;
[0395] The memory 402 stores a program, and the processor 401 calls the program to execute any of the above data recognition methods based on feature vector matching.
[0396] The present application also relates to a computer-readable storage medium, on which a program is stored. When the program is run on a computer, the computer executes any of the above methods.
[0397] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0398] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0399] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0400] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0401] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, read-only memory), random access memory (RAM, random access memory), disk or optical disk and other media that can store program code.
Claims
1. A data recognition method based on feature vector matching, characterized in that: The method comprises: Based on the trusted execution environment TEE of the data provider, obtain database description information, table description information and field description information of the data to be detected, where the data to be detected is text data; Based on the trusted execution environment TEE of the data provider, the data to be detected is segmented to obtain a feature vector of the data to be detected; Loading a feature vector library of predefined core data and important data into the TEE; For each feature vector to be matched in the feature vector library, the Euclidean distance between the feature vectors and all predefined core data and important data feature vectors is calculated through matrix operations; Determine the number of feature vectors that match the feature vector to be matched according to the calculation result; Calculate the weight value of the data according to the feature vector matching results in the description information, table description information and field description information of the database, wherein the weight of the description information is 0.2, the weight of the table description information is 0.5, and the weight of the field description information is 0.3; According to the matching results and the data size, the probability value of the data to be detected belonging to the core data or important data is calculated, and the probability value is obtained by weighted calculation based on the number of matches and the data size; Determining whether the probability value exceeds a predetermined threshold, and if so, determining that the data to be detected is core data or important data; The weight value is calculated by the following formula: Weight value=(W_db*N_db+W_table*N_table+W_field*N_field)*S; Wherein, W_db represents the weight of database description information, N_db represents the matching number of feature vectors of database description information, W_table represents the weight of table description information, N_table represents the matching number of feature vectors of table description information, W_field represents the weight of field description information, N_field represents the matching number of feature vectors of field description information, and S represents the data scale factor, which represents the influencing factor of data scale and is used to adjust the calculation result according to the scale of the data to be detected; The calculation of the probability value of the data to be detected belonging to the core data or important data based on the matching result and the data scale, wherein the probability value is obtained by weighted calculation based on the number of matches and the data scale, includes: Obtain matching results, including the number of matches M_total and N_db, N_table, and N_field; Get the number of records N_records and the number of fields N_fields of the data to be tested; The S is calculated by the following formula: ; Among them, N_records-ref is the number of reference records, indicating the preset standard size, and N_fields-ref is the number of reference fields, indicating the preset standard size; Obtaining the weight value; The weight value is adjusted according to the M_total by the following formula: ; Among them, M_ref is the reference value of the number of matches, which represents the total number of matches of typical core data or important data; W_adjusted represents the adjusted weight value; The adjusted weight value is combined with the data scale factor through the following formula to calculate the final probability value: P_final=W_adjusted*S; Where P represents the final probability value.
2. The data recognition method based on feature vector matching according to claim 1, characterized in that: The table description information is metadata at the database level, including the name, type, and structure of the database; The database description information is the metadata of the data table, including the table name, field information, and index status; The field description information is a description of each field, including field name, data type, field length, primary key or foreign key attributes; The step of obtaining database description information, table description information and field description information of the data to be detected includes: Determine a matching driver according to the database type, the driver being JDBC or ODBC; Configure the database address, port, user name and password; Obtain the database description information through the system table; Extract table_name, table_type, and table_comment fields; Query the information_schema.COLUMNS system table to obtain the field description information.
3. The data recognition method based on feature vector matching according to claim 1, characterized in that: The calculation of the Euclidean distance between each feature vector to be matched and all predefined core data and important data feature vectors through matrix operations includes: Calculate using the following formula: ; Among them, v represents the feature vector to be matched, and the matrix M represents the shape of (m,n), where m is the number of predefined feature vectors and n is the vector dimension; The Euclidean distance between v and each row in M is calculated by the above operation.
4. The data recognition method based on feature vector matching according to claim 3 is characterized in that: Determining the number of feature vectors matching the feature vector to be matched according to the calculation result includes: Determine the distance threshold threshold; Construct all the output Euclidean distances d into a distance list; For each distance value d in the distance list, determine whether d≤threshold is satisfied; If satisfied, it is determined that the corresponding feature vector matches the predefined feature vector; Count the number of eigenvectors that meet the conditions.
5. The data recognition method based on feature vector matching according to claim 1, characterized in that: The determining that the probability value exceeds a predetermined threshold, and if so, determining that the data to be detected is core data or important data includes: Determine a preset first threshold and a second threshold, wherein the first threshold is greater than the second threshold; If the probability value satisfies that the probability value is greater than the first threshold, determining that the data to be detected is core data; If the probability value satisfies the condition that it is greater than the second threshold value and less than the first threshold value, it is determined that the data to be detected is important data; If the probability value satisfies and is smaller than the second threshold, it is determined that the data to be detected is important data.
6. A data recognition system based on feature vector matching, characterized in that: For executing the method according to any one of claims 1 to 5, the system comprises: An acquisition unit, used to acquire database description information, table description information and field description information of the data to be detected based on the trusted execution environment TEE of the data provider; A word segmentation unit is used to perform word segmentation processing on the data to be detected based on the trusted execution environment TEE of the data supplier to obtain a feature vector of the data to be detected; A loading unit, used for loading a feature vector library of predefined core data and important data into the TEE; A distance calculation unit, used for calculating the Euclidean distance between each feature vector to be matched in the feature vector library and all predefined core data and important data feature vectors through matrix operation; A quantity determination unit, used to determine the quantity of feature vectors matching the feature vector to be matched according to the calculation result; A weight value calculation unit, used to calculate the weight value of the data according to the feature vector matching results in the description information, table description information and field description information of the database, wherein the weight of the description information is 0.2, the weight of the table description information is 0.5, and the weight of the field description information is 0.3; A probability value calculation unit, used to 1 calculate the probability value of the data to be detected belonging to the core data or important data according to the matching result and the data scale, wherein the probability value is obtained by weighted calculation based on the number of matches and the data scale; The judging unit is used to judge whether the probability value exceeds a predetermined threshold value, and if so, determine that the data to be detected is core data or important data.
7. A device for data recognition based on feature vector matching, characterized in that: The device comprises: Processor, memory, input-output unit, and bus; The processor is connected to the memory, the input and output unit, and the bus; The memory stores a program, and the processor calls the program to execute the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a program stored thereon, wherein the program, when executed on a computer, performs the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Rubbish text recognition method and system
CN101477544A
Face image recognition method and device, computing equipment and medium
CN112784823A
Image retrieval method based on multi-scale feature and space attention mechanism fusion
CN119046493A