Data identification method and device, electronic equipment, storage medium and program product
By obtaining the data to be identified and its sensitive types and subject vocabulary of the industry to which it belongs, and using a large language model combined with the subject vocabulary recognition correlation, the problem of low accuracy of data sensitive types in the prior art is solved, and higher recognition accuracy and adaptability are achieved.
Patent Information
- Application Number
- CN202510104134.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art has low accuracy in identifying sensitive types of data, making it difficult to adapt to the diversity and complexity of data types, especially in the face of new data patterns.
By obtaining the sensitive types and subject vocabulary of the data to be identified and its industry to which it belongs, using a preset large language model combined with the subject vocabulary of sensitive types, the correlation between the data to be identified and each sensitive type is determined, thereby identifying the target sensitive type.
It improves the accuracy of data sensitive type identification, and can more accurately identify the sensitive type of data to be identified, adapting to changes in different industries and data types.
Smart Images

Figure CN120011568A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of data security technology, and in particular, to data identification methods, devices, electronic devices, storage media, and program products. Background Art
[0002] In the field of data security, the identification of sensitive data is a prerequisite for protecting sensitive data from being leaked. However, the diversity and complexity of data pose significant challenges to the identification of sensitive data. For example, data can exist in many forms, including structured database records, unstructured text files, emails, etc., and the types and definitions of sensitive data may vary by organization, region, and industry. This diversity further increases the complexity of sensitive data identification.
[0003] At present, predefined rules and other methods are often used to identify sensitive types of data. However, due to the diversity of data types, predefined rules cannot cover all situations and are difficult to adapt to new data. For example, new data may contain sensitive information that cannot be identified by predefined rules, which leads to the current technical problem of low accuracy in identifying sensitive types of data.
[0004] The above contents are only used to assist in understanding the technical solutions of the embodiments of the present application, and do not constitute an admission that the above contents are prior art. Summary of the invention
[0005] The main purpose of the embodiments of the present application is to provide a data identification method, device, electronic device, storage medium and program product, aiming to solve the technical problem of low accuracy in identifying sensitive types of data.
[0006] To achieve the above object, an embodiment of the present application provides a data identification method, the method comprising:
[0007] Acquire the data to be identified, each sensitive type of the industry to which the data to be identified belongs, and a subject word of each sensitive type, wherein the subject word of the sensitive type is determined based on sample data of the industry of the sensitive type;
[0008] Inputting each sensitive type of the industry, the subject vocabulary of each sensitive type, and the data to be identified into a preset large language model, and determining the relevance of the data to be identified with each sensitive type through the preset large language model based on the subject vocabulary of each sensitive type;
[0009] The sensitive type whose correlation is greater than a preset correlation threshold is used as the target sensitive type of the data to be identified.
[0010] In one embodiment, the method further comprises:
[0011] Obtaining each sensitive type determined based on a preset industry specification of the industry, and obtaining industry sample data of each of the sensitive types;
[0012] For each of the sensitive types, a candidate vocabulary set is extracted from the industry sample data of the sensitive type, and based on the candidate vocabulary set, a subject vocabulary of the sensitive type is determined.
[0013] In one embodiment, the step of extracting a candidate vocabulary set from the sensitive type industry sample data and determining the sensitive type subject vocabulary based on the candidate vocabulary set includes:
[0014] Performing word segmentation processing on the sensitive type of industry sample data, and determining a candidate vocabulary set for the industry sample data according to the word segmentation processing result;
[0015] Based on the candidate vocabulary set, construct a vocabulary adjacency matrix;
[0016] Calculating the resonance value of each candidate word in the candidate word set according to a preset hyperbolic tangent function, a preset temperature control adjustment parameter and the word adjacency matrix;
[0017] In the candidate vocabulary set, a preset number of target vocabulary with a higher resonance value ranking are determined, and each of the target vocabulary is used as a subject vocabulary of the sensitive type, wherein the resonance value of the candidate vocabulary ranked higher is greater than the resonance value of the candidate vocabulary ranked lower.
[0018] In one embodiment, the word segmentation processing result includes an initial vocabulary, and the industry sample data includes a plurality of sub-sample data;
[0019] The step of performing word segmentation processing on the sensitive type of industry sample data and determining a candidate vocabulary set for the industry sample data according to the word segmentation processing result comprises:
[0020] For each of the sub-sample data, performing word segmentation processing on the sub-sample data according to a preset word segmentation tool, and filtering out stop words in the sub-sample data to obtain an initial vocabulary of the sub-sample data;
[0021] By using a preset keyword extraction algorithm and the dependency relationships corresponding to the initial words, the dependency weights of the initial words are calculated, and the dependency weights are sorted in reverse order, and the initial words with dependency weights ranked in the front preset percentile are used as candidate words for the sub-sample data;
[0022] The candidate words corresponding to each of the sub-sample data are aggregated to obtain a candidate word set for the industry sample data.
[0023] In one embodiment, the step of constructing a vocabulary adjacency matrix based on the candidate vocabulary set includes:
[0024] Determining a vocabulary vector for each candidate vocabulary in the candidate vocabulary set;
[0025] For each target candidate word, based on the word vectors corresponding to each of the candidate words, calculate the word similarity between the target candidate word and each other candidate word in the candidate word set;
[0026] Based on the vocabulary similarities corresponding to each of the target candidate vocabulary, a vocabulary adjacency matrix is constructed;
[0027] The target candidate word is any candidate word in the candidate word set, and the other candidate words are any candidate words in the candidate word set except the target candidate word.
[0028] In one embodiment, the step of determining the vocabulary vector of each candidate vocabulary in the candidate vocabulary set comprises:
[0029] Obtain pre-trained industry vectorization models;
[0030] For each of the candidate words, the candidate word is input into the industry vectorization model, and the candidate word is mapped to the vector space of the industry vectorization through the industry vectorization model to obtain the word vector of the candidate word.
[0031] In one embodiment, the step of calculating the resonance value of each candidate word in the candidate word set according to the preset hyperbolic tangent function, the preset temperature control adjustment parameter and the word adjacency matrix includes:
[0032] For each candidate word, obtaining the similarity of each word corresponding to the candidate word in the word adjacency matrix;
[0033] For each of the vocabulary similarities, calculating a sub-resonance value of the vocabulary similarity according to a preset hyperbolic tangent function and a preset temperature control adjustment parameter;
[0034] For each of the candidate words, the sub-resonance values corresponding to the candidate words are accumulated to obtain the resonance value of the candidate word.
[0035] In addition, to achieve the above-mentioned purpose, the embodiment of the present application provides a data identification device, the device comprising:
[0036] An acquisition module, used to acquire the data to be identified, each sensitive type of the industry to which the data to be identified belongs, and a subject word of each sensitive type, wherein the subject word of the sensitive type is determined based on sample data of the industry of the sensitive type;
[0037] A relevance determination module, used to input each sensitive type of the industry, a subject word of each sensitive type, and the data to be identified into a preset large language model, and determine the relevance of the data to be identified with each sensitive type respectively according to the subject words of each sensitive type through the preset large language model;
[0038] The target sensitive type determination module is used to take the sensitive type whose correlation is greater than a preset correlation threshold as the target sensitive type of the data to be identified.
[0039] In addition, to achieve the above-mentioned purpose, an embodiment of the present application also provides an electronic device, which includes: a memory, a processor, and a program of the data identification method stored in the memory and executable on the processor. When the program of the data identification method is executed by the processor, the steps of the data identification method as described above can be implemented.
[0040] In addition, to achieve the above-mentioned purpose, an embodiment of the present application also provides a computer-readable storage medium, on which a program for implementing the data identification method is stored. When the program of the data identification method is executed by a processor, the steps of the data identification method as described above are implemented.
[0041] In addition, to achieve the above-mentioned purpose, an embodiment of the present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned data identification method when executed by a processor.
[0042] One or more technical solutions proposed in the embodiments of the present application have at least the following technical effects: the present application can obtain the data to be identified, the various sensitive types of the industry described in the data to be identified, and the subject vocabulary of the sensitive type, and the subject vocabulary of the sensitive type is determined based on the industry sample data of the sensitive type, and the industry sample data of the sensitive type can reflect the industry data commonly used by the sensitive type, so the present application can more accurately determine the subject vocabulary of the sensitive type through the industry sample data of the sensitive type. Then, the data to be identified, the various sensitive types, and the subject vocabulary of each sensitive type can be input into a preset large language model, and through the powerful semantic understanding ability of the preset large language model itself, combined with the subject vocabulary corresponding to each sensitive type, the correlation between the data to be identified and each sensitive type is determined, so that the sensitive type with the correlation greater than the preset correlation threshold is used as the target sensitive type of the data to be identified, realizing the identification of the sensitive type of the data to be identified and improving the accuracy of the sensitive type identification of the data. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the embodiments of the present application, and together with the description are used to explain the principles of the embodiments of the present application.
[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0045] Figure 1 This is a flow chart of an embodiment of a data identification method according to an embodiment of the present application;
[0046] Figure 2 A schematic diagram of the vector space corresponding to the industry vectorization model in the data identification method of the embodiment of the present application;
[0047] Figure 3 A schematic diagram of a process for determining a subject word in a data identification method according to an embodiment of the present application;
[0048] Figure 4 A schematic diagram of a process for identifying sensitive types of data to be identified in a data identification method according to an embodiment of the present application;
[0049] Figure 5 This is a schematic diagram of a flow chart of presetting a large language model to identify data to be identified in the data identification method of an embodiment of the present application;
[0050] Figure 6 This is a schematic diagram of the module structure of the data identification device according to an embodiment of the present application;
[0051] Figure 7 Schematic diagram of the device structure of the hardware operating environment involved in the data identification method in the embodiment of the present application.
[0052] The purpose, features and advantages of the embodiments of the present application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0053] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the embodiments of the present application and are not used to limit the embodiments of the present application.
[0054] In order to better understand the technical solutions of the embodiments of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0055] With the rapid development of information technology, data has become one of the most valuable assets of enterprises. However, data security issues have become increasingly prominent, and data leakage and abuse incidents have occurred frequently, bringing huge risks to enterprises and individuals. How to ensure data security and prevent improper exposure of sensitive information is an urgent problem that needs to be solved. In the field of data security, the identification of sensitive data is a prerequisite for protecting data. However, the diversity and complexity of data have brought significant challenges to the identification of sensitive data. Data can exist in many forms, including structured database records, unstructured text files, emails, etc., and the types and definitions of sensitive data may vary by organization, region, and industry. This difference further increases the complexity of sensitive data identification.
[0056] Although there are many methods for identifying sensitive types of data, there is still the problem of inaccurate identification of sensitive types of data. For the method of using predefined rules to identify sensitive types of data, the predefined rules are too fixed and difficult to cover all situations. They are also not universal across industries. Predefined rules need to be customized for different industries, but they cannot cover all types of sensitive data corresponding to the industry. Therefore, predefined rules are also difficult to accurately identify sensitive types of data.
[0057] For machine learning models, they may misidentify due to bias or insufficiency of training data. When using algorithms such as deep neural networks, a large amount of labeled data is often required for training, and they may perform poorly when faced with unseen data patterns. In addition, as regulations and business needs change, the definition of sensitive data is constantly being updated, and predefined rules and machine learning models are difficult to adapt to changes in new data, and thus it is difficult to accurately identify the sensitive types of new data.
[0058] To this end, the embodiment of the present application provides a data identification method, in which the data to be identified, the various sensitive types of the industry of the data to be identified, and the subject vocabulary of the sensitive type can be obtained, and the subject vocabulary of the sensitive type is determined based on the industry sample data of the sensitive type, and the industry sample data of the sensitive type can reflect the industry data commonly used by the sensitive type, so the present application can more accurately determine the subject vocabulary of the sensitive type through the industry sample data of the sensitive type. Then, the data to be identified, the various sensitive types, and the subject vocabulary of each sensitive type can be input into a preset large language model, and the target subject associated with the data to be identified in each subject vocabulary can be determined through the powerful semantic understanding ability of the preset large language model itself, so that the target sensitive type of the data to be identified can be determined based on the sensitive type to which the target subject belongs, thereby improving the accuracy of sensitive type identification of the data.
[0059] Based on this, the present application embodiment provides a data identification method, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the data identification method of the present application. The data identification method includes steps S10 to S30:
[0060] Step S10, obtaining the data to be identified, each sensitive type of the industry to which the data to be identified belongs, and a subject word of each sensitive type, wherein the subject word of the sensitive type is determined based on sample data of the industry of the sensitive type;
[0061] It should be noted that the data to be identified is data that needs to be identified as sensitive types. The data to be identified can be documents, texts, etc., and this embodiment does not specifically limit this. Since the sensitive types of data corresponding to different industries are different, and the importance of data is also different, it is necessary to determine the industry to which the data to be identified belongs, and then obtain the sensitive type of the industry and the subject vocabulary corresponding to the sensitive type. This facilitates the subsequent accurate identification of the sensitive type of the data to be identified.
[0062] Each sensitive type may correspond to multiple subject words, which are words related to the sensitive type. For example, when the sensitive type is traffic, the subject words corresponding to the sensitive type may be traffic jam, car accident, walking, and crowds of people. The above is only one example, and this embodiment does not specifically limit this. The sensitive type can be pre-set based on the industry, and each industry can correspond to multiple sensitive types.
[0063] Each sensitive type has its own corresponding industry sample data. The sensitive type can be bound to the industry sample data, and the subject vocabulary corresponding to the sensitive type can be determined from the industry sample data.
[0064] Exemplarily, the data to be identified, each sensitive type of the industry to which the data to be identified belongs, and the subject vocabulary of each sensitive type are obtained.
[0065] In a feasible embodiment, step S10 further includes steps S11 to S12:
[0066] Step S11, obtaining various sensitive types determined based on preset industry specifications of the industry, and obtaining industry sample data of each sensitive type;
[0067] It should be noted that each industry can have its own corresponding preset industry specifications, which can be determined in advance based on the corresponding industry policies of the industry. Different types of data are defined in the industry policies, such as transaction type data, customer data, etc., and the corresponding security level is also specified for each type in the industry policy. The preset industry specifications can be used to describe the data level, data security, etc. of different data in the industry. The data level can be used to characterize the sensitivity and value of the data. For example, the higher the data level, the higher the sensitivity and value of the data level, indicating that the data at this level may be a type that cannot be leaked. Different industries have different preset industry specifications. For example, there may be industries such as finance or medical care, and this embodiment does not make specific limitations on this. The preset industry specifications can also be determined based on actual conditions, and this embodiment does not make specific limitations on this.
[0068] Each sensitive type of an industry can be determined based on preset industry specifications. For example, based on preset industry specifications, different levels of sensitive types corresponding to an industry can be constructed. For example, the same level can correspond to multiple sensitive types, and the sensitivity of sensitive types at different levels may be different. The higher the level, the higher the sensitivity may be. The higher the sensitivity, the more important the data is and the less likely it is to be leaked. The sensitive type may be, for example, an ID card, a bank card, a medical record, etc., which is not specifically limited in this embodiment.
[0069] Industry sample data is data associated with sensitive types. Industry sample data can also be data such as documents and texts, which are not specifically limited in this embodiment. Industry sample data represents a set of data with common characteristics or attributes. Each sensitive type has its own corresponding industry sample data, and data in industry sample data of the same sensitive type have common characteristics or attributes. For example, sensitive types can also be types such as financial reports, case information, technical manuals, legal documents, and health records.
[0070] Exemplarily, each sensitive type determined in advance based on preset industry specifications and the industry sample data corresponding to each sensitive type are obtained, and the sensitive type and the industry sample data of the sensitive type can be bound. This embodiment obtains the industry sample data of the sensitive type, thereby facilitating the subsequent determination of the subject vocabulary of the sensitive type from the industry sample data.
[0071] Step S12, extracting a candidate vocabulary set from the sensitive type industry sample data, and determining a sensitive type subject vocabulary based on the candidate vocabulary set.
[0072] It should be noted that the industry sample data may include multiple sub-sample data, and the sub-sample data may be data of sample documents, etc. In this embodiment, the corresponding subject vocabulary can be automatically determined based on the industry sample data. In other embodiments, the subject vocabulary may also be manually determined based on the industry sample data. When manually determining the subject vocabulary based on the industry sample data, it is necessary to determine multiple representative sample documents corresponding to the sensitive type to ensure that they can cover the key features and diversity of the created sensitive type. By analyzing the sample documents, the key concepts, terms and topics in each sample document are identified, and the keywords or phrases that can represent the sample documents are extracted as the subject vocabulary, so that an accurate, consistent and effective subject vocabulary can be created for the defined sensitive type, which is helpful for the subsequent preset large language model to understand the data to be identified.
[0073] The step of automatically determining corresponding subject vocabulary based on industry sample data includes automatically extracting a candidate vocabulary set from sensitive type industry sample data, and automatically determining sensitive type subject vocabulary based on the candidate vocabulary set, thereby automatically determining sensitive type subject vocabulary and improving the efficiency of determining sensitive type subject vocabulary.
[0074] Exemplarily, for each sensitive type, keywords or phrases that can represent the industry sample data and are manually extracted from the industry sample data of the sensitive type can be obtained as the subject vocabulary, and / or, a candidate vocabulary set can be automatically extracted from the industry sample data of the sensitive type, and the subject vocabulary of the sensitive type can be determined based on the candidate vocabulary set. In the process of automatically determining the subject vocabulary, the subject vocabulary of the sensitive type will be determined based on the similarity between the candidate vocabulary in the candidate vocabulary set, so that the subject vocabulary can represent the industry text data, etc.
[0075] In a feasible embodiment, step S12 further includes steps S121 to S124:
[0076] Step S121, performing word segmentation processing on the sensitive type of industry sample data, and determining a candidate vocabulary set for the industry sample data according to the word segmentation processing result;
[0077] It should be noted that the word segmentation process for industry sample data can be to divide the industry sample data into semantic words, and then a candidate word set for the industry sample data can be determined. The candidate word set includes multiple candidate words corresponding to the industry sample data. The industry sample data includes multiple sub-sample data, and the sub-sample data can be data such as documents or texts, and this embodiment does not make specific limitations on this.
[0078] Exemplarily, for each sub-sample data in the industry sample data, the word segmentation process is performed on each sub-sample data to obtain a word segmentation result, and the candidate word set is determined based on the word segmentation result. The word segmentation result can include words existing in the sub-sample data and the like.
[0079] In a feasible embodiment, the word segmentation result includes initial words, the industry sample data includes multiple sub-sample data, and step S121 further includes steps S1211 to S1213:
[0080] Step S1211, for each sub-sample data, perform word segmentation on the sub-sample data according to a preset word segmentation tool and filter out the stop words in the sub-sample data to obtain the initial words of the sub-sample data;
[0081] Step S1212, through a preset keyword extraction algorithm and the dependency relationship between each initial word, calculate the dependency weight of each initial word, perform a reverse order sorting on the dependency weights, and take the initial words whose dependency weights are ranked in the top preset percentile as the candidate words of the sub-sample data;
[0082] Step S1213, aggregate the candidate words corresponding to each of the sub-sample data to obtain the candidate word set of the industry sample data.
[0083] It should be noted that the preset word segmentation tool is used to perform word segmentation on the sub-sample. The preset word segmentation tool can be Jieba (a Chinese word segmentation tool), or other tools that can perform Chinese word segmentation. This embodiment does not make specific limitations on this. And during the process of performing word segmentation on the sub-sample data, the stop words in the sub-sample data will also be filtered out, so that the initial words of the sub-sample data can be obtained. Stop words are words that can be ignored. For example, stop words refer to words that are considered to have no actual meaning in text processing. For example, stop words can be: "de", "zai", "le", etc. This embodiment does not make specific limitations on this, and stop words can be determined in advance. Initial words are the words remaining in the sub-sample data after word segmentation and filtering out stop words.
[0084] The preset keyword extraction algorithm is used to extract candidate words that appear frequently in the sub-sample data. For example, the preset keyword extraction algorithm can determine the important words in the sub-sample data, that is, the candidate words, from the initial words of the sub-sample data. The preset keyword extraction algorithm can be a TextRank algorithm (text ranking algorithm) or other algorithms for extracting keywords, which is not specifically limited in this embodiment. Each initial word has its own corresponding dependency relationship. For the target initial word, the dependency relationship of the target initial word is the dependency relationship between the target initial word and other initial words in the sub-sample data. The target initial word can be any initial word in the sub-sample data, and other initial words are not equal to the target initial word.
[0085] For example, there are three words ABC, the dependency of A can be, A depends on B, A depends on C, the dependency of B can be, A depends on B, and the dependency of C can be, A depends on C. For example, I eat apples, A can be eat, B is apple, and C is me. The above examples are provided for easy understanding, and this embodiment does not specifically limit the initial words and the dependency of the initial words.
[0086] A preset keyword extraction algorithm can be used to calculate the dependency weight of each initial vocabulary and sort the dependency weights in reverse order. For example, the ones with larger dependency weights are placed in front. The preset percentile can be determined based on actual conditions. For example, the preset percentile can be 10%, or 12%. For example, if the preset percentile is 10%, all initial vocabulary with dependency weights in the top 10% can be used as candidate vocabulary.
[0087] Exemplarily, each sub-sample data may be considered as a document or a row of text content, etc., and this embodiment does not make specific limitations on this. For each sub-sample data, the sub-sample data is segmented according to a preset segmentation tool, and stop words in the sub-sample data are filtered out to obtain the initial vocabulary of the sub-sample data; for each sub-sample data, the dependency weight of each initial vocabulary in the sub-sample data is calculated through a preset keyword extraction algorithm and the dependency relationship corresponding to each initial vocabulary in the sub-sample data, and each dependency weight is sorted in reverse order, and the initial vocabulary with the dependency weight ranked in the front preset percentile is used as the candidate vocabulary of the sub-sample data; the candidate vocabulary corresponding to each sub-sample data is used as the candidate vocabulary set of the industry sample data.
[0088] Step S122, constructing a vocabulary adjacency matrix based on the candidate vocabulary set;
[0089] It should be noted that the total number of candidate words in the candidate word set is n, and the word adjacency matrix is an N×N matrix. The word adjacency matrix can reflect the word similarity between each candidate word in the candidate word set. Word similarity is used to describe the degree of correlation between two candidate words. The higher the word similarity, the higher the correlation between the two candidate words corresponding to the word similarity. The higher the correlation, the closer the semantics between the two candidate words.
[0090] Exemplarily, a vocabulary adjacency matrix is constructed based on vocabulary similarities between candidate vocabulary in the candidate vocabulary set.
[0091] In a feasible embodiment, step S122 further includes steps S1221 to S1223:
[0092] Step S1221, determining the vocabulary vector of each candidate vocabulary in the candidate vocabulary set;
[0093] It should be noted that the vocabulary vector can be used to describe the semantics of the candidate vocabulary, and can also be used to describe the context information corresponding to the candidate vocabulary. Each candidate vocabulary in the candidate vocabulary set has its own corresponding vocabulary vector.
[0094] For example, a preset industry vectorization model may be used to determine the vocabulary vector of each candidate word.
[0095] In a feasible embodiment, step S1221 further includes steps A10 to A20:
[0096] Step A10, obtaining a pre-trained industry vectorization model;
[0097] Step A20: for each candidate word, the candidate word is input into the industry vectorization model, and the candidate word is mapped to the industry vectorization vector space through the industry vectorization model to obtain the word vector of the candidate word.
[0098] It should be noted that the industry vectorization model can be pre-trained. The industry vectorization model has a corresponding vector space. By mapping the candidate vocabulary in the vector space, the semantics and context information of the candidate vocabulary can be reflected in the vector space. The semantics and context information of the candidate vocabulary are expressed through the vocabulary vector. The context information is the possible context vocabulary corresponding to the candidate vocabulary in the industry.
[0099] Exemplarily, the candidate vocabulary is input into the industry vectorization model, and the corresponding position of the candidate vocabulary in the vector space is determined by the industry vectorization model, so as to map the candidate vocabulary to the vector space and obtain the vocabulary vector of the candidate vocabulary. The industry vectorization model can be obtained by training a preset model to be trained using the industry expectation library of the industry.
[0100] This embodiment determines the vocabulary vector of the candidate vocabulary through the industry vectorization model, so that it is convenient to describe the semantics and contextual information of the candidate vocabulary through the vocabulary vector, and then it is convenient for the subsequent preset large language model to understand the semantics and contextual information of the vocabulary, and then it is convenient to more accurately identify sensitive types.
[0101] Among them, the step of obtaining a pre-trained industry vectorization model includes: obtaining an industry corpus of the industry; training a preset model to be trained through each corpus word in the industry corpus and each corresponding context information to obtain a trained industry vectorization model; wherein, the training process of the model to be trained includes: mapping each corpus word in the industry corpus to the training vector space of the model to be trained based on the context information corresponding to each corpus word in the industry corpus.
[0102] It should be noted that the industry corpus includes data within the industry, such as document data, text data, etc. Corpus words are words that appear in documents or texts in the industry corpus. The context information of the corpus words may include the previous information and / or the following information of the corpus words. The previous information may be a phrase or a word, and the following information may also be a phrase or a word, etc. This embodiment does not make specific restrictions on this. The preset model to be trained may be a neural network model, or other machine learning models, which is not specifically limited in this embodiment. The model to be trained has a corresponding training vector space, and the size of the training vector space can be set in advance based on actual conditions. This embodiment does not make specific restrictions on this. After the training of the model to be trained is completed, an industry vectorization model with a vector space can be obtained. By training the model to be trained, the model to be trained can learn the context information corresponding to each corpus word in the industry, etc.
[0103] Exemplarily, the model to be trained can be trained through each corpus word in the industry corpus and each corresponding context information to obtain a trained industry vectorization model. For example, the model to be trained can be subjected to unsupervised training, etc., which is not specifically limited in this embodiment. The training process of the model to be trained can specifically include, based on each corpus word and each corresponding context information, mapping each corpus word in the industry corpus to the training vector space of the model to be trained to achieve training of the model to be trained. For example, in the training process of the model to be trained, semantically similar words can also be mapped to similar positions in the training vector space, so that semantically similar words can be associated in the vector space, so that after the vocabulary vector is generated using the industry vectorization model, it is convenient to calculate the vocabulary similarity between each candidate vocabulary through the vocabulary vector.
[0104] This embodiment trains the model to be trained through each corpus word and the corresponding context information of the industry corpus, so that the model to be trained can learn the context information corresponding to each corpus word in the industry, thereby facilitating the subsequent generation of vocabulary vectors of candidate words, so that the vocabulary similarity between candidate words can be calculated through the vocabulary vectors.
[0105] For a better understanding of this embodiment, please refer to Figure 2 , Figure 2 It is a schematic diagram of the vector space corresponding to the industry vectorization model. Figure 2 Where X and Y can be expressed as the horizontal and vertical coordinates of the vector space, Figure 2 G1~Gn in the above may be n candidate words in the vector space. For example, G1 may be a candidate word for queen, etc. This embodiment does not specifically limit this. [0.3, 0.4, ..., 0.7] corresponding to G1 is the word vector of G1. The word vector can be represented by multiple numerical values, among which 0.3, 0.4, ..., 0.7 may represent the position of words related to G1. For example, 0.3 may represent the position of a word in the vector space, and 0.4 and 0.7 may also represent the position of a word in the vector space.
[0106] Step S1222, for each target candidate word, based on the word vectors corresponding to each candidate word, calculate the word similarity between the target candidate word and each other candidate word in the candidate word set;
[0107] Step S1223, constructing a vocabulary adjacency matrix based on the vocabulary similarities corresponding to each target candidate vocabulary;
[0108] The target candidate word is any candidate word in the candidate word set, and the other candidate words are any candidate words in the candidate word set except the target candidate word.
[0109] It should be noted that the vocabulary similarity is used to characterize the degree of correlation between two candidate vocabulary. Each target candidate vocabulary has corresponding multiple vocabulary similarities. For the same target candidate vocabulary, there is a corresponding vocabulary similarity between the target candidate vocabulary and each other candidate vocabulary in the candidate vocabulary set.
[0110] For example, when the number of candidate words in the candidate word set is n, then the target candidate word corresponds to n-1 other candidate words, and the number of vocabulary similarities between the target candidate word and other candidate words is also n-1. A vocabulary adjacency matrix can be constructed based on the vocabulary similarities corresponding to all target candidate words in the candidate word set, and the vocabulary similarity between candidate words can be reflected by the vocabulary adjacency matrix. The vocabulary similarity can be the cosine similarity between two vocabulary vectors.
[0111] Exemplarily, the total number of candidate words in the candidate vocabulary set is n. Based on the total number of candidate words in the candidate vocabulary set n, an n×n adjacency vocabulary matrix is constructed. Referring to Table 1, the vocabulary adjacency matrix is represented in a tabular form. Table 1 is:
[0112]
[0113]
[0114] Among them, S1~Sn represent each candidate word in the candidate vocabulary set, there are n candidate words, there are n vocabulary vectors, Cij is the vocabulary similarity between candidate word Si and candidate word Sj, for example, C12 in Table 1 is the vocabulary similarity between candidate word S1 and candidate word S2, and the vocabulary similarity between S1 and S1 is 1. i is less than or equal to n, and i is greater than or equal to 1, j is less than or equal to n, and j is greater than or equal to 1. Si represents the i-th candidate word in the candidate vocabulary set, and Sj represents the j-th candidate word in the candidate vocabulary set. The vocabulary vector of Si can be expressed as Ai, and the cosine similarity between the two vocabulary vectors can be used as the vocabulary similarity. Formula 1 for calculating cosine similarity (lexical similarity) is:
[0115]
[0116] Among them, v is the vocabulary similarity between Si and Sj, A i is the vocabulary vector of Si, A j is the vocabulary vector of Sj, and || represents the second-order norm.
[0117] Step S123, calculating the resonance value of each candidate word in the candidate word set according to a preset hyperbolic tangent function, a preset temperature control adjustment parameter and a word adjacency matrix;
[0118] It should be noted that the preset hyperbolic tangent function is used to smooth the vocabulary similarities corresponding to the candidate vocabulary, so as to reduce the influence of the extreme vocabulary similarities in the candidate vocabulary on the overall vocabulary similarities corresponding to the candidate vocabulary. The overall similarity corresponding to the candidate vocabulary includes the vocabulary similarities between the candidate vocabulary and other candidate vocabulary respectively. The resonance value can be used to describe the mutual influence between the candidate vocabulary and other candidate vocabulary in the candidate vocabulary set. The larger the resonance value, the stronger the correlation between the candidate vocabulary and the industry sample document, and the more it can represent the characteristics of the industry sample document of the sensitive type. The preset temperature control adjustment parameters can be set based on actual conditions, and this embodiment does not make specific limitations on this.
[0119] Exemplarily, for each sample word, the word similarity corresponding to the sample word can be obtained from the word adjacency matrix, and the resonance value of the candidate word is calculated based on the word similarity corresponding to the sample word, a preset hyperbolic tangent function, and a preset temperature control adjustment parameter. By determining the resonance value of the candidate word, this embodiment can avoid extreme candidate words in the candidate word set from having a greater impact on the entire candidate word set, so that a more accurate topic word can be determined more accurately from the candidate word set later.
[0120] In a feasible embodiment, step S123 further includes steps B10 to B30:
[0121] Step B10, for each candidate word, obtaining the similarity of each word corresponding to the candidate word in the word adjacency matrix;
[0122] Step B20, for each vocabulary similarity, calculating the sub-resonance value of the vocabulary similarity according to a preset hyperbolic tangent function and a preset temperature control adjustment parameter;
[0123] Step B30: for each candidate word, the sub-resonance values corresponding to the candidate word are accumulated to obtain the resonance value of the candidate word.
[0124] It should be noted that the higher the preset temperature control adjustment parameter, the smoother the sub-resonance value of the vocabulary similarity, which can reduce the influence of extreme vocabulary similarity on the resonance value of the candidate vocabulary, and the lower the preset temperature control adjustment parameter, the higher the influence of vocabulary similarity on the resonance value of the candidate vocabulary can be amplified, that is, the lower preset temperature control adjustment parameter can amplify the contribution of higher vocabulary similarity to the resonance value of the candidate vocabulary, and the higher preset temperature control adjustment parameter can suppress the influence of extreme vocabulary similarity on the resonance value of the candidate vocabulary. For example, the extreme vocabulary similarity is too high and close to 1, or it can be a small vocabulary similarity, close to 0, and this embodiment does not make specific limitations on this.
[0125] For example, the preset hyperbolic tangent function can be expressed as Formula 2, and the sub-resonance value of the vocabulary similarity can be calculated by Formula 2:
[0126]
[0127] Wherein, T is a preset temperature control adjustment parameter, tanh(v / T) is a calculated sub-resonance value, v is a vocabulary similarity, and e is the base of a natural logarithm. This embodiment adjusts the vocabulary similarity by presetting a hyperbolic tangent function and a preset temperature control adjustment parameter, thereby facilitating the subsequent calculation of the resonance value of the candidate vocabulary, so as to select the candidate vocabulary with the highest resonance value as the subject vocabulary.
[0128] Exemplarily, for each candidate word, the formula for calculating the resonance value of the candidate word is Formula 3:
[0129]
[0130] Among them, RS i is the resonance value of the candidate word, v ij is the vocabulary similarity between the i-th candidate word and the j-th candidate word, j is not equal to i, j can start from 1 to n. tanh(v ij / T) is represented by the sub-resonance value corresponding to the candidate word.
[0131] This embodiment calculates the resonance value of each candidate word, so that the resonance value can reflect the relevance of the candidate word in the industry sample data, thereby facilitating the subsequent screening of the subject words that best represent the sensitive type of industry sample data through the resonance value, thereby facilitating the improvement of the accuracy of the subsequent identification of the sensitive type of data.
[0132] Step S124, in the candidate vocabulary set, determine a preset number of target vocabulary with a higher resonance value, and use each target vocabulary as a sensitive type of topic vocabulary, wherein the resonance value of the candidate vocabulary with a higher resonance value is greater than the resonance value of the candidate vocabulary with a lower resonance value.
[0133] It should be noted that the resonance values may be sorted in descending order, and the preset number may be set based on actual conditions. For example, the preset number may be 3, 4, 5, etc., and this embodiment does not specifically limit this.
[0134] For example, by sorting the resonance values in descending order, a target word ranked first by a preset number of candidate words can be selected as a topic word of the sensitive type. The target word is a candidate word ranked first by a preset number of resonance values. There are a preset number of target words.
[0135] This embodiment calculates the lexical similarity between candidate words and determines the resonance value of each candidate word, so that the subject words with high relevance in the industry sample data can be screened out through the resonance value, and the characteristics of the industry sample data can be better expressed through the subject words, so that the subsequent preset large language model can accurately calculate the correlation between the sensitive type and the data to be identified.
[0136] For a better understanding of this embodiment, please refer to Figure 3, the process of determining the subject vocabulary in this embodiment is briefly described: Step F10: sensitive type C is bound to the corresponding industry sample data, sensitive type C is bound to the corresponding industry sample data of sensitive type C, the subject vocabulary can be customized, or step F20 can be executed to generate self-generated subject vocabulary, step F20 can include steps F21 to F25, step F21: determine candidate vocabulary from industry sample data. The implementation of step F21 can use the preset keyword extraction algorithm referred to by F21a to determine candidate vocabulary in the industry sample data. Step F22: vectorize the candidate words. Specifically, the industry corpus referred to by F22a can be used to train the industry vector model referred to by F22b. After the candidate words are vectorized, the vocabulary vectors of the candidate words can be used to calculate the vocabulary similarity corresponding to the candidate words, and then execute step F23: construct a vocabulary adjacency matrix; execute step F24: calculate the resonance value of the candidate words; the topic words can be screened based on each resonance value, and execute step F25: sensitive type topic words C1, C2...Ck, C1~Ck are the topic words corresponding to the sensitive type C, and k is a positive integer.
[0137] Step S20, inputting each sensitive type of the industry, the subject words of each sensitive type, and the data to be identified into a preset large language model, and determining the relevance of the data to be identified with each sensitive type based on the subject words of each sensitive type through the preset large language model;
[0138] It should be noted that the preset large language model may be an LLM (Large Language Model) model, etc., which is not specifically limited in this embodiment. The preset large language model may be used to determine the correlation between the data to be identified and each sensitive type. The correlation is used to characterize the degree of correlation between the data to be identified and the sensitive type. The correlation is positively correlated with the correlation. The correlation may be represented by a correlation score.
[0139] Exemplarily, a preset prompt word template can be used to first determine each sensitive type, the subject vocabulary of each sensitive type, and the prompt sentence corresponding to the data to be identified, and the prompt word can be input into a preset large language model to output the correlation between the data to be identified and each sensitive type through the preset large language model.
[0140] In a feasible embodiment, step S20 may also include generating a sensitive identification prompt sentence for the data to be identified based on a preset prompt word template, the data to be identified, each sensitive type, and the subject vocabulary corresponding to each sensitive type; inputting the sensitive identification prompt sentence into a preset large language model, and calculating the correlation between the data to be identified and each sensitive type respectively through the preset large language model according to the subject vocabulary corresponding to each sensitive type.
[0141] It should be noted that the preset prompt word template can be pre-set based on actual conditions, and this embodiment does not specifically limit this. The sensitive identification prompt sentence is used to prompt the preset large language model to calculate the relevance of the data to be identified with each sensitive type based on each subject vocabulary.
[0142] The correlation can be represented by a correlation score. The higher the correlation score corresponding to the sensitive type, the more likely the sensitive type is to be associated with the data to be identified, and the more the sensitive type can reflect the data to be identified. The range of the correlation score can be a preset range, which can be 0 to 1. It can be understood that the correlation score corresponding to each sensitive type will not exceed 1 or be less than 0. The preset range can be set based on actual conditions, and this embodiment does not specifically limit this.
[0143] For example, the data to be identified can be text to be identified. When there is a text to be identified, the sensitive types of the industries corresponding to the text to be identified are transportation, National Day, and medical health. The subject words corresponding to each sensitive type are transportation: traffic jam, huge crowds, unable to walk, car accident; National Day: celebration, national flag, national anthem, activity, National Day; Medical health: medical expenses, expensive medicine, life, poverty, difficulty. According to the preset prompt word template, the corresponding sensitive identification prompt sentence can be determined as:
[0144] “Please judge the relevance of the following text content to each sensitive type. Each sensitive type has some subject words. If relevant, please give a relevance score of 0-1 for each sensitive type (only sensitive types are considered when outputting). When scoring, both the sensitive type itself and its subject words should be fully considered, and the relevance of subject words cannot be considered only. Output in json format is required. Sensitive types and their subject words are as follows:
[0145] Traffic: traffic jams, huge crowds, inability to move, car accidents;
[0146] National Day: celebration, national flag, national anthem, activities, national day;
[0147] Health care: medical expenses, expensive medicines, life, poverty, and difficulties;
[0148] Text content: At 3 pm on October 1, 2023, Tianfu Avenue was congested. The reason was a serious car accident on Tianfu Avenue in Chengdu, resulting in 3 deaths and 1 injury. "
[0149] Among them, the text content in the sensitive recognition prompt sentence is the content of the text to be recognized. After the above sensitive recognition prompt sentence is input into the preset large language model, the answer output by the preset large language model can be: "{"Traffic": 1.0, / / The text mentioned the congestion and traffic accident on Tianfu Avenue, which is closely related to the traffic theme; "National Day": 0.2, / / Although the date mentioned in the text is National Day, the text content is mainly unrelated to the National Day celebrations; "Health Care": 0.1 / / The text mentioned the casualties in the car accident, but did not mention the keywords of medical and health themes such as medical expenses, drug prices or life}".
[0150] Among them, the preset large language model outputs the correlation between each sensitive type and the data to be identified, such as the correlation corresponding to traffic is 1, the correlation corresponding to National Day is 0.2, and the correlation corresponding to medical health is 0.1. In addition to outputting the correlation of each sensitive type, the preset large language model also explains the association between each sensitive type and the data to be identified, so as to facilitate users to understand the correlation between sensitive types and the data to be identified, and improve the interpretability of the correlation corresponding to sensitive types.
[0151] Step S30: taking the sensitive type with a correlation greater than a preset correlation threshold as the target sensitive type of the data to be identified.
[0152] It should be noted that the preset correlation threshold can be set based on actual conditions. For example, the preset correlation threshold can be 0.9, etc., which is not specifically limited in this embodiment. The correlation of the sensitive type is greater than the preset correlation threshold, indicating that the data to be identified is more relevant to the sensitive type. The data to be identified can correspond to one or more target sensitive types, which is not specifically limited in this embodiment.
[0153] Exemplarily, a set M may be used to represent the target sensitive type of the data to be identified. For example, M may be expressed as Formula 4:
[0154] M={f|f∈F and R(f,d)≥θ}(Formula 4)
[0155] Among them, ∈ means belonging, | means "under the condition of...", f is a sensitive type, F is a set of sensitive types in the industry, d is the data to be identified, θ is a preset correlation threshold, R(f,d)≥θ is a target sensitive type whose correlation with the data to be identified is greater than or equal to the preset correlation threshold θ, and Formula 4 means that the set M consists of all sensitive types t that belong to the set F and whose correlation with the data to be identified d is at least θ. For example, continuing with the previous example, if θ is 0.9, the sensitive type set of the industry is F = {transportation, National Day, medical health}, the correlation of transportation output by the preset large language model is 1, the correlation of National Day is 0.1, and the correlation of medical health is 0.1, then the target sensitive type of the data to be identified is M = {transportation}.
[0156] Sensitive data generated during the operation of an enterprise, such as customer personal information and financial data, may cause huge economic losses and reputation risks to the enterprise once it is leaked or abused. Identifying sensitive types of data can help enterprises detect potential security threats in a timely manner and strengthen data security protection measures.
[0157] The embodiment of the present application can obtain the data to be identified, each sensitive type of the industry of the data to be identified, and the subject vocabulary of the sensitive type, and the subject vocabulary of the sensitive type is determined based on the industry sample data of the sensitive type, and the industry sample data of the sensitive type can reflect the industry data commonly used by the sensitive type, so the embodiment of the present application can more accurately determine the subject vocabulary of the sensitive type through the industry sample data of the sensitive type. Then, the data to be identified, each sensitive type and the subject vocabulary of each sensitive type can be input into the preset large language model, and the correlation between the data to be identified and each sensitive type can be determined by combining the subject vocabulary corresponding to each sensitive type through the powerful semantic understanding ability of the preset large language model itself, so as to use the sensitive type with a correlation greater than the preset correlation threshold as the target sensitive type of the data to be identified, thereby realizing the identification of the sensitive type of the data to be identified and improving the accuracy of the sensitive type identification of the data.
[0158] For a better understanding of this embodiment, please refer to Figure 4 , an example is given to illustrate the process of identifying the sensitive type of the data to be identified: Step Y10: obtain the preset industry specification of the industry to which the data to be identified belongs; Step Y20: create the sensitive type corresponding to the industry; Step Y30: bind each sensitive type to the corresponding industry sample data; obtain the industry sample data of each sensitive type, and bind the sensitive type to the corresponding industry sample data. There are two methods for determining the subject vocabulary of the sensitive type, including step Y40 and step Y50. Step Y40: customize the subject vocabulary of the sensitive type based on the industry sample data of the sensitive type; Step Y50: filter out candidate vocabulary from the industry sample data, calculate the resonance value of the candidate vocabulary, and filter out the subject vocabulary of the sensitive type based on the resonance value; Step Y60: input each sensitive type and the corresponding subject vocabulary, Step Y70: input the data to be identified, Step Y60 and Step Y70 can be performed simultaneously, for example, the sensitive type, the subject vocabulary corresponding to the sensitive type and the data to be identified can be determined, and the sensitive identification prompt sentence input into the preset large language model can be determined. Execute step Y80: preset a large language model to evaluate the relevance of the data to be identified with each sensitive type; step Y90: determine the sensitive type of the data to be identified. For example, the sensitive type with a relevance greater than a preset relevance threshold is used as the target sensitive type of the data to be identified.
[0159] Further, see Figure 5, the specific identification process of the sensitive type of the data to be identified is illustrated by example, Z10 refers to the industry sample data of the sensitive type, which may include sensitive type A and sensitive type B. The industry sample data of sensitive type A may be sample document 1, sample document 2, sample document m, etc., where m is a positive integer, and this embodiment does not make specific limitations on this. The industry sample data of sensitive type B may be sample document 1, sample document 2, sample document m, etc., where m is a positive integer, and this embodiment does not make specific limitations on this. After determining the industry sample data corresponding to sensitive type A and sensitive type B, step Z20 may be executed, in which the subject vocabulary may be customized or self-generated, and the process of self-generating subject vocabulary may refer to steps S1211 to S1214. Through step Z20, the subject vocabulary of sensitive type A and the subject vocabulary of sensitive type B may be generated. The subject vocabulary of sensitive type A and the subject vocabulary of sensitive type B may be input into the preset large language model. Z40 corresponds to the data to be identified, and the data to be identified may include multiple target documents, for example, target document 1, target document 2 to target document p, etc., where p is a positive integer, and this embodiment does not specifically limit this. The data to be identified referred to in Z40 can also be input into a preset large language model. Z50 refers to a preset large language model. The preset large language model can identify the data to be identified, for example, it can identify the correlation between the data to be identified and the sensitive type. In this example, for each target document in the data to be identified, the preset large language model can identify the correlation between the target document and each sensitive type, thereby determining the target sensitive type of the target document. The data to be identified may include one target document or multiple target documents, and this embodiment does not specifically limit this. For example, the recognition result of the preset large language model is shown in Z60, and the recognition result includes the recognition results of target document 1, target document 2, and target document p. For example, a represents the recognition result of target document 1, a can be the correlation between target document 1 and each sensitive type, b can be the correlation between target document 2 and each sensitive type, etc.
[0160] The present application also provides a data identification device, please refer to Figure 6 , the device comprises:
[0161] The acquisition module 10 is used to acquire the data to be identified, each sensitive type of the industry to which the data to be identified belongs, and a subject word of each sensitive type, wherein the subject word of the sensitive type is determined based on the industry sample data of the sensitive type;
[0162] A relevance determination module 20 is used to input each sensitive type of the industry, a subject word of each sensitive type, and the data to be identified into a preset large language model, and determine the relevance of the data to be identified with each sensitive type according to the subject words of each sensitive type through the preset large language model;
[0163] The target sensitive type determination module 30 is used to take the sensitive type whose correlation is greater than a preset correlation threshold as the target sensitive type of the data to be identified.
[0164] The data identification device provided in the embodiment of the present application adopts the data identification method in the above embodiment, aiming to solve the technical problem of low accuracy in identifying sensitive types of data. Compared with the prior art, the beneficial effects of the data identification method provided in the embodiment of the present application are the same as those of the data identification method provided in the above embodiment, and other technical features in the data identification device are the same as those disclosed in the above embodiment method, which will not be described in detail here.
[0165] An embodiment of the present application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the data identification method in the above-mentioned embodiment.
[0166] Reference below Figure 7 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure. The electronic device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0167] like Figure 7 As shown, the electronic device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM) 1004. Various programs and data required for the operation of the electronic device are also stored in the RAM 1004. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus.
[0168] Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 1009. The communication device can allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows an electronic device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or have alternatively.
[0169] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 1009, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment of the present disclosure are executed.
[0170] The electronic device provided in the embodiment of the present application adopts the data identification method in the above embodiment 1 to solve the technical problem of low accuracy in identifying sensitive types of data. Compared with the prior art, the beneficial effects of product flow data allocation provided in the embodiment of the present application are the same as the beneficial effects of the data identification method provided in the above embodiment, and other technical features in the data identification device are the same as the features disclosed in the above embodiment method, which will not be repeated here.
[0171] It should be understood that the various parts of the present disclosure can be implemented with hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0172] The above are only specific implementations of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or replacements within the technical scope disclosed in the embodiments of the present application, which should be included in the protection scope of the embodiments of the present application. Therefore, the protection scope of the embodiments of the present application shall be based on the protection scope of the claims.
[0173] This embodiment provides a computer-readable storage medium having computer-readable program instructions stored thereon, and the computer-readable program instructions are used to execute the data identification method in the above-mentioned embodiment 1.
[0174] The computer-readable storage medium provided in the embodiment of the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor devices, equipment or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable EPROM (Electrical Programmable Read Only Memory, read-only memory) or flash memory, an optical fiber, a portable compact disk CD-ROM (compact disc read-only memory, read-only memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution device, device or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency, radio frequency), etc., or any suitable combination of the above.
[0175] The computer-readable storage medium may be included in the electronic device, or may exist independently without being installed in the electronic device.
[0176] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by an electronic device, the electronic device: obtains the data to be identified, each sensitive type of the industry to which the data to be identified belongs, and the subject vocabulary of each sensitive type, wherein the subject vocabulary of the sensitive type is determined based on the industry sample data of the sensitive type; inputs each sensitive type of the industry, the subject vocabulary of each sensitive type and the data to be identified into a preset large language model, and determines the correlation between the data to be identified and each sensitive type respectively according to the subject vocabulary of each sensitive type through the preset large language model; and uses the sensitive type whose correlation is greater than a preset correlation threshold as the target sensitive type of the data to be identified.
[0177] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a LAN (local area network) or WAN (Wide Area Network), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0178] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the equipment, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based device that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0179] The modules involved in the embodiments described in the present disclosure may be implemented by software or hardware, wherein the name of the module does not limit the unit itself in some cases.
[0180] The computer-readable storage medium provided in the embodiment of the present application stores computer-readable program instructions for executing the above-mentioned data identification method, aiming to solve the technical problem of low accuracy in identifying sensitive types of data. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in the embodiment of the present application are the same as the beneficial effects of the data identification method provided in the above-mentioned embodiment, and will not be elaborated here.
[0181] An embodiment of the present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned data identification method when executed by a processor.
[0182] The computer program product provided in the embodiment of the present application is intended to solve the technical problem of low accuracy in identifying sensitive types of data. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiment of the present application are the same as the beneficial effects of the data identification method provided in the above embodiment, which will not be repeated here.
[0183] The above are only preferred embodiments of the embodiments of the present application, and are not intended to limit the patent scope of the embodiments of the present application. Any equivalent structure or equivalent process transformation made using the description and drawings of the embodiments of the present application, or directly or indirectly applied in other related technical fields, are also included in the patent processing scope of the embodiments of the present application.
Claims
1. A data identification method, characterized in that: The method includes: Acquire the data to be identified, each sensitive type of the industry to which the data to be identified belongs, and a subject word of each sensitive type, wherein the subject word of the sensitive type is determined based on sample data of the industry of the sensitive type; Inputting each sensitive type of the industry, the subject vocabulary of each sensitive type, and the data to be identified into a preset large language model, and determining the relevance of the data to be identified with each sensitive type through the preset large language model based on the subject vocabulary of each sensitive type; The sensitive type whose correlation is greater than a preset correlation threshold is used as the target sensitive type of the data to be identified.
2. The data identification method according to claim 1, characterized in that: The method further comprises: Obtaining each sensitive type determined based on a preset industry specification of the industry, and obtaining industry sample data of each of the sensitive types; For each of the sensitive types, a candidate vocabulary set is extracted from the industry sample data of the sensitive type, and based on the candidate vocabulary set, a subject vocabulary of the sensitive type is determined.
3. The data identification method according to claim 2, characterized in that: The step of extracting a candidate vocabulary set from the sensitive type industry sample data and determining the sensitive type subject vocabulary based on the candidate vocabulary set comprises: Performing word segmentation processing on the sensitive type of industry sample data, and determining a candidate vocabulary set for the industry sample data according to the word segmentation processing result; Based on the candidate vocabulary set, construct a vocabulary adjacency matrix; Calculating the resonance value of each candidate word in the candidate word set according to a preset hyperbolic tangent function, a preset temperature control adjustment parameter and the word adjacency matrix; In the candidate vocabulary set, a preset number of target vocabulary with a higher resonance value ranking are determined, and each of the target vocabulary is used as a subject vocabulary of the sensitive type, wherein the resonance value of the candidate vocabulary ranked higher is greater than the resonance value of the candidate vocabulary ranked lower.
4. The data identification method according to claim 3, characterized in that: The word segmentation processing result includes an initial vocabulary, and the industry sample data includes a plurality of sub-sample data; The step of performing word segmentation processing on the sensitive type of industry sample data and determining a candidate vocabulary set for the industry sample data according to the word segmentation processing result comprises: For each of the sub-sample data, performing word segmentation processing on the sub-sample data according to a preset word segmentation tool, and filtering out stop words in the sub-sample data to obtain an initial vocabulary of the sub-sample data; By using a preset keyword extraction algorithm and the dependency relationships corresponding to the initial words, the dependency weights of the initial words are calculated, and the dependency weights are sorted in reverse order, and the initial words with dependency weights ranked in the front preset percentile are used as candidate words for the sub-sample data; The candidate words corresponding to each of the sub-sample data are aggregated to obtain a candidate word set for the industry sample data.
5. The data identification method according to claim 3, characterized in that: The step of constructing a vocabulary adjacency matrix based on the candidate vocabulary set comprises: Determining a vocabulary vector for each candidate vocabulary in the candidate vocabulary set; For each target candidate word, based on the word vectors corresponding to each of the candidate words, calculate the word similarity between the target candidate word and each other candidate word in the candidate word set; Based on the vocabulary similarities corresponding to each of the target candidate vocabulary, a vocabulary adjacency matrix is constructed; The target candidate word is any candidate word in the candidate word set, and the other candidate words are any candidate words in the candidate word set except the target candidate word.
6. The data identification method according to claim 5, characterized in that: The step of determining the vocabulary vector of each candidate vocabulary in the candidate vocabulary set comprises: Obtain pre-trained industry vectorization models; For each of the candidate words, the candidate word is input into the industry vectorization model, and the candidate word is mapped to the vector space of the industry vectorization through the industry vectorization model to obtain the word vector of the candidate word.
7. The data identification method according to claim 3, characterized in that: The step of calculating the resonance value of each candidate word in the candidate word set according to the preset hyperbolic tangent function, the preset temperature control adjustment parameter and the word adjacency matrix comprises: For each candidate word, obtaining the similarity of each word corresponding to the candidate word in the word adjacency matrix; For each of the vocabulary similarities, calculating a sub-resonance value of the vocabulary similarity according to a preset hyperbolic tangent function and a preset temperature control adjustment parameter; For each of the candidate words, the sub-resonance values corresponding to the candidate words are accumulated to obtain the resonance value of the candidate word.
8. A data identification device, characterized in that: The device comprises: An acquisition module, used to acquire the data to be identified, each sensitive type of the industry to which the data to be identified belongs, and a subject word of each sensitive type, wherein the subject word of the sensitive type is determined based on sample data of the industry of the sensitive type; A relevance determination module, used to input each sensitive type of the industry, a subject word of each sensitive type, and the data to be identified into a preset large language model, and determine the relevance of the data to be identified with each sensitive type respectively according to the subject words of each sensitive type through the preset large language model; The target sensitive type determination module is used to take the sensitive type whose correlation is greater than a preset correlation threshold as the target sensitive type of the data to be identified.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the steps of the data identification method described in any one of claims 1 to 7.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, on which is stored a program for implementing the data identification method, and the program for implementing the data identification method is executed by a processor to implement the steps of the data identification method as described in any one of claims 1 to 7.
11. A program product, characterized in that The program product is a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of the data identification method according to any one of claims 1 to 7 are implemented.