Vector signature method, device and system for preventing information leakage
By obtaining the word frequency characteristics and semantic characteristics of text data, cluster analysis is carried out to generate vector signatures, which solves the problem of missing semantic correlation in the prior art, and realizes efficient identification of text tampering and information protection against leakage.
Patent Information
- Application Number
- CN202510828363.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-20
AI Technical Summary
The existing text hash value generation and signature mechanism lacks semantic correlation, which causes attackers to implement collision attacks through synonyms, and cannot recognize hidden tampering at the semantic level, resulting in the risk of sensitive information leakage.
By obtaining the word frequency feature values and semantic feature distances of the text data to be signed, clustering analysis is performed, vector signatures are generated, and asymmetric encryption is used to bind vector graphics with the user's public key to improve tamper recognition capabilities.
It improves the sensitivity to identification of text tampering, reduces the risk of information leakage, and enhances the security of vector signatures.
Smart Images

Figure CN120354434A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data security technology, and in particular, to a vector signature method, device, and system for preventing information leakage. Background Art
[0002] The core value of vector signature lies in constructing a mathematical binding relationship between "content - signature". By converting an electronic signature into a vector graph with coordinate invariance and unique geometric features (such as official seals composed of Bezier curves, handwritten signature paths, etc.), the verification of document integrity can be transformed into a "feature value comparison" process: when the text content is illegally tampered with (such as deleting confidential clauses, replacing key data fields, adjusting the core semantic logic), the text feature value and the mapping information embedded in the signature will no longer match, and the system can accurately identify the risk of information leakage by comparing the feature value differences in real time. This technology plays a key role in scenarios with extremely high requirements for data security such as electronic contract signing. Taking financial contracts as an example, vector signatures can ensure the immutability of sensitive fields such as "amount", "repayment term", and "liability for breach of contract" through feature value binding, effectively preventing legal disputes and asset losses caused by signature forgery or text tampering, and becoming a basic security barrier for information leakage prevention in the digital age.
[0003] The currently widely used single - dimensional mechanism of "generating signatures based on text hash values" has a significant defect of lacking semantic relevance. Traditional hash calculations only rely on bit - by - bit matching of character sequences and do not establish a deep - level logical mapping with text semantics, resulting in attackers being able to use the characteristics of natural language to carry out "collision attacks", that is, by means of synonym replacement, constructing tampered texts with different semantics but the same character sequence hash values. In such attack scenarios, electronic signatures may not be able to recognize hidden tampering at the semantic level, allowing sensitive information to be illegally replaced or blocked, and thus posing a serious risk of information leakage. Summary of the Invention
[0004] In order to solve the above - mentioned technical problems, the purpose of this application is to provide a vector signature method, device, and system for preventing information leakage, and the specific technical solutions adopted are as follows: In the first aspect, an embodiment of this application provides a vector signature method for preventing information leakage, and the method includes the following steps: Obtain all text data to be signed; Divide the text in all text data to be signed into highly sensitive text and low-sensitive text; obtain the frequency statistics eigenvalue of each word according to the word frequency of each word in the text data to be signed where it is located and the word frequency in all text data to be signed; obtain the feature weight of a single category of sensitive text according to the ratio of the average level of the frequency statistics eigenvalues of all words in the single category of sensitive text to the sum of the average levels of the frequency statistics eigenvalues of all words in the two categories of sensitive text; obtain the word frequency eigenvalue of each word according to the frequency statistics eigenvalue of each word and the feature weight of the sensitive text of its corresponding category. Obtain the semantic feature distance and word frequency feature distance of each word respectively according to the distance between the corresponding word vectors of each word and every other word in the text data to be signed where it is located and the difference in word frequency eigenvalues, and map each word to a two-dimensional coordinate system with the semantic feature distance and word frequency feature distance as the abscissa and ordinate of each word; cluster the words in the two-dimensional coordinate system according to the distance between any two words in the two-dimensional coordinate system and the distance between their corresponding word vectors, and divide all words into multiple clusters; select some words from all text data to be signed for chunking according to the average level of the semantic feature distance and the average level of the word frequency feature distance in each cluster, and then generate a vector signature.
[0005] Preferably, the process of obtaining the frequency statistics eigenvalue of each word is as follows: use the word frequency of each word in all text data to be signed as the first eigenvalue of each word; use the word frequency of each word in the text data to be signed where it is located as the second eigenvalue of each word; use the mean value of the first eigenvalue and the second eigenvalue of each word as the frequency statistics eigenvalue of each word.
[0006] Preferably, the process of obtaining the feature weight of a single category of sensitive text is as follows: denote the mean value of the frequency statistics eigenvalues of all words in the low-sensitive text as the first feature coefficient of the low-sensitive text, and denote the mean value of the frequency statistics eigenvalues of all words in the high-sensitive text as the second feature coefficient of the high-sensitive text; calculate the sum of the first feature coefficient and the second feature coefficient, denote the ratio of the first feature coefficient to the sum value as the feature weight of the low-sensitive text, and denote the ratio of the second feature coefficient to the sum value as the feature weight of the high-sensitive text.
[0007] Preferably, the word frequency eigenvalue of each word is the product of the frequency statistics eigenvalue of each word and the feature weight of the sensitive text of its corresponding category.
[0008] Preferably, the process of obtaining the semantic feature distance and word frequency feature distance of each word is as follows: use the mean value of the DTW distances between the corresponding word vectors of each word and every other word in the text data to be signed where it is located as the semantic feature distance of each word; use the mean value of the absolute differences in word frequency eigenvalues between each word and every other word in the text data to be signed where it is located as the word frequency feature distance of each word.
[0009] Preferably, the specific process of dividing all the words into multiple clustering clusters is as follows: according to the distance between any two words in the two-dimensional coordinate system and the distance between their corresponding word vectors, obtain the similar replacement distance between the corresponding two words; take the coordinates of all the words in the two-dimensional coordinate system as the input of the clustering algorithm, and take the similar replacement distance between each pair of words as the metric distance of the clustering algorithm to obtain the clustering results of all the words; wherein, the calculation formula of the similar replacement distance between the two words is: ; in the formula, represents the th and the th similar replacement distances between words, represents the Euclidean distance between the th and the th words, represents the DTW distance between the word vectors corresponding to the th and the th words.
[0010] Preferably, the specific process of selecting some words from all the text data to be signed and sealed for chunking is as follows: Calculate the priority coefficients of each clustering cluster: ; in the formula, represents the priority coefficient of the jth clustering cluster, represents the mean value of the semantic feature distances of all the words in the jth clustering cluster, represents the mean value of the word frequency feature distances of all the words in the jth clustering cluster; Determine the selection order of the clustering clusters in the order of the priority coefficients of all the clustering clusters from large to small, and the selection order of the words within the clustering clusters is selected in the order of the product of the semantic feature distance and the word frequency feature distance corresponding to each word from large to small, and the selected words are sequentially placed into a preset number of chunks until the size of each chunk meets the preset fixed chunk size.
[0011] Preferably, the specific process of generating the vector signature is as follows: perform iterative hashing calculation on the words within the chunk after chunking to generate a hash value; implicitly embed the hash value into the vector signature template; use AES-256 symmetric encryption and user public key asymmetric encryption to encrypt the text data to be signed and sealed, bind the vector graph embedded with the hash value to the encrypted text, and generate a visual vector signature through position anchoring and transparency control.
[0012] In a second aspect, an embodiment of the present application provides a vector signature device for preventing information leakage, and the load flexible adjustment device includes: a data collection module, a semantic word frequency analysis module, and a vector signature generation module.
[0013] A data acquisition module for obtaining all text data to be signed and sealed. A semantic word frequency analysis module for obtaining the word frequency feature values of each word based on the semantic sensitivity features and word frequency features of each word. A vector signature generation module for classifying each word into multiple clustering clusters based on the semantic features and word frequency features of each word, selecting some words for chunking according to the priority order of each clustering cluster and the priority order of each word in each clustering cluster, and then generating a vector signature based on the words within the chunks after chunking.
[0014] Thirdly, an embodiment of the present application also provides a vector signature system for preventing information leakage. The system includes a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of the method for preventing information leakage of a vector signature described in any one of the above are implemented.
[0015] The present application has at least the following beneficial effects: The present application fully considers the semantic features of different texts in the text data to be signed and sealed, analyzes the word frequency features in different text data to be signed and sealed, constructs word frequency feature values, evaluates the influence of synonym replacement on the hash mapping result of the text data, and highlights the sensitivity features of different word semantic replacements; based on the analysis results of the sensitivity features, comprehensively considering the differences in semantic features and word frequency features between different texts in the text data chunking process, adjusts the text data chunking in the hash mapping process, preferentially allocates the words with greater sensitivity to the chunks. On the premise of calculating based on the round function, if synonym replacement occurs, it will have a greater impact on the hash calculation result, thereby improving the sensitivity of tampering recognition of the text data to be signed and sealed, enhancing the security of the vector signature, and reducing the risk of text leakage. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 It is a flowchart of the steps of a method for preventing information leakage of a vector signature provided by an embodiment of the present application; Figure 2 It is a schematic structural diagram of a device for preventing information leakage of a vector signature provided by an embodiment of the present application. Detailed Embodiments
[0018] To further elaborate on the technical means and effects adopted by this application to achieve the intended invention purpose, the following provides a detailed description of a vector signature method, device, and system for preventing information leakage proposed according to this application, including its specific implementation, structure, features, and effects, in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0019] Unless otherwise specified and limited, terms such as "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, such that a circuit structure, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the article or device including the said element. Additionally, the term "and / or" used herein includes any and all combinations of one or more of the related listed items. All technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs.
[0020] The following specifically describes the specific solutions of a vector signature method, device, and system for preventing information leakage provided by this application in combination with the accompanying drawings.
[0021] Please refer to Figure 1 , which shows a step flowchart of a vector signature method for preventing information leakage provided by an embodiment of this application. The method includes the following steps: Step 1: Obtain all text data to be signed.
[0022] In this application, the OA system and the electronic signature platform are docked through the API to obtain the text data to be signed, including enterprise contracts and approval documents, with a focus on collecting texts that require handwritten signatures, official seals, and seam seals; it should be noted that the enterprise contracts and approval documents are obtained under authorization for the implementation and analysis of the method in this application.
[0023] Specifically, the LTP tool is used for Chinese word segmentation, part-of-speech tagging, and named entity recognition, and word vectors are generated through Word2Vec; random characters and duplicate texts are filtered through regular expressions, the character encoding is unified to UTF-8, and format parameters such as paragraph spacing are standardized. The SimHash algorithm is used to remove semantically duplicate texts, and finally, the preprocessed word segmentation, tagging, and recognition results of each text data to be signed are obtained.
[0024] Step 2: Divide the text in all text data to be signed into highly sensitive text and low-sensitive text; obtain the word frequency statistical feature value of each word according to the word frequency of each word in the text data to be signed and the word frequency in all text data to be signed; obtain the feature weight of a single type of sensitive text according to the ratio of the average level of the word frequency statistical feature values of all words in a single type of sensitive text to the sum of the average levels of the word frequency statistical feature values of all words in two types of sensitive text; obtain the word frequency feature value of each word according to the word frequency statistical feature value of each word and the feature weight of the corresponding category of sensitive text.
[0025] To enhance the security of vector signatures, this application extracts the text data feature values and embeds them into the vector path of the vector signature to achieve information leakage prevention without affecting the appearance of the signature. Using the text data to be signed as the associated information carrier of the vector signature, if the text data is tampered with by means of synonym replacement, etc., it will directly affect the verification and recognition results. Aiming at the unstable features in the text that are vulnerable to malicious modification, by establishing a logical mapping between the signature and the text semantics, the recognition ability for semantic tampering is improved. The specific analysis and processing process is as follows: First, in the text data to be signed, this application marks all the text in the text data to be signed as two categories by the BIO annotation method, namely highly sensitive text and low-sensitive text; for example, the fixed expressions of legal provisions and industry standard terms in contract documents are classified as low-sensitive text because of their fixed expressions and low possibility of replacement and tampering; while the text with a high modification frequency such as amount and signer information is used as highly sensitive text.
[0026] Taking all the preprocessed text data to be signed as the input, use the TF-IDF algorithm to perform word frequency statistics on the word segmentation results to obtain the word frequency of each word in all the text data to be signed, which is used as the first feature value of each word to reflect the overall word frequency feature of each word in all the text data to be signed. To further refine the sensitive features in different text data, the word frequency of each word in the text data to be signed where the word is located is used as the second feature value of each word to reflect the local word frequency feature of the text data to be signed where the word is located. For each word in a single text data to be signed, the average value of the first feature value and the second feature value of each word is used as the word frequency statistical feature value of each word. The larger the word frequency statistical feature value, the more significant the word frequency feature of the corresponding word in the current text and historical texts, and the smaller the possibility of synonym replacement modification of the word; on the contrary, the greater the possibility of synonym replacement modification of the word.
[0027] Based on the above calculation results, the mean value of the word frequency statistical feature values of all words in the low-sensitivity text is denoted as the first feature coefficient of the low-sensitivity text, and the mean value of the word frequency statistical feature values of all words in the high-sensitivity text is denoted as the second feature coefficient of the high-sensitivity text. To highlight the difference in word frequency features between the two types of texts and reduce the interference of synonym replacement on sensitive feature analysis during the block division process, calculate the sum of the first feature coefficient and the second feature coefficient, and denote the ratio of the first feature coefficient to the sum value as the feature weight of the low-sensitivity text, and denote the ratio of the second feature coefficient to the sum value as the feature weight of the high-sensitivity text.
[0028] Based on the weights determined above, perform weighted adjustment on the word frequency features affected by synonym replacement for each word; specifically, calculate the product of the word frequency statistical feature value corresponding to each word and the feature weight of the sensitive text of its corresponding category, and use the product as the word frequency feature value of each word; the larger the word frequency feature value, the greater the impact of the synonym replacement of the corresponding word vector under the hash mapping of the text data to be signed, that is, according to the comprehensive analysis of semantic features and word frequency features, the greater the possibility of its synonym replacement affecting the hash calculation result.
[0029] Step 3: According to the distance between the corresponding word vectors of each word and each other word in the text data to be signed and the difference in word frequency feature values, obtain the semantic feature distance and word frequency feature distance of each word respectively, and map each word to a two-dimensional coordinate system with them as the horizontal and vertical coordinates of each word; according to the distance between any two words in the two-dimensional coordinate system and the distance between their corresponding word vectors, cluster the words in the two-dimensional coordinate system, and divide all words into multiple clustering clusters; according to the average level of the semantic feature distance and the average level of the word frequency feature distance of all words in each clustering cluster, select some words from all the text data to be signed for block division, and then generate a vector signature.
[0030] Further, in the process of forming a hash value by performing a hash mapping on the text data to be signed, compared with the words with relatively large differences in semantic features and word frequency features in the overall text data, the greater the sensitivity of the text recognition is affected by synonym replacement. If synonym replacement occurs, considering the differences in word frequency features and semantic features, the replaced word will have relatively large feature changes compared with the overall text. Therefore, in this application, to highlight the sensitivity characteristics of different words to the hash mapping change of the text data to be signed, the DTW distance between the word vectors corresponding to each word and each other word in the text data to be signed where it is located is calculated, and the mean value of all DTW distances is used as the semantic feature distance of each word. And the absolute difference between the word frequency feature values of each word and each other word in the text data to be signed where it is located is calculated, and the mean value of all the absolute values is used as the word frequency feature distance of each word. The greater the semantic feature distance and the word frequency feature distance, the greater the sensitivity of the semantic feature change of the current word in the text data to be signed, and then the more likely the word is to be tampered with due to synonym replacement.
[0031] Construct a two-dimensional coordinate system with the above semantic feature distance and word frequency feature distance as the abscissa and ordinate respectively, and map all words to the constructed two-dimensional coordinate system according to the semantic feature distance and word frequency feature distance of all words. To accurately reflect the semantic similarity between different words, as well as the differences in semantic features and word frequency features between different words compared with the overall text data, in this embodiment, the similarity replacement distance of different words is calculated based on the mapping result. The specific calculation formula is: ; where represents the similarity replacement distance between the -th and the -th words, represents the Euclidean distance between the -th and the -th words, represents the DTW distance between the word vectors corresponding to the -th and the -th words.
[0032] If the calculated similarity replacement distance is larger, it indicates that the semantics between different words are more similar, and are more similar to the semantic features and word frequency features of the overall text data to be signed. Therefore, the mapping results of all words in the two-dimensional coordinate system are used as the input of the CURE (Clustering using representative) clustering algorithm, the similarity replacement distance between each word is used as the metric distance, and the size of the contraction factor is set to 0.5 to obtain the clustering results of all words. In this application, the similarity replacement distance between words is used as the metric distance in the clustering division process to reduce the risk that an attacker performs a synonymous replacement on the text to be signed and then tampers with the text. The specific clustering division process of the CURE algorithm is well-known to those skilled in the art, and the detailed process will not be elaborated here.
[0033] Furthermore, to further highlight the sensitive features of text content replacement in the text hashing mapping process, in this embodiment, based on the above clustering division results, a comprehensive analysis is performed on the quantity and word frequency features between different clustering clusters. If the relative semantic features and relative word frequency features of the words in the clustering cluster after division are more significant, then the words in this clustering cluster should be preferentially selected when performing word chunking.
[0034] Specifically, calculate the priority coefficient of each clustering cluster, and its specific calculation formula is: ; in the formula, represents the priority coefficient of the jth clustering cluster, represents the mean value of the semantic feature distances of all words in the jth clustering cluster, represents the mean value of the word frequency feature distances of all words in the jth clustering cluster. That is, when the semantic feature differences and word frequency feature differences of the corresponding words in the jth clustering cluster are larger, it indicates that the words in this clustering cluster are more likely to be tampered with due to synonymous replacement. Then, more attention should be paid to the words in this clustering cluster during the hashing mapping process.
[0035] Hashing mapping performs the calculation of the round function based on data chunking, that is, the calculation result of the previous chunk result is used as the input vector for the calculation of the next chunk. Therefore, in this application, to improve the sensitivity of identifying synonymous replacements in the hashing mapping process, based on the priority coefficients of all clustering clusters, the selection order of the clustering clusters is determined in descending order of the priority coefficients. The selection order of the words within the clustering cluster is selected in descending order of the product of the semantic feature distance and the word frequency feature distance corresponding to each word, and the selected words are sequentially placed into a preset number of chunks until the size of each chunk meets the preset fixed chunk size.
[0036] It should be noted that the words that have a greater impact on the sensitivity of the hash mapping are used as the priority selection data for chunking. The purpose is to preferentially allocate the words with greater sensitivity to the chunks. On the premise of calculating based on the round function, when synonym replacement occurs during the hash calculation process, it will have a greater impact on the hash calculation result, thereby improving the sensitivity of tampering recognition for the text data to be signed.
[0037] In this embodiment of the hash mapping algorithm, the SHA3-512 algorithm is used for hash mapping calculation. Its fixed chunk size is 512 bit, and the size of the output hash value is 512 bit. First, the SHA3-512 algorithm is used to perform iterative hash calculation on the words within the chunks after chunking to generate a 512-bit hash value, that is, the hash value. Its mapping process enhances the semantic tampering detection and recognition ability by fusing the lexical sensitivity feature values; subsequently, the hash value is converted into a 128-bit floating-point number sequence, and the hash value is implicitly embedded by fine-tuning the coordinate parameters of the vector signature template path (such as the last four digits after the decimal point of the control point coordinates of the Bezier curve). For example, every 4 bits of the hash value are mapped to the decimal number of the coordinate suffix (for example, the hash value "a1b2" is converted to "0.1234"), and at the same time, non-visible paths (such as the seal auxiliary line) are used for redundant storage to avoid affecting the visual effect; finally, combining AES-256 symmetric encryption and user public key asymmetric encryption, the text data to be signed is encrypted, the vector graph embedded with the hash value is bound to the encrypted text, and the visual vector signature is generated through position anchoring and transparency control to ensure the concealment of the hash value embedding and the integrity of the signature appearance.
[0038] Please refer to Figure 2 , Figure 2 is a schematic structural diagram of a vector signature device for preventing information leakage provided by an embodiment of the present application. In this embodiment, each unit included in the terminal is used to execute each step in the corresponding embodiment of a vector signature method for preventing information leakage. Refer to Figure 2 , the vector signature device includes: a data acquisition module, a semantic word frequency analysis module, and a vector signature generation module.
[0039] The data acquisition module is used to obtain all the text data to be signed; The semantic word frequency analysis module is used to obtain the word frequency feature values of each word based on the semantic sensitivity features and word frequency features of each word; The vector signature generation module is used to divide each word into multiple clustering clusters based on the semantic features and word frequency features of each word, and select some words for chunking according to the priority order of each clustering cluster and the priority order of each word in each clustering cluster, and then generate a vector signature according to the words within the chunks after chunking.
[0040] Based on the same inventive concept as the above method, an embodiment of the present application further provides a vector signature system for preventing information leakage, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of any one of the above-mentioned vector signature methods for preventing information leakage are implemented.
[0041] The embodiments in the present application are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments.
[0042] It should be noted that unless otherwise specified and limited, terms such as "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a circuit structure, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the article or device including the element. In addition, the term "and / or" used herein includes any and all combinations of one or more of the related listed items.
[0043] Those skilled in the art will readily think of other implementation schemes of the present application after considering the specification and practicing the invention herein. The present application is intended to cover any variations, uses or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not invented by the present application.
[0044] It should be understood that the present application is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope.
Claims
1. A vector signature method for preventing information leakage, characterized in that, The method includes the following steps: Obtain all text data to be signed; Divide the text in all text data to be signed into highly sensitive text and low-sensitive text; According to the word frequency of each vocabulary in the text data to be signed where it is located and the word frequency in all text data to be signed, obtain the word frequency statistical characteristic value of each vocabulary; According to the ratio of the average level of the word frequency statistical characteristic values of all vocabularies in a single type of sensitive text to the sum of the average levels of the word frequency statistical characteristic values of all vocabularies in two types of sensitive text, obtain the characteristic weight of the single type of sensitive text; According to the word frequency statistical characteristic value of each vocabulary and the characteristic weight of the sensitive text of its corresponding category, obtain the word frequency characteristic value of each vocabulary. According to the distance between the corresponding word vectors of each vocabulary and every other vocabulary in the text data to be signed where it is located and the difference in word frequency characteristic values, respectively obtain the semantic characteristic distance and word frequency characteristic distance of each vocabulary, and use them as the horizontal and vertical coordinates of each vocabulary to map each vocabulary into a two-dimensional coordinate system; According to the distance between any two vocabularies in the two-dimensional coordinate system and the distance between their corresponding word vectors, cluster the vocabularies in the two-dimensional coordinate system, and divide all vocabularies into multiple clustering clusters; According to the average level of the semantic characteristic distance and the average level of the word frequency characteristic distance of all vocabularies in each clustering cluster, select some vocabularies from all text data to be signed for chunking, and then generate a vector signature.
2. The vector signature method for preventing information leakage according to claim 1, wherein The process of obtaining the word frequency statistical characteristic value of each vocabulary is as follows: Take the word frequency of each vocabulary in all text data to be signed as the first characteristic value of each vocabulary; Take the word frequency of each vocabulary in the text data to be signed where it is located as the second characteristic value of each vocabulary; Take the mean of the first characteristic value and the second characteristic value of each vocabulary as the word frequency statistical characteristic value of each vocabulary.
3. A method for preventing information leakage in vector signature as claimed in claim 1, characterized in that The process of obtaining the characteristic weight of a single type of sensitive text is as follows: Denote the mean of the word frequency statistical characteristic values of all vocabularies in the low-sensitive text as the first characteristic coefficient of the low-sensitive text, and denote the mean of the word frequency statistical characteristic values of all vocabularies in the high-sensitive text as the second characteristic coefficient of the high-sensitive text; Calculate the sum of the first characteristic coefficient and the second characteristic coefficient, and denote the ratio of the first characteristic coefficient to the sum value as the characteristic weight of the low-sensitive text, and denote the ratio of the second characteristic coefficient to the sum value as the characteristic weight of the high-sensitive text.
4. A vector signature method for preventing information leakage according to claim 1, characterized in that, The word frequency characteristic value of each vocabulary is the product of the word frequency statistical characteristic value of each vocabulary and the characteristic weight of the sensitive text of its corresponding category.
5. A vector signature method for preventing information leakage according to claim 1, characterized in that The processes of obtaining the semantic characteristic distance and word frequency characteristic distance of each vocabulary are as follows: Take the mean of the DTW distances between the corresponding word vectors of each vocabulary and every other vocabulary in the text data to be signed where it is located as the semantic characteristic distance of each vocabulary; Take the mean of the absolute differences in word frequency characteristic values between each vocabulary and every other vocabulary in the text data to be signed where it is located as the word frequency characteristic distance of each vocabulary.
6. The vector signature method for preventing information leakage according to claim 1, characterized in that, The specific process of dividing all words into multiple clustering clusters is as follows: According to the distance between any two words in the two-dimensional coordinate system and the distance between their corresponding word vectors, obtain the similarity replacement distance between the corresponding two words; Use the coordinates of all words in the two-dimensional coordinate system as the input of the clustering algorithm, and use the similarity replacement distance between each word as the metric distance of the clustering algorithm to obtain the clustering result of all words; Among them, the calculation formula for the similarity replacement distance between the two words is: ; In the formula, represents the -th and the -th similarity replacement distance between words, represents the Euclidean distance between the -th and the -th words, represents the DTW distance between the word vectors corresponding to the -th and the -th words.
7. A vector signature method for preventing information leakage according to claim 1, characterized in that, The specific process of selecting some vocabularies from all text data to be signed for chunking is as follows: Calculate the priority coefficient of each clustering cluster: ; In the formula, represents the priority coefficient of the j-th clustering cluster, represents the mean of the semantic feature distances of all words in the j-th clustering cluster, represents the mean of the word frequency feature distances of all words in the j-th clustering cluster; Determine the selection order of clustering clusters in descending order of the priority coefficients of all clustering clusters. The selection order of the words within a clustering cluster is determined in descending order of the product of the semantic feature distance and the word frequency feature distance corresponding to each word. The selected words are sequentially placed into a preset number of blocks until the size of each block meets the preset fixed block size.
8. A vector signature method for preventing information leakage according to claim 1, characterized in that, The specific process of generating the vector signature is as follows: perform iterative hashing calculation on the words within the block after segmentation to generate a hash value; implicitly embed the hash value into the vector signature template; use AES-256 symmetric encryption and user public key asymmetric encryption to encrypt the text data to be signed, bind the vector graph with the embedded hash value to the encrypted text, and generate a visual vector signature through position anchoring and transparency control.
9. A vector signature device for preventing information leakage, characterized in that, Implement a vector signature method for preventing information leakage as described in any one of claims 1-8. The vector signature device includes: A data acquisition module for obtaining all text data to be signed; A semantic word frequency analysis module for obtaining the word frequency feature values of each word based on the semantic sensitivity features and word frequency features of each word; A vector signature generation module for dividing each word into multiple clustering clusters based on the semantic features and word frequency features of each word, selecting some words for segmentation according to the priority order of each clustering cluster and the priority order of each word in each clustering cluster, and then generating a vector signature based on the words within the block after segmentation.
10. A vector signature system for preventing information leakage, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements a vector signature method for preventing information leakage as described in any one of claims 1-8.
Citation Information
Patent Citations
Label generation method, system and device based on semantic similarity model and medium
CN114443850A
Multi-feature fusion double-target self-supervision medical problem text clustering method and system
CN116543406A
Image generation method, system and device based on text data
CN119049051A
Text classification method, apparatus and device, and computer-readable storage medium
WO2020207167A1