A High-Security Hierarchical Encryption Method and System for Financial Data

By hierarchical encryption of financial data, combined with the correlation factors and differential analysis of structured and unstructured data, the problems of excessive computing resource usage and time-consuming encryption and decryption in the existing technology are solved, and efficient and secure data processing is achieved.

CN119918091BActive Publication Date: 2025-07-11SHENGYIN CONSUMER FINANCE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510397140.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-11
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

In the prior art, the structured data and unstructured data encryption methods of financial data are single, resulting in excessive computing resources occupied and a long time to encrypt and decryption process, and the use of unstructured data is limited.

Method used

Financial data is divided into structured data and unstructured data. Through vectorized processing and correlation factor analysis, hierarchical encryption is performed based on the semantic correlation and differences of user data.

Benefits of technology

实现了在保证安全性的同时,减少计算资源需求,优化加解密过程,提高数据处理效率,并提供个性化的数据保护方案。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119918091B_ABST
    Figure CN119918091B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of financial data encryption technology, and particularly relates to a high-security hierarchical encryption method and system for financial data. First, by analyzing the single correlation factor between the unstructured data and structured data of a user, and analyzing the semantic correlation factor between the unstructured data of the user and its neighborhood data, the correlation degree between the unstructured data and structured data of the user is obtained; then, by analyzing the difference degree between the associated unstructured data of the user and other users, and combining the correlation degree, the importance of the associated unstructured data of the user is obtained; finally, based on the single correlation factor and importance, the unstructured data is classified into different importance levels, and the hierarchical encryption of the unstructured data is completed. The present invention effectively reduces the occupation of computing resources while ensuring the security of financial data, improves the data processing efficiency, and realizes the refined management of data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of financial data encryption, and particularly relates to a high-security hierarchical encryption method and system for financial data. Background Art

[0002] Financial data refers to various types of data generated in financial activities, including but not limited to transaction records, user information, account balances, contract information, etc. These financial data usually have high value and high sensitivity, involving personal privacy, corporate secrets, etc. The financial industry is an information-intensive industry, and a large amount of financial data needs to be stored and managed. These financial data can be classified into three categories according to the type at the time of storage: structured data, unstructured data, and semi-structured data. Among them: structured data refers to data stored in the form of a table, and each data has a fixed field and data type; unstructured data refers to data without a fixed format and structure, such as text files, audio files, video files, etc.; semi-structured data is a data type between structured data and unstructured data, having a certain structure, but not as strict as structured data.

[0003] Due to the fact that structured data generally directly represents the amount value and various values directly related to the amount value, its importance is extremely high; at the same time, because the format of structured data is fixed and the scale is small, the computing resources required for encrypting it are also less; therefore, when encrypting financial data in the prior art, structured data is often encrypted as a whole and at the highest encryption level. The data types of unstructured data and semi-structured data are relatively diverse, and their data volume is also large, and they often contain more non-important information. If they are encrypted at a complete high level, it will consume more resources and cause the encryption and decryption processes to take a long time, thus severely restricting the use of unstructured data and semi-structured data. Summary of the Invention

[0004] In order to solve the above technical problems, the purpose of the present invention is to provide a high-security hierarchical encryption method and system for financial data.

[0005] According to the first aspect of the embodiments of the present invention, a high-security hierarchical encryption method for financial data is provided, and the technical solution adopted is specifically as follows:

[0006] Collect financial data, and divide the financial data into structured data and unstructured data;

[0007] Perform vectorization processing on the structured data and the unstructured data to obtain a structured data vector and an unstructured data vector;

[0008] Analyze a single association factor between the unstructured data vector of the current user and the structured data vector to obtain an associated unstructured data vector;

[0009] Analyze the semantic association factor between the associated unstructured data vector of the current user and its neighborhood data vector, and combine the single association factor to obtain the association degree between the associated unstructured data vector of the current user and the structured data vector;

[0010] Analyze the difference degree between the associated unstructured data vectors of the current user and other users, and combine the association degree to obtain the importance of each associated unstructured data vector of all users;

[0011] Based on the single association factor and the importance, perform an importance level division on the unstructured data vector to complete the hierarchical encryption of the unstructured data.

[0012] In some embodiments of the present invention, analyzing a single association factor between the unstructured data vector of the current user and the structured data vector to obtain an associated unstructured data vector includes:

[0013] Based on the cross-attention mechanism, analyze the first attention score between the unstructured data vector of the current user and the structured data vector to obtain a single association factor between any unstructured data vector of the current user and all structured data vectors;

[0014] Set an association factor threshold, normalize the single association factor to obtain a normalized single association factor, and the unstructured data vector corresponding to the normalized single association factor greater than the association factor threshold is the associated unstructured data vector.

[0015] In some embodiments of the present invention, analyzing the semantic association factor between the associated unstructured data vector of the current user and its neighborhood data vector includes:

[0016] Divide the unstructured data vector of the current user into a text data vector and a non-text data vector;

[0017] Divide the text data vector into modules according to punctuation marks to obtain text data modules;

[0018] Extract the characterization feature values of all the text data modules;

[0019] Based on the self-attention mechanism, obtain the second attention score between any two of the characterization feature values and the second self-attention weight of each characterization feature value;

[0020] Based on the second attention score and the second self-attention weight, obtain the third attention score between the text data vector corresponding to the associated unstructured data vector of the current user and its neighborhood data vector, and the third self-attention weight of the text data vector corresponding to the associated unstructured data vector and its neighborhood data vector;

[0021] According to the third attention score and the third self-attention weight, and in combination with the distance between the text data vector corresponding to the associated unstructured data vector and its neighborhood data vector, obtain the semantic association factor between the associated unstructured data vector and its neighborhood data vector in the text data vector of the current user.

[0022] In some embodiments of the present invention, extract the characterization feature values of all the text data modules, including:

[0023] Based on the self-attention mechanism, obtain the initial self-attention weights of all the text data vectors included in the text data module;

[0024] Take the eigenvalue of the text data vector corresponding to the maximum value of the initial self-attention weight in the text data module as the characterization feature value of the text data module.

[0025] In some embodiments of the present invention, the method further includes:

[0026] Based on the self-attention mechanism, obtain the fourth attention score between the non-text data vectors, and the fourth self-attention weight of each non-text data vector;

[0027] According to the fourth attention score and the fourth self-attention weight, and in combination with the distance between the non-text data vector corresponding to the associated unstructured data vector and its neighborhood data vector, obtain the semantic association factor between the associated unstructured data vector and its neighborhood data vector in the non-text data vector.

[0028] In some embodiments of the present invention, analyze the degree of difference between the associated unstructured data vectors of the current user and other users, including:

[0029] Analyze the matching relationship between the associated unstructured data vector of the current user and the associated unstructured data vectors of other users, and obtain the matching vector corresponding to the associated unstructured data vector of the current user in the associated unstructured data vectors of other users;

[0030] Analyze the eigenvalue difference between the associated unstructured data vector of the current user and the corresponding matching vector of other users, and obtain the degree of difference between the associated unstructured data vectors of the current user and other users.

[0031] In some embodiments of the present invention, analyzing the eigenvalue difference between the associated unstructured data vector of the current user and the matching vectors corresponding to other users to obtain the difference degree between the associated unstructured data vectors of the current user and other users includes:

[0032] Calculating the absolute value of the eigenvalue difference between the associated unstructured data vector of the current user and the matching vectors corresponding to other users;

[0033] Calculating the mean value of the absolute values of the eigenvalue differences between all the associated unstructured data vectors of the current user and the matching vectors corresponding to other users to obtain the average difference;

[0034] Combining the absolute value of the eigenvalue difference and the average difference to obtain the difference degree between the unstructured data vectors of the current user and other users.

[0035] In some embodiments of the present invention, based on the single association factor and the importance, performing an important level division on the unstructured data vector to complete the hierarchical encryption of the unstructured data, including:

[0036] Dividing the unstructured data vectors corresponding to the normalized single association factor less than or equal to the association factor threshold into the lowest important level;

[0037] Presetting an importance threshold, and performing an important level division on the associated unstructured data vector according to the importance of the unstructured data vector;

[0038] Setting different lengths of secret keys for the unstructured data vectors according to the important levels to complete the hierarchical encryption of the unstructured data.

[0039] According to the second aspect of the embodiments of the present invention, a high-security hierarchical encryption system for financial data is provided, including: a memory and a processor, wherein:

[0040] The memory is used for storing program codes;

[0041] The processor is used for reading the program codes stored in the memory and executing the method described in the first aspect of the embodiments of the present invention.

[0042] In some embodiments of the present invention, the processor includes:

[0043] A financial data acquisition module, configured to acquire financial data, divide the financial data into structured data and unstructured data; and perform feature vectorization processing on the structured data and the unstructured data to obtain structured data vectors and unstructured data vectors;

[0044] An association degree analysis module, which is used to analyze a single association factor between the unstructured data vector of the current user and the structured data vector to obtain an associated unstructured data vector; and analyze a semantic association factor between the associated unstructured data vector and its neighborhood data vector, and combine the single association factor to obtain the association degree between the associated unstructured data vector of the current user and the structured data vector;

[0045] An importance analysis module, which is used to analyze the difference degree between the associated unstructured data vectors of the current user and other users, and combine the association degree to obtain the importance of each associated unstructured data vector of all users;

[0046] An importance level classification module, which is used to classify the importance level of the unstructured data vector based on the single association factor and the importance, and complete the hierarchical encryption of the unstructured data.

[0047] Compared with the prior art, a high-security hierarchical encryption method and system for financial data provided by the present invention have the following beneficial effects:

[0048] 1. By dividing financial data into structured data and unstructured data, and adopting different encryption methods for different data types, where the unstructured data is hierarchically encrypted, the hierarchical encryption method effectively reduces the demand for computing resources while ensuring the security of financial data; and due to the optimization of the encryption and decryption processes, the delay during use is reduced, and the data processing efficiency is improved.

[0049] 2. By combining the association degree between the unstructured data and the structured data of the current user and the difference degree between the associated unstructured data vectors of the current user and other users, the importance of the unstructured data of the current user is analyzed, and the unstructured data is classified according to the importance level, realizing the refined management of financial data.

[0050] 3. By combining the single association factor between the unstructured data and the structured data of the current user and the semantic association factor within the unstructured data, the association degree between the unstructured data vector of the current user and the structured data vector is analyzed, which not only considers the association between the unstructured data and the structured data, but also takes into account the structural characteristics within the unstructured data, improves the accuracy of the association degree, and provides a reliable data basis for subsequent importance analysis.

[0051] 4. In some embodiments of the present invention, by dividing the unstructured data vector into a text data vector and a non-text data vector, and adopting different methods for analyzing the internal relevance of the unstructured data for the text data vector and the non-text data vector, wherein for the text data vector, module division is performed according to punctuation marks, and relevance analysis is performed based on the modules, the problem of ignoring the weak association relationship between paragraphs due to the strong internal association within the paragraphs can be effectively avoided.

[0052] 5. In some embodiments of the present invention, according to different importance levels of the unstructured data, keys of different lengths are set for the unstructured data, providing flexible encryption options and a more personalized data protection scheme, which can adapt to different levels of security requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0054] Figure 1 It is a schematic diagram of the basic process of a high-security hierarchical encryption method for financial data provided by an embodiment of the present invention;

[0055] Figure 2 It is a schematic diagram of the basic composition of a high-security hierarchical encryption system for financial data provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following will, in combination with the drawings and preferred embodiments, elaborate in detail on a high-security hierarchical encryption method and system for financial data proposed according to the present invention, its specific implementation manners, structures, features and effects. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. In addition, the specific features, structures or characteristics in one or more embodiments can be combined in any suitable form.

[0057] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this invention belongs. Terms such as "including", "comprising", or any other variant thereof are intended to cover non-exclusive inclusion, such that a circuit structure, article, or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of additional identical elements in the article or device including the element.

[0058] The following specifically describes the specific solution of a high-security hierarchical encryption method for financial data provided by the present invention in conjunction with the accompanying drawings.

[0059] Please refer to Figure 1 , which shows the basic process of a high-security hierarchical encryption method for financial data provided by an embodiment of the present invention.

[0060] As Figure 1 shown, a high-security hierarchical encryption method for financial data provided by an embodiment of the present invention specifically includes:

[0061] S100: Collect financial data and divide the financial data into structured data and unstructured data.

[0062] Financial data refers to various types of data generated in financial activities, including but not limited to transaction records, user information, account balances, contract information, etc. Financial storage data types refer to different ways and technologies for storing and managing data in the financial industry. Financial data can be divided into structured data, unstructured data, semi-structured data, etc. according to the storage data type, and different data types have different characteristics and uses.

[0063] Structured data refers to data stored in tabular form, and each structured data has fixed fields and data types; in the financial industry, structured data is commonly used to store user information, transaction records, financial statements, etc.; the storage methods of structured data can be relational databases, data warehouses, etc.

[0064] Unstructured data refers to data without a fixed format and structure, such as text files, audio files, video files, etc.; in the financial industry, unstructured data is commonly used to store contract files, market reviews, user feedback, etc.; the storage methods of unstructured data can be file systems, object storage, etc.

[0065] Semi-structured data is a data type that lies between structured data and unstructured data. It has a certain structure but is not as strict as structured data. In the financial industry, semi-structured data is often used to store emails, log files, reports, etc. The storage methods of semi-structured data can be document databases, SQL databases, etc.

[0066] Extract structured financial data from the database; obtain unstructured financial data from the file system or memory. For semi-structured data, it should be split into structured data and unstructured data according to structure when stored, so it is not analyzed separately in the embodiments of the present invention. That is, the financial data is divided into structured data and unstructured data.

[0067] In financial data, structured data refers to data stored in tabular form, most of which are numerical information. Structured data generally directly represents the amount value and various values directly related to the amount value. Therefore, the importance of structured data is extremely high. At the same time, because the format of structured data is fixed and the scale is small, the computing resources required for encrypting it are also less. Therefore, when encrypting financial data, structured data is often encrypted as a whole and at the highest encryption level.

[0068] The data types of unstructured data in financial data are relatively diverse and the data volume is also large, which often contains a lot of unimportant information. Therefore, performing complete high-level encryption on it will consume more resources and affect its use. In the unstructured data of financial data, due to the high similarity of financial operations, there is a strong overlap in the financial data of different users. Therefore, the importance can be judged by analyzing the information overlap (similarity) of different users. In addition, the structured data of a user can often represent the key information of that user. Therefore, the higher the similarity between the unstructured data and the structured data of a user, the stronger the importance of the unstructured data. Therefore, the importance can be judged by analyzing the similarity between the structured data and the unstructured data of a user. Specifically, it includes steps S200 to S500.

[0069] S200: Perform vectorization processing on the structured data and the unstructured data to obtain a structured data vector and an unstructured data vector.

[0070] For the convenience of subsequent data use, first perform vectorization processing on the structured data and the unstructured data to obtain a structured data vector and an unstructured data vector. The specific implementation method is:

[0071] For structured data, it is feature-vectorized through word embedding to obtain the structured data vector corresponding to the structured data, and all the structured data vectors of the same user are arranged in the order of position or time before and after to obtain the structured data vector set. , where represents the structured data vector set of the th user, represents the th structured data vector of the th user, represents the number of structured data vectors included in the th user.

[0072] For unstructured data, since unstructured data is usually in the formats of text data, video data, audio data, etc., feature extraction is first performed on the unstructured data according to existing feature extraction technologies, where:

[0073] Text data: The text is converted into a vector form through word embedding techniques (such as Word2Vec or BERT), and semantic features of key sentences and phrases in the text are extracted through an encoder, such as Transformer, to obtain the feature values of different texts;

[0074] Video data: Technologies such as Convolutional Neural Network (CNN), Region Proposal Network (RPN), and Feature Pyramid Network (FPN) are used to extract local features of the image, perform scene detection and recognition, and obtain the feature values at different moments of the video data;

[0075] Audio data: It is converted into text data through audio recognition technology, and feature extraction is performed through text processing methods. Therefore, audio data can be classified as text data type in subsequent processing.

[0076] Then, based on the extracted features of the unstructured data, the unstructured data is converted into unstructured data vectors, and all the unstructured data vectors of the same user are arranged in the order of position or time before and after to obtain the unstructured data vector set , where represents the unstructured data vector set of the th user, represents the th unstructured data vector of the th user, represents the The number of unstructured data vectors included in a user.

[0077] S300: Analyze the single correlation factor between the unstructured data vector and the structured data vector of the current user to obtain the correlated unstructured data vector.

[0078] The structured data of a user can often represent the key information of the user. The higher the similarity between the customer's unstructured data and the structured data, the stronger the importance of the unstructured data. Therefore, the importance can be judged by analyzing the similarity between the customer's structured data and unstructured data, that is, the stronger the correlation between the unstructured data and the structured data, the more important the unstructured data.

[0079] Based on the above analysis, in the embodiments of the present invention, the correlated unstructured data vector is obtained by analyzing the single correlation factor between the unstructured data vector and the structured data vector of the current user. The specific implementation method is as follows:

[0080] First, obtain the eigenvalues of each structured data vector in the structured data vector set of the user, and arrange all the eigenvalues in the order of arrangement in the structured data vector set to obtain the structured eigenvalue vector , where represents the structured eigenvalue vector of the th user, represents the th structured eigenvalue of the th user (i.e., the eigenvalue corresponding to the structured data vector ). And obtain the eigenvalues of each unstructured data vector in the unstructured data vector set of the user, and arrange all the eigenvalues in the order of arrangement in the unstructured data vector set to obtain the unstructured eigenvalue vector , where represents the unstructured eigenvalue vector of the th user, represents the th unstructured eigenvalue of the th user (i.e., the eigenvalue corresponding to the unstructured data vector

[0081] Then, based on the cross-attention mechanism, analyze the first attention score between the unstructured data vector and the structured data vector of the current user to obtain the single correlation factor between any unstructured data vector of the current user and all structured data vectors. Further elaboration is as follows: Based on the cross-attention mechanism, analyze the The unstructured eigenvalue vector of a user In the th unstructured data vector and the th structured eigenvalue vector of a user In the th structured data vector, the first attention score is obtained, and the th single association factor between the th unstructured data vector of a user and the th structured data vector is obtained.

[0082] Finally, an association factor threshold is set, which can take a value of 0.5. The single association factor is normalized by maximum and minimum values to obtain a normalized single association factor. The unstructured data vector corresponding to the normalized single association factor greater than the association factor threshold is an associated unstructured data vector.

[0083] S400: Analyze the semantic association factor between the associated unstructured data vector of the current user and its neighborhood data vector, and combine the single association factor to obtain the association degree between the associated unstructured data vector of the current user and the structured data vector.

[0084] What is obtained in step S300 is the association between single elements between the structured data and unstructured data of the same user. This association is only obtained by comparing eigenvalue vectors, and it cannot highlight the association between different data in the unstructured data and cannot accurately represent the importance of the structure in the unstructured data. It is also necessary to analyze the association of the structure of the unstructured data itself by combining the self-attention mechanism.

[0085] Based on the above analysis, in the embodiment of the present invention, by analyzing the semantic association factor between the associated unstructured data vector of the current user and its neighborhood data vector, and combining the single association factor, the association degree between the associated unstructured data vector of the current user and the structured data vector is obtained.

[0086] Among them, based on the above analysis, in the embodiment of the present invention, by analyzing the semantic association factor between the associated unstructured data vector of the current user and its neighborhood data vector, the specific implementation method is as follows:

[0087] The unstructured data vector of the current user is divided into a text data vector and a non-text data vector. For example, documents, audio, etc. are divided into text data; pictures, videos, etc. are divided into non-text data.

[0088] For text data vectors, since text data can be segmented by punctuation marks therein and the content in a paragraph is usually strongly correlated, using self-attention to analyze the correlation will cause the correlation to be more concentrated on the correlation between two punctuation marks and ignore the correlation between text segments. Therefore, in some embodiments of the present invention, the specific implementation manner of performing semantic correlation factor analysis on text data vectors is as follows:

[0089] First, divide the text data vector of the current user into modules according to punctuation marks to obtain text data modules, where each text data module contains one or more text data vectors (unstructured data vectors), and arrange all the text data modules of the current user in the order of their positions to obtain a text data module set.

[0090] Then, extract the characterization feature values of all text data modules in the text data module set of the current user. A more specific method for extracting the characterization feature values is as follows: Based on the self-attention mechanism, obtain the initial self-attention weights of all text data vectors included in the text data module; use the feature value of the text data vector corresponding to the maximum initial self-attention weight in the text data module as the characterization feature value of the text data module, obtain the characterization feature values corresponding to all text data modules, and arrange all the characterization feature values in the order in the text data module set to obtain a characterization feature value vector. It should be noted that other feature value extraction methods can also be used here as long as they can represent the feature values between local paragraphs.

[0091] Then, based on the self-attention mechanism, obtain the second attention scores between any two characterization feature values of the current user, and the second self-attention weights of each characterization feature value. Further elaboration is as follows: Based on the self-attention mechanism, obtain the second attention scores between any two characterization feature values in the characterization feature value vector of the th user, and obtain the second self-attention weights of each characterization feature value.

[0092] Then, based on the second attention scores and the second self-attention weights, obtain the third attention scores between the text data vector corresponding to the associated unstructured data vector of the current user and its neighborhood data vectors, and the third self-attention weights of the text data vector corresponding to the associated unstructured data vector and its neighborhood data vectors; where the neighborhood data vector is defined as other text data vectors whose distance from the text data vector corresponding to the associated unstructured data vector is within 32. More specifically: Denote the second attention score between the text data vector corresponding to the associated unstructured data vector of the th user and the characterization feature value of the text data module corresponding to its neighborhood data vector as the The third attention score between the text data vector corresponding to the associated unstructured data vector of the th user and its neighborhood data vector; the second self-attention weight of the representation eigenvalue of the text data module corresponding to the text data vector corresponding to the associated unstructured data vector of the th user and the text data module corresponding to its neighborhood data vector is denoted as the third self-attention weight of the text data vector corresponding to the associated unstructured data vector of the

[0093] Finally, based on the third attention score and the third self-attention weight, combined with the distance between the text data vector corresponding to the associated unstructured data vector and its neighborhood data vector, the semantic association factor between the associated unstructured data vector and its neighborhood data vector in the text data vector of the current user is obtained.

[0094] Meanwhile, for the non-text data vector, based on the self-attention mechanism, the fourth attention score between the non-text data vectors and the fourth self-attention weight of each non-text data vector are obtained; based on the fourth attention score and the fourth self-attention weight, combined with the distance between the non-text data vector corresponding to the associated unstructured data vector and its neighborhood data vector, the semantic association factor between the associated unstructured data vector and its neighborhood data vector in the non-text data vector is obtained.

[0095] Integrating the text data vector and the non-text data vector, the calculation formula for the semantic association factor between the associated unstructured data vector and its neighborhood data vector in the text data vector of the th user is as follows:

[0096]

[0097] In the formula: represents the semantic association factor between the th associated unstructured data vector and the th neighborhood data vector of the th user, where the th associated unstructured data vector is a text data vector or a non-text data vector; represents the self-attention weight (the third self-attention weight or the fourth self-attention weight) of the th associated unstructured data vector of the th user; represents the self-attention weight (the third self-attention weight or the fourth self-attention weight) of the th neighborhood data vector corresponding to the th associated unstructured data vector of the th user; Indicates the th user's th attention score (the third attention score or the fourth attention score) between the associated unstructured data vector and its th neighborhood data vector; Indicates the th user's th distance between the associated unstructured data vector and its th neighborhood data vector; Indicates the set of neighborhood data vectors of the th user's th associated unstructured data vector, that is, it represents the set of other data vectors within a distance of 32 from the th user's th associated unstructured data vector; +0.1 is to avoid the denominator of the fraction being zero, resulting in the fraction being invalid.

[0098] Attention score The larger the value, the stronger the correlation between the associated unstructured data vector and its neighborhood data vector, indicating a larger semantic correlation factor; the self-attention weight of the associated unstructured data vector The larger the value, the more important the semantics of the data vector, indicating a larger semantic correlation factor; the self-attention weight of the neighborhood data vector The smaller it is, the less important it is in semantics, and its correlation with the associated unstructured data vector will be weaker. Therefore, the attention score is adjusted by it.

[0099] After obtaining the semantic correlation factor between the associated unstructured data vector of the current user and its neighborhood data vector, combined with the single correlation factor, the correlation degree between the associated unstructured data vector of the current user and the structured data vector is obtained. Construct the th user's th associated unstructured data vector and the th structured data vector, and the calculation formula for the correlation degree is:

[0100]

[0101] In the formula: Indicates the correlation degree between the th user's th associated unstructured data vector and the th structured data vector; Indicates the th user's th associated unstructured data vector and the A single association factor between structured data vectors; Indicating the th user's th associated unstructured data vector and its th semantic association factor between the neighborhood data vectors.

[0102] Single association factor The larger the value, the greater the correlation between the associated unstructured data vector and the structured data vector, and the greater the degree of association between the associated unstructured data vector and the structured data vector; Semantic association factor The larger the value, the greater the correlation between the associated unstructured data vector and its neighborhood data vector, indicating that the structure between the associated unstructured data vector and its neighborhood data vector is stronger, and the greater the degree of association between the associated unstructured data vector and the structured data vector.

[0103] S500: Analyze the degree of difference between the associated unstructured data vectors of the current user and other users, and combine with the degree of association to obtain the importance of each associated unstructured data vector of all users.

[0104] In the unstructured data of financial data, due to the high similarity of financial operations, there is a strong overlap in the financial data of different users, and the importance can be judged by the information overlap (similarity) of different users.

[0105] Based on the above analysis, in the embodiments of the present invention, the importance of each associated unstructured data vector of all users is obtained by analyzing the degree of difference between the associated unstructured data vectors of the current user and other users and combining with the degree of association.

[0106] By analyzing the degree of difference between the associated unstructured data vectors of the current user and other users, the specific implementation method is:

[0107] First, obtain the set of unstructured data vectors of the current user and other users and .

[0108] Then, through DTW, the set of unstructured data vectors and are globally matched to obtain the matching relationship between the associated unstructured data vector of the current user and the associated unstructured data vectors of other users, and further obtain the matching vectors corresponding to the associated unstructured data vector of the current user in the associated unstructured data vectors of other users. For example, the th user's An associated unstructured data vector At the th user, the matching vector among all associated unstructured data vectors is .

[0109] Finally, analyze the eigenvalue difference between the associated unstructured data vector of the current user and the matching vectors corresponding to other users to obtain the difference degree between the associated unstructured data vectors of the current user and other users. More specifically, calculate the absolute value of the eigenvalue difference between the associated unstructured data vector of the current user and the matching vectors corresponding to other users; calculate the average value of the absolute values of the eigenvalue differences between all associated unstructured data vectors of the current user and the matching vectors corresponding to other users to obtain the average difference; combine the absolute value of the eigenvalue difference and the average difference to obtain the difference degree between the associated unstructured data vectors of the current user and other users. Construct the th user's th associated unstructured data vector and the th user's corresponding matching vector, and the difference degree calculation formula is:

[0110]

[0111] In the formula: Represents the difference degree between the th user's th associated unstructured data vector and the th user's corresponding matching vector; Represents the eigenvalue of the th user's th associated unstructured data vector ; Represents the eigenvalue of the th user's th associated unstructured data vector at the th user's corresponding matching vector is ; Represents the average difference between all associated unstructured data vectors of the th user and the th user's corresponding matching vector; Is to prevent the denominator from being 0.

[0112] Represents the absolute value of the eigenvalue difference between the associated unstructured data vector of the current user and the matching vectors corresponding to other users. The larger this value is, the greater the difference between the two data vectors; the average difference is smaller, indicating that the overall difference between the two users is smaller. Through Pairwise weighting, that is, when the overall difference between two users is small, the difference between two data vectors is relatively more prominent.

[0113] In financial data, the stronger the difference in unstructured data vectors between different users, the more the data vectors can represent the personalized information of users, and the more important the data vectors are. The stronger the correlation between the unstructured data vectors and structured data vectors of the same user, the stronger the correlation between the unstructured data vectors of the user and the core data vectors, indicating the greater importance of the unstructured data vectors.

[0114] Therefore, after obtaining the degree of difference between the associated unstructured data vectors of the current user and those of other users, further combining with the degree of association, the importance of each associated unstructured data vector of all users is obtained. The importance calculation formula for the th associated unstructured data vector of the

[0115]

[0116] th user is: Indicates the importance of the th associated unstructured data vector of the th user; Indicates the degree of association between the th associated unstructured data vector of the th user and the th structured data vector; Indicates the number of structured data vectors of the th user; Indicates the degree of difference between the th associated unstructured data vector of the th user and the matching vector corresponding to the

[0117] th user;

[0118] Indicates the number of other users; Indicates the linear normalization function.

[0117] The greater the degree of association between the associated unstructured data vector of a user and all structured data vectors, the greater the importance of the associated unstructured data vector; the greater the degree of difference between the associated unstructured data vector of a user and the matching vectors corresponding to all other users, the more the associated unstructured data vector can represent the personalized information of the user, indicating the greater importance of the associated unstructured data vector.

[0118] S600: Based on the single correlation factor and the importance, perform an importance level classification on the unstructured data vector to complete the hierarchical encryption of the unstructured data.

[0119] Based on the single correlation factor and the importance, perform an importance level classification on the unstructured data vector to complete the hierarchical encryption of the unstructured data. The specific implementation method is as follows: Since the unstructured data vector includes a correlated unstructured data vector (the unstructured data vector corresponding to the normalized single correlation factor greater than the correlation factor threshold) and an uncorrelated unstructured data vector (the unstructured data vector corresponding to the normalized single correlation factor less than or equal to the correlation factor threshold), therefore, the unstructured data vector corresponding to the normalized single correlation factor less than or equal to the correlation factor threshold is classified as the lowest importance level, that is, the unstructured data vector is classified as the lowest importance level; at the same time, for the correlated unstructured data vector, a preset importance threshold is set, and the values can be 0.5 and 0.7. When the importance of the correlated unstructured data vector is greater than or equal to 0.7, the correlated unstructured data vector is classified as a high importance level; when the importance of the correlated unstructured data vector is greater than or equal to 0.5 and less than 0.7, the correlated unstructured data vector is classified as a medium importance level; when the importance of the correlated unstructured data vector is less than 0.5, the correlated unstructured data vector is classified as a low importance level.

[0120] The importance levels of the unstructured data vectors of users are different, and the security requirements for the unstructured data of users are different. Therefore, according to the importance levels of the unstructured data vectors, keys with different lengths are set for encryption, so as to complete the hierarchical encryption of the unstructured data in the financial data.

[0121] Based on the same inventive concept as the above method, this embodiment also provides a high-security hierarchical encryption system for financial data.

[0122] Please refer to Figure 2 , which shows the basic composition of a high-security hierarchical encryption system for financial data provided by an embodiment of the present invention.

[0123] As Figure 2 shown, a high-security hierarchical encryption system for financial data includes: a memory 10 and a processor 20, where:

[0124] The memory 10 is used to store program codes;

[0125] A processor 20, configured to read the program code stored in a memory 10 and execute operations including collecting financial data, classifying the financial data into structured data and unstructured data; performing vectorization processing on the structured data and the unstructured data to obtain a structured data vector and an unstructured data vector; analyzing a single correlation factor between the unstructured data vector of the current user and the structured data vector to obtain a correlated unstructured data vector; analyzing a semantic correlation factor between the correlated unstructured data vector of the current user and its neighborhood data vector, and combining the single correlation factor to obtain the correlation degree between the correlated unstructured data vector of the current user and the structured data vector; analyzing the difference degree between the correlated unstructured data vectors of the current user and other users, and combining the correlation degree to obtain the importance of each correlated unstructured data vector of all users; and performing an importance level classification on the unstructured data vector based on the single correlation factor and the importance to complete the hierarchical encryption of the unstructured data.

[0126] Further, the processor includes: a financial data collection module 21, a correlation degree analysis module 22, an importance analysis module 23, and an importance level classification module 24, where:

[0127] The financial data collection module 21 is configured to collect financial data, classify the financial data into structured data and unstructured data; and perform feature vectorization processing on the structured data and the unstructured data to obtain a structured data vector and an unstructured data vector;

[0128] The correlation degree analysis module 22 is configured to analyze a single correlation factor between the unstructured data vector of the current user and the structured data vector to obtain a correlated unstructured data vector; and analyze a semantic correlation factor between the correlated unstructured data vector and its neighborhood data vector, and combine the single correlation factor to obtain the correlation degree between the correlated unstructured data vector of the current user and the structured data vector;

[0129] The importance analysis module 23 is configured to analyze the difference degree between the correlated unstructured data vectors of the current user and other users, and combine the correlation degree to obtain the importance of each correlated unstructured data vector of all users;

[0130] The importance level classification module 24 is configured to perform an importance level classification on the unstructured data vector based on the single correlation factor and the importance.

[0131] It should be noted that: the above sequence of the embodiments of the present invention is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0132] In the present specification, each embodiment is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key points of each embodiment are the differences from other embodiments.

Claims

1. A high-security hierarchical encryption method for financial data, characterized in that, The method includes: Collect financial data and divide the financial data into structured data and unstructured data; Perform vectorization processing on the structured data and the unstructured data to obtain a structured data vector and an unstructured data vector; Analyze a single association factor between the unstructured data vector of the current user and the structured data vector to obtain an associated unstructured data vector; Analyze a semantic association factor between the associated unstructured data vector of the current user and its neighborhood data vector, and combine the single association factor to obtain the association degree between the associated unstructured data vector of the current user and the structured data vector; Analyze the difference degree between the associated unstructured data vectors of the current user and other users, and combine the association degree to obtain the importance of each associated unstructured data vector of all users; Based on the single association factor and the importance, perform an important level division on the unstructured data vector to complete the hierarchical encryption of the unstructured data; Analyzing a single association factor between the unstructured data vector of the current user and the structured data vector includes: Based on the cross-attention mechanism, analyze the first attention score between the unstructured data vector of the current user and the structured data vector to obtain a single association factor between any unstructured data vector of the current user and all structured data vectors; Analyzing the semantic association factor between the associated unstructured data vector of the current user and its neighborhood data vector includes: Divide the unstructured data vector of the current user into a text data vector and a non-text data vector; Divide the text data vector according to punctuation marks to obtain text data modules; Extract the characterization feature values of all the text data modules; Based on the self-attention mechanism, obtain the second attention score between any two of the characterization feature values and the second self-attention weight of each characterization feature value; Based on the second attention score and the second self-attention weight, obtain the third attention score between the text data vector corresponding to the associated unstructured data vector of the current user and its neighborhood data vector, and the third self-attention weight of the text data vector corresponding to the associated unstructured data vector and its neighborhood data vector; Based on the third attention score and the third self-attention weight, and combining the distance between the text data vector corresponding to the associated unstructured data vector and its neighborhood data vector, obtain the semantic association factor between the associated unstructured data vector and its neighborhood data vector in the text data vector of the current user; construct the th user's th associated unstructured data vector and the th structured data vector, and the calculation formula for the degree of association is: In the formula: represents the degree of association between the th user's th associated unstructured data vector and the th structured data vector; represents the single association factor between the th user's th associated unstructured data vector and the th structured data vector; represents the semantic association factor between the th user's th associated unstructured data vector and its th neighborhood data vector; Analyzing the difference degree between the associated unstructured data vectors of the current user and other users includes: Analyze the matching relationship between the associated unstructured data vector of the current user and the associated unstructured data vectors of other users to obtain the matching vector corresponding to the associated unstructured data vector of the current user in the associated unstructured data vectors of other users; Analyze the eigenvalue difference between the associated unstructured data vector of the current user and the corresponding matching vector of other users to obtain the difference degree between the associated unstructured data vectors of the current user and other users; Construct the th importance calculation formula for the associated unstructured data vector of the user is: In the formula: Indicates the th importance of the associated unstructured data vector of the user; Represents the th degree of association between the associated unstructured data vector of the user and the th structured data vector; Represents the number of user structured data vectors; Represents the th degree of difference between the associated unstructured data vector of the user and the th matching vector corresponding to the user; Represents the number of other users; Represents the linear normalization function; Extracting the characterization feature values of all the text data modules includes: Based on the self-attention mechanism, obtain the initial self-attention weights of all the text data vectors included in the text data module; Use the eigenvalue of the text data vector corresponding to the maximum initial self-attention weight in the text data module as the representative eigenvalue of the text data module.

2. The high-security hierarchical encryption method for financial data according to claim 1, wherein Analyze the single correlation factor between the unstructured data vector of the current user and the structured data vector to obtain the correlated unstructured data vector, including: Set a correlation factor threshold, normalize the single correlation factor to obtain a normalized single correlation factor, and the unstructured data vector corresponding to the normalized single correlation factor greater than the correlation factor threshold is the correlated unstructured data vector.

3. The high-security hierarchical encryption method for financial data according to claim 1, characterized in that, The method further includes: Based on the self-attention mechanism, obtain the fourth attention score between the non-text data vectors and the fourth self-attention weight of each non-text data vector; According to the fourth attention score and the fourth self-attention weight, and in combination with the distance between the non-text data vector corresponding to the correlated unstructured data vector and its neighborhood data vector, obtain the semantic correlation factor between the correlated unstructured data vector and its neighborhood data vector in the non-text data vectors.

4. The high-security hierarchical encryption method for financial data according to claim 1, characterized in that, Analyze the eigenvalue difference between the correlated unstructured data vector of the current user and the matching vector corresponding to other users to obtain the difference degree between the correlated unstructured data vectors of the current user and other users, including: Calculate the absolute value of the eigenvalue difference between the correlated unstructured data vector of the current user and the matching vector corresponding to other users; Calculate the mean value of the absolute values of the eigenvalue differences between all the correlated unstructured data vectors of the current user and the matching vectors corresponding to other users to obtain the average difference; Combine the absolute value of the eigenvalue difference and the average difference to obtain the difference degree between the unstructured data vectors of the current user and other users.

5. The high-security hierarchical encryption method for financial data according to claim 2, wherein Based on the single correlation factor and the importance, perform an important level classification on the unstructured data vectors to complete the hierarchical encryption of the unstructured data, including: Classify the unstructured data vectors corresponding to the normalized single correlation factor less than or equal to the correlation factor threshold into the lowest important level; Preset an importance threshold, and perform an important level classification on the correlated unstructured data vectors according to the importance of the unstructured data vectors; According to the important level, set different lengths of secret keys for the unstructured data vectors to complete the hierarchical encryption of the unstructured data.

6. A high-security hierarchical encryption system for financial data, characterized in that, The system includes: a memory and a processor, wherein: The memory is used to store program codes; The processor is used to read the program codes stored in the memory and execute the method according to any one of claims 1 to 5.

7. The high-security hierarchical encryption system for financial data according to claim 6, characterized in that, The processor includes: A financial data collection module, which is used to collect financial data, divide the financial data into structured data and unstructured data; and perform feature vectorization processing on the structured data and the unstructured data to obtain structured data vectors and unstructured data vectors; The correlation degree analysis module is used to analyze a single correlation factor between the unstructured data vector of the current user and the structured data vector to obtain a correlated unstructured data vector; and analyze the semantic correlation factor between the correlated unstructured data vector and its neighborhood data vector, and combine the single correlation factor to obtain the correlation degree between the correlated unstructured data vector of the current user and the structured data vector; The importance analysis module is used to analyze the difference degree between the correlated unstructured data vectors of the current user and other users, and combine the correlation degree to obtain the importance of each correlated unstructured data vector of all users; The importance level division module is used to divide the importance levels of the unstructured data vectors based on the single correlation factor and the importance, and complete the hierarchical encryption of the unstructured data.

Citation Information

Patent Citations

  • Method and system for processing unstructured data

    CN103761337A

  • Data classification and grading safety protection system suitable for power industry

    CN112364377A