Data conformal encryption parallel processing transmission method
By identifying sensitive fields and conformal encryption of encrypted data, combining distributed scheduling of logical sharding and hash digest, the parallel processing and distributed transmission of multi-field conformal encryption in large-scale concurrent scenarios is solved, and efficient and stable data transmission and consistency verification are achieved.
Patent Information
- Application Number
- CN202510823173.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2045-06-19
AI Technical Summary
The prior art is difficult to realize parallel processing and distributed efficient transmission of multi-field conformal encryption in large-scale concurrency scenarios, and fails to effectively handle field hierarchical sensitivity differences, the coupling mechanism between encrypted ciphertext logical structure and shard scheduling, resulting in delays and disorders of data reconstruction and consistency verification.
Through sensitive field recognition and conformal structure analysis, the field is encrypted using a conformal encryption algorithm, and combined with logical sharding and hash digest, the encrypted ciphertext shard is sent to multiple nodes using a distributed scheduling strategy, supporting shard reorganization and consistency verification.
It significantly improves the transmission efficiency and end-to-end integrity guarantee capabilities after data conformal encryption, reduces transmission delay and risk of disorder, and improves the parallelism and stability of the system.
Smart Images

Figure CN120434035A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of encrypted transmission, and in particular to a data shape-preserving encrypted parallel processing and transmission method. Background Art
[0002] With the rapid development of data-driven businesses, especially in data-sensitive scenarios such as finance, healthcare, and e-commerce, privacy protection and efficient transmission have become key issues in information system design. Format-Preserving Encryption (FPE), because it can encrypt data while preserving its original structure (such as length and format), has been widely used to encrypt sensitive data such as database fields, log information, and transaction records. However, with the continued growth of data volumes and the increasing demand for real-time processing, ensuring the security of format-preserving encryption while supporting parallel encryption processing and distributed, efficient transmission has become a key bottleneck in current research and engineering implementation.
[0003] In recent years, scholars and practitioners at home and abroad have proposed various solutions to the encryption and transmission of structured sensitive data. In terms of conformal encryption, the FF1 and FF3 algorithms proposed by NIST have become mainstream implementations in the industry. Their core principle is to preserve the format and semantic boundaries of the original data based on the Feistel network, and they are widely used to encrypt and protect financial fields such as bank card numbers and ID card numbers. However, traditional conformal encryption algorithms primarily target static data processing and lack systematic support for multi-field conformal coupled encryption, structural perturbation control, and coordinated transmission of heterogeneous nodes in large-scale concurrent scenarios. Existing research, such as patent CN108881257B, "A Distributed Search Cluster Encrypted Transmission Method and Encrypted Transmission Distributed Search Cluster," proposes an encrypted transmission method for distributed search scenarios. This method obtains distributed cluster node attribute information, generates node certificates by the management node, and implements encrypted data transmission based on the node certificates, effectively improving identity authentication and security during distributed transmission. However, this solution fails to conduct a fine-grained analysis of the data structure and ignores the impact of field-level differences on transmission efficiency, decryption order, and data reconstruction accuracy.
[0004] In addition, traditional methods have the following problems: they do not model the sensitivity differences at the field level in the data structure, making it difficult to meet the requirements of conformal encryption under structure-awareness; they do not consider the coupling mechanism between the logical structure of the encrypted ciphertext and shard scheduling, resulting in delays and confusion risks during data reconstruction and consistency verification.
[0005] Therefore, there is an urgent need for a data conformal encryption parallel processing and transmission method that integrates field sensitive identification, conformal encryption, parallel scheduling transmission and integrity verification mechanism to enhance the structure preservation, processing parallelism and transmission stability of encrypted data in a distributed environment. Summary of the Invention
[0006] In view of this, the present invention provides a data conformal encryption parallel processing and transmission method, which maintains the original field format characteristics through conformal structure analysis and encryption operations, and at the same time constructs a logical sharding mechanism based on position and hash and a hash summary generation method, and uses a parallel transmission method to send encrypted ciphertext shards to multiple nodes. This method supports shard reorganization and consistency verification, significantly enhancing the transmission efficiency and end-to-end integrity assurance capabilities of data after conformal encryption.
[0007] To achieve the above objectives, the present invention provides a data conformal encryption parallel processing and transmission method, comprising the following steps: S1: Identify sensitive fields in the data to be encrypted, obtain the sensitive fields in the data to be encrypted, and perform conformal structure analysis on the sensitive fields to obtain the field boundaries and conformal encryption factors of the sensitive fields; S2: Extract the encrypted fields corresponding to the sensitive fields based on the field boundaries, use the conformal encryption algorithm combined with the conformal encryption factor to conformally encrypt the fields to be encrypted, and construct the encrypted ciphertext of the data to be encrypted based on the conformal encryption results of the fields to be encrypted; S3: Logically shard the encrypted ciphertext to obtain shard data of the encrypted ciphertext, along with hash digest and shard verification information. The shard data with the accompanying information is sent to the distributed nodes using a distributed scheduling strategy. S4: The distributed node uploads the shard data with accompanying information to the receiving end. The receiving end extracts the logical shard order from the shard verification information attached to the shard data, reorganizes and decrypts the shard data with accompanying information, and compares the hash summary to verify data integrity and consistency.
[0008] Optionally, sensitive fields of the encrypted data are identified, including: Obtain the user's data to be encrypted, use the pre-built field dictionary to identify and separate the fields in the data to be encrypted, and obtain the field sequence of the data to be encrypted; Use the rule engine and regular expressions to build a rule matching pattern, match the fields in the field sequence, and extract the successfully matched fields as candidate fields; The context fields of the candidate fields are extracted to form a candidate field sequence. The word vector model is used to encode the fields in the candidate field sequence. The lightweight sensitive field identification module is used to identify sensitive fields in the candidate field sequence after word vector encoding. The probability that the candidate field in the candidate field sequence is a sensitive field is obtained, and the candidate field with a probability higher than the preset sensitive probability threshold is regarded as a sensitive field.
[0009] Optionally, performing conformal structure analysis on the sensitive field includes: The conformal structure analysis includes identifying field boundaries and generating conformal encryption factors; Extract the location of the sensitive field, initially construct a boundary window containing the sensitive field, calculate the information entropy of each character in the boundary window, calculate the sensitive field boundary window score of the boundary window based on the information entropy, and fine-tune the boundary window until the boundary window length is less than the preset window length threshold. The boundary window containing the sensitive field and with the highest sensitive field boundary window score is used as the sensitive boundary window of the sensitive field, and the positions at both ends of the sensitive boundary window are used as the field boundaries of the sensitive field. The sensitive field boundary window score calculation formula of the boundary window is: ; in, Indicates that sensitive fields are included The bounding window, Represents the bounding window Sensitive field boundary window score, Indicates the sensitive fields output by the lightweight sensitive field identification module is the probability of sensitive fields, Represents the bounding window Information entropy of the frequency distribution of the inner field; The formula for generating the conformal encryption factor of the sensitive field is: ; in, represents the conformal encryption factor of the sensitive field y, Indicates the ID of the user associated with the data to be encrypted. The timestamp representing the generation of the conformal encryption factor of the sensitive field y, Indicates splicing processing, Represents a hash operation.
[0010] Optionally, conformal encryption is performed on the field to be encrypted using a conformal encryption algorithm combined with a conformal encryption factor, including: Extract the character sequence of the sensitive field in the corresponding sensitive boundary window as the field to be encrypted; Using a word vector model to perform word vector encoding on the data to be encrypted, obtaining a word vector data sequence corresponding to the data to be encrypted, and extracting the word vector encoding result of the field to be encrypted; The word vector encoding result of the field to be encrypted is divided into a left half and a right half, and the word vector encoding result is round-encrypted using a Feistel encryption round. After the round encryption is completed, the left half and the right half of the last round of encryption output are spliced together to obtain the conformal encryption result of the field to be encrypted; The conformal encryption result of the field to be encrypted is used to replace the word vector encoding result of the field to be encrypted in the word vector data sequence to obtain the encrypted ciphertext corresponding to the data to be encrypted.
[0011] Optionally, logically sharding the encrypted ciphertext to obtain shard data of the encrypted ciphertext includes: Identify the conformal encryption result of the field to be encrypted in the encrypted ciphertext, add a binary fragmentation mark code, a conformal encryption result length code and a fragmentation sequence number code before the conformal encryption result, use the character position before the fragmentation mark code and the fragmentation sequence number code as the logical fragmentation position, logically fragment the encrypted ciphertext, and obtain the fragmentation data of the encrypted ciphertext.
[0012] Optionally, the shard data is accompanied by a hash digest and shard verification information, including: Performing a hash digest operation on the shard data to obtain a hash digest of the shard data; Using the fragment mark code, the conformal encryption result length code and the fragment sequence number code in the fragment data as fragment verification information; The hash summary is attached to the fragment data to obtain the fragment data with the hash summary and fragment verification information.
[0013] Optionally, a distributed scheduling strategy is used to send the sharded data with accompanying information to distributed nodes, including: The fragmented data with accompanying information is fragmented data with accompanying hash summary and fragment verification information; Calculate the scheduling score between the fragmented data of the accompanying information and the distributed nodes in turn, and send the fragmented data of the accompanying information to the distributed node with the highest scheduling score as the distributed scheduling strategy of the fragmented data of the accompanying information, where the fragmented data of the accompanying information and the distributed nodes are The scheduling score between them is: , ; ; ; ; in, Shard data and distributed nodes that represent accompanying information The scheduling score value between Indicates the nth distributed node, N indicates the total number of distributed nodes, represents the scheduling weight coefficient, Represents a distributed node The load factor, The shard data data that represents the accompanying information is distributed on the node The space capacity ratio in Shard data and distributed nodes that represent accompanying information Communication quality of the communication link between them; Both represent the scheduling resource weights, Represents a distributed node The transmission speed of the fragmented data to the receiving end, Indicates the preset maximum transmission speed. Represents a distributed node The number of fragmented data to be sent; The number of bits of the fragment data that represents the accompanying information, Represents a distributed node The number of bits of space available; Shard data and distributed nodes representing the calculation of incidental information When scheduling the score value between the shard data sender and the distributed node The transmission delay between Indicates the preset maximum transmission delay.
[0014] Optionally, the receiving end extracts the logical sharding order from the sharding checksum information included with the sharding data, reassembles and decrypts the sharding data, and compares the hash digest to verify data integrity and consistency, including: The receiving end receives the fragmented data of the accompanying information, extracts the logical fragmentation order from the fragmentation verification information, and sorts the fragmented data of the accompanying information according to the logical fragmentation order; Extract the hash summary from the shard data with the accompanying information, remove the hash summary and shard verification information from the shard data with the accompanying information, and recalculate the hash summary of the shard data. If the recalculated result is consistent with the extracted hash summary, it means that the data integrity and consistency verification has passed, and the shard data is decrypted.
[0015] Compared with the prior art, the present invention has the following beneficial effects: First, by calculating the feature information of candidate fields in multiple dimensions, this application systematically solves the core problems of traditional sensitive field identification methods in complex data environments, such as poor accuracy, high missed recognition rate, and insufficient semantic understanding.
[0016] At the same time, the distributed scheduling strategy proposed in the present invention introduces a composite scoring mechanism through a scheduling scoring function, comprehensively processes the three core dimensions of computing load, storage capacity and network delay, and uses the load factor to dynamically disperse the scheduling pressure of high-load nodes, thereby improving the overall response speed of the system. The spatial capacity ratio is used to reflect the system's adaptability to the actual resource carrying capacity of the node. This design enables sharded data to be scheduled to nodes with sufficient spatial resources first, reducing the sharding failure rate and the secondary scheduling overhead caused by insufficient resources. By reversely measuring the communication delay, the system is more inclined to select low-latency link nodes for scheduling, reducing transmission congestion and improving the real-time transmission guarantee capability of sharded data. It is particularly suitable for scheduling tasks in time-sensitive or frequently interactive scenarios, and improves the parallelism of sharded data transmission. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A flowchart of a data conformal encryption and parallel processing and transmission method provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0018] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0019] The present invention provides a method for parallel processing and transmission of data with conformal encryption. The method can be executed by at least one of a server, a terminal, or other electronic device capable of executing the method provided by the present invention. In other words, the method can be executed by software or hardware installed on a terminal or server device, where the software can be a blockchain platform. The server can include, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.
[0020] Reference Figure 1 , embodiment 1 of the present invention is: A data shape-preserving encryption parallel processing and transmission method comprises the following steps: S1: Identify sensitive fields in the data to be encrypted, obtain the sensitive fields in the data to be encrypted, and perform conformal structure analysis on the sensitive fields to obtain the field boundaries and conformal encryption factors of the sensitive fields.
[0021] Identify sensitive fields in encrypted data, including: Obtain the user's data to be encrypted, use the pre-built field dictionary to identify and separate the fields in the data to be encrypted, and obtain the field sequence of the data to be encrypted; A rule matching pattern is constructed using a rule engine and regular expressions to match fields in a field sequence, and successfully matched fields are extracted as candidate fields. Specifically, the rule matching pattern includes multiple sets of matching rules, each set of matching rules includes field length restrictions, field type restrictions, and typical positions, where field types include numeric fields, Chinese character fields, special symbol characters, and English character fields. Fields in the field sequence are input into the rule matching pattern, and the matching confidence between the fields and the matching rules in the rule matching pattern is calculated. If the matching confidence is higher than a preset confidence threshold, it indicates a successful match. The formula for calculating the matching confidence between field Y and matching rule rule is: ; in, Indicates the matching confidence between field Y and matching rule rule. Indicates the field length of field Y, The fields length limit, field type limit, and typical position of the matching rule are represented in sequence. The typical position represents the preferred field position of the sensitive field associated with the matching rule. Indicates field Y and field type restrictions The fit, Indicates field Y and typical position The fit, Indicates that field Y meets the field type restrictions The number of characters, Indicates the position number of field Y in the field sequence. Represents position control term, exp( ) represents an exponential function with a natural constant as its base.
[0022] As an embodiment of the present invention, the matching rules include identification card number matching rules, mobile phone number matching rules, and email address matching rules, among which the identification card number matching rules are: field type: numeric field + special English characters (can be at the end of X / x), field length limit: 18 bits; the mobile phone number matching rules are: field type: numeric field + starting with 1 + the second digit is 3 to 9, field length limit: 11 bits; the email address matching rules are: field type: English character field + special symbol field, field length limit: ≥6 and ≤50; Extract the context fields of the candidate fields to form a candidate field sequence. Use the word vector model to encode the fields in the candidate field sequence. Combined with the lightweight sensitive field identification module, perform sensitive field identification on the candidate field sequence after word vector encoding. Obtain the probability that the candidate field in the candidate field sequence is a sensitive field, and select the candidate field with a probability higher than the preset sensitive probability threshold as a sensitive field. Specifically, the sensitive field identification formula is: ; in, Indicates candidate fields is the probability of sensitive fields, Represents the candidate fields output by the lightweight sensitive field identification module is the context-aware probability of the sensitive field, Indicates candidate fields The regularized confidence score of Indicates candidate fields The contextual semantic similarity of Both represent probability weights; Indicates candidate fields The corresponding candidate field sequence after word vector encoding, is the weight matrix parameter, is the bias parameter, is the Sigmoid activation function, ReLU represents the activation function, i.e. the rectified linear unit; Indicates candidate fields The confidence level of the match with the matching rule. Indicates rule matching mode, max{ } means taking the maximum value of the set, Represents taking a set The maximum value in ; Indicates candidate fields The word vector encoding result of the context field, Word vector encoding results representing common sensitive words; express The maximum cosine similarity between them, the common sensitive words include ID card number, bank card, etc.
[0023] It should be noted that the context-aware probability of the present invention can capture the implicit semantic connection between contexts from the candidate field sequence, especially when the field name is missing and the field structure is vague, it can still maintain a high recognition accuracy, making up for the limitations of regular matching means in processing fuzzy data. Regular confidence scoring does not simply determine whether the rule is hit, but introduces factors such as format integrity and character type conformity to quantify the matching quality of the field with the predefined format, thereby supporting reasonable scoring when the field is truncated, nested or slightly deformed. Furthermore, the typical position matching degree uses a Gaussian position window function to model the expected position of the field in the structural domain, so that the system can learn and preferentially identify rules such as the identity card number usually appearing in the second field, which significantly improves the sensitivity to field positions in structured or semi-structured data. Context semantic similarity extracts whether the context field contains sensitive words such as identity card and bank, and models them through word vector similarity to further confirm the sensitivity of the field at the semantic level, solving the situation where the field structure is insensitive but the context has strong prompts.
[0024] This application systematically solves the core problems of traditional sensitive field identification methods in complex data environments, such as poor accuracy, high missed recognition rate, and insufficient semantic understanding, by calculating the feature information of candidate fields in multiple dimensions. The sensitive field identification formula not only integrates the probability estimation of sensitive fields based on contextual semantics, but also introduces format fidelity evaluation (such as regular rule matching score, character type consistency), typical position matching confidence evaluation (such as position window Gaussian weighted function), and semantic consistency analysis of field names and their context words, thereby constructing a highly robust and compatible sensitive field identification formula. The sensitive fields are subjected to conformal structural analysis, including: The conformal structure analysis includes identifying field boundaries and generating conformal encryption factors; Extract the position of the sensitive field, initially construct a boundary window containing the sensitive field, calculate the information entropy of each character in the boundary window, calculate the sensitive field boundary window score of the boundary window based on the information entropy, and fine-tune the boundary window until the boundary window length is less than the preset window length threshold, the boundary window containing the sensitive field and the highest sensitive field boundary window score is obtained, and the positions at both ends of the sensitive boundary window are used as the field boundaries of the sensitive field. The sensitive field boundary window score calculation formula of the boundary window is: ; in, Indicates that sensitive fields are included The bounding window, Represents the bounding window Sensitive field boundary window score, Indicates the sensitive fields output by the lightweight sensitive field identification module is the probability of sensitive fields, Represents the bounding window Information entropy of the frequency distribution of the internal field; specifically: ; in, Represents the bounding window The frequency of occurrence of the cth field in the field sequence of the data to be encrypted; Generate a conformal encryption factor for the sensitive field, where the formula for generating the conformal encryption factor is: ; in, represents the conformal encryption factor of the sensitive field y, Indicates the ID of the user associated with the data to be encrypted. The timestamp representing the generation of the conformal encryption factor of the sensitive field y, Indicates splicing processing, represents a hash operation; as an embodiment of the present invention, the hash operation adopts the SHA-256 hash function.
[0025] It should be noted that by combining the contextual information entropy of sensitive fields with the recognition probability of sensitive fields, multi-scale adaptive perception of field boundaries is achieved. This method can autonomously screen the most credible field boundary areas in fuzzy boundary scenarios (such as embedded text and logs without field names), thereby improving the accuracy and stability of boundary recognition.
[0026] This application gradually scores and fine-tunes the initial boundary window by introducing the information entropy of the field frequency distribution. It can accurately capture the true semantic boundaries of sensitive fields while keeping the window length not exceeding the set threshold, avoiding field truncation or misidentification caused by traditional fixed-pattern or regular extraction methods. This ensures that the encryption operation acts on the complete field semantic unit, so that the structural distribution of the original field can be maintained after conformal encryption, ensuring the structural consistency of the data before and after encryption. Moreover, through the optimization driven by high-sensitivity entropy values, the selected sensitive boundary window is highly consistent in semantics and has the characteristics of high information entropy concentration and prominent changes, which helps the subsequent encryption module to accurately act on the sensitive data carrier and reduce the risk of structural drift. This method can effectively compress the range of the field to be encrypted to a highly sensitive area, reduce the redundant data included in the encryption range, thereby reducing the number of subsequent rounds of encryption execution, reducing the key expansion burden, and improving the overall encryption processing speed, which is conducive to the system's parallel deployment and real-time processing needs.
[0027] S2: Extract the encrypted fields corresponding to the sensitive fields based on the field boundaries, use the conformal encryption algorithm combined with the conformal encryption factor to conformally encrypt the fields to be encrypted, and construct the encrypted ciphertext of the data to be encrypted based on the conformal encryption result of the fields to be encrypted.
[0028] The fields to be encrypted are encrypting in a conformal manner using a conformal encryption algorithm combined with a conformal encryption factor, including: Extract the character sequence of the sensitive field in the corresponding sensitive boundary window as the field to be encrypted; Using a word vector model to perform word vector encoding on the data to be encrypted, obtaining a word vector data sequence corresponding to the data to be encrypted, and extracting the word vector encoding result of the field to be encrypted; specifically, the word vector model is a bag-of-words model, and the word vector encoding result is a binary code; The word vector encoding result of the field to be encrypted is divided into a left half and a right half. The word vector encoding result is encrypted using the Feistel encryption round. After the round encryption is completed, the left half and the right half of the last round encryption output are spliced to obtain the conformal encryption result of the field to be encrypted. The round encryption process of each round is as follows: The number of rounds of the current encryption, the conformal encryption factor of the sensitive field contained in the field to be encrypted, and the current right half are concatenated into an intermediate variable; The intermediate variable is encrypted using the AES encryption method combined with the system encryption key to obtain the ciphertext fragment corresponding to the intermediate variable; Take the modulus of the ciphertext fragment and perform an XOR operation with the current left half to generate a new right half; Exchange the new right half with the current left half to obtain the encrypted left half and right half of the current round.
[0029] As an embodiment of the present invention, the number of rounds of encryption is 10; The conformal encryption result of the field to be encrypted is used to replace the word vector encoding result of the field to be encrypted in the word vector data sequence to obtain the encrypted ciphertext corresponding to the data to be encrypted.
[0030] It should be noted that the present application implements conformal encryption based on the Feistel network structure, performs multiple rounds of encryption on the word vector encoding of the encrypted field, and enhances the binding of semantics and format information by introducing conformal encryption factors and the intermediate variable construction method of the original right half. In each round of encryption, AES encryption converts the intermediate variable into a nonlinear mapping output, avoiding the possibility of fixed pattern matching. The advantage of the Feistel network structure is that it is highly reversible and the round function can be freely designed, and it can be decrypted without additional inverse operation logic. At the same time, since the word vector contains context, position and field semantics when encoding, it is represented by a fixed-length vector, and the entire encryption process maintains the consistency of the input and output dimensions. By introducing conformal encryption factors in the round function, it is ensured that even if the word vector features of different format fields are similar, their encryption paths will be significantly different, and the final encryption result will still maintain the same dimension as the original word vector. Therefore, in the decoding or data display stage, it can be directly mapped back to the ciphertext field consistent with the original field structure, ensuring that the encrypted data can still be embedded in the conformal characteristics of the original data stream. In addition, this method introduces nonlinear and diffusion characteristics by alternating operations on left and right vector fragments, so that the encryption result of sensitive fields not only depends on the word vector itself, but also is jointly associated with its semantic label, number of rounds, and structural context, thereby achieving the dual encryption goals of structural preservation and semantic perturbation, effectively improving the security and anti-correlation of conformal encryption.
[0031] S3: Logically shard the encrypted ciphertext to obtain the shard data of the encrypted ciphertext, along with the hash summary and shard verification information. The shard data with the attached information is sent to the distributed nodes using a distributed scheduling strategy.
[0032] The encrypted ciphertext is logically fragmented to obtain fragmented data of the encrypted ciphertext, including: Identify the conformal encryption result of the to-be-encrypted field in the encrypted ciphertext, add a binary fragmentation mark code, a conformal encryption result length code, and a fragmentation sequence number code before the conformal encryption result, use the character positions before the fragmentation mark code and the fragmentation sequence number code as logical fragmentation positions, and logically fragment the encrypted ciphertext to obtain fragmentation data of the encrypted ciphertext. Specifically, the fragmentation sequence number code is the number of fragmentation mark codes currently added.
[0033] The shard data is accompanied by a hash summary and shard verification information, including: Performing a hash digest operation on the shard data to obtain a hash digest of the shard data; Using the fragment mark code, the conformal encryption result length code and the fragment sequence number code in the fragment data as fragment verification information; Attach the hash digest to the shard data to obtain the shard data with the hash digest and shard verification information; Specifically, the fragment marking code consists of a special and fixed fragment positioning code and a matching rule code indicating a successful match of sensitive fields contained in the fragment data, and is used to indicate the sensitive semantic type associated with the current fragment data, such as ID card, mobile phone number, email address, etc. The hash summary operation method of the shard data includes a first-level hash summary generation operation and a second-level hash summary generation operation, wherein the first-level hash summary generation operation uses the SHA-256 hash function to generate a summary of the shard data that does not contain the shard mark code and the shard sequence number code to obtain a first-level hash summary of the shard data for basic integrity verification, and the first-level hash summary is spliced with the shard sequence number code in the shard data. The HMAC hash function based on the system-level hash key is used to perform a hash operation on the splicing result to generate a second-level hash summary to detect whether the shard data is replaced or misplaced, and the second-level hash summary is used as the hash summary of the shard data; Furthermore, if an attacker replaces the content of the shard data and forges the first-level hash digest, the attacker cannot forge the second-level hash digest because he does not know the system-level hash key. The second-level hash digest can then be used to detect whether the content has been replaced. If the shard data is placed in the wrong position, that is, the attacker attacks the shard checksum information, the part of the second-level hash digest that is used to encode the shard sequence number will be abnormal, and the identification and reassembly sequence will be abnormal. By introducing a sensitive semantic label embedding mechanism and attaching identifiable matching rule codes to each shard data, the interpretability of fields in distributed processing and semantic auditing is enhanced, and semantic shard reconstruction and conflict detection are supported. The proposed dual-layer hash verification strategy (SHA-256 + HMAC) not only implements integrity verification, but also supports refined position consistency detection through structural nested verification, effectively addressing various transmission risks such as tampering, replacement, and dislocation of shard data. The overall solution takes into account conformality, structural identification, and consistency verification, and can provide a basic guarantee for highly reliable and secure data encryption and transmission.
[0034] A distributed scheduling strategy is used to send shard data with accompanying information to distributed nodes, including: Calculate the scheduling score between the fragmented data of the accompanying information and the distributed nodes in turn, and send the fragmented data of the accompanying information to the distributed node with the highest scheduling score as the distributed scheduling strategy of the fragmented data of the accompanying information, where the fragmented data of the accompanying information and the distributed nodes are The scheduling score between them is: , ; ; ; ; in, Shard data and distributed nodes that represent accompanying information The scheduling score value between Indicates the nth distributed node, N indicates the total number of distributed nodes, represents the scheduling weight coefficient, Represents a distributed node The load factor, The shard data data that represents the accompanying information is distributed on the node The space capacity ratio in Shard data and distributed nodes that represent accompanying information Communication quality of the communication link between them; Both represent the scheduling resource weights, Represents a distributed node The transmission speed of the fragmented data to the receiving end, Indicates the preset maximum transmission speed. Represents a distributed node The number of fragmented data to be sent; The number of bits of the fragment data that represents the accompanying information, Represents a distributed node The number of bits of space available; Shard data and distributed nodes representing the calculation of incidental information When scheduling the score value between the shard data sender and the distributed node The transmission delay between Indicates the preset maximum transmission delay.
[0035] It should be noted that traditional scheduling functions mostly focus on single-dimensional resource indicators or fixed node distribution strategies, which are difficult to cope with the scheduling challenges brought about by dynamic loads and data structure differences. This scheduling scoring function introduces a composite scoring mechanism to comprehensively process the three core dimensions of computing load, storage capacity and network delay. It uses load factors to dynamically disperse the scheduling pressure of high-load nodes, improves the overall response speed of the system, and uses the spatial capacity ratio to reflect the system's adaptability to the actual resource carrying capacity of the nodes. This design allows sharded data to be scheduled to nodes with sufficient spatial resources first, reducing the sharding failure rate and the secondary scheduling overhead caused by insufficient resources. By reversely measuring the communication delay, the system is more inclined to select low-latency link nodes for scheduling, reducing transmission congestion and improving the real-time transmission guarantee capability of sharded data. It is particularly suitable for scheduling tasks in time-sensitive or frequently interactive scenarios, and improves the parallelism of sharded data transmission.
[0036] S4: The distributed node uploads the shard data with the accompanying information to the receiving end. The receiving end extracts the logical shard order from the shard verification information attached to the shard data, reorganizes and decrypts the shard data with the accompanying information, and compares the hash summary to verify the data integrity and consistency. This includes: The receiving end receives the fragmented data of the accompanying information, extracts the logical fragmentation order from the fragmentation check information, and sorts the fragmented data of the accompanying information according to the logical fragmentation order; specifically, the logical fragmentation order is the fragmentation sequence code in the fragmentation check information; Extract the hash summary from the shard data with the accompanying information, remove the hash summary and shard verification information from the shard data with the accompanying information, and recalculate the hash summary of the shard data. If the recalculated result is consistent with the extracted hash summary, it means that the data integrity and consistency verification has passed, and the shard data is decrypted.
[0037] It should be noted that the decryption process of the present invention includes extracting the conformal encryption result in the sliced data based on the conformal encryption result length code in the slice verification information, decrypting the conformal encryption result using the inverse algorithm of the conformal encryption algorithm, obtaining the word vector encoding result corresponding to the sliced data, arranging the word vector encoding results in the order of the sliced data, and obtaining the word vector data sequence corresponding to the data to be encrypted as the decryption result; the conformal encryption result is in the front row of the sliced data, and the position of the conformal encryption result is located based on the conformal encryption result length code.
[0038] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.
[0039] It should be noted that the serial numbers of the above-mentioned embodiments of the present invention are for descriptive purposes only and do not represent the advantages or disadvantages of the embodiments. In addition, the terms "including", "comprising" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, device, article or method comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, device, article or method. In the absence of further restrictions, an element defined by the sentence "including a ..." does not exclude the presence of other identical elements in the process, device, article or method comprising the element.
[0040] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0041] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. A data shape-preserving encryption parallel processing and transmission method, characterized in that: The method comprises: S1: Identify sensitive fields in the data to be encrypted, obtain the sensitive fields in the data to be encrypted, and perform conformal structure analysis on the sensitive fields to obtain the field boundaries and conformal encryption factors of the sensitive fields; S2: Extract the encrypted fields corresponding to the sensitive fields based on the field boundaries, use the conformal encryption algorithm combined with the conformal encryption factor to conformally encrypt the fields to be encrypted, and construct the encrypted ciphertext of the data to be encrypted based on the conformal encryption results of the fields to be encrypted; S3: Logically shard the encrypted ciphertext to obtain shard data of the encrypted ciphertext, along with hash digest and shard verification information. The shard data with the accompanying information is sent to the distributed nodes using a distributed scheduling strategy. S4: The distributed node uploads the shard data with accompanying information to the receiving end. The receiving end extracts the logical shard order from the shard verification information attached to the shard data, reorganizes and decrypts the shard data with accompanying information, and compares the hash summary to verify data integrity and consistency.
2. The data conformal encryption parallel processing and transmission method according to claim 1, characterized in that: Identify sensitive fields in encrypted data, including: Obtain the user's data to be encrypted, use the pre-built field dictionary to identify and separate the fields in the data to be encrypted, and obtain the field sequence of the data to be encrypted; Use the rule engine and regular expressions to build a rule matching pattern, match the fields in the field sequence, and extract the successfully matched fields as candidate fields; The context fields of the candidate fields are extracted to form a candidate field sequence. The word vector model is used to encode the fields in the candidate field sequence. The lightweight sensitive field identification module is used to identify sensitive fields in the candidate field sequence after word vector encoding. The probability that the candidate field in the candidate field sequence is a sensitive field is obtained, and the candidate field with a probability higher than the preset sensitive probability threshold is regarded as a sensitive field.
3. The data shape-preserving encryption parallel processing and transmission method according to claim 2, characterized in that: Performing conformal structural analysis on the sensitive fields, including: The conformal structure analysis includes identifying field boundaries and generating conformal encryption factors; Extract the position of the sensitive field, initially construct a boundary window containing the sensitive field, calculate the information entropy of each character in the boundary window, calculate the sensitive field boundary window score of the boundary window based on the information entropy, and fine-tune the boundary window until the boundary window length is less than the preset window length threshold. The boundary window containing the sensitive field and with the highest sensitive field boundary window score is used as the sensitive boundary window of the sensitive field; the positions at both ends of the sensitive boundary window are used as the field boundaries of the sensitive field. The sensitive field boundary window score calculation formula of the boundary window is: ; in, Indicates that sensitive fields are included The bounding window, Represents the bounding window Sensitive field boundary window score, Indicates the sensitive fields output by the lightweight sensitive field identification module is the probability of sensitive fields, Represents the bounding window Information entropy of the frequency distribution of the inner field; The formula for generating the conformal encryption factor of the sensitive field is: ; in, represents the conformal encryption factor of the sensitive field y, Indicates the ID of the user associated with the data to be encrypted. The timestamp representing the generation of the conformal encryption factor of the sensitive field y, Indicates splicing processing, Represents a hash operation.
4. The data shape-preserving encryption parallel processing and transmission method according to claim 3, characterized in that: The fields to be encrypted are encrypting in a conformal manner using a conformal encryption algorithm combined with a conformal encryption factor, including: Extract the character sequence of the sensitive field in the corresponding sensitive boundary window as the field to be encrypted; The word vector model is used to perform word vector encoding on the data to be encrypted to obtain a word vector data sequence corresponding to the data to be encrypted, and the word vector encoding result of the field to be encrypted is extracted; the word vector encoding result of the field to be encrypted is divided into a left half and a right half, and the word vector encoding result is round-encrypted using a Feistel encryption round. After the round encryption is completed, the left half and the right half of the last round encryption output are spliced to obtain the conformal encryption result of the field to be encrypted; The conformal encryption result of the field to be encrypted is used to replace the word vector encoding result of the field to be encrypted in the word vector data sequence to obtain the encrypted ciphertext corresponding to the data to be encrypted.
5. The data shape-preserving encryption parallel processing and transmission method according to claim 4, characterized in that: The encrypted ciphertext is logically fragmented to obtain fragmented data of the encrypted ciphertext, including: Identify the conformal encryption result of the field to be encrypted in the encrypted ciphertext, add a binary fragmentation mark code, a conformal encryption result length code and a fragmentation sequence number code before the conformal encryption result, use the character position before the fragmentation mark code and the fragmentation sequence number code as the logical fragmentation position, logically fragment the encrypted ciphertext, and obtain the fragmentation data of the encrypted ciphertext.
6. The data shape-preserving encryption parallel processing and transmission method according to claim 5, characterized in that: The shard data is accompanied by a hash summary and shard verification information, including: Performing a hash digest operation on the shard data to obtain a hash digest of the shard data; Using the fragment mark code, the conformal encryption result length code and the fragment sequence number code in the fragment data as fragment verification information; The hash summary is attached to the fragment data to obtain the fragment data with the hash summary and fragment verification information.
7. The data shape-preserving encryption parallel processing and transmission method according to claim 6, characterized in that: A distributed scheduling strategy is used to send shard data with accompanying information to distributed nodes, including: The fragmented data with accompanying information is fragmented data with accompanying hash summary and fragment verification information; Calculate the scheduling score between the fragmented data of the accompanying information and the distributed nodes in turn, and send the fragmented data of the accompanying information to the distributed node with the highest scheduling score as the distributed scheduling strategy of the fragmented data of the accompanying information, where the fragmented data of the accompanying information and the distributed nodes are The scheduling score between them is: , ; ; ; ; in, Shard data and distributed nodes that represent accompanying information The scheduling score value between Indicates the nth distributed node, N indicates the total number of distributed nodes, represents the scheduling weight coefficient, Represents a distributed node The load factor, The shard data data that represents the accompanying information is distributed on the node The space capacity ratio in Shard data and distributed nodes that represent accompanying information The communication quality of the communication link between Indicates the preset maximum number of pending sends; Both represent the scheduling resource weights, Represents a distributed node The transmission speed of the fragmented data to the receiving end, Indicates the preset maximum transmission speed. Represents a distributed node The number of fragmented data to be sent; The number of bits of the fragment data that represents the accompanying information, Represents a distributed node The number of bits of space available; Shard data and distributed nodes representing the calculation of incidental information When scheduling the score value between the shard data sender and the distributed node The transmission delay between Indicates the preset maximum transmission delay.
8. The data shape-preserving encryption parallel processing and transmission method according to claim 1, characterized in that: The receiving end extracts the logical sharding order from the sharding checksum information included with the sharding data, reorganizes and decrypts the sharding data, and compares the hash digest to verify data integrity and consistency, including: The receiving end receives the fragmented data of the accompanying information, extracts the logical fragmentation order from the fragmentation verification information, and sorts the fragmented data of the accompanying information according to the logical fragmentation order; Extract the hash summary from the shard data with the accompanying information, remove the hash summary and shard verification information from the shard data with the accompanying information, and recalculate the hash summary of the shard data. If the recalculated result is consistent with the extracted hash summary, it means that the data integrity and consistency verification has passed, and the shard data is decrypted.
Citation Information
Patent Citations
Financial privacy data conformal encryption and decryption method and system
CN120086897A
Weight management method and system for neural network processing, and neural network processor
US20200019843A1