Data transmission method for casualty accident processing

By analyzing the coding differences of the one-hot coding sequence in the K-anonymous model and screening target positions, and determining the total number of equivalent categories based on the visitor's access permissions, the problem of poor generalization effect in the existing technology is solved, and more efficient data transmission of casualties and accidents is achieved.

CN119046985BActive Publication Date: 2025-05-16平湖市公安局交通警察大队
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411169592.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-24
Publication Date
2025-05-16
Estimated Expiration
2044-08-24

AI Technical Summary

Technical Problem

The existing K-anonymous model has poor results when generalizing the casualty data, affecting the quality of data transmission.

Method used

By analyzing the encoding differences of each position in the generalization region in the single-hot coding sequence, the target positions with high independence are screened for generalization operations, and the total number of equivalent classes is adaptively determined based on the visitor's access rights, and multiple divisions are performed to filter the optimal division results.

Benefits of technology

It achieves better generalization effect at a lower cost of distortion, and improves the transmission quality and safety of casualties data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119046985B_ABST
    Figure CN119046985B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data transmission, and in particular to a data transmission method for handling casualty accidents. The method comprises: obtaining a one-hot coding sequence corresponding to the data of persons responsible for several casualty accidents; analyzing the difference between the coding of each position in the generalized area and the coding of other positions in the one-hot coding sequence to select the target position; combining the number of types of coding of the target position in all the one-hot coding sequences and the access rights of the visitor, determining the total number of equivalence classes, and dividing the one-hot coding sequence for a preset number of times, and selecting the optimal division result based on the degree of confusion of the coding information in the non-generalized area and the degree of confusion of the coding information of the target position in the division result; performing a generalization operation on the coding of the target position according to the optimal division result, and transmitting the data after the generalization operation if it meets the K-anonymity rule of the K-anonymity model. The present invention improves the generalization effect and transmission quality of casualty accident data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data transmission, and in particular to a data transmission method for handling casualty accidents. Background Art

[0002] In the handling of casualties such as traffic accidents and industrial accidents, it is often necessary to register the private data of all parties involved in the accident. This private data involves personal privacy. If it is leaked during the transmission process, it may cause security risks. Therefore, in the transmission process of casualty accident handling data, it is usually necessary to desensitize the data to protect personal privacy, avoid network attacks, and ensure transmission security.

[0003] The K-anonymity model is an effective data desensitization method, which has a good data protection effect when facing data attacks such as differential attacks. The model first generalizes the private data of a person responsible for the accident, so that network attackers cannot use differential attacks to identify the information of the person responsible for the accident. When performing data generalization operations, the traditional K-anonymity model randomly selects the data of multiple people responsible for the accident to form an equivalence class, and generalizes the quasi-identifiers of the people responsible for the accident in the equivalence class to achieve the desensitization effect. This algorithm uses random selection to construct equivalence classes, but does not select according to the data characteristics of the person responsible for the accident, which will make the generalization operation less effective, thereby affecting the transmission quality of casualty accident data. Summary of the invention

[0004] In order to solve the problem that the generalization effect of the existing methods is poor when generalizing the casualty accident data, the purpose of the present invention is to provide a data transmission method for casualty accident processing, and the technical solution adopted is as follows:

[0005] The present invention provides a data transmission method for handling casualty accidents, the method comprising the following steps:

[0006] Obtaining a one-hot encoding sequence corresponding to data of persons responsible for a number of casualty accidents; the one-hot encoding sequence includes a generalized region and a non-generalized region;

[0007] Analyze the difference between the encoding of each position in the generalization area and the encoding of other positions in all one-hot encoding sequences to screen the target position; combine the number of encoding types of the target position in all one-hot encoding sequences and the access rights of the visitor to determine the total number of equivalence classes;

[0008] Divide all the one-hot coding sequences a preset number of times, wherein the number of categories of each division is the total number of the equivalence classes; based on the degree of confusion of the coding information in the non-generalized area and the degree of confusion of the coding information at the target position in the division result of each division, select the optimal division result;

[0009] The encoding of the target position is generalized according to the optimal division result to determine whether the data after the generalization operation conforms to the K-anonymity rule of the K-anonymity model. If so, the data after the generalization operation and the corresponding fuzzy coding table are transmitted.

[0010] Preferably, the analysis of the difference between the encoding of each position in the generalization region and the encoding of other positions in all one-hot encoding sequences to screen the target position includes:

[0011] The encodings of the candidate positions of the generalized region in all the one-hot encoding sequences constitute the one-hot encoding sequence corresponding to the candidate positions; the candidate position is any position of the generalized region in the one-hot encoding sequence;

[0012] According to the comprehensive difference between the one-hot encoding sequence corresponding to the candidate position and the one-hot encoding sequences corresponding to all other positions, an independence weight corresponding to the candidate position is obtained, wherein the comprehensive difference is positively correlated with the independence weight;

[0013] Filter target locations based on the independence weight corresponding to each location.

[0014] Preferably, the screening of target locations based on the independence weight corresponding to each location includes:

[0015] The position corresponding to the maximum independence weight is determined as the target position.

[0016] Preferably, the method of combining the number of types of encodings of target positions in all one-hot encoding sequences and the access rights of visitors to determine the total number of equivalence classes includes:

[0017] The total number of equivalence classes is determined by rounding up the ratio of the number of types of encodings at the target position in all one-hot encoding sequences to the visitor's access characteristic value;

[0018] Among them, the visitor's access characteristic value is negatively correlated with the access permission.

[0019] Preferably, the screening of the optimal division result based on the confusion degree of the coding information in the non-generalized area and the confusion degree of the coding information at the target position in the division result of each division includes:

[0020] For the division result of any division: for the cth equivalence class, the difference between the degree of confusion of all codes in the non-generalized area of ​​the cth equivalence class and the degree of confusion of all target positions in the cth equivalence class is recorded as the division evaluation value of the cth equivalence class; based on the overall distribution of the division evaluation values ​​of all equivalence classes in the division result of this division, the generalization evaluation value of this division is obtained;

[0021] The optimal partitioning result is selected according to the generalization evaluation value of each partitioning.

[0022] Preferably, the screening of the optimal partitioning result according to the generalization evaluation value of each partitioning includes:

[0023] The partition result corresponding to the largest partition evaluation value is taken as the optimal partition result.

[0024] Preferably, obtaining the confusion degree of all codes in the non-generalized region in the cth equivalence class includes:

[0025] The frequency of occurrence of each code in the non-generalized area in the cth equivalence class is counted, and the frequency is substituted into the calculation formula of the information entropy to obtain the degree of confusion of all codes in the non-generalized area in the cth equivalence class.

[0026] Preferably, the step of obtaining the generalization evaluation value of the sub-division based on the overall distribution of the division evaluation values ​​of all equivalence classes in the division result of the sub-division includes:

[0027] The sum of the partition evaluation values ​​of all equivalent classes in the partition result of this partition is taken as the generalization evaluation value of this partition.

[0028] Preferably, after determining whether the data after the generalization operation conforms to the K-anonymity rule of the K-anonymity model, the method further includes:

[0029] If the data after the generalization operation does not conform to the K-anonymity rule of the K-anonymity model, the positions in the generalization area of ​​all one-hot coding sequences except the target position are recorded as positions to be analyzed; the difference between the encoding of each position to be analyzed and the encoding of other positions in the generalization area of ​​all one-hot coding sequences is analyzed to screen new target positions; the total number of new equivalence classes is determined by combining the number of types of encodings of the new target positions in all one-hot coding sequences and the access rights of the visitor;

[0030] Based on the total number of new equivalence classes, all one-hot encoding sequences are divided a preset number of times, and the encoding of the new target position is generalized to determine whether the data after the generalization operation conforms to the K-anonymity rule of the K-anonymity model. If so, the data after the generalization operation and the corresponding fuzzy coding table are transmitted.

[0031] Preferably, the transmitting of the data after the generalization operation and the corresponding fuzzy coding table includes:

[0032] An asymmetric encryption algorithm is used to encrypt the data after the generalization operation and the corresponding fuzzy coding table, and the encrypted data is transmitted.

[0033] The present invention has at least the following beneficial effects:

[0034] The present invention first analyzes the difference between the code of each position in the generalization area and the code of other positions in all one-hot coding sequences, selects the code position with high independence as the target position for generalization operation, and generalizes the code of the target position in the one-hot coding sequence, which can achieve better generalization effect with less distortion. According to the number of types of the code of the target position in all one-hot coding sequences and the access rights of visitors, the total number of equivalence classes is adaptively determined, and then the one-hot coding sequence is divided for multiple times. Based on the confusion degree of the code information in the non-generalization area and the confusion degree of the code information of the target position in the division result of each division, the optimal division result is selected from all the division results, and then the code of the target position in each equivalence class is generalized based on the division result, and combined with the K-anonymity rule of the K-anonymity model, the data after the generalization operation and the corresponding fuzzy coding table are transmitted. The method provided by the present invention can make the data of the person responsible for the casualty accident achieve better generalization effect with less distortion cost, and improve the data transmission quality under the premise of ensuring the safety of the data transmission of the casualty accident. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0036] Figure 1 A flow chart of a data transmission method for handling casualty accidents provided by an embodiment of the present invention;

[0037] Figure 2 A flowchart of a method for obtaining the total number of equivalence classes provided in an embodiment of the present invention;

[0038] Figure 3 This is a structural block diagram of a data transmission system for handling casualty accidents provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0039] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the data transmission method for handling casualty accidents proposed by the present invention is described in detail below in conjunction with the accompanying drawings and preferred embodiments.

[0040] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0041] The specific scheme of the data transmission method for handling casualty accidents provided by the present invention is described in detail below with reference to the accompanying drawings.

[0042] Embodiment of a data transmission method for handling casualty accidents:

[0043] The specific scenario targeted by this embodiment is: in the process of transmitting the data of the person responsible for the casualty accident, since this part of the data often involves the privacy of the person responsible, it is necessary to desensitize the data of the casualty accident in combination with the visitor's access rights, and then transmit the desensitized data to complete the transmission of the data of the person responsible for the casualty accident.

[0044] This embodiment proposes a data transmission method for handling casualty accidents, such as Figure 1 As shown, the data transmission method for handling casualty accidents in this embodiment includes the following steps:

[0045] Step S1, obtaining a one-hot encoding sequence corresponding to data of persons responsible for several casualty accidents; the one-hot encoding sequence includes a generalized region and a non-generalized region.

[0046] Since casualties often involve responsible persons related to the accident, and these data often involve the privacy of the responsible persons, in order to prevent privacy leakage during the transmission process, it is necessary to perform desensitization and generalization operations on them. First, collect the responsible person data of all casualties in the current time period, where one piece of data corresponds to one responsible person for the accident, including multiple attributes of the responsible person for the accident. In this embodiment, the attributes of the responsible person for the accident are divided into three categories according to the K-anonymity model. The K-anonymity model is a commonly used technology in the field of data transmission. The three types of attributes include: identifier: the unique identification code of the responsible person for the accident, including: ID number, telephone number, address, name; quasi-identifier: multiple data cross-verification is required to determine the specific responsible person for the accident corresponding to the data, including: gender, age, vehicle color, driver's license status, accident location; public sensitive attributes: sensitive information that needs to be clearly transmitted and cannot identify the specific responsible person for the accident, including: accident weather, accident cause, accident location, accident time. Therefore, one piece of data of the responsible person for the accident is composed of three types of attributes, and multiple pieces of data of the responsible person for the accident constitute the data that needs to be transmitted in this embodiment. It should be noted that: this embodiment is explained by taking traffic accident as an example, that is, the collected data of responsible persons is the data of responsible persons of traffic accident. As other implementations, it can also be other types of accident data, such as industrial accident data. The current time period is a set of all historical moments whose time interval with the current moment is less than or equal to the preset time length. In this embodiment, the preset time length is 1 month. In specific applications, the implementer can set it according to the specific situation.

[0047] The collected data of casualties are stored in the data storage module, and the data requester sends an access instruction to the data center. The access instruction includes the identity authentication code of the data sender and the requested data range. In this embodiment, the data request range is: all casualties data within xx hours, xx minutes, and xx seconds of the accident time. After receiving the access instruction, the access permission module confirms the identity information in the access instruction, identifies the specific identity of the access instruction, and determines the visitor's visitor rights. In this embodiment, the visitor's access characteristic value is set manually. The value of the visitor's access characteristic value in this embodiment is [1, 5]. The visitor's access characteristic value is negatively correlated with the access permission. The negative correlation indicates that the dependent variable will decrease as the independent variable increases, and the dependent variable will increase as the independent variable decreases; the smaller the visitor's access characteristic value, the greater the corresponding access permission, the lower the generalization degree of the data received by the visitor, and the higher the data information content. In specific applications, the implementer sets the access characteristic value according to specific circumstances.

[0048] The access feature value is used to determine whether the visitor has access rights. When the visitor has access rights, in order to protect the privacy of the person responsible for the accident, the data desensitization submodule is enabled, and the identifier part of the accident responsible person's data cannot be transmitted. In the transmitted data, a piece of accident responsible person data includes two attributes: quasi-identifier and public privacy attribute. All accident responsible person data within the data request range constitutes the transmitted data. Next, the transmitted data will be desensitized to obtain the final transmitted data.

[0049] This embodiment adopts the K-anonymity model to perform data desensitization processing. Since the traditional K-anonymity model generalizes by randomly selecting generalization attributes, resulting in poor generalization efficiency and loss of generalized data information, this embodiment extracts features from the transmitted data and selects desensitized quasi-identifiers based on the features, thereby improving the generalization efficiency and generalization effect while ensuring data security.

[0050] Since the lengths of different responsible person data may be different, in this embodiment, for different responsible person data: the data with the same attributes are right-aligned, and 0 is added to the empty positions, that is, all responsible person data are aligned so that the lengths of all responsible person data are equal. It should be noted that the responsible person data mentioned later are all pre-processed data. For any responsible person data: perform one-hot encoding on the responsible person data to obtain the one-hot encoding sequence corresponding to the responsible person data. Using this method, the one-hot encoding sequence corresponding to each responsible person data can be obtained, that is, multiple one-hot encoding sequences are obtained, and the lengths of all one-hot encoding sequences are equal. One-hot encoding is a prior art and will not be described in detail here.

[0051] In the one-hot encoding sequence, the encoding corresponding to the quasi-identifier is recorded as the generalization area, and the encoding other than the generalization area is recorded as the non-generalization area.

[0052] So far, using the method provided in this embodiment, multiple one-hot encoding sequences are obtained, and the generalization region and the non-generalization region in the one-hot encoding sequence are divided.

[0053] Step S2, analyzing the difference between the encoding of each position in the generalization area and the encoding of other positions in all one-hot encoding sequences, and screening the target position; combining the number of types of encodings of the target position in all one-hot encoding sequences and the access rights of the visitor, and determining the total number of equivalence classes.

[0054] In order to prevent the leakage of casualty accident data, it is necessary to perform a fuzzification operation on the casualty accident data to be transmitted. The essence of the fuzzification operation process is the process of increasing the entropy of the one-hot coded information. The faster the entropy increase, the greater the generalization degree. However, due to the mutual correlation between different one-hot codes, the original content of the fuzzy code position can be determined by other code positions, affecting the generalization efficiency and generalization effect of the one-hot code. In this embodiment, the code position with a high degree of independence is first selected as the position to be generalized, and then the total number of equivalence classes is adaptively determined according to the visitor's access rights, thereby improving the division result of the one-hot code sequence and ensuring the generalization effect of the data.

[0055] Step S21, screening the target position according to the difference between the encoding of each position in the generalization region and the encoding of other positions in all one-hot encoding sequences.

[0056] Next, a position of the generalization region in the one-hot encoding sequence is taken as an example for description. The method provided in this embodiment can be used to process other positions of the generalization region in the one-hot encoding sequence.

[0057] Specifically, any position in the generalization region in the one-hot coding sequence is recorded as a candidate position. Since the lengths of all one-hot coding sequences are equal, there is a candidate position in each one-hot coding sequence, and the codes of the candidate positions in the generalization region in all one-hot coding sequences constitute the one-hot coding sequence corresponding to the candidate position. Using this method, the one-hot coding sequence corresponding to each position in the generalization region can be obtained, and then the independence weight corresponding to the candidate position is obtained based on the comprehensive difference between the one-hot coding sequence corresponding to the candidate position and the one-hot coding sequences corresponding to all other positions, and the comprehensive difference is positively correlated with the independence weight.

[0058] Among them, the comprehensive difference between two coding sequences can be characterized by a variety of indicators, such as: the dynamic time warping distance between the two coding sequences, the negative correlation mapping value of the cosine similarity between the two coding sequences, the editing distance between the two coding sequences, etc. In this embodiment, the editing distance between the two coding sequences is used to characterize the comprehensive difference between the two, that is, the editing distance between the two coding sequences is taken as the comprehensive difference between the two coding sequences.

[0059] A positive correlation means that the dependent variable will increase as the independent variable increases, and the dependent variable will decrease as the independent variable decreases. The specific relationship can be a multiplication relationship, an addition relationship, or the power of an exponential function with a base greater than 1.

[0060] In this embodiment, a specific calculation method for the independence weight corresponding to the candidate position is given: the cumulative sum of the comprehensive differences between the one-hot coding sequence corresponding to the candidate position and the one-hot coding sequences corresponding to all other positions in the generalization area except the candidate position is used as the independence weight corresponding to the candidate position. The greater the comprehensive difference between the one-hot coding sequence corresponding to the candidate position and the one-hot coding sequences corresponding to other positions, the smaller the data correlation between the coding of the candidate position of the one-hot coding sequence and the coding of other positions, the greater the independence weight of the coding of the candidate position, and the better the generalization effect that can be achieved by selecting the coding of the candidate position for fuzzification operation during generalization.

[0061] By using the above method, the independence weight corresponding to each position in the generalization area can be obtained. The larger the independence weight, the better the effect of fuzzy processing on the encoding of the candidate position. Therefore, in this embodiment, the position corresponding to the maximum independence weight is determined as the target position.

[0062] Step S22, combining the number of types of encodings of target positions in all one-hot encoding sequences and the access rights of visitors, to determine the total number of equivalence classes.

[0063] After determining the target position of the one-hot encoding sequence, the K-anonymity model needs to classify all responsible person data and generalize the responsible person data belonging to the same category to achieve the purpose of desensitization.

[0064] When classifying the responsible person data, since the data in the non-generalized area will not be fuzzified, in order to avoid the attacker from performing differential attacks and inferring the original data through the data in the non-generalized area, it is necessary to ensure that the data in the non-generalized area in the equivalence class is as evenly distributed as possible during classification. In essence, it is to avoid differential attacks by increasing the degree of confusion of the encoding of the non-generalized area in the equivalence class. This embodiment will determine the total number of equivalence classes by combining the number of encoding types of the target position in the one-hot encoding sequence and the access rights of the visitor.

[0065] Specifically, the number of encoding types of the target position in all one-hot encoding sequences is counted, and the result of rounding up the ratio between the number of encoding types of the target position in all one-hot encoding sequences and the visitor's access characteristic value is determined as the total number of equivalence classes.

[0066] So far, using the above method, the total number of equivalence classes is obtained, such as Figure 2 As shown, the figure is a flow chart of the method for obtaining the total number of equivalence classes.

[0067] Step S3, dividing all the one-hot coding sequences a preset number of times, wherein the number of categories for each division is the total number of the equivalence classes; based on the degree of confusion of the coding information in the non-generalized area and the degree of confusion of the coding information at the target position in the division result of each division, selecting the optimal division result.

[0068] When dividing the equivalence classes of the K-anonymity model, the uniformity of the non-generalized area is usually used as the optimization target of the intelligent optimization algorithm to obtain the equivalence class division method with the highest information entropy in the non-generalized area. In this embodiment, the one-hot encoding sequence is divided into equivalence classes multiple times, and the optimal division result is screened out from all the division results, and then generalization processing is performed to improve the generalization effect of the data and reduce the risk of accident data leakage.

[0069] Specifically, the preset number of times is set manually. In this embodiment, the preset number of times is set to 100 times. In specific applications, the implementer can set it according to the specific situation. The intelligent optimization algorithm is used to divide all the unique hot coding sequences multiple times, and the total number of divisions is 100 times, and the division results of each division are obtained. In this embodiment, the intelligent optimization algorithm selected is the particle swarm optimization algorithm. As other implementation methods, the genetic algorithm or the ant colony algorithm can also be selected for processing. The specific division process in this embodiment is as follows: set the division parameter of the equivalence class to C, construct an M-dimensional search space, where the value range of m is a positive integer between 1 and M, and the value range of the m-th dimension data is an integer between 1 and C, representing that the m-th responsible person data belongs to the c-th equivalence class; where M is the number of one-hot encoding sequences, and C is the total number of equivalence classes; for a particle in the search space, it contains the equivalence class of each responsible person data, that is, a particle corresponds to an equivalence class division method; using the M-dimensional space as the search space, the particle swarm optimization algorithm is used for processing, and the necessary parameters are: the number of search iterations is 100; the output is a particle, which corresponds to an equivalence class division method. The particle swarm optimization algorithm is a prior art and will not be described in detail here.

[0070] For the division result of any division: for the cth equivalence class, the difference between the degree of confusion of all codes in the non-generalized area of ​​the cth equivalence class and the degree of confusion of the codes of all target positions in the cth equivalence class is recorded as the division evaluation value of the cth equivalence class; wherein, the method for obtaining the degree of confusion of all codes in the non-generalized area of ​​the cth equivalence class is: counting the frequency of occurrence of each code in the non-generalized area of ​​the cth equivalence class, substituting the frequency of occurrence of all codes in the non-generalized area of ​​the cth equivalence class into the calculation formula of information entropy, and taking the obtained information entropy as the degree of confusion of all codes in the non-generalized area of ​​the cth equivalence class. The larger the information entropy, the larger the amount of information in the non-generalized area of ​​the cth equivalence class, the more chaotic the distribution of codes in the non-generalized area of ​​the cth equivalence class, and the more public sensitive information in the cth equivalence class is generalized, the faster the entropy of the sensitive information part in the entire information can be increased, and the less susceptible to differential attacks. The method for obtaining the degree of confusion of the encoding of all target positions in the cth equivalence class is as follows: the frequency of occurrence of each encoding of the target position in all one-hot encoding sequences in the cth equivalence class is counted, the frequency of occurrence of each encoding of the target position in all one-hot encoding sequences in the cth equivalence class is substituted into the calculation formula of information entropy, and the obtained information entropy is used as the degree of confusion of the encoding of all target positions in the cth equivalence class. The calculation formula of information entropy is an existing calculation formula, which will not be repeated here. The method for obtaining the difference between the degree of confusion of all encodings in the non-generalized area in the cth equivalence class and the degree of confusion of the encoding of all target positions in the cth equivalence class is as follows: the degree of confusion of all encodings in the non-generalized area in the cth equivalence class is subtracted from the degree of confusion of the encoding of all target positions in the cth equivalence class, and the result is used as the difference between the degree of confusion of all encodings in the non-generalized area in the cth equivalence class and the degree of confusion of the encoding of all target positions in the cth equivalence class. By adopting this method, the division evaluation value of each equivalence class in the division result can be obtained. Next, the generalization evaluation value of the partition is obtained based on the overall distribution of the partition evaluation values ​​of all equivalence classes in the partition result; the overall distribution of the partition evaluation values ​​of all equivalence classes can be represented by the cumulative sum of the partition evaluation values ​​of all equivalence classes, can be represented by the average value of the partition evaluation values ​​of all equivalence classes, or can be represented by the mode of the partition evaluation values ​​of all equivalence classes; in this embodiment, the generalization evaluation value of the partition is represented by the cumulative sum of the partition evaluation values ​​of all equivalence classes in the partition result, that is, the cumulative sum of the partition evaluation values ​​of all equivalence classes in the partition result is used as the generalization evaluation value of the partition.

[0071] By adopting the above method, the generalization evaluation value of each division can be obtained. The larger the generalization evaluation value, the better the current division result. After performing the generalization operation based on the current division result, a better desensitization effect can be achieved at a lower distortion cost. Based on this, this embodiment takes the division result corresponding to the maximum value of the generalization evaluation value as the optimal division result.

[0072] Step S4, generalize the code of the target position according to the optimal division result, and determine whether the data after the generalization operation conforms to the K-anonymity rule of the K-anonymity model. If so, transmit the data after the generalization operation and the corresponding fuzzy coding table.

[0073] In the above steps, the codes of the target positions in all the one-hot coding sequences are divided into C equivalence classes. For any equivalence class in the optimal division result: generalization operation is performed on the codes of all target positions in the equivalence class to obtain the data after generalization operation, and a fuzzy coding table is constructed. The generalization operation and the construction process of the fuzzy coding table are both existing technologies and will not be described in detail here.

[0074] The K-anonymity rule of the K-anonymity model is used to judge the data after the generalization operation to determine whether the transmitted data after the generalization operation conforms to the K-anonymity rule. If it conforms, the asymmetric encryption algorithm is used to encrypt the data after the generalization operation and the corresponding fuzzy coding table, and the encrypted data is transmitted. The asymmetric encryption algorithm is a prior art and will not be described in detail here. If it does not conform, the positions except the target position in the generalization area of ​​all one-hot coding sequences are recorded as positions to be analyzed; the difference between the encoding of each position to be analyzed in the generalization area of ​​all one-hot coding sequences and the encoding of other positions is analyzed to screen new target positions; combined with all one-hot coding sequences The number of types of encodings of the new target position in the column and the access rights of the visitor are used to determine the total number of new equivalence classes; based on the total number of new equivalence classes, all one-hot encoding sequences are divided into a preset number of times using an intelligent optimization algorithm, and the encoding of the new target position is generalized, and the data after the generalization operation is further judged whether it meets the K-anonymity rules of the K-anonymity model. If it does, the data after the generalization operation and the corresponding fuzzy coding table are transmitted; that is, when the transmitted data does not meet the K-anonymity rules, return to step S2, re-screen the target position, and then re-generalize using the method provided in this embodiment until the transmitted data meets the K-anonymity rules of the K-anonymity model. The K-anonymity rules of the K-anonymity model are existing rules and will not be described in detail here.

[0075] So far, the method provided in this embodiment has been used to complete the generalization and transmission processing of the casualty accident data.

[0076] This embodiment first analyzes the difference between the code of each position in the generalization area and the code of other positions in all one-hot coding sequences, selects the code position with a high degree of independence as the target position for generalization operation, and generalizes the code of the target position in the one-hot coding sequence, which can achieve better generalization effect with less distortion. According to the number of types of codes of the target position in all one-hot coding sequences and the access rights of visitors, the total number of equivalence classes is adaptively determined, and then the one-hot coding sequence is divided multiple times. Based on the confusion degree of the code information in the non-generalization area and the confusion degree of the code information of the target position in the division result of each division, the optimal division result is selected from all division results, and then the code of the target position in each equivalence class is generalized based on the division result. Combined with the K-anonymity rule of the K-anonymity model, the data after the generalization operation and the corresponding fuzzy coding table are transmitted. The method provided in this embodiment can achieve better desensitization effect of the data of the person responsible for the casualty accident at a lower distortion cost, and improve the data transmission rate and transmission quality under the premise of ensuring the security of data transmission of the casualty accident.

[0077] Embodiment of a data transmission system for handling casualty accidents:

[0078] See also Figure 3 , which shows a structural block diagram of a data transmission system for handling casualty accidents provided by an embodiment of the present invention. The system may include a data acquisition module, an equivalence class total number determination module, a division result screening module, and a data transmission module.

[0079] The data collection module is used to obtain a one-hot coding sequence corresponding to the data of the persons responsible for several casualties; the one-hot coding sequence includes a generalized area and a non-generalized area;

[0080] The module for determining the total number of equivalence classes is used to analyze the difference between the encoding of each position in the generalization region and the encoding of other positions in all one-hot encoding sequences, and screen the target position; the total number of equivalence classes is determined by combining the number of types of encoding of the target position in all one-hot encoding sequences and the access rights of the visitor;

[0081] A partition result screening module is used to partition all one-hot coding sequences for a preset number of times, wherein the number of categories for each partition is the total number of equivalence classes; based on the degree of confusion of the coding information in the non-generalized area and the degree of confusion of the coding information at the target position in the partition result of each partition, the optimal partition result is screened;

[0082] The data transmission module is used to generalize the encoding of the target position according to the optimal division result, and determine whether the data after the generalization operation conforms to the K-anonymity rule of the K-anonymity model. If so, the data after the generalization operation and the corresponding fuzzy coding table are transmitted.

[0083] It should be understood that Figure 3 The structural block diagram of the data transmission system for handling casualties and its modules shown can be implemented in various ways. For example, in some embodiments, the system and its modules can be implemented by hardware, software, or a combination of software and hardware. Among them, the hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or a dedicated design hardware. Those skilled in the art will understand that the above-mentioned method and system can be implemented using computer executable instructions and / or included in a processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. Such code is provided on the system and its modules of this specification. Not only can the hardware circuits such as ultra-large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc. be implemented, it can also be implemented by software executed by various types of processors, and it can also be implemented by a combination of the above-mentioned hardware circuits and software (for example, firmware).

[0084] For more details about the above modules, please refer to other places in this manual (for example Figure 1 , Figure 2 part and its related description), which will not be repeated here.

[0085] In other embodiments, a device is also provided, including a memory and a processor. The memory is used to store executable program code, and the processor is used to call and run the executable program code from the memory, so that the device executes the above-mentioned data transmission method for handling casualties. The device can be specifically a chip, a component or a module, and the chip may include a connected processor and a memory; wherein the memory is used to store instructions, and when the processor calls and executes the instructions, the chip can execute the data transmission method for handling casualties provided in the above-mentioned embodiment.

[0086] In other embodiments, a computer program product is also provided. When the computer program product is run on a computer, the computer is caused to execute the above-mentioned related steps to implement the data transmission method for handling casualty accidents provided in the above-mentioned embodiments.

[0087] In other embodiments, a computer-readable storage medium is also provided, in which a computer program code is stored. When the computer program code is executed on a computer, the computer executes the above-mentioned related method steps to implement the data transmission method for handling casualty accidents provided in the above-mentioned embodiment.

[0088] Among them, the provided systems, devices, computer program products, and computer-readable storage media are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above and will not be repeated here.

[0089] It should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A data transmission method for handling casualty accidents, characterized in that: The method comprises the following steps: Obtaining a one-hot encoding sequence corresponding to data of persons responsible for a number of casualty accidents; the one-hot encoding sequence includes a generalized region and a non-generalized region; Analyze the difference between the encoding of each position in the generalization area and the encoding of other positions in all one-hot encoding sequences to screen the target position; combine the number of encoding types of the target position in all one-hot encoding sequences and the access rights of the visitor to determine the total number of equivalence classes; Divide all the one-hot coding sequences a preset number of times, wherein the number of categories of each division is the total number of the equivalence classes; based on the degree of confusion of the coding information in the non-generalized area and the degree of confusion of the coding information at the target position in the division result of each division, select the optimal division result; Perform generalization operation on the encoding of the target position according to the optimal division result, and judge whether the data after generalization operation conforms to the K-anonymity rule of the K-anonymity model. If so, transmit the data after generalization operation and the corresponding fuzzy coding table; The screening target location includes: The encodings of the candidate positions of the generalized region in all the one-hot encoding sequences constitute the one-hot encoding sequence corresponding to the candidate positions; the candidate position is any position of the generalized region in the one-hot encoding sequence; According to the comprehensive difference between the one-hot encoding sequence corresponding to the candidate position and the one-hot encoding sequences corresponding to all other positions, an independence weight corresponding to the candidate position is obtained, wherein the comprehensive difference is positively correlated with the independence weight; Filter target locations based on the independence weight corresponding to each location.

2. The data transmission method for handling casualty accidents according to claim 1, characterized in that: The screening of the target location based on the independence weight corresponding to each location includes: The position corresponding to the maximum independence weight is determined as the target position.

3. The data transmission method for handling casualty accidents according to claim 1, characterized in that: The method combines the number of types of encodings of the target positions in all one-hot encoding sequences and the access rights of the visitor to determine the total number of equivalence classes, including: The total number of equivalence classes is determined by rounding up the ratio of the number of types of encodings at the target position in all one-hot encoding sequences to the visitor's access characteristic value; Among them, the visitor's access characteristic value is negatively correlated with the access permission.

4. The data transmission method for handling casualty accidents according to claim 1, characterized in that: The method of selecting the optimal division result based on the confusion degree of the coding information in the non-generalized area and the confusion degree of the coding information at the target position in the division result of each division includes: For the division result of any division: for the cth equivalence class, the difference between the degree of confusion of all codes in the non-generalized area of ​​the cth equivalence class and the degree of confusion of all target positions in the cth equivalence class is recorded as the division evaluation value of the cth equivalence class; based on the overall distribution of the division evaluation values ​​of all equivalence classes in the division result of this division, the generalization evaluation value of this division is obtained; The optimal partitioning result is selected according to the generalization evaluation value of each partitioning.

5. The data transmission method for handling casualty accidents according to claim 4, characterized in that: The step of selecting the optimal partitioning result according to the generalization evaluation value of each partitioning includes: The partition result corresponding to the largest partition evaluation value is taken as the optimal partition result.

6. The data transmission method for handling casualty accidents according to claim 4, characterized in that: The acquisition of the confusion degree of all encodings in the non-generalized area in the cth equivalence class includes: The frequency of occurrence of each code in the non-generalized area in the cth equivalence class is counted, and the frequency is substituted into the calculation formula of the information entropy to obtain the degree of confusion of all codes in the non-generalized area in the cth equivalence class.

7. The data transmission method for handling casualty accidents according to claim 4, characterized in that: The generalization evaluation value of the sub-division is obtained based on the overall distribution of the division evaluation values ​​of all equivalent classes in the division result of the sub-division, including: The sum of the partition evaluation values ​​of all equivalent classes in the partition result of this partition is taken as the generalization evaluation value of this partition.

8. The data transmission method for handling casualty accidents according to claim 1, characterized in that: After judging whether the data after the generalization operation conforms to the K-anonymity rule of the K-anonymity model, it also includes: If the data after the generalization operation does not conform to the K-anonymity rule of the K-anonymity model, the positions in the generalization area of ​​all one-hot coding sequences except the target position are recorded as positions to be analyzed; the difference between the encoding of each position to be analyzed and the encoding of other positions in the generalization area of ​​all one-hot coding sequences is analyzed to screen new target positions; the total number of new equivalence classes is determined by combining the number of types of encodings of the new target positions in all one-hot coding sequences and the access rights of the visitor; Based on the total number of new equivalence classes, all one-hot encoding sequences are divided a preset number of times, and the encoding of the new target position is generalized to determine whether the data after the generalization operation conforms to the K-anonymity rule of the K-anonymity model. If so, the data after the generalization operation and the corresponding fuzzy coding table are transmitted.

9. The data transmission method for handling casualty accidents according to claim 1, characterized in that: The transmitting of the data after the generalization operation and the corresponding fuzzy coding table includes: An asymmetric encryption algorithm is used to encrypt the data after the generalization operation and the corresponding fuzzy coding table, and the encrypted data is transmitted.

Citation Information

Patent Citations

  • K-anonymous data processing method and system based on dual coding and cluster mapping

    CN113378223A

  • Recursive block partitioning

    IN201647025455A