Data Processing Method, Apparatus, Server, and Storage Medium
By receiving data access requests from user terminals, obtaining access permission levels, finding matching text fragments from the annotated text stored in the blockchain, and performing permission control operations, the problem of privacy data leakage in the existing technology is solved and data security is improved.
Patent Information
- Application Number
- CN202111256884.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-27
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-10-27
AI Technical Summary
In the prior art, directly returning the data content accessed by users can easily lead to privacy data leakage and low data security.
By receiving data access requests from the user terminal, obtaining the access permission level, finding matching text fragments from the annotation text stored in the blockchain, and performing permission control operations, returning data content that meets the permission level, and at the same time, format conversion and sensitive level annotation processing of the original medical data.
It realizes that data content that meets the access permission level is returned to users while ensuring data security, improving data security.
Smart Images

Figure CN114003929B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a data processing method, apparatus, server, and storage medium. Background Art
[0002] In recent years, with the continuous development of big data technology, big data technology has brought great convenience to enterprises and users in all walks of life. For example, in the medical industry, doctors can analyze the changes in patients' conditions based on big data analysis results to better carry out relevant follow-up work. In practical applications, while big data brings convenience to people, it may also bring some problems, such as the problem of privacy data leakage.
[0003] In related technologies, usually the data content accessed by the user is directly returned to the user, which easily causes the leakage of privacy data and results in low data security. Summary of the Invention
[0004] In view of this, embodiments of the present application provide a data processing method, apparatus, server, and storage medium to solve the problem in related technologies that directly returning the data content accessed by the user to the user easily causes the leakage of privacy data and results in low data security.
[0005] The first aspect of the embodiments of the present application provides a data processing method, including:
[0006] Receiving a data access request sent by a user terminal corresponding to a target account, and obtaining the access permission level of the target account. The data access request includes access requirement description information, and the target account is pre-assigned an access permission level, where the access permission level corresponds to the sensitivity level of the accessed content;
[0007] Searching in at least one annotated text stored in the blockchain for an annotated text that matches the access requirement description information, where each text segment in the annotated text is annotated with a sensitivity level;
[0008] Performing a permission control operation on text segments in the found annotated text whose sensitivity level does not match the access permission level of the target account to obtain an access text, and sending the access text to the user terminal. The permission control operation is used to control at least one of the following permissions: editing permission, visibility permission.
[0009] Further, the method further includes:
[0010] Obtaining original medical data, and performing format conversion on the original medical data to obtain a target text in text format;
[0011] Split the target text into multiple text segments, and determine the segment type and corresponding sensitivity level of each text segment according to the content of each text segment;
[0012] Perform information annotation processing on the target text according to the sensitivity level of each text segment to obtain an annotated text, and store the annotated text in the blockchain.
[0013] Further, split the target text into multiple text segments, and determine the segment type and corresponding sensitivity level of each text segment according to the content of each text segment, including:
[0014] Perform word segmentation on the target text to obtain multiple segmented words and the word segmentation position information of each segmented word in the target text;
[0015] Determine the word segmentation type of each segmented word according to a preset keyword set, and split the target text into multiple text segments according to the word segmentation type of each segmented word, where the preset keywords in the preset keyword set correspond to keyword types;
[0016] Determine the segment type of each text segment and the segment position information of each text segment in the target text according to the word segmentation type and word segmentation position information of the segmented words included in each text segment, and determine the sensitivity level of each text segment according to the segment type of each text segment.
[0017] Further, determine the word segmentation type of each segmented word according to the preset keyword set, including:
[0018] For each segmented word, calculate the similarity between the segmented word and each preset keyword in the preset keyword set, determine the preset keyword in the preset keyword set whose corresponding similarity meets the preset similarity condition as the preset keyword matched with the segmented word, and determine the keyword type corresponding to the preset keyword matched with the segmented word as the word segmentation type of the segmented word.
[0019] Further, split the target text into multiple text segments according to the word segmentation type of each segmented word, including:
[0020] Traverse each segmented word in the target text. If the word segmentation type of the currently accessed segmented word is the same as that of the previous segmented word, divide the currently accessed segmented word into the text segment to which the previous segmented word belongs;
[0021] If the word segmentation type of the currently accessed segmented word is different from that of the previous segmented word, divide the currently accessed segmented word into a new text segment different from the text segment to which the previous segmented word belongs, and so on, until the text segment division of each segmented word is completed to obtain multiple text segments.
[0022] Further, storing the annotated text in the blockchain includes:
[0023] Generating a first key pair for the annotated text, where the first key pair includes a first private key and a first public key;
[0024] Encrypting the annotated text according to the first private key and storing the encrypted annotated text in the blockchain;
[0025] Generating a second key pair according to the account information of the target account, where the second key pair includes a second private key and a second public key;
[0026] Encrypting the first private key according to the second public key of the target account and storing the encrypted first private key.
[0027] Further, finding the annotated text matching the access requirement description information from at least one annotated text stored in the blockchain includes:
[0028] If the access requirement description information includes a text identifier, finding the encrypted first private key corresponding to the text identifier and finding the encrypted annotated text corresponding to the text identifier from at least one encrypted annotated text stored in the blockchain;
[0029] Decrypting the encrypted first private key according to the second public key of the target account to obtain the first private key, and decrypting the found encrypted annotated text according to the obtained first private key to obtain the annotated text matching the access requirement description information.
[0030] The second aspect of the embodiments of the present application provides a data processing device, including:
[0031] A request receiving unit, configured to receive a data access request sent by a user terminal corresponding to a target account and obtain the access permission level of the target account. The data access request includes access requirement description information, and the target account is pre-assigned an access permission level, where the access permission level corresponds to the sensitivity level of the accessed content;
[0032] A text searching unit, configured to find the annotated text matching the access requirement description information from at least one annotated text stored in the blockchain, where each text segment in the annotated text is marked with a sensitivity level;
[0033] A data control unit, configured to perform a permission control operation on the text segments whose sensitivity levels do not match the access permission level of the target account from the found annotated text to obtain the access text, and send the access text to the user terminal. The permission control operation is used to control at least one of the following permissions: editing permission, visibility permission.
[0034] Further, the device further includes a text storage unit. The text storage unit includes a format conversion module, a level determination module, and a storage execution module.
[0035] The format conversion module is used to obtain the original medical data, perform format conversion on the original medical data, and obtain the target text in text format.
[0036] The level determination module is used to split the target text into multiple text segments, and determine the segment type and corresponding sensitivity level of the corresponding text segment according to the content of each text segment.
[0037] The storage execution module is used to perform information annotation processing on the target text according to the sensitivity level of each text segment to obtain the annotated text, and store the annotated text in the blockchain.
[0038] Further, the level determination module is specifically used for:
[0039] Perform word segmentation on the target text to obtain multiple segmented words and the word segmentation position information of each segmented word in the target text.
[0040] Determine the word segmentation type of each segmented word according to the preset keyword set, and split the target text into multiple text segments according to the word segmentation type of each segmented word. The preset keywords in the preset keyword set correspond to keyword types.
[0041] Determine the segment type of the corresponding text segment and the segment position information of the corresponding text segment in the target text according to the word segmentation type and word segmentation position information of the segmented words included in each text segment, and determine the sensitivity level of the corresponding text segment according to the segment type of each text segment.
[0042] Further, in the level determination module, determining the word segmentation type of each segmented word according to the preset keyword set includes:
[0043] For each segmented word, calculate the similarity degree between the segmented word and each preset keyword in the preset keyword set, determine the preset keyword in the preset keyword set whose corresponding similarity degree meets the preset similarity condition as the preset keyword matched with the segmented word, and determine the keyword type corresponding to the preset keyword matched with the segmented word as the word segmentation type of the segmented word.
[0044] Further, in the level determination module, splitting the target text into multiple text segments according to the word segmentation type of each segmented word includes:
[0045] Traverse each segmented word in the target text. If the word segmentation type of the currently accessed segmented word is the same as that of the previous segmented word, divide the currently accessed segmented word into the text segment to which the previous segmented word belongs.
[0046] If the word segmentation type of the currently accessed segmented word is inconsistent with that of the previous segmented word, then the currently accessed segmented word is divided into a new text segment different from the text segment to which the previous segmented word belongs, and so on, until the text segment division of each segmented word is completed, obtaining multiple text segments.
[0047] Further, in the storage execution module, storing the annotated text in the blockchain includes:
[0048] Generating a first key pair for the annotated text, where the first key pair includes a first private key and a first public key;
[0049] Performing encryption processing on the annotated text according to the first private key, and storing the encrypted annotated text in the blockchain;
[0050] Generating a second key pair according to the account information of the target account, where the second key pair includes a second private key and a second public key;
[0051] Encrypting the first private key according to the second public key of the target account, and storing the encrypted first private key.
[0052] Further, the text search unit is specifically configured to:
[0053] If the access requirement description information includes a text identifier, then searching for the encrypted first private key corresponding to the text identifier, and searching for the encrypted annotated text corresponding to the text identifier from at least one encrypted annotated text stored in the blockchain;
[0054] Decrypting the encrypted first private key according to the second public key of the target account to obtain the first private key, and decrypting the found encrypted annotated text according to the obtained first private key to obtain the annotated text matching the access requirement description information.
[0055] The third aspect of the embodiments of the present application provides a server, including a memory, a processor, and a computer program stored in the memory and executable on the server. When the processor executes the computer program, it implements the steps of the data processing method provided in the first aspect.
[0056] The fourth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps of the data processing method provided in the first aspect.
[0057] Implementing the data processing method, apparatus, server, and storage medium provided by the embodiments of the present application has the following beneficial effects: By pre-assigning access privilege levels to each target account, when a user logs in to the target account through a user terminal to access the stored labeled text, partial content that conforms to the access privilege level can be returned to the user, which can ensure data security. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the accompanying drawings required for use in the embodiments or related technologies. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0059] Figure 1 is the implementation flowchart of a data processing method provided by the embodiments of the present application;
[0060] Figure 2 is the implementation flowchart of another data processing method provided by the embodiments of the present application;
[0061] Figure 3 is the implementation flowchart of storing labeled text in a blockchain provided by the embodiments of the present application;
[0062] Figure 4 is the structural block diagram of a data processing apparatus provided by the embodiments of the present application;
[0063] Figure 5 is the structural block diagram of a server provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] To make the objectives, technical solutions, and advantages of the present application clearer and more understandable, the following further details the present application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0065] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use the knowledge to obtain the best results in theory, methods, technologies, and application systems.
[0066] The basic technologies of artificial intelligence generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technologies, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0067] In the embodiments of the present application, based on artificial intelligence technology, it is to return data content that meets the user's access permission level to the user to ensure data security.
[0068] The data processing method involved in the embodiments of the present application can be executed by a server. When the data processing method is executed by the server, the execution entity is the server.
[0069] It should be noted that the above-mentioned server may include, but is not limited to, a server, a mobile phone, a tablet, or a wearable intelligent device, etc. In addition, the above-mentioned server may be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0070] Please refer to Figure 1 , Figure 1 which shows the implementation flowchart of a data processing method provided by the embodiments of the present application, including:
[0071] Step 101, receive a data access request sent by a user terminal corresponding to a target account, and obtain the access permission level of the target account.
[0072] Among them, the data access request includes access requirement description information, and the target account is pre-assigned an access permission level. Among them, the access permission level corresponds to the sensitivity level of the accessed content. The above-mentioned target account is usually a registered account.
[0073] In practice, the access requirement description information can be the text identifier of the annotated text. For example, it can be "C001", or it can be the title of the annotated text.
[0074] Among them, the above-mentioned access permission level is usually information used to indicate specific access permissions. One access permission level can correspond to multiple sensitivity levels or one sensitivity level. For example, if the access permission level is A level, the corresponding sensitivity levels can be level 1, level 2, and level 3.
[0075] Here, the user terminal can send a data access request to the above-mentioned execution entity. In this way, the above-mentioned execution entity can receive the data access request, and can use the account information of the target account logged in by the user terminal to find the access permission level of the target account from the pre-stored correspondence between account information and access permission levels.
[0076] In practice, before the above-mentioned step 101, the above data processing method may further include the following steps: in response to meeting the preset permission allocation condition, allocate an access permission level to the target account.
[0077] Among them, the above-mentioned preset permission allocation condition is usually a condition preset for triggering the allocation of access permission levels.
[0078] In practice, the preset permission allocation condition may include at least one of the following three items.
[0079] The first item, it is detected that a new target account is successfully registered. Here, when a new target account is successfully registered, an access permission level can be allocated to the new target account.
[0080] The second item, a permission change request sent by the target terminal corresponding to the target account is received. Among them, the above-mentioned permission change request is usually information for requesting a change in the access permission level. For example, it can request to change the access permission level from level A to level B. The above-mentioned target terminal corresponding to the target account usually refers to the terminal device logged in to the target account. Here, when the permission change request sent by the target terminal is received, the above-mentioned execution entity can change the access permission level corresponding to the target account to be consistent with the level requested by the permission change request.
[0081] The third item, a permission change request sent by the management terminal is received. Here, the management terminal is usually the terminal of the management staff. After the above-mentioned execution entity receives the permission change request sent by the management terminal, it can change the access permission level corresponding to the target account to be consistent with the level requested by the permission change request.
[0082] In practice, when the current preset permission allocation condition is met, the above-mentioned execution entity can allocate an access permission level to the target account. For example, an access permission level of level A can be allocated to the doctor's account, and an access permission level of level B can be allocated to the account of the technical personnel using medical data for research and development.
[0083] Step 102, find the annotation text that matches the access requirement description information from at least one annotation text stored in the blockchain.
[0084] Among them, each text segment in the annotated text is marked with a sensitivity level. The annotated text generally refers to the text in which each included text segment is marked with a sensitivity level. The sensitivity level is generally information used to indicate the sensitivity degree of the content of the text segment. For example, it can be level 1. In practice, there are usually multiple text segments in the annotated text, and for each text segment, a sensitivity level can be marked.
[0085] Here, the above-mentioned execution entity can use the access requirement description information to find the annotated text that matches the access requirement description information from at least one stored annotated text. For example, if the access requirement description information includes the title of the annotated text, the above-mentioned execution entity can use the title included in the access requirement description information to find the annotated text corresponding to the title from the blockchain.
[0086] Step 103: Perform a permission control operation on the text segment whose sensitivity level does not match the access permission level of the target account from the found annotated text to obtain the access text, and send the access text to the user terminal.
[0087] Among them, the permission control operation is used to control at least one of the following permissions: edit permission, visible permission.
[0088] Among them, the above-mentioned text segment whose sensitivity level does not match the access permission level of the target account generally refers to the text segment whose corresponding sensitivity level does not belong to the sensitivity level corresponding to the access permission level of the target account. For example, if the access permission level of target account A is level A, and the sensitivity levels corresponding to level A are level 1, level 2, and level 3. If there are 3 text segments in the annotated text, namely X, Y, and Z, and the sensitivity level corresponding to X is level 1, the sensitivity level corresponding to Y is level 2, and the sensitivity level corresponding to Z is level 4, at this time, the text segment in the annotated text whose sensitivity level does not match the access permission level of the target account is text segment Z.
[0089] In practice, if the permission control operation is used to control the edit permission, the permission control operation can include: deleting the text segment whose sensitivity level does not match the access permission level of the target account. If the permission control operation is used to control the visible permission, the permission control operation can include: setting the edit state of the text segment whose sensitivity level does not match the access permission level of the target account to the non-editable state.
[0090] Here, the above-mentioned execution entity can perform a permission control operation on the found annotated text to implement the processing of the content of the corresponding text segment.
[0091] The method provided in this embodiment can ensure data security by pre-assigning access privilege levels to each target account, so that when a user logs in to the target account through a user terminal to access the stored annotated text, only a part of the content that meets the access privilege level can be returned to the user.
[0092] Please refer to Figure 2 , Figure 2 which is a flowchart of the implementation of a data processing method provided in an embodiment of this application. The data processing method provided in this embodiment may include the following steps:
[0093] Step 201: Obtain the original medical data, and perform format conversion on the original medical data to obtain the target text in text format.
[0094] Among them, the original medical data is usually the data generated during the medical process. The original medical data may have a voice part or a text part.
[0095] Among them, the target text is usually the original medical data in text form.
[0096] Here, the above-mentioned execution entity may obtain the original medical data locally or from other devices connected by communication. Then, the obtained original medical data is converted into text format to obtain the target text.
[0097] Step 202: Split the target text into multiple text segments, and determine the segment type and the corresponding sensitivity level of each text segment according to the content of each text segment.
[0098] Here, the above-mentioned execution entity may split the target text into multiple text segments based on the segmentation in the target text. In practice, since each paragraph in the text usually expresses the same theme, the above-mentioned execution entity may take each paragraph as a text segment. Then, for each text segment, the above-mentioned execution entity may analyze the text segment, such as semantic analysis, to determine the segment type of the text segment, and find the corresponding sensitivity level from the pre-stored segment type-sensitivity level correspondence table.
[0099] It should be noted that in the field of medical data, the segment types are usually relatively fixed. For example, they may be name type, gender type, ID number type, diagnosis result type, etc.
[0100] Step 203: Perform information annotation processing on the target text according to the sensitivity levels of the text segments to obtain the annotated text, and store the annotated text in the blockchain.
[0101] Among them, information annotation processing is usually used to mark the sensitivity level of text fragments at the corresponding positions of the text fragments. In this way, it is possible to quickly find the text fragments corresponding to the sensitivity level in the marked text, which helps to improve the data processing efficiency.
[0102] Here, the above-mentioned execution entity can mark the sensitivity level corresponding to the text fragment at the position of the text fragment to obtain the target text after marking, denoted as the marked text. Then, the marked text can be stored in the blockchain. It should be noted that due to the anti-tampering feature of the blockchain, storing the marked text in the blockchain can ensure the security and reliability of the stored data.
[0103] Step 204, receive a data access request sent by the user terminal corresponding to the target account, and obtain the access permission level of the target account.
[0104] Among them, the data access request includes access requirement description information, and the target account is pre-assigned an access permission level, where the access permission level corresponds to the sensitivity level of the content to be accessed.
[0105] Step 205, search for the marked text that matches the access requirement description information from at least one marked text stored in the blockchain.
[0106] Among them, each text fragment in the marked text is marked with a sensitivity level.
[0107] Step 206, perform a permission control operation on the text fragments whose sensitivity levels do not match the access permission level of the target account from the found marked text to obtain the access text, and send the access text to the user terminal.
[0108] Among them, the permission control operation is used to control at least one of the following permissions: editing permission, visible permission.
[0109] In this embodiment, the specific operations of steps 204-206 are Figure 1 substantially the same as the operations of steps 101-103 in the embodiment shown, and will not be elaborated here.
[0110] This embodiment can process the original medical data to obtain the corresponding marked text, and store the obtained marked text in the blockchain. Due to the anti-tampering feature of the blockchain, storing the marked text in the blockchain can ensure the security and reliability of the stored data.
[0111] In some alternative implementation manners of this embodiment, splitting the target text into multiple text fragments, and determining the fragment type and the corresponding sensitivity level of the corresponding text fragment according to the content of each text fragment may include the following steps 1 to 3.
[0112] Step 1: Segment the target text to obtain multiple segmented words and the position information of each segmented word in the target text.
[0113] Here, the above-mentioned execution entity can segment the target text in multiple ways. For example, the execution entity can segment the target text using the shortest path segmentation method (N-Short Path). For another example, the execution entity can also segment the target text using the maximum probability segmentation method (Maximum Probability). For still another example, the execution entity can also segment the target text using the maximum matching method (Maximum Matching). Here, after the execution entity segments the target text, at least one segmented word in the target text can be obtained. Among them, the above-mentioned segmented word is the word obtained after segmenting the target text.
[0114] In practice, the above-mentioned execution entity usually inputs the target text into a pre-trained segmentation model to obtain multiple segmented words and the position information of each segmented word in the target text, denoted as the segmentation position information. Among them, the segmentation model is used to represent the corresponding relationship between the target text, the segmented words, and the position information of the segmented words in the target text. As an example, the segmentation model can be a model obtained by training an initial model (such as a convolutional neural network (Convolutional Neural Network, CNN), a residual network (ResNet), etc.) using machine learning methods based on training samples.
[0115] Step 2: Determine the segmentation type of each segmented word according to the preset keyword set, and segment the target text into multiple text segments according to the segmentation type of each segmented word.
[0116] Among them, the preset keywords in the preset keyword set correspond to keyword types.
[0117] Here, for each segmented word, the above-mentioned execution entity can find the same preset keyword as it in the preset keyword set, and then determine the keyword type corresponding to the found preset keyword as the segmentation type of the segmented word. Then, the execution entity can use the combination of multiple consecutive segmented words in the target text with the same segmentation type as a text segment, so as to obtain multiple text segments.
[0118] Optionally, determining the word segmentation types of each segmented word according to the preset keyword set may include: for each segmented word, calculating the similarity degree between the segmented word and each preset keyword in the preset keyword set, determining the preset keyword in the preset keyword set whose corresponding similarity degree meets the preset similarity condition as the preset keyword matching the segmented word, and determining the keyword type corresponding to the preset keyword matching the segmented word as the word segmentation type of the segmented word.
[0119] Among them, the above-mentioned preset similarity condition is usually a condition set in advance. For example, the preset similarity condition can be that the similarity degree is greater than 80%, or the similarity degree is the largest.
[0120] Here, for each segmented word, the execution entity can find the preset keyword with a relatively high similarity degree to the segmented word from the preset keyword set, and then determine the keyword type corresponding to the found preset keyword as the word segmentation type of the segmented word.
[0121] Optionally, segmenting the target text into multiple text segments according to the word segmentation types of each segmented word may include: first, traversing each segmented word in the target text. If the word segmentation type of the currently accessed segmented word is the same as that of the previous segmented word, then divide the currently accessed segmented word into the text segment to which the previous segmented word belongs. Then, if the word segmentation type of the currently accessed segmented word is different from that of the previous segmented word, then divide the currently accessed segmented word into a new text segment different from the text segment to which the previous segmented word belongs, and so on, until the text segment division of each segmented word is completed to obtain multiple text segments.
[0122] Here, the execution entity can use two adjacent segmented words with different corresponding word segmentation types as the separation points to separate the target text into multiple text segments, which can realize dividing each adjacent segmented word with the same corresponding word segmentation type into the same text segment.
[0123] Step 3: Determine the segment type of the corresponding text segment and the segment position information of the corresponding text segment in the target text according to the word segmentation type and the word segmentation position information of the segmented words included in each text segment, and determine the sensitivity level of the corresponding text segment according to the segment type of each text segment.
[0124] Here, the most frequently occurring word segmentation type in the text segment can be determined as the segment type of the text segment. And the position interval formed by the position of the first segmented word and the position of the last segmented word included in the text segment is determined as the position of the text segment. Here, the sensitivity level of the text segment can be found from the pre-stored segment type - sensitivity level correspondence table.
[0125] Please refer to Figure 3 ,Figure 3 This is a flowchart for implementing storing annotated text in a blockchain provided by an embodiment of the present application, which may include the following steps:
[0126] Step 301: Generate a first key pair for the annotated text.
[0127] Among them, the first key pair includes a first private key and a first public key.
[0128] Here, the above-mentioned execution entity may use a key generation algorithm to generate a first key pair for the annotated text.
[0129] Step 302: Encrypt the annotated text according to the first private key and store the encrypted annotated text in the blockchain.
[0130] Here, the above-mentioned execution entity may use the first private key in the first key pair to encrypt the annotated text and store the encrypted annotated text in the blockchain.
[0131] Step 303: Generate a second key pair according to the account information of the target account.
[0132] Among them, the second key pair includes a second private key and a second public key.
[0133] Here, the above-mentioned execution entity may use a key generation algorithm to generate a second key pair for the target account.
[0134] Step 304: Encrypt the first private key according to the second public key and store the encrypted first private key.
[0135] Here, the above-mentioned execution entity may use the second public key of the target account to encrypt the first private key of the annotated text to obtain the encrypted first private key.
[0136] It should be noted that encrypting the private key of the annotated text with the public key of the user can decrypt the private key of the annotated text with the private key of the user when accessing the annotated text to obtain the private key of the annotated text. Then, decrypt the annotated text with the private key of the annotated text. Realizing further confidentiality of the stored annotated data helps to further improve data security.
[0137] In some alternative implementation manners, searching for an annotated text matching the access requirement description information from at least one annotated text stored in the blockchain may include:
[0138] First, if the access requirement description information includes a text identifier, search for the corresponding encrypted first private key and search for the corresponding encrypted annotated text from at least one encrypted annotated text stored in the blockchain.
[0139] Here, when the access requirement description information includes a text identifier, the above-mentioned execution entity may use the text identifier to find the encrypted first private key corresponding to the text identifier from the pre-stored correspondence between the text identifier and the encrypted first private key. Then, the above-mentioned execution entity may find the encrypted annotated text corresponding to the text identifier from multiple encrypted annotated texts stored in the blockchain. It should be noted that the correspondence between the text identifier and the encrypted first private key may be stored in the blockchain, may be stored locally, or may be stored in other devices communicatively connected to the execution entity.
[0140] Then, decrypt the encrypted first private key according to the second public key of the target account to obtain the first private key, and decrypt the found encrypted annotated text according to the obtained first private key to obtain the annotated text that matches the access requirement description information.
[0141] Here, the above-mentioned execution entity may use the second public key of the target account to decrypt the obtained encrypted first private key to obtain the first private key. Then, the above-mentioned execution entity may use the obtained first private key to decrypt the obtained encrypted annotated text to obtain the decrypted annotated text required by the user.
[0142] It should be noted that when a user accesses an annotated text, decrypting the private key of the annotated text with the user's private key can obtain the private key of the annotated text, and then decrypting the annotated text with the private key of the annotated text can further improve data security.
[0143] Please refer to Figure 4 , Figure 4 which is a structural block diagram of a data processing device 400 provided by an embodiment of the present application. In this embodiment, each unit included in the data processing device is used to execute Figures 1 - 3 the corresponding steps in the corresponding embodiment. Specifically, please refer to Figures 1 - 3 and Figures 1 - 3 the relevant descriptions in the corresponding embodiments. For the sake of convenience of description, only the parts related to this embodiment are shown. Refer to Figure 4 , the data processing device 400 includes:
[0144] A request receiving unit 401, configured to receive a data access request sent by a user terminal corresponding to a target account, and obtain the access permission level of the target account. The data access request includes access requirement description information, and the target account is pre-assigned an access permission level, where the access permission level corresponds to the sensitivity level of the accessed content;
[0145] A text searching unit 402, configured to search for an annotated text that matches the access requirement description information from at least one annotated text stored in the blockchain, where each text segment in the annotated text is annotated with a sensitivity level;
[0146] The data control unit 403 is configured to perform a permission control operation on a text segment whose sensitivity level does not match the access permission level of the target account from the found annotated text, obtain an access text, and send the access text to the user terminal. The permission control operation is used to control at least one of the following permissions: edit permission, visibility permission.
[0147] As an embodiment of the present application, the device further includes a text storage unit (not shown in the figure). The text storage unit includes a format conversion module, a level determination module, and a storage execution module.
[0148] The format conversion module is configured to obtain the original medical data, perform format conversion on the original medical data, and obtain a target text in text format.
[0149] The level determination module is configured to split the target text into multiple text segments, and determine the segment type and corresponding sensitivity level of the corresponding text segment according to the content of each text segment.
[0150] The storage execution module is configured to perform information annotation processing on the target text according to the sensitivity level of each text segment, obtain an annotated text, and store the annotated text in the blockchain.
[0151] As an embodiment of the present application, the level determination module is specifically configured to:
[0152] Perform word segmentation processing on the target text to obtain multiple segmented words and the word segmentation position information of each segmented word in the target text.
[0153] Determine the word segmentation type of each segmented word according to the preset keyword set, and split the target text into multiple text segments according to the word segmentation type of each segmented word. The preset keywords in the preset keyword set correspond to keyword types.
[0154] Determine the segment type of the corresponding text segment and the segment position information of the corresponding text segment in the target text according to the word segmentation type and word segmentation position information of the segmented words included in each text segment, and determine the sensitivity level of the corresponding text segment according to the segment type of each text segment.
[0155] As an embodiment of the present application, in the level determination module, determining the word segmentation type of each segmented word according to the preset keyword set includes:
[0156] For each segmented word, calculate the similarity degree between the segmented word and each preset keyword in the preset keyword set, determine the preset keyword in the preset keyword set whose corresponding similarity degree meets the preset similarity condition as the preset keyword matching the segmented word, and determine the keyword type corresponding to the preset keyword matching the segmented word as the word segmentation type of the segmented word.
[0157] In one embodiment of the present application, in the level determination module, according to the word segmentation types of each segmented word, the target text is segmented into multiple text segments, including:
[0158] Traverse each segmented word in the target text. If the word segmentation type of the currently accessed segmented word is the same as that of the previous segmented word, then the currently accessed segmented word is classified into the text segment to which the previous segmented word belongs;
[0159] If the word segmentation type of the currently accessed segmented word is different from that of the previous segmented word, then the currently accessed segmented word is classified into a new text segment different from the text segment to which the previous segmented word belongs, and so on, until the text segment division of each segmented word is completed, obtaining multiple text segments.
[0160] In one embodiment of the present application, in the storage execution module, storing the annotated text in the blockchain includes:
[0161] Generate a first key pair for the annotated text, where the first key pair includes a first private key and a first public key;
[0162] Encrypt the annotated text according to the first private key, and store the encrypted annotated text in the blockchain;
[0163] Generate a second key pair according to the account information of the target account, where the second key pair includes a second private key and a second public key;
[0164] Encrypt the first private key according to the second public key, and store the encrypted first private key.
[0165] As one embodiment of the present application, the text search unit 402 is specifically configured to:
[0166] If the access requirement description information includes a text identifier, then search for the encrypted first private key corresponding to the text identifier, and search for the encrypted annotated text corresponding to the text identifier from at least one encrypted annotated text stored in the blockchain;
[0167] Decrypt the encrypted first private key according to the second public key of the target account to obtain the first private key, and decrypt the found encrypted annotated text according to the obtained first private key to obtain the annotated text matching the access requirement description information.
[0168] The device provided in this embodiment can, by pre-assigning access permission levels to each target account, when a user logs in to the target account through a user terminal to access the stored annotated text, return partial content that meets the access permission level to the user, which can ensure data security.
[0169] It should be understood that Figure 4In the structural block diagram of the data processing device shown, each unit is used to execute Figures 1 - 3 each step in the corresponding embodiment, and for Figures 1 - 3 each step in the corresponding embodiment has been explained in detail in the above embodiments. For details, please refer to Figures 1 - 3 and Figures 1 - 3 the relevant descriptions in the corresponding embodiments, which will not be elaborated here.
[0170] Figure 5 Figure 12 is a structural block diagram of a server provided in another embodiment of the present application. As Figure 5 shown, the server 500 in this embodiment includes: a processor 501, a memory 502, and a computer program 503 stored in the memory 502 and executable on the processor 501, such as a program for the data processing method. When the processor 501 executes the computer program 503, it implements the steps in each of the above-mentioned embodiments of various data processing methods, such as Figure 1 the steps 101 to 103 shown in Figure 15. Alternatively, when the processor 501 executes the computer program 503, it implements the functions of each unit in the above-mentioned Figure 4 corresponding embodiment. For example, Figure 4 the functions of the units 401 to 403 shown in Figure 19. For details, please refer to Figure 4 the relevant descriptions in the corresponding embodiment, which will not be elaborated here.
[0171] Exemplarily, the computer program 503 can be divided into one or more units. One or more units are stored in the memory 502 and executed by the processor 501 to complete the present application. One or more units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 503 in the server 500. For example, the computer program 503 can be divided into a request receiving unit, a text searching unit, and a data control unit. The specific functions of each unit are as described above.
[0172] The server may include, but is not limited to, the processor 501 and the memory 502. Those skilled in the art can understand that Figure 5 Figure 28 is only an example of the server 500 and does not constitute a limitation on the server 500. It may include more or fewer components than shown, or combine certain components, or different components. For example, the turntable device may further include input / output devices, network access devices, a bus, etc.
[0173] The so-called processor 501 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0174] The memory 502 may be an internal storage unit of the server 500, such as the hard disk or memory of the server 500. The memory 502 may also be an external storage device of the server 500, such as a plug-in hard disk equipped on the server 500, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 502 may also include both the internal storage unit of the server 500 and the external storage device. The memory 502 is used to store computer programs and other programs and data required by the turntable device. The memory 502 may also be used to temporarily store data that has been output or is to be output.
[0175] In addition, in each embodiment of the present application, the functional units may be integrated in one processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.
[0176] When an integrated module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium can be non-volatile or volatile. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of this application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.
[0177] The above embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of this application, and should all be included in the protection scope of this application.
Claims
1. A data processing method, characterized in that, The method includes: Receiving a data access request sent by a user terminal corresponding to a target account, and obtaining the access privilege level of the target account. The data access request includes access requirement description information, and the target account is pre-assigned an access privilege level, where the access privilege level corresponds to the sensitivity level of the accessed content; Searching, from at least one annotated text stored in the blockchain, for an annotated text that matches the access requirement description information, where each text segment in the annotated text is annotated with a sensitivity level; Performing a privilege control operation on the text segments whose sensitivity levels do not match the access privilege level of the target account from the found annotated text to obtain an access text, and sending the access text to the user terminal. The privilege control operation is used to control at least one of the following privileges: editing privilege, visibility privilege; The method further includes: Obtaining original medical data, performing format conversion on the original medical data to obtain a target text in text format; Performing word segmentation processing on the target text to obtain a plurality of segmented words and the word segmentation position information of each segmented word in the target text; Determining the word segmentation type of each segmented word according to a preset keyword set, and splitting the target text into a plurality of text segments according to the word segmentation type of each segmented word, where the preset keywords in the preset keyword set correspond to keyword types; Determining the segment type of the corresponding text segment and the segment position information of the corresponding text segment in the target text according to the word segmentation type and the word segmentation position information of the segmented words included in each text segment, and determining the sensitivity level of the corresponding text segment according to the segment type of each text segment; Performing information annotation processing on the target text according to the sensitivity level of each text segment to obtain an annotated text, and storing the annotated text in the blockchain.
2. The data processing method according to claim 1, characterized in that The determining the word segmentation type of each segmented word according to the preset keyword set includes: For each segmented word, calculating the similarity degree between the segmented word and each preset keyword in the preset keyword set, determining the preset keyword in the preset keyword set that corresponds to a similarity degree satisfying a preset similarity condition as the preset keyword matching the segmented word, and determining the keyword type corresponding to the preset keyword matching the segmented word as the word segmentation type of the segmented word.
3. The data processing method according to claim 1, wherein The splitting the target text into a plurality of text segments according to the word segmentation type of each segmented word includes: Traversing each segmented word in the target text. If the word segmentation type of the currently accessed segmented word is the same as that of the previous segmented word, then divide the currently accessed segmented word into the text segment to which the previous segmented word belongs; If the word segmentation type of the currently accessed segmented word is different from that of the previous segmented word, then divide the currently accessed segmented word into a new text segment different from the text segment to which the previous segmented word belongs, and so on, until the text segment division of each segmented word is completed to obtain a plurality of text segments.
4. The data processing method according to any one of claims 1 to 3, characterized in that, The storing the annotated text in the blockchain includes: Generating a first key pair for the annotated text, where the first key pair includes a first private key and a first public key; Encrypt the annotated text according to the first private key, and store the encrypted annotated text in the blockchain; Generate a second key pair according to the account information of the target account, where the second key pair includes a second private key and a second public key; Encrypt the first private key according to the second public key, and store the encrypted first private key.
5. The data processing method according to claim 4, wherein Searching for an annotated text matching the access requirement description information from at least one annotated text stored in the blockchain includes: If the access requirement description information includes a text identifier, search for the encrypted first private key corresponding to the text identifier, and search for the encrypted annotated text corresponding to the text identifier from at least one encrypted annotated text stored in the blockchain; Decrypt the encrypted first private key according to the second public key of the target account to obtain the first private key, and decrypt the found encrypted annotated text according to the obtained first private key to obtain the annotated text matching the access requirement description information.
6. A data processing device, characterized in that, The device includes: A request receiving unit, configured to receive a data access request sent by a user terminal corresponding to a target account, and obtain the access permission level of the target account. The data access request includes access requirement description information, and the target account is pre-assigned an access permission level, where the access permission level corresponds to the sensitivity level of the accessed content; A text searching unit, configured to search for an annotated text matching the access requirement description information from at least one annotated text stored in the blockchain, where each text segment in the annotated text is marked with a sensitivity level; A data control unit, configured to perform a permission control operation on a text segment whose sensitivity level does not match the access permission level of the target account from the found annotated text to obtain an access text, and send the access text to the user terminal. The permission control operation is used to control at least one of the following permissions: editing permission, visibility permission; A format conversion module, configured to obtain original medical data, and perform format conversion on the original medical data to obtain a target text in text format; A level determination module, configured to perform word segmentation processing on the target text to obtain a plurality of segmented words and the word segmentation position information of each segmented word in the target text; determine the word segmentation type of each segmented word according to a preset keyword set, and segment the target text into a plurality of text segments according to the word segmentation type of each segmented word, where the preset keywords in the preset keyword set correspond to keyword types; determine the segment type of the corresponding text segment and the segment position information of the corresponding text segment in the target text according to the word segmentation type and the word segmentation position information of the segmented words included in each text segment, and determine the sensitivity level of the corresponding text segment according to the segment type of each text segment; A storage execution module, configured to perform information annotation processing on the target text according to the sensitivity level of each text segment to obtain an annotated text, and store the annotated text in the blockchain.
7. A server, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 5 are implemented.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method and device for data verification
CN107172081A
Private data access method and device and electronic equipment
CN111400765A
Data management method and device based on block chain
CN113407954A