Regional medical information sharing method and system based on artificial intelligence

Through artificial intelligence-based methods, the separation and encryption of medical information is solved, and the problem of privacy leakage and tampering in medical information sharing is achieved, and safe and efficient information sharing is achieved.

CN120260772AActive Publication Date: 2025-07-04WUHAN SHENGBOHUI INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510325994.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-04
Estimated Expiration
2045-03-19

AI Technical Summary

Technical Problem

Existing medical information sharing technology is difficult to achieve effective sharing on the premise of ensuring information security, especially the risk of patient privacy leakage and information tampering.

Method used

Using an artificial intelligence-based method, text information is extracted through the object detection algorithm, and the text information is divided into key text and regular text using the key information classification model, and encrypted separately. The non-text information is encrypted by combining vectorization processing and group encryption algorithm, and finally uploaded to the medical information sharing terminal through a cloud server.

Benefits of technology

It realizes dual encryption of key text and non-text information during the medical information sharing process, guarantees privacy and security, prevents information leakage and tampering, adapts to the access needs of different medical roles, and simplifies the key allocation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260772A_ABST
    Figure CN120260772A_ABST
Patent Text Reader

Abstract

The invention discloses a regional medical information sharing method based on artificial intelligence, and relates to the field of information processing, and the method comprises the steps: obtaining regional medical information; extracting text information in the regional medical information; dividing the text information into a key text and a conventional text; text encryption of the key text is completed, and text encryption information is obtained; performing information compression encryption on the non-text information to obtain non-text encrypted information; performing global encryption on the conventional text, the text encryption information and the non-text encryption information to obtain medical encryption information; and uploading the medical encrypted information to a medical information sharing terminal. According to the method and the device, the medical information can be effectively shared in a state that the medical information is not leaked.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of information processing, and in particular, to a method and system for sharing regional medical information based on artificial intelligence. Background Art

[0002] With the increasing popularity of data informatization, more and more medical information is stored in each medical institution. This medical information is an important basis for medical research, clinical diagnosis, and treatment. If medical information sharing can be achieved between different medical institutions, the medical quality can be greatly improved and the development of medicine can be promoted. However, various problems are faced in the process of realizing medical information sharing. For example, since medical information belongs to sensitive data, a sharing mode without proper encryption protection may lead to the leakage of patients' privacy, resulting in serious consequences.

[0003] Existing medical information sharing technologies include directly uploading and sharing after hiding the privacy information in medical information through anonymization technology. Although this method encrypts the privacy information of patients, it does not encrypt other non-privacy information, and it is still possible to infer the privacy information of patients by combining non-privacy information, resulting in the leakage of patients' information. Summary of the Invention

[0004] Embodiments of the present application provide a method and system for sharing regional medical information based on artificial intelligence, which are used to solve the problem in the prior art that it is difficult to ensure information sharing of medical information without information leakage.

[0005] To achieve the above object, the embodiments of the present application adopt the following technical solutions:

[0006] In a first aspect, a method for sharing regional medical information based on artificial intelligence is provided. The method includes:

[0007] Obtain the regional medical information of the target area;

[0008] Locate and extract the text information in the regional medical information based on a target detection algorithm;

[0009] Input the text information into a key information classification model, output the information classification result of the text information through the key information classification model, and divide the text information into key text and regular text according to the information classification result. The key information classification model is constructed based on a recurrent neural network model;

[0010] Complete the text encryption of the key text based on a text encryption algorithm to obtain text encryption information;

[0011] Vectorize the non-text information other than the text information in the regional medical information, and perform information compression encryption on the vectorized non-text information to obtain non-text encryption information;

[0012] Based on the group encryption algorithm, globally encrypt the regular text, text encrypted information, and non-text encrypted information in the regional medical information to obtain medical encrypted information;

[0013] Upload the medical encrypted information to a pre-constructed medical information sharing terminal through a cloud server.

[0014] Optionally, the key information classification model includes a feature extraction module, an element marking module, a key degree module, and an information classification module. The steps for outputting the information classification result of the text information through the key information classification model are as follows:

[0015] Segment the text information to obtain multiple text information sequences;

[0016] For any text information sequence, input the text information sequence into the feature extraction module, and extract the text information features of the text information sequence through the feature extraction module;

[0017] Input the text information features into the element marking module, mark the element types of the text information sequence through the element marking module, and output the text element sequence of the text information sequence according to the element type marking result;

[0018] Calculate the information weight of the text element sequence through the key degree module, and output the key degree vector of the text element sequence according to the information weight;

[0019] Input all the text information features and the key degree vector into the information classification module in sequence for information classification, and output the information classification result of the text information through the information classification module.

[0020] Optionally, the steps for extracting the text information features of the text information sequence through the feature extraction module are as follows:

[0021] Convert the text information sequence into an information vector sequence through the feature extraction module;

[0022] Use the feature extraction module to complete the context semantic analysis of the information vector sequence, and extract the text semantic features of all character vectors in the information vector sequence according to the context semantic analysis result;

[0023] Based on the feature extraction module, analyze the semantic correlation degree between all character vectors, and assign weights to all text semantic features according to the semantic correlation degree;

[0024] Weightedly fuse all the text semantic features with weights assigned through the feature extraction module, and output the text information features.

[0025] Optionally, the step of calculating the information weight of the text element sequence by the key degree module and outputting the key degree vector of the text element sequence according to the information weight includes the following steps:

[0026] Calculate the sequence self-information of all text elements in the text element sequence through the key degree module and using the self-information formula, and calculate the information entropy weight of all text elements based on the sequence self-information;

[0027] Calculate the sequence mutual information between all text elements through the key degree module and using the mutual information formula, and calculate the mutual information weight of all text elements based on the sequence mutual information;

[0028] Integrate all the information entropy weights and mutual information weights to obtain the information weight of the text element sequence;

[0029] Construct a text key degree matrix of the text element sequence based on the information weight;

[0030] Vectorize the text key degree matrix, and output the key degree vector of the text element sequence through the key degree module.

[0031] Optionally, the step of vectorizing non-text information other than text information in the medical area medical information and performing information compression and encryption on the vectorized non-text information to obtain non-text encrypted information includes the following steps:

[0032] Vectorize the non-text information to obtain a non-text information vector;

[0033] Decompose the non-text information vector based on the empirical mode decomposition algorithm to obtain multiple non-text sub-informations;

[0034] For any non-text sub-information, perform lossless compression on the non-text sub-information based on the coding technology to obtain non-text compressed information;

[0035] Use the chaos system to encrypt the non-text compressed information to obtain non-text encrypted sub-information;

[0036] Integrate all non-text encrypted sub-informations to obtain non-text encrypted information.

[0037] Optionally, the step of decomposing the non-text information vector based on the empirical mode decomposition algorithm to obtain multiple non-text sub-informations includes the following steps:

[0038] Extract the extreme value interval of the non-text information vector based on the empirical mode decomposition algorithm;

[0039] Calculate the local average value of the non-text information vector according to the extreme value interval and using the integral average formula, and generate an information mean curve of the non-text information vector based on the local average value;

[0040] Iteratively extract all information components of the non-text information vector based on the information mean curve;

[0041] After reconstructing all information components, obtain multiple vector information matrices;

[0042] Construct multiple non-text sub-informations based on all vector information matrices.

[0043] Optionally, the steps for lossless compression of non-text sub-information based on coding technology to obtain non-text compressed information are as follows:

[0044] For any information element in the non-text sub-information, generate a predicted value of the information element according to multiple adjacent elements of the information element and through a linear prediction algorithm;

[0045] Calculate the absolute value of the difference between the element value of the information element and the predicted value to obtain the predicted element error.

[0046] Integrate all predicted element errors to obtain the predicted error non-text;

[0047] Decompose the predicted error non-text based on bit-plane decomposition technology to obtain multiple error bit-planes;

[0048] For any error bit-plane, divide the error bit-plane into multiple non-overlapping error blocks;

[0049] Traverse all non-overlapping error blocks to extract elements, and synthesize element planes according to the element extraction results;

[0050] Use coding technology to complete double coding of all element planes, and complete element compression of all element planes based on the double coding results to obtain non-text compressed information.

[0051] Optionally, the steps for chaotic encryption of non-text compressed information using a chaotic system to obtain non-text encrypted sub-information include the following:

[0052] Drive the chaotic system through a preset initial key to complete chaotic mapping of the non-text compressed information and generate a chaotic non-text sequence;

[0053] Quantize the chaotic non-text sequence and generate an integer index sequence according to the quantization result;

[0054] Construct a Latin square based on the integer index sequence and using a dynamic filling rule;

[0055] Scramble the non-text compressed information based on the Latin square to obtain the scrambled non-text;

[0056] Perform modulo addition operation on the scrambled non-text and the diffusion key, and obtain the intermediate non-text, where the diffusion key is generated based on the Latin square;

[0057] Perform a diffusion operation on the intermediate non-text according to the initial key to obtain the diffused non-text;

[0058] Use the diffused non-text as the new non-text compression information, and repeat the above scrambling and diffusion steps until the preset maximum encryption times are reached, and output the non-text encrypted sub-information.

[0059] Optionally, the method further includes the following steps:

[0060] Receive the medical encrypted information and extract the pre-stored reference hash information in the medical encrypted information;

[0061] Generate a hash path according to the medical encrypted information, and calculate the root hash of the medical encrypted information according to the hash path;

[0062] Integrate the hash path and the root hash to obtain the verification hash information;

[0063] Verify whether the reference hash information is consistent with the verification hash information. If the verification hash information is consistent with the verification hash information, it is determined that the medical encrypted information passes the verification;

[0064] If the verification hash information is inconsistent with the verification hash information, it is determined that the medical encrypted information fails the verification.

[0065] In a second aspect, the present application provides an artificial intelligence-based regional medical information sharing system, which is characterized by including:

[0066] A memory configured to store instructions; and

[0067] A processor configured to call instructions from the memory and be able to implement a method for sharing regional medical information based on artificial intelligence according to any one of the first aspect.

[0068] Through the above technical solutions, the text information in the regional medical information is screened out, and the text information is divided into key text and regular text by extracting the text information features and the element types included therein. The key text is encrypted locally separately, which improves the confidentiality of the key information and prevents the leakage of the key text during the sharing of medical information, and at the same time can effectively save computing power. Then, the non-text information other than the text information in the regional medical information is encrypted. After the local encryption of the key text and the non-text information is completed, the regional medical information is globally encrypted. Through the above method, double encryption of the key text and the non-text information can be achieved, ensuring its privacy and security, and ensuring that there will be no information leakage and information tampering during the sharing of medical information.

[0069] Other features and advantages of the embodiments of the present application will be described in detail in the subsequent specific implementation part. Brief Description of the Drawings

[0070] Figure 1 It is a schematic flow chart of a method for sharing regional medical information based on artificial intelligence provided by an embodiment of the present application;

[0071] Figure 2 It is a schematic structural diagram of a regional medical information sharing system provided by an embodiment of the present application;

[0072] Figure 3 It is a schematic flow chart of a process for compressing and encrypting non-text provided by an embodiment of the present application;

[0073] Figure 4 It is an example diagram of an error bit plane and an element plane provided by an embodiment of the present application;

[0074] Figure 5 It is a schematic diagram of an information compression process provided by an embodiment of the present application;

[0075] Figure 6 It is a schematic flow chart of a process based on Latin square chaotic encryption provided by an embodiment of the present application. Detailed Description of the Embodiments

[0076] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. It should be understood that the specific embodiments described herein are only used to illustrate and explain the embodiments of the present application, and are not used to limit the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0077] It should be noted that if there are directional indications (such as up, down, left, right, front, back...) involved in the embodiments of the present application, the directional indications are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly.

[0078] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of this application, such descriptions of "first", "second", etc. are only for descriptive purposes and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one such feature. Additionally, the technical solutions between various embodiments may be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0079] Figure 1 Schematically shows a flowchart of a method for sharing regional medical information based on artificial intelligence according to an embodiment of the present application. As Figure 1 shown, the embodiments of the present application provide a method for sharing regional medical information based on artificial intelligence, and the method may include the following steps:

[0080] S101. Obtain the regional medical information of the target area;

[0081] In this embodiment, the regional medical information refers to various document information generated during the medical diagnosis of patients in medical institutions, including but not limited to electronic medical records, payment bills, inspection reports (blood routine reports, urine routine reports, etc.), imaging reports (X-ray films, CT images, etc.), registration forms, etc. Since the above regional medical information contains information involving patient privacy such as patient identity identifiers (such as name, ID number), disease details (such as HIV test results), and biometric data (such as gene sequences), in order to prevent the leakage of the above privacy information during the process of information sharing, it is necessary to encrypt it before sharing the information. Additionally, when patients undergo disease diagnosis and treatment in different medical institutions, relevant medical staff may refer to the regional medical information of the patient in other hospitals. If the regional medical information is not encrypted during the process of cross-institutional information sharing, there may be a risk of tampering with the patient's regional medical information. Once the medical staff refers to the tampered regional medical information of the patient during the process of disease diagnosis and treatment for the patient, it may lead to misdiagnosis and thus affect the patient's physical condition. Based on the above, in order to ensure that the regional medical information will not be leaked or tampered with during the process of information sharing, it is necessary to encrypt the regional medical information before sharing the information to ensure the information security of the regional medical information.

[0082] S102. Locate and extract the text information in the regional medical information based on the target detection algorithm;

[0083] In this embodiment, the text information area in the regional medical information is quickly framed using an object detection algorithm, and then the text information area is text-verified to check whether it only contains text information. If the verification passes, the segmented text information is output as the final result, that is, the extraction of the text information is completed. If the verification fails, the object detection algorithm is used to perform secondary text framing on the text information area until the verification passes.

[0084] Specifically, common object detection algorithms include YOLOv3, Faster R-CNN, etc. Taking YOLOv3 as an example, YOLOv3 uses Darknet-53 (a convolutional neural network model) as the backbone network, which contains 53 convolutional layers. The residual connection is used to balance the calculation efficiency and feature extraction ability. By predicting bounding boxes on 3 feature layers of different scales (13×13, 26×26, 52×52), large, medium, and small-sized text areas are captured respectively. At the same time, the anchor box mechanism is introduced. For the aspect ratio of the text information (such as the long strip-shaped test item name, square-shaped inspection result), 9 anchor boxes are predefined (3 for each scale) to improve the positioning accuracy of dense text. CIoU Loss (Complete Intersection over Union Loss) is used as its loss function to improve the text positioning accuracy by increasing the overlap degree between the predicted box and the ground truth box and the matching degree of their geometric relationship. Finally, binary cross-entropy is used to determine whether each anchor box contains text information to avoid interference from complex backgrounds (such as the organ contours in medical images). The steps of using YOLOv3 to frame the text information in the regional medical information include: first, preprocess the medical information such as denoising and geometric correction, and then unify the sizes of all information in the regional medical information to facilitate subsequent text detection and positioning. The Feature Pyramid Network (FPN) mechanism of YOLOv3 is used to perform multi-scale feature extraction and fusion on the regional medical information, that is, by capturing large-range text (such as the report title), medium-sized text (such as the inspection item name), and small-sized text (such as the drug dosage) in the regional medical information on 3 feature layers of different scales (13×13, 26×26, 52×52) respectively, and then splicing and merging the texts of different scales to complete the framing and positioning of the text information. Finally, the area containing the text information is segmented using the bounding box coordinates output by YOLOv3, and the text information area is output.

[0085] OCR technology can be used to verify the text information area, and the steps include: first, grayscale the text information area, denoise it (such as removing interference such as ink stains and creases), and perform geometric correction (such as correcting the tilt angle of the text line, etc.). Then, split the text information into individual characters, send the split characters into the OCR engine, use pattern recognition and machine learning algorithms for recognition, and output the recognition results. For example, use the CRNN model to output the confidence of each character. If the average confidence of each character in the text information area is greater than or equal to the preset confidence threshold, it is determined that the text information area passes the verification. On the contrary, if the average confidence of each character in the text information area is less than the preset confidence threshold, it is determined that the text information area fails the verification. The CRNN model (Convolutional Recurrent Neural Network) is a deep learning model that combines the advantages of convolutional neural networks (CNN) and recurrent neural networks (RNN). It includes a convolutional layer, a recurrent layer, and a transcription layer, can handle sequences of arbitrary length, and does not depend on character segmentation or horizontal scale normalization. It is suitable for variable image text recognition tasks. In addition, CRNN can also be applied to other image sequence recognition tasks, such as music symbol recognition, etc., showing good versatility.

[0086] Since regional medical information contains multiple types of information such as text information and image information, the encryption levels and appropriate encryption methods for different types of information are different. Therefore, it is necessary to distinguish different types of text information. In addition, since the text information contains various types of data, these data include key texts with a relatively high degree of sensitivity such as patient ID numbers and patient gene test results, and also include conventional texts with a relatively low degree of sensitivity such as gender and general medical terms (such as upper respiratory tract infection, cold). Therefore, it is necessary to separately extract the text information in the regional medical information, classify the extracted text information, and encrypt the text information with different sensitivities to different degrees to ensure the information security during the sharing of medical information.

[0087] S103. Input the text information into the key information classification model, output the information classification result of the text information through the key information classification model, and divide the text information into key texts and conventional texts according to the information classification result. The key information classification model is constructed based on the recurrent neural network model;

[0088] In this embodiment, the feature extraction module is constructed based on a bidirectional gated recurrent unit and is used to extract the text information features of a text information sequence. The bidirectional gated recurrent unit is a variant of the recurrent neural network model, and its core lies in the gate mechanism. The gate mechanism includes an update gate, a retention gate, and a candidate state. The element tagging module is constructed based on the conditional random field (CRF) model. The CRF model is a discriminative probability model commonly used for annotating or analyzing sequence data. The CRF makes predictions by learning the conditional probability distribution between label sequences under the condition of a given observation sequence. Its core idea is to train the model by maximizing the conditional probability. The key degree module is constructed based on a neural network model (which can be a fully connected neural network model) and is used to output the key probability that all text information sequences are key texts based on the text information features and the key degree vector. The text information sequences with a key probability greater than the preset probability threshold can be regarded as key texts, and the text information sequences with a key probability less than or equal to the preset probability threshold can be regarded as regular texts.

[0089] To facilitate subsequent feature extraction of text information, it is necessary to preprocess the text information by standardizing it, unifying non-standard abbreviations (such as "myocardial infarction" and "heart attack") and spelling variants (such as "fever" and "pyrexia") in the text information into medical standard terms. The UMLS (Unified Medical Language System) can be used to standardize the text information. In addition, since the volume of text information is large and not conducive to the processing of the key information classification model, the text information is first segmented into multiple text information sequences according to a preset length, and the text information sequences are sequentially input into the feature extraction module for feature extraction.

[0090] Specifically, the input text information first passes through the embedding layer of the feature extraction module. First, the text information sequence is tokenized, and then each word is converted into a vector of a fixed dimension using word embedding technology to form a character vector. The character vectors are integrated to obtain an information vector sequence. The forward GRU in the encoding layer of the gated recurrent unit processes the information vector sequence from left to right to capture the historical information of the information vector sequence, and the reverse GRU processes the information vector sequence from right to left to capture the future information of the information vector sequence. The captured historical information and future information are concatenated to obtain a hidden state sequence containing bidirectional context information, that is, the text semantic feature.

[0091] Since the information vector sequence may be long, it is necessary to perform global feature extraction on the text information sequence. By introducing an attention mechanism, the feature extraction module can better obtain the influence between character vectors that are far apart. The semantic correlation degree between all character vectors can be calculated through a similarity formula. The higher the semantic correlation degree between character vectors, the higher the assigned weight. Finally, all text semantic features with weight assignment are weighted and fused to obtain text information features, which are used as the output of the feature extraction module.

[0092] Taking the text information features as the input of the element tagging module, the element tagging module is constructed based on the CRF model. By modeling the relationship between the observation features and text elements and the transition probability between text elements, the globally optimal text element sequence is finally output. First, map the text information features to a low-dimensional feature space with the dimension of the number of entity categories to obtain low-dimensional features. The number of entity categories refers to the number of information element types in the text information, and the information element type refers to the type of text elements included in the text information. The text elements include identity information, disease names, disease types, etc. The text elements can be used as the labels of the text information. Use the pre-set inference algorithm (such as the Viterbi algorithm) in the CRF model to calculate the scores between the low-dimensional features and different text element sequences, and take the text element sequence with the highest score as the output.

[0093] Count the occurrence times of each type of text element according to the text element sequence, and calculate the occurrence probability of each type of text element in all text element sequences according to the occurrence times of the text elements. Then, use the self-information formula to calculate the sequence self-information of all text elements in each text element sequence of the text information sequence, and input the calculated sequence self-information into the calculation formula of the information entropy weight in turn to calculate the information entropy weight of all text elements. Use the mutual information formula to calculate the sequence mutual information between all text elements, and input the calculated sequence mutual information into the calculation formula of the mutual information weight in turn to calculate the mutual information weight of all text elements. Integrate all the information entropy weights and mutual information weights to obtain the information weight of the text element sequence, and construct the text key degree matrix of the text element sequence based on the information weight. Vectorize the text key degree matrix, and output the key degree vector of the text element sequence through the key degree module. The text key degree matrix can be transformed into a key degree vector by using the flattening method, that is, splicing the matrix into a vector in the order of rows or columns, and directly retaining all the original data.

[0094] Finally, input all the text information features and the key degree vector into the information classification module for information classification in turn, and output the information classification result of the text information through the information classification module. The key degree module is constructed based on a neural network model (which can be a fully connected neural network model).

[0095] Since the text information contains various types of data, these data include key texts with a relatively high degree of sensitivity and also include regular texts with a relatively low degree of sensitivity. Due to the large amount of information in the text information, in order to save computing power and improve the encryption efficiency, the text information in the regional medical information is extracted separately, and the extracted text information is classified. According to the different degrees of sensitivity, the text information is divided into key texts and regular texts, and different levels of encryption are performed on the two, which not only ensures the information security of the regional medical information, but also saves computing power and reduces the encryption cost of the regional medical information.

[0096] S104. Perform text encryption on the key text based on the text encryption algorithm to obtain text encryption information;

[0097] In this embodiment, common text encryption algorithms include attribute-based encryption (ABE), homomorphic encryption, lattice-based encryption, etc. Taking the lattice-based encryption algorithm as an example, using the lattice-based encryption algorithm to perform text encryption on the key text to obtain text encryption information includes the following steps:

[0098] 1.1. Preset the encryption parameters for lattice encryption. The encryption parameters include the polynomial dimension N, the modulus q, and the modulus p. Define the polynomial ring and define the value range of the polynomial coefficients (such as -1, 0, 1). Define the polynomial ring. The modulus q must be much larger than the modulus p, and the modulus q and the modulus p are relatively prime. For example, N = 443, p = 3, and q = 677 can be set;

[0099] 1.2. Generate the private key (f, g). Randomly generate two polynomials, namely polynomial f and polynomial g, whose coefficients are randomly distributed among -1, 0, 1, and polynomial f has an inverse element under the modulus q and the modulus p. The polynomial f and the polynomial g form the private key (f, g), where the inverse element refers to in modular arithmetic, if there exists a number b such that a×b = 1 (mod c), then b is the inverse element of a under the modulus c;

[0100] 1.3. Generate the public key h and calculate the public key polynomial: is the inverse element of polynomial f under the modulus q, and mod is the modular operation;

[0101] 1.4. Divide the key text into multiple key sub-texts according to a fixed length, and encode all the key sub-texts according to a preset encoding rule to obtain multiple data information sequences. The encoding rule can be an encoding rule such as ASCII or Unicode;

[0102] 1.5. Map all data information sequences into the polynomial coefficient space, that is, convert all data information sequences into multiple information polynomials respectively. For example, if the data information sequence is [10, 20, 30], assuming p = 3 and N = 5, then each number in the data information sequence modulo 3 gives [1, 2, 0]. Since N = 5, two zeros need to be filled, resulting in the polynomial 1 + 2x + 0x 2 + 0x 3 + 0x 4 , and integrating gives 1 + 2x, that is, the data information sequence [10, 20, 30] is converted into the polynomial 1 + 2x;

[0103] 1.6. Randomly generate a noise polynomial r, whose coefficient value range is (-1, 0, 1), which is used to ensure that different ciphertexts are generated during the encryption of text information, to improve the security of text encrypted information and increase the difficulty of cracking text encrypted information;

[0104] 1.7. Perform an encryption operation to encrypt the information polynomial. The encryption formula is: s = r × h + t mod q, where t is the information polynomial and s is the encrypted information polynomial. After integrating and splicing all the encrypted information polynomials, the text encrypted information is obtained.

[0105] Since the lattice-based encryption algorithm has a high complexity and can resist traditional computing power attacks (such as brute force cracking and side-channel attacks), it can effectively guarantee the information security of key texts. In addition, due to the importance of key texts to patients, they often need to be stored for a long time. The lattice-based encryption algorithm can resist attacks from future quantum computers and can avoid the risk of leakage of text encrypted information caused by breakthroughs in quantum computing technology. Therefore, the lattice-based encryption algorithm can ensure the long-term security of key text information. In addition, precisely because the lattice-based encryption algorithm has a high complexity, in order to reduce the encryption computing power burden and reduce the encryption time, it is necessary to screen out the key texts in the text information and only perform lattice-based encryption on the key texts. While reducing the encryption computing power and the encryption time, it can also adapt to the access needs of different medical staff roles and simplify the key distribution process. For example, for medical staff with lower access rights, only the key to the medical encrypted information is provided, and the key to the medical encrypted information can only access regular texts. Only medical staff with higher access rights can provide the key to the key text. Through the above method, the leakage risk of key texts can be further reduced and the access process of regional medical information can be standardized.

[0106] S105. Vectorize the non-text information in the regional medical information other than the text information, and perform information compression and encryption on the vectorized non-text information to obtain non-text encrypted information;

[0107] In this embodiment, the non-text information mainly refers to image information, such as X-ray films, CT images, etc. Image information is usually stored in the form of a two-dimensional matrix, such as a single-channel matrix of a grayscale image or an RGB three-channel matrix of a color image, etc. Therefore, the image matrix of the non-text information can be unfolded into a one-dimensional vector in a specific order (such as row-major) to obtain a non-text information vector.

[0108] Then, the extreme points of the non-text information vector are extracted, and the extreme value intervals of the non-text information vector are divided according to the extreme points. The integral average value of the pixel values in each extreme value interval is calculated using the integral average formula, and the integral average values of adjacent extreme value intervals are connected by a straight line to form a continuous curve, that is, the information mean curve. The non-text information vector is subtracted from the information mean curve to obtain an information candidate component, and this process is repeated until the information candidate component meets the IMF condition to obtain the first IMF (information component). The remaining non-text information vector is used as the new non-text information vector to repeat the EMD decomposition step until no more information components can be extracted, and finally several information components and corresponding residuals are obtained. Each information component and the corresponding residual are rearranged into a two-dimensional matrix according to the size of the non-text information (such as M×N), that is, a vector information matrix. After normalizing and denoising the vector information matrix, a sub-image, that is, non-text sub-information, is generated.

[0109] For any non-text sub-information, the pixel values of multiple adjacent elements of any information element are used for linear prediction to obtain the predicted value of the information element. Then, the absolute value of the difference between the element value and the predicted value is calculated to obtain the prediction element error, that is, the predicted error value. After calculating the prediction element errors of all information elements, the prediction element errors are used as the pixel values of the image, and all pixel values are integrated to obtain the prediction error non-text, which can also be called the prediction error image.

[0110] The prediction error image is decomposed into multiple error bit planes. For any error bit plane, first, the error bit plane is divided into non-overlapping blocks according to a preset fixed size to obtain multiple error non-overlapping blocks. Then, in the raster scan order, the binary pixel values in each error non-overlapping block are combined into a string of bitstreams to complete element extraction, and the bitstreams are converted into multiple decimal pixel values. After all decimal pixel values are combined into an element plane. After all error bit planes are converted into element planes, each pixel plane is double-encoded using two different coding techniques, which can be Huffman coding and arithmetic coding, to complete the compression of the element plane and obtain non-text compressed information.

[0111] The initial key is converted into the control parameters and initial values of the chaotic system through a key expansion algorithm. The control parameters and initial values are used to drive the chaotic system to generate a high-entropy chaotic sequence, i.e., a chaotic non-text sequence. The floating-point numbers in the chaotic non-text sequence are mapped to an integer interval to obtain an integer index sequence. The integer index sequence is used as the index for constructing a Latin square. After obtaining multiple different Latin squares using the same method, the non-text compression information is divided into sub-blocks matching the order of the Latin square. Different Latin squares are used to perform permutation operations on different sub-blocks. According to the element values of the Latin square, the row positions and column positions of the data within the sub-blocks are permuted. The permuted sub-blocks are spliced according to the arrangement positions before division to obtain the permuted non-text. The diffusion parameters are extracted from the Latin square, and a diffusion key is generated through a non-linear function. Then, a modulo addition operation is performed on the permuted non-text and the diffusion key to obtain the intermediate non-text. Finally, a modulo addition operation is performed on the intermediate non-text and the initial key to complete the diffusion operation and obtain the diffused non-text. The diffused non-text is used as the information to be encrypted, and the permutation and diffusion steps are repeated until the preset maximum number of encryption times (which can be three times) is reached to obtain the non-text encrypted sub-information. All the non-text encrypted sub-information is spliced to obtain the non-text encrypted information.

[0112] S106. Based on the group encryption algorithm, perform global encryption on the regular text, text encrypted information, and non-text encrypted information in the regional medical information to obtain the medical encrypted information;

[0113] In this embodiment, common encryption algorithms include AES (Advanced Encryption Standard), SM4 (National Cryptography Algorithm), RSA algorithm, etc. To ensure the security of regular text, text encryption information, and non-text encryption information, AES and SM4 can be combined as group encryption algorithms to encrypt regular text, text encryption information, and non-text encryption information. First, use the AES algorithm to encrypt the information to be encrypted, where the information to be encrypted includes regular text, text encryption information, and non-text encryption information. Specifically, first select the key length of AES (128 bits, 192 bits, or 256 bits), and then select a suitable encryption mode (such as ECB, CBC, GCM, etc.); use a secure random number generator to generate an AES key, and then uniformly encode the information to be encrypted to obtain encoded information, which can be a byte array or a string. Additionally, if the length of the encoded information is not a multiple of 16 bytes, use a padding algorithm (such as PKCS#7) to pad the data to meet the block size requirements of AES; arrange the encoded information that meets the block size requirements of AES into a 4x4 matrix to obtain an encoded information matrix; arrange the AES key into a 4x4 matrix to obtain a key matrix; then perform an exclusive OR operation on the encoded information matrix and the key matrix to complete the initial round key encryption and obtain a first-level encrypted matrix; perform multiple rounds (which can be 8 rounds) of composite transformation operations such as byte substitution, row shift, column confusion, and key encryption on the first-level encrypted matrix to obtain medical encrypted information;

[0114] Among them, byte substitution means mapping each byte (matrix element) in the first-level encrypted matrix to a new value through an S-box (S substitution box) non-linear substitution table to break the statistical characteristics of the data; row shift means performing a cyclic shift operation on each row of the first-level encrypted matrix, where the first row remains unchanged, the second row is shifted left by one position, the third row is shifted left by two positions, and the fourth row is shifted left by three positions to increase the diffusion effect; column confusion means performing a matrix multiplication operation on each column of the first-level encrypted matrix to achieve confusion between columns to ensure the encryption strength; key encryption means performing an exclusive OR operation on the first-level encrypted matrix that has undergone byte substitution, row shift, and column confusion with an extended key matrix, where the extended key matrix is a matrix obtained by converting the AES key through a key expansion algorithm, and the key expansion algorithm is an algorithm that expands a shorter key into a longer, more complex, and more secure key, and a longer key can be generated by hashing the shorter key through an encryption function.

[0115] After the AES algorithm completes the encryption of the information to be encrypted, the SM4 algorithm can also be used to encrypt the AES key of the AES algorithm. The encryption steps of the SM4 algorithm include: First, select an elliptic curve and a base point G, select a random private key d, whose length can be 256 bits; use the private key d to calculate the public key P = d * G, and the public key P is used for encryption; select a random number k, whose length is the same as the order n of the elliptic curve; calculate the point Q on the elliptic curve, and the calculation formula is Q = k * G; calculate the shared secret S = H(M||Q), where H is a hash function, M is the message to be encrypted, that is, the AES key, and || is the OR operation; calculate the temporary public key T = S * G, and use the public key P and the temporary public key T to encrypt the message M to obtain the encrypted AES key.

[0116] The AES encryption algorithm has advantages such as fast encryption speed and high security. Especially for information with a large amount of data such as regional medical information, using the AES encryption algorithm can not only ensure the security of regional medical information, but also have a faster encryption speed, reduce the computing power occupation of the server or terminal device, and reduce the overall operating cost of the medical information system. However, as a symmetric encryption algorithm, the AES encryption algorithm requires both the information sender and the information receiver to share the key. However, directly transmitting the AES key has the risk of being intercepted. Since the text encrypted information and the non-text encrypted information have already undergone one round of encryption, even if the AES key is intercepted, the leakage risk of the two is still relatively small. However, for conventional texts, once the AES key is maliciously intercepted, it is very likely that the conventional text will be leaked or tampered with. Although the information contained in the conventional text is not critical information, if it is tampered with, it will still affect the medical decisions of medical staff. For example, tampering with conventional diagnostic data (such as disease classification codes) may distort the epidemiological statistics results and lead to incorrect prevention and control strategies formulated by relevant departments. Therefore, it is necessary to ensure that large-scale leakage events or tampering events do not occur in conventional texts. The SM2 encryption algorithm protects the AES key transmission through asymmetric encryption, ensuring that only the receiving party holding the private key can decrypt the AES key. This mechanism makes up for the key distribution defect of symmetric encryption. In addition, since only short data (256-bit AES key) needs to be asymmetrically encrypted, it greatly saves computing power and reduces the overall operating cost of the medical information system. In addition, by encrypting the information in multiple different methods and rounds, while ensuring the information security of regional medical information, it can also adapt to the access needs of different medical staff roles and simplify the key distribution process. For example, for medical staff with lower access rights, only the key to the medical encrypted information is provided, and only medical staff with higher access rights can provide the key to the critical text. Through the above methods, the leakage risk of critical texts can be further reduced, and the access process of regional medical information can be standardized.

[0117] S107. Upload the encrypted medical information to a pre-built medical information sharing terminal via a cloud server.

[0118] In this embodiment, a reference hash path is generated according to the medical encryption information, and a reference root hash is calculated according to the reference hash path, and the reference hash path and the reference root hash are integrated into the reference hash information. Specifically, the medical encryption information to be uploaded is divided into multiple data blocks according to a preset fixed size (such as 1MB), ensuring that each data block is independently verifiable, and applying a hash algorithm (such as SHA-256 or national encryption SM3) to each data block to generate a unique hash value as the leaf node of the Merkle tree. The hash values ​​of two adjacent leaf nodes are spliced ​​to obtain a spliced ​​hash value, and the hash algorithm is used again for the spliced ​​hash value to calculate the hash value of the spliced ​​hash value, and the hash value of the spliced ​​hash value is used as the parent node hash, and the steps of generating the parent node hash are repeated until a unique root hash is generated, and the unique root hash is used as the reference root hash for subsequent verification of the integrity of the uploaded medical encryption information. The reference hash path refers to the hash path used as the verification benchmark for subsequent information integrity verification, and the hash path contains all parent node hashes and spliced ​​hash values.

[0119] Reference Figure 2 , connected to the cloud server through SSL / TLS protocol or VPN, the client uploads the medical encrypted information to the medical information sharing terminal through the cloud server. When the medical encrypted information is uploaded, it is accompanied by the corresponding benchmark hash information and metadata for subsequent terminal verification. The metadata includes the upload time of the medical encrypted information, the encryption method of the medical encrypted information, the key index of the medical encrypted information, the identity of the uploader and other information. The key index points to the location of a specific key stored in the key management system (KMS). The key only exists in KMS and memory, and is not stored on the disk or transmitted to avoid being intercepted. Even if the metadata is leaked, the attacker cannot obtain the key based on the index alone (additional breakthrough of KMS permission control is required), which greatly ensures the security of the key.

[0120] When a target user (usually a medical staff member) needs to query regional medical information, they need to first log in to the medical information sharing terminal using a personal identity credential (such as a work number + dynamic password, biometric identification, or digital certificate). The medical information sharing terminal verifies the identity information of the medical staff member, determines the query permissions of the medical staff member, and retrieves the corresponding key index according to the query permissions. For example, if the permissions of this medical staff member are relatively low, the medical information sharing terminal will only retrieve the key index of the medical encrypted information. After the medical information sharing terminal sends the key index of the medical encrypted information to the Key Management System (KMS), the KMS will only retrieve the key of the medical encrypted information. After obtaining the key of the medical encrypted information, this medical staff member can only decrypt the medical encrypted information and can only view the regular text, but cannot view other key texts and non-text information. Only by obtaining the keys of the key texts and non-text information can other key texts and non-text information be viewed, and obtaining the keys of the key texts and non-text information requires higher query permissions. Based on this, a hierarchical encryption method is used to encrypt regional medical information. While ensuring the security of regional medical information, it is also conducive to the management of keys and the permission allocation of medical staff members. Different keys are allocated to medical staff members with different permissions, ensuring the security of regional medical information and standardizing the permission allocation and access process of regional medical information.

[0121] In one embodiment, the key information classification model includes a feature extraction module, an element marking module, a key degree module, and an information classification module. The steps for outputting the information classification result of the text information through the key information classification model are as follows:

[0122] Segment the text information to obtain multiple text information sequences;

[0123] For any text information sequence, input the text information sequence into the feature extraction module, and extract the text information features of the text information sequence through the feature extraction module;

[0124] Input the text information features into the element marking module, mark the element types of the text information sequence through the element marking module, and output the text element sequence of the text information sequence according to the element type marking result;

[0125] Calculate the information weight of the text element sequence through the key degree module, and output the key degree vector of the text element sequence according to the information weight;

[0126] Input all the text information features and the key degree vector into the information classification module in sequence for information classification, and output the information classification result of the text information through the information classification module.

[0127] In this embodiment, the feature extraction module is constructed based on a bidirectional gated recurrent unit and is used to extract the text information features of a text information sequence. The bidirectional gated recurrent unit is a variant of the recurrent neural network model, and its core lies in the gate mechanism. The gate mechanism includes an update gate, a retention gate, and a candidate state. The element tagging module is constructed based on the conditional random field (CRF) model. The CRF model is a discriminative probability model commonly used for annotating or analyzing sequence data. The CRF makes predictions by learning the conditional probability distribution between label sequences under the condition of a given observation sequence. Its core idea is to train the model by maximizing the conditional probability. The key degree module is constructed based on a neural network model (which can be a fully connected neural network model) and is used to output the key probability that all text information sequences are key texts based on the text information features and the key degree vector. The text information sequences with key probabilities greater than the preset probability threshold can be regarded as key texts, and the text information sequences with key probabilities less than or equal to the preset probability threshold can be regarded as regular texts.

[0128] To facilitate subsequent feature extraction of text information, it is necessary to preprocess the text information by standardization, unifying non-standard abbreviations (such as "myocardial infarction" and "heart attack") and spelling variants (such as "fever" and "pyrexia") in the text information into medical standard terms. The Unified Medical Language System (UMLS) can be used to standardize the text information. Additionally, since the volume of text information is large and not conducive to the processing of the key information classification model, the text information is first segmented into multiple text information sequences according to a preset length, and the text information sequences are sequentially input into the feature extraction module for feature extraction.

[0129] Specifically, the input text information first passes through the embedding layer of the feature extraction module. First, the text information sequence is tokenized, and then each word is converted into a vector of a fixed dimension using word embedding technology to form character vectors. The character vectors are integrated to obtain an information vector sequence. The forward GRU in the encoding layer of the gated recurrent unit processes the information vector sequence from left to right to capture the historical information of the information vector sequence, that is, to capture the relationship between the current character vector and the previous text. The reverse GRU processes the information vector sequence from right to left to capture the future information of the information vector sequence, that is, to capture the relationship between the current character vector and the subsequent text. The captured historical information and future information are concatenated to obtain a hidden state sequence containing bidirectional context information, that is, the text semantic feature.

[0130] Since the information vector sequence may be relatively long, it is necessary to extract global features from the text information sequence. By introducing an attention mechanism, the feature extraction module can better capture the influence between character vectors that are far apart. The semantic correlation degree between all character vectors can be calculated through a similarity formula. Commonly used similarity formulas include the cosine similarity formula, the Euclidean distance formula, etc. The character vectors are input into the similarity formula in pairs, and the similarity between the character vectors is calculated as the semantic correlation degree. The higher the semantic correlation degree between the character vectors, the higher the assigned weight. This is to identify the keyword phrases in the text information and give higher attention to the keyword phrases, which can effectively filter out irrelevant information and make the finally obtained text information features better reflect the true features of the text information sequence. Finally, all the text semantic features with weight assignment are weighted and fused to obtain the text information features, which are used as the output of the feature extraction module.

[0131] The text information features are used as the input of the element tagging module, which is constructed based on the CRF model. By modeling the relationship between the observation features and the text elements and the transition probability between text elements, the globally optimal text element sequence is finally output. First, the text information features are mapped to a low-dimensional feature space with the dimension of the number of entity categories to obtain low-dimensional features. The number of entity categories refers to the number of information element types in the text information, and the information element type refers to the type of text elements included in the text information. The text elements include identity information, disease names, disease types, etc. The text elements can be used as the labels of the text information. The scores between the low-dimensional features and different text element sequences are calculated using the pre-set inference algorithm (such as the Viterbi algorithm) in the CRF model, and the text element sequence with the highest score is used as the output. Specifically, the emission probability between the low-dimensional features and different text elements is calculated. The emission probability represents the likelihood that a certain text element generates the low-dimensional feature given the low-dimensional feature. The state transition probability between text elements is calculated, which represents the likelihood of transitioning from one text element to another. Finally, the scores between the low-dimensional features and different text element sequences are calculated based on the emission probability and the state transition probability, and the text element sequence with the highest score is selected as the output. Tagging the text elements of the text information is to facilitate the subsequent calculation of the information entropy weight and mutual information weight of different text information in the text information, so that the information entropy weight and mutual information weight can be used as one of the core factors for screening key texts.

[0132] Count the occurrence times of each type of text element according to the text element sequence, calculate the occurrence probability of each type of text element in the entire text element sequence based on the occurrence times of the text elements, then use the self-information formula to calculate the sequence self-information of all text elements in each text element sequence of the text information sequence, and input the calculated sequence self-information into the calculation formula of the information entropy weight in turn to calculate the information entropy weight of all text elements. Calculate the sequence mutual information between all text elements using the mutual information formula, and input the calculated sequence mutual information into the calculation formula of the mutual information weight in turn to calculate the mutual information weight of all text elements. Integrate all the information entropy weights and mutual information weights to obtain the information weight of the text element sequence, and construct the text key degree matrix of the text element sequence based on the information weight. Vectorize the text key degree matrix, and output the key degree vector of the text element sequence through the key degree module. The text key degree matrix can be transformed into a key degree vector using the flattening method, that is, the matrix is concatenated into a vector in the order of rows or columns, directly retaining all the original data.

[0133] Finally, input all the text information features and the key degree vector into the information classification module for information classification in turn, and output the information classification result of the text information through the information classification module. The key degree module is constructed based on a neural network model (which can be a fully connected neural network model). Taking the fully connected layer neural network model as an example, through the fully connected layer of the key degree module and using activation functions (such as ReLU, Sigmoid, Tanh) to introduce non-linear features, perform feature fusion on the text information features and the key degree vector and then map them to obtain high-order comprehensive features. Finally, output the probability value (0 - 1) of the text information being the key text through the Sigmoid activation function, and use the key text probability value as the information classification result.

[0134] In addition, the key information classification model has been pre-trained. It is trained through a pre-constructed training set, which contains historical medical text information with different text lengths and different information densities. Input the training set into the key information classification model that has not been trained yet, output the predicted value through the key information classification model that has not been trained yet, compare the predicted value with the true label, calculate the loss value (such as cross-entropy loss), and continuously optimize the key information classification model that has not been trained yet according to the loss value until the model metrics (accuracy, precision, recall, etc.) of the key information classification model that has not been trained yet reach the preset metric threshold or the number of training times reaches the preset maximum training times, and the training is completed.

[0135] Since the text information contains various types of data, these data include key texts with a relatively high degree of sensitivity such as patient ID numbers and patient gene test results, and also include conventional texts with a relatively low degree of sensitivity such as gender and general medical terms (such as upper respiratory tract infection, cold). Due to the large amount of information in the text information, in order to save computing power and improve encryption efficiency, the text information in the regional medical information is extracted separately, and the extracted text information is classified. According to the different degrees of sensitivity, the text information is divided into key texts and conventional texts, and different levels of encryption are performed on the two, which not only ensures the information security of the regional medical information, but also saves computing power and reduces the encryption cost of the regional medical information.

[0136] In one embodiment, the steps of extracting the text information features of the text information sequence by the feature extraction module are as follows:

[0137] Convert the text information sequence into an information vector sequence through the feature extraction module;

[0138] Use the feature extraction module to complete the context semantic analysis of the information vector sequence, and extract the text semantic features of all character vectors in the information vector sequence according to the context semantic analysis results;

[0139] Based on the feature extraction module, analyze the semantic correlation degree between all character vectors, and assign weights to all text semantic features according to the semantic correlation degree;

[0140] Weightedly fuse all text semantic features with weight assignment through the feature extraction module, and output text information features.

[0141] In this embodiment, the feature extraction module is constructed based on a bidirectional gated recurrent unit and an attention mechanism. The bidirectional gated recurrent unit can learn the features of time series data in two directions, forward and backward, and obtain forward and backward hidden state representations respectively. Then, the gated recurrent unit integrates and filters the information according to these two hidden states to obtain the final representation result. This structure can better capture the long-term dependence relationship in time series data and improve the performance and generalization ability of the model. The input text information first passes through the embedding layer of the feature extraction module. First, the text information sequence is tokenized, and then each word is converted into a vector of a fixed dimension using word embedding technology to form a character vector. The character vectors are integrated to obtain an information vector sequence. Word embedding maps words or phrases to vectors in the real number domain, and each word or phrase has a corresponding vector representation. This representation method not only reduces the dimension of the data, but also can capture the semantic relationship between words.

[0142] The forward GRU encoded in the gated recurrent unit processes the sequence of information vectors from left to right, capturing the historical information of the sequence of information vectors, that is, capturing the relationship between the current character vector and the previous context. The backward GRU processes the sequence of information vectors from right to left, capturing the future information of the sequence of information vectors, that is, capturing the relationship between the current character vector and the following context. The captured historical information and future information are concatenated to obtain a hidden state sequence containing bidirectional context information, that is, the text semantic feature. The text semantic feature contains both the local semantics (such as specific symptoms or terms) at each position and also integrates the global context dependence.

[0143] Since the sequence of information vectors may be long, it is necessary to perform global feature extraction on the text information sequence. By introducing the attention mechanism, the feature extraction module can better obtain the influence between character vectors that are far apart. The semantic correlation degree between all character vectors can be calculated through a similarity formula. Commonly used similarity formulas include the cosine similarity formula, the Euclidean distance formula, etc. The character vectors are input into the similarity formula in pairs, and the similarity between the character vectors is calculated as the semantic correlation degree. The principle is that character vectors with higher correlation intensity generally have closer positions in the medical text. For example, the words "diabetes" and "insulin", so when training the word embedding model, their context environments are similar, resulting in their vectors being close in space. At this time, whether using the cosine similarity or the Euclidean distance, this proximity can be detected, thereby reflecting the semantic correlation degree between them. The higher the semantic correlation degree between character vectors, the higher the assigned weight. This is to be able to identify the keyword words in the text information and give higher attention to the keyword words, which can effectively filter out irrelevant information, so that the finally obtained text information feature can better reflect the true feature of the text information sequence. Finally, the weighted fusion of all text semantic features with completed weight assignment is performed to obtain the text information feature, and the text information feature is used as the output of the feature extraction module.

[0144] Identifying the semantic features of text information is also to be able to more accurately screen out key texts. This is because text information usually contains a large number of complex terms, redundant descriptions, and implicit logics. Directly screening key texts based on surface features (such as keyword frequency) is likely to miss semantic associations or introduce noise. Extracting the semantic features of text information maps the text to a high-dimensional semantic space, capturing the deep associations between words, phrases, and paragraphs, thereby accurately positioning key texts.

[0145] In one embodiment, calculating the information weight of the text element sequence through the key degree module and outputting the key degree vector of the text element sequence according to the information weight includes the following steps:

[0146] Calculate the sequence self-information of all text elements in the text element sequence through the key degree module and using the self-information formula, and calculate the information entropy weight of all text elements based on the sequence self-information;

[0147] Calculate the sequence mutual information between all text elements through the key degree module and using the mutual information formula, and calculate the mutual information weight of all text elements based on the sequence mutual information;

[0148] Integrate all the information entropy weights and mutual information weights to obtain the information weight of the text element sequence;

[0149] Construct the text key degree matrix of the text element sequence based on the information weight;

[0150] Vectorize the text key degree matrix and output the key degree vector of the text element sequence through the key degree module.

[0151] In this embodiment, the text element sequence refers to the text elements included in the text information sequence, and the text element refers to the information type with a relatively high occurrence frequency in the historical regional medical information. For example, identity information, disease name, disease type, drug information, etc. are all text elements, and different types of text elements are numbered. For example, the number of identity information is A1, the disease name is A2, and the text element sequence is a sequence containing text element numbers. For example, a sub-text in the text information sequence contains identity information A1, mild disease A4, blood information A6, and drug information A9, then the text element sequence K of this sub-text = (A1, A4, A6, A9).

[0152] Count the occurrence times of each type of text element according to the text element sequence, calculate the occurrence probability of each type of text element in all text element sequences according to the occurrence times of the text element, and then use the self-information formula to calculate the sequence self-information of all text elements in each text element sequence of the text information sequence. The self-information formula is as follows:

[0153]

[0154] where p(x i ) is the occurrence probability of the i-th type of text element in all text element sequences, and n is the number of text element types.

[0155] Input the calculated all sequence self-informations into the calculation formula of the information entropy weight in turn to calculate the information entropy weights of all text elements. The calculation formula of the information entropy weight is as follows:

[0156]

[0157] Calculate the sequence mutual information between all text elements using the mutual information formula. The mutual information formula is as follows:

[0158]

[0159] Among them, p(x i , x j ) represents the occurrence probability of the i-th type of text element and the j-th type of text element in the same text element sequence among all text element sequences, and p(x j ) represents the occurrence probability of the j-th type of text element in all text element sequences.

[0160] Input the calculated mutual information of all sequences into the calculation formula of the mutual information weight in turn to calculate the mutual information weight of all text elements. The calculation formula of the mutual information weight is as follows:

[0161]

[0162] Integrate all information entropy weights and mutual information weights to obtain the information weight of the text element sequence, and construct a text key degree matrix of the text element sequence based on the information weight. For example, R ab is an element in the text key degree matrix. When a is equal to b, R ab is the information entropy weight of the text element. When a is not equal to b, R ab is the mutual information weight between text elements.

[0163] Vectorize the text key degree matrix and output the key degree vector of the text element sequence through the key degree module. The text key degree matrix can be transformed into a key degree vector by using the flattening method, that is, the matrix is concatenated into a vector in the order of rows or columns, directly retaining all original data.

[0164] Calculating the sequence self-information can reflect the independent information volume of text elements. The higher the sequence self-information, the more dispersed the distribution of the text element, the greater the information volume it contains, the higher the privacy leakage risk, and the higher its key degree. By calculating the self-information of text elements, their potential importance can be revealed. The sequence mutual information represents the correlation strength between two text elements. A higher mutual information indicates a strong correlation between text elements, and the combined leakage risk surges. Therefore, it shows that the key degrees of both are also higher and need to be strictly encrypted. The information entropy weight and the mutual information weight represent the key degree of a single text element among all text elements. Therefore, according to the key degree vector of each text information sequence, the importance of this text information sequence compared with other text information sequences can be measured. Therefore, the key degree vector is used as one of the factors for screening key texts.

[0165] In one of the embodiments, refer to Figure 3, vectorize non-text information other than text information in the medical information of the medical area, and perform information compression and encryption on the non-text information after the vectorization process to obtain non-text encrypted information, including the following steps:

[0166] Vectorize the non-text information to obtain a non-text information vector;

[0167] Decompose the non-text information vector based on the empirical mode decomposition algorithm to obtain multiple non-text sub-informations;

[0168] For any non-text sub-information, perform lossless compression on the non-text sub-information based on the coding technology to obtain non-text compressed information;

[0169] Use the chaos system to perform chaos encryption on the non-text compressed information to obtain non-text encrypted sub-information;

[0170] Integrate all non-text encrypted sub-informations to obtain non-text encrypted information.

[0171] In this embodiment, the non-text information mainly refers to image information, for example, X-ray films, CT images, etc. Image information is usually stored in the form of a two-dimensional matrix, such as a single-channel matrix of a grayscale image or an RGB three-channel matrix of a color image, etc. Therefore, the image matrix of the non-text information can be expanded into a one-dimensional vector in a specific order (such as row-first) to obtain a non-text information vector. For example, image information with a size of M×N can be transformed into a vector with a length of M×N, and each element corresponds to the pixel value of the pixel. This process preserves the spatial relationship of the pixels, which is convenient for generating non-text sub-informations after the EMD decomposition of the non-text information vector, that is, performing EMD decomposition on the image information to obtain multiple sub-images.

[0172] First, check the pixel values of each data point in the non-text information vector one by one from left to right, extract the extreme points of the non-text information vector, and divide the extreme value intervals of the non-text information vector according to the extreme points. Calculate the integral average value of the pixel values in each extreme value interval using the integral average formula. The integral average value is the local average value of the non-text information vector. After calculating the integral average values of all extreme value intervals, connect the integral average values of adjacent extreme value intervals with a straight line to form a continuous curve, that is, the information mean curve. Subtract the information mean curve from the non-text information vector to obtain an information candidate component. Repeat this process until the information candidate component meets the IMF condition to obtain the first IMF (information component). Use the remaining non-text information vector as a new non-text information vector and repeat the EMD decomposition steps until no more information components can be extracted. Finally, obtain several information components and the corresponding residuals. Rearrange each information component and the corresponding residual into a two-dimensional matrix according to the size of the non-text information (such as M×N), that is, a vector information matrix. After normalizing and denoising the vector information matrix, generate sub-images, that is, non-text sub-informations.

[0173] For any non-text sub-information, the predicted value of an information element is obtained by linearly predicting the pixel values of multiple adjacent elements of any one of the information elements. Then, the absolute value of the difference between the element value and the predicted value is calculated to obtain the predicted element error, that is, the predicted error value. After calculating the predicted element errors of all information elements, the predicted element errors are used as the pixel values of the image, and all pixel values are integrated to obtain the predicted error non-text, which can also be called the predicted error image.

[0174] The predicted error image is decomposed into multiple error bit planes, and the error bit planes can be divided into high-order planes and low-order planes. By representing the pixel values (gray values or pixel brightness) of the predicted error image in binary, the binary pixel values are obtained, and then the binary pixel values are arranged in order from high to low to form the error bit planes. For any error bit plane, first, the error bit plane is divided into non-overlapping blocks of a preset fixed size in the order from left to right and top to bottom to obtain multiple error non-overlapping blocks. Then, in the raster scan order, the binary pixel values in each error non-overlapping block are combined into a bit stream to complete element extraction, and the bit stream is converted into multiple decimal pixel values, and all decimal pixel values are combined into an element plane. After all error bit planes are converted into element planes, each pixel plane is double-coded using two different coding techniques, which can be Huffman coding and arithmetic coding, to complete the compression of the element plane and obtain the non-text compressed information.

[0175] The initial key is preset (such as a 128-bit binary string). The initial key is converted into the control parameters and initial values of the chaotic system through the key expansion algorithm. The control parameters and initial values are used to drive the chaotic system to generate a high-entropy chaotic sequence, that is, the chaotic non-text sequence. The floating-point numbers in the chaotic non-text sequence are mapped to the integer interval to obtain an integer index sequence. The integer index sequence is used as the index for constructing the Latin square, that is, the Latin square is constructed using the integer index sequence. Specifically, first, an empty initial matrix is initialized, and then the initial matrix is filled row by row. For example, if the integer index sequence is {Q1, Q2, Q3,..., Qn}, and the value of each element in the integer index sequence belongs to {0, 1, 2,..., n - 1}, the first row of the initial matrix is filled according to the integer index sequence to obtain (0, 1, 2,..., n - 1), and the arrangement of the i-th row (i is greater than or equal to 2) of the initial matrix is the (i - 1)-th row circularly shifted to the right by Qi times until the number of rows is the same as the number of columns.

[0176] After obtaining multiple different Latin squares using the same method, divide the non-text compressed information into sub-blocks that match the order of the Latin square, and use different Latin squares to perform permutation operations on different sub-blocks. Permute the row positions and column positions of the data within the sub-blocks according to the element values of the Latin square. Concatenate the permuted sub-blocks in the arrangement position before division to obtain the permuted non-text. Extract diffusion parameters (such as row sum, column XOR value) from the Latin square, generate a diffusion key through a non-linear function, then perform modulo addition operation on the permuted non-text and the diffusion key to obtain the intermediate non-text, achieving bit-level diffusion. Finally, perform modulo addition operation on the intermediate non-text and the initial key to complete the diffusion operation and obtain the diffused non-text. Use the diffused non-text as the information to be encrypted and repeat the permutation and diffusion steps until the preset maximum encryption times (which can be three times) are reached to obtain the non-text encrypted sub-information, and concatenate all the non-text encrypted sub-information to obtain the non-text encrypted information.

[0177] In one embodiment, decomposing the non-text information vector based on the empirical mode decomposition algorithm to obtain multiple non-text sub-informations includes the following steps:

[0178] Extract the extreme value intervals of the non-text information vector based on the empirical mode decomposition algorithm;

[0179] Calculate the local average value of the non-text information vector according to the extreme value intervals and through the integral average formula, and generate the information mean curve of the non-text information vector based on the local average value;

[0180] Iteratively extract all the information components of the non-text information vector based on the information mean curve;

[0181] After reconstructing all the information components, obtain multiple vector information matrices;

[0182] Construct multiple non-text sub-informations based on all the vector information matrices.

[0183] In this embodiment, since the empirical mode decomposition algorithm is based on the time-scale characteristics of the signal itself and does not require prior setting of basis functions, it has stronger self-adaptability than Fourier analysis and wavelet analysis. In theory, it can be applied to the decomposition of any type of signal. The basic principle of the EMD algorithm is to decompose a real-valued signal into a series of intrinsic mode functions by separating the local mean. In this application, a chaotic encryption algorithm is used to encrypt non-text information. Among them, the chaotic system used for encryption includes a low-dimensional chaotic information and a high-dimensional chaotic system. However, the chaotic sequence generated by the low-dimensional chaotic system has disadvantages such as short cipher period, low precision, and small key space. Although the chaotic sequence generated by the high-dimensional chaotic system can overcome the above deficiencies, the increase in the dimension of the chaotic system will inevitably increase the computational complexity, affect the computational accuracy, and the encryption algorithm becomes very complex, which may lead to difficulties in later decryption. Therefore, in order to overcome the above problems, it is necessary to use the empirical mode decomposition algorithm to decompose the non-text information into several non-text sub-information, and then compress and encrypt the non-text sub-information.

[0184] First, check the pixel values of each data point in the non-text information vector one by one from left to right. If the pixel value of a certain data point is greater than the pixel values of its left and right adjacent data points, then mark this data point as a maximum point. On the contrary, if the pixel value of a certain data point is less than the pixel values of its left and right adjacent data points, then mark this data point as a minimum point. In addition, for the data points at the starting point and the ending point in the non-text information vector, only the pixel values of the one-sided adjacent data points need to be compared. By arranging the maximum points and minimum points in the marked order, the non-text information vector can be divided into multiple extreme value intervals. The extreme value interval refers to the region between adjacent maximum points and minimum points.

[0185] Use the integral average formula to calculate the integral average value of the pixel values in each extreme value interval. The integral average value is the local average value of the non-text information vector. The integral average formula is:

[0186]

[0187] where, j k+1 is the pixel value corresponding to the k + 1th maximum point or minimum point, j k is the pixel value corresponding to the kth maximum point or minimum point, and F(j) is the non-text information vector.

[0188] After calculating the integral average value of all extreme value intervals, connect the integral average values of adjacent extreme value intervals with straight lines to form a continuous curve, that is, the information mean curve. The traditional envelope method calculates the mean through the upper and lower envelope lines, while the integral average method directly divides the intervals based on the extreme points and uses the integral to reflect the overall energy of the pixel values within the intervals, avoiding the overshoot or undershoot errors (such as the oscillation of cubic spline interpolation at the endpoints) that may be introduced by envelope interpolation, and can more intuitively reflect the statistical characteristics of the pixel values within the extreme value intervals. Subtract the information mean curve from the non-text information vector to obtain the information candidate component, and repeat this process until the information candidate component meets the IMF conditions. The IMF conditions refer to that the number of extreme points (maximum points or minimum points) is close to the number of zero-crossing points, and the envelope mean is zero, to obtain the first IMF (information component). Take the remaining non-text information vector as the new non-text information vector and repeat the EMD decomposition steps until no more information components can be extracted, and finally obtain several information components and the corresponding residuals.

[0189] Rearrange each information component and the corresponding residual into a two-dimensional matrix according to the size of the non-text information (such as M×N), that is, the vector information matrix. After normalizing and denoising the vector information matrix, generate a sub-image, that is, the non-text sub-information. This is because the amplitude ranges of each information component may be different, and normalization can ensure consistent contrast during visualization. In addition, high-frequency information components may have noise, so methods such as Gaussian filtering and median filtering need to be used for denoising processing to make the finally obtained non-text sub-information of higher quality.

[0190] In one embodiment, the lossless compression of the non-text sub-information is completed based on the coding technology, and the non-text compressed information is obtained as follows:

[0191] For any information element in the non-text sub-information, generate a predicted value of the information element according to multiple adjacent elements of the information element through a linear prediction algorithm;

[0192] Calculate the absolute value of the difference between the element value of the information element and the predicted value to obtain the predicted element error;

[0193] Integrate all the predicted element errors to obtain the predicted error non-text;

[0194] Decompose the predicted error non-text based on the bit-plane decomposition technology to obtain multiple error bit-planes;

[0195] For any error bit-plane, divide the error bit-plane into multiple non-overlapping error blocks;

[0196] Traverse all the non-overlapping error blocks for element extraction, and synthesize the element plane according to the element extraction results;

[0197] Complete the dual coding of all element planes using coding techniques, and based on the dual coding results, compress the elements of all element planes to obtain non-text compression information.

[0198] In this embodiment, the non-text information is used as image information, which contains a large number of pixel information. There is often a correlation between adjacent pixels, and their pixel values are mostly close in most cases. Based on this characteristic, the pixel values of multiple adjacent elements of the information element are linearly predicted to obtain the predicted value of the information element. Since the non-text information is image information, the non-text sub-information is a sub-image of the image information, the information element and the adjacent elements are the pixels of the sub-image, the predicted value is the predicted pixel value, the element value of the information element refers to the original pixel value of the pixel in the sub-image, and the pixel value refers to the color or brightness information contained in each pixel in the image. For example, the pixel values of a grayscale image are usually represented by integers from 0 to 255, 0 represents black, 255 represents white, and the intermediate values represent different grayscales. The linear prediction algorithm is a method for predicting future values through linear regression analysis. For example, the mean value of the pixel values between all adjacent elements of a certain information element is calculated as the predicted element of the information element. Then, the absolute value of the difference between the element value and the predicted value is calculated to obtain the predicted element error, that is, the predicted error value. The histogram of the predicted element error calculated by this method usually presents a Laplacian-like distribution, with more redundant space available for compression. After calculating the predicted element errors of all information elements, the predicted element errors are used as the pixel values of the image, and all pixel values are integrated to obtain the predicted error non-text, which can also be called the predicted error image.

[0199] The prediction error image is decomposed into multiple error bit planes, which can be divided into high-order planes and low-order planes. By representing the pixel values (gray values or pixel brightness) of the prediction error image in binary, the binary pixel values are obtained, and then the binary pixel values are arranged in order from the highest bit to the lowest bit to form the error bit planes. Each error bit plane contains the bit information at the same position of all pixels in the prediction error image. The importance of these bit planes gradually decreases because the information in the high-order bits is usually more important than that in the low-order bits. For example, each pixel value of an 8-bit gray image can be decomposed into 8 binary bits, forming 8 bit planes from the lowest bit to the highest bit. Each bit plane only contains the 0 or 1 value of the corresponding bit position. Since there are still many 0s in the high-order planes obtained by decomposing the prediction error image, and the low-order planes show a disorderly situation of 0s and 1s, aiming at this characteristic of the bit planes, a method of synthesizing pixels by dividing bit planes is designed to aggregate the information of each bit plane, compress the 0 values in the high-order planes and the disordered data in the low-order planes as much as possible, so as to facilitate making more embeddable space. This is because more embeddable space allows non-text encrypted information to be dispersed in multiple dimensions (such as pixel values, frequency domain coefficients, bit planes, etc.), avoiding attackers from obtaining the complete key through local cracking, and can also better support more complex chaotic encryption operations in the follow-up, improving the security of non-text encrypted information.

[0200] Referring to Figure 4 , for any error bit plane, first divide the error bit plane into non-overlapping blocks according to a preset fixed size in the order from left to right and from top to bottom, obtaining multiple error non-overlapping blocks. Non-overlapping block division means that in image processing, a large-sized image is divided into several non-overlapping sub-blocks. This division method makes each sub-block independent in space and does not interfere with each other, thus simplifying the processing process. Then, through the raster scan order, the binary pixel values in each error non-overlapping block are combined into a string of bitstreams to complete element extraction. The bitstreams are converted into multiple decimal pixel values, and all the decimal pixel values are combined into an element plane. Among them, a bitstream refers to a data stream composed of a string of binary numbers, and the raster scan order refers to the order of scanning from left to right, from top to bottom, scanning one row first and then moving to the starting position of the next row to continue scanning. The element plane can also be called a pixel plane, that is, a planar image containing pixel values.

[0201] Referring to Figure 5, after converting all error bit planes into element planes, each pixel plane is double-coded using two different coding techniques, which can be Huffman coding and arithmetic coding. Specifically, the occurrence frequency of pixel values in each element plane is counted to obtain the pixel frequency, and a corresponding Huffman dictionary is generated according to the pixel frequency. Short codes are assigned to high-frequency pixel values, and long codes are assigned to low-frequency values. This method can ensure that the compressed data has a smaller volume while maintaining the integrity of the original data. The element plane corresponding to the high-order bit plane has more continuously repeated pixel values, and Huffman coding can significantly shorten the code length of high-frequency symbols (such as long strings of 0 or 1), significantly reducing its pixel amount. The element plane corresponding to the low-order bit plane may have fewer continuously repeated pixel values, but there may be repeated values locally. Huffman coding can compress these limited repeated pixel values. Since there may still be continuously identical pixel values (such as long strings of 0 or 1) in the element plane corresponding to the high-order bit plane after Huffman coding, arithmetic coding is used to further eliminate the repeated pixel values in the element plane corresponding to the high-order bit plane, completing the further compression of the element plane and obtaining non-text compressed information.

[0202] Huffman coding is an entropy coding algorithm for lossless data compression. Its basic principle is to construct a Huffman tree (also known as an optimal binary tree) and use the frequency of character occurrence to assign different coding lengths to achieve data compression. Arithmetic coding is also a lossless data compression method and belongs to a type of entropy coding. Its core idea is to encode the entire input message as a decimal number between 0 and 1 according to the probabilities of different symbols appearing in the input message, and select a representative decimal number within the final mapping interval as the coding output. This decimal number represented in binary form is the final coding result.

[0203] Through the above method, the compression of each non-text sub-information can be achieved. After compression, non-text compressed information is obtained, and its data volume is significantly reduced compared to the non-text sub-information, which can effectively improve the efficiency of subsequent chaotic encryption. At the same time, compression eliminates the statistical redundancy between pixels (such as continuously identical pixel values), making chaotic encryption more focused on effective information and avoiding redundant data from diluting the diffusion property of the chaotic encryption algorithm. In addition, compression destroys the spatial correlation between pixels, and the diffusion property of chaotic encryption can further disrupt the intra-block structure. Under the dual measures, it is more difficult to infer the plaintext pattern from the ciphertext, improving the security of the non-text encrypted information.

[0204] In one embodiment, chaotic encryption of non-text compressed information using a chaotic system to obtain non-text encrypted sub-information includes the following steps:

[0205] Drive the chaotic system through a pre-set initial key to complete the chaotic mapping of the non-text compressed information and generate a chaotic non-text sequence;

[0206] Quantize the chaotic non-text sequence and generate an integer index sequence according to the quantization result;

[0207] Construct a Latin square based on the integer index sequence and using the dynamic filling rule;

[0208] Scramble the non-text compressed information with the Latin square as a reference to obtain the scrambled non-text;

[0209] Perform modulo addition operation on the scrambled non-text and the diffusion key to obtain the intermediate non-text, and the diffusion key is generated based on the Latin square;

[0210] Perform diffusion operation on the intermediate non-text according to the initial key to obtain the diffused non-text;

[0211] Use the diffused non-text as the new non-text compressed information and repeat the above scrambling and diffusion steps until the preset maximum encryption times are reached, and output the non-text encrypted sub-information.

[0212] In this embodiment, the principle of chaotic encryption is based on the non-linearity of the chaotic system and the sensitivity to the initial conditions. By generating complex random sequences for data confusion and diffusion, encryption is achieved. The chaotic system has inherent randomness and extreme sensitivity to the initial conditions, which makes the chaotic sequence show good randomness and complexity in encryption and is difficult to be predicted and reconstructed.

[0213] Refer to Figure 6 , the initial key is preset (such as a 128-bit binary string). The initial key is converted into the control parameters and initial values of the chaotic system (such as the μ value of the Logistic map, the σ value of the Lorenz system, etc.) through the key expansion algorithm. The control parameters and initial values are used to drive the chaotic system to generate a high-entropy chaotic sequence, that is, the chaotic non-text sequence. Since the chaotic non-text sequence is a floating-point number sequence and the floating-point number is realized by the combination of the mantissa and the exponent, linear quantization or non-linear quantization is performed on the floating-point numbers in the chaotic non-text sequence to obtain an integer index sequence. This process is a process of mapping the floating-point numbers in the chaotic non-text sequence to the integer interval. Use the integer index sequence as the index for constructing the Latin square. Specifically, first initialize an empty initial matrix, and then fill the initial matrix row by row. For example, the integer index sequence is {Q1, Q2, Q3,..., Qn}, and the value of each element in the integer index sequence belongs to {0, 1, 2,..., n - 1}. Fill the first row of the initial matrix according to the integer index sequence to get (0, 1, 2,..., n - 1). The arrangement of the i-th row (i is greater than or equal to 2) of the initial matrix is the (i - 1)-th row circularly shifted to the right Qi times until the number of rows is the same as the number of columns. For example, the finally obtained Latin square can be:

[0214]

[0215] After obtaining multiple different Latin squares using the same method, divide the non-text compression information into sub-blocks that match the order of the Latin square. Use different Latin squares to perform permutation operations on different sub-blocks. According to the element values of the Latin square, permute the row positions and column positions of the data within the sub-block, that is, rearrange the arrangement order of the data in the i-th row within the sub-block according to the arrangement order of the elements in the i-th row of the Latin square, and rearrange the arrangement order of the data in the i-th column within the sub-block according to the arrangement order of the elements in the i-th column of the Latin square. For example, if the third row of the Latin square is [3, 0, 2, 1], then the original row arrangement order of the sub-block is permuted from [a, b, c, d] to [b, d, a, c]. Concatenate the scrambled sub-blocks according to the arrangement positions before division to obtain the scrambled non-text. Extract diffusion parameters (such as row sum, column XOR value) from the Latin square, and generate a diffusion key through a non-linear function. The non-linear function can be an S-box or a Feistel network, etc. Taking the S-box as an example, divide the diffusion parameter block into S-box input units (such as 4 bits or 8 bits), and perform look-up table replacement through a pre-defined or dynamically generated S-box to obtain the diffusion key. Then perform modulo addition operation on the scrambled non-text and the diffusion key to obtain the intermediate non-text, that is, the image obtained during the encryption process, to achieve bit-level diffusion. Finally, perform modulo addition operation on the intermediate non-text and the initial key to complete the diffusion operation and obtain the diffused non-text. Use the diffused non-text as the information to be encrypted and repeat the scrambling and diffusion steps until the preset maximum encryption times (which can be three times) are reached to obtain the non-text encrypted sub-information. The modulo addition operation means d = e + f (mod n), where d is the result of the modulo addition operation, and mod represents the modulo operation, that is, the operation of finding the remainder.

[0216] Implement multi-level scrambling and modulo addition diffusion through dynamically generated Latin squares to destroy the correlation of adjacent pixels and enhance the ability to resist statistical attacks. At the same time, for the high redundancy characteristics of non-text information, quickly break the continuity of same-color pixels through chaotic scrambling to avoid the low efficiency problem of traditional encryption algorithms. In short, the complexity of chaotic mapping and Latin square permutation operations is low, which is suitable for encrypting large amounts of data of image information (non-text information) in regional medical information. While ensuring the security of non-text information, the encryption speed is also fast.

[0217] In one of the embodiments, the method further includes the following steps:

[0218] Receive medical encryption information and extract the pre-stored reference hash information in the medical encryption information;

[0219] Generate a hash path according to the medical encryption information and calculate the root hash of the medical encryption information according to the hash path;

[0220] Integrate the hash path and the root hash to obtain the verification hash information;

[0221] Verify whether the reference hash information is consistent with the verification hash information. If the verification hash information is consistent with the verification hash information, it is determined that the medical encrypted information passes the verification;

[0222] If the verification hash information is inconsistent with the verification hash information, it is determined that the medical encrypted information fails the verification.

[0223] In this embodiment, the reference hash information is pre-calculated, which includes a reference hash path and a reference root hash. The calculation methods of the hash path and the root hash are exactly the same as those of the reference hash information. First, generate a hash path based on the medical encrypted information, and calculate the root hash according to the hash path. Integrate the hash path and the root hash into the reference hash information. Specifically, divide the uploaded medical encrypted information into multiple data blocks according to a preset fixed size (such as 1MB), ensure that each data block is independently verifiable, and apply a hash algorithm (such as SHA-256 or national cipher SM3) to each data block to generate a unique hash value as the leaf node of the Merkle tree. Concatenate the hash values of adjacent two leaf nodes to obtain a concatenated hash value, use the hash algorithm on the concatenated hash value again to calculate the hash value of the concatenated hash value, and use the hash value of the concatenated hash value as the parent node hash. Repeat the step of generating the parent node hash until a unique root hash is generated. The hash path is the set of all parent node hashes and concatenated hash values. Compare whether the root hash is consistent with the reference hash, and verify whether the hash path is consistent with the reference hash path. If both are consistent, it is determined that the medical encrypted information passes the verification. If either the root hash or the hash path is inconsistent with the reference hash or the reference hash path in the reference hash information, it is determined that the medical encrypted information fails the integrity verification and may be tampered with, and the uploader of the medical encrypted information needs to re-encrypt and upload the regional medical information.

[0224] This application also discloses an artificial intelligence-based regional medical information sharing system, which is characterized by including:

[0225] A memory configured to store instructions; and

[0226] A processor configured to call instructions from the memory and be able to implement a method for artificial intelligence-based regional medical information sharing according to any one of the above.

[0227] Among them, the processor can adopt a central processing unit (CPU). Of course, according to the actual usage situation, other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. can also be adopted. The general-purpose processor can adopt a microprocessor or any conventional processor, etc. This application does not make any restrictions on this.

[0228] Among them, the memory can be an internal storage unit of the computer device, for example, the hard disk or memory of the computer device, or an external storage device of the computer device, for example, a plug-in hard disk, a smart memory card (SMC), a secure digital card (SD), or a flash memory card (FC) equipped on the computer device. Moreover, the memory can also be a combination of the internal storage unit and the external storage device of the computer device. The memory is used to store computer programs and other programs and data required by the computer device. The memory can also be used to temporarily store the data that has been output or will be output. This application does not make any restrictions on this.

[0229] The embodiment of the present application also provides a machine-readable storage medium, on which instructions are stored, and these instructions are used to make the machine execute the above-mentioned method for sharing regional medical information based on artificial intelligence.

[0230] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0231] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0232] These computer program instructions can also be stored in a computer-readable memory that can guide the computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device realizes the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0233] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 steps of the functions specified in one block or multiple blocks.

[0234] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0235] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.

[0236] Computer-readable media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory media, such as modulated data signals and carrier waves.

[0237] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, commodity or device including the element.

[0238] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A method for sharing regional medical information based on artificial intelligence, characterized in that , The method includes the following steps: Obtain the regional medical information of the target area; Based on the target detection algorithm, locate and extract the text information in the regional medical information; Input the text information into the key information classification model, and output the information classification result of the text information through the key information classification model. Divide the text information into key text and regular text according to the information classification result. The key information classification model is constructed based on the recurrent neural network model; Complete the text encryption of the key text based on the text encryption algorithm to obtain the text encryption information; Vectorize the non-text information in the regional medical information other than the text information, and perform information compression encryption on the vectorized non-text information to obtain the non-text encryption information; Complete the global encryption of the regular text, text encryption information and non-text encryption information in the regional medical information based on the group encryption algorithm to obtain the medical encryption information; Upload the medical encryption information to the pre-constructed medical information sharing terminal through the cloud server.

2. The method according to claim 1, characterized in that, The key information classification model includes a feature extraction module, an element marking module, a key degree module and an information classification module. The steps for outputting the information classification result of the text information through the key information classification model are as follows: Segment the text information to obtain multiple text information sequences; For any text information sequence, input the text information sequence into the feature extraction module, and extract the text information features of the text information sequence through the feature extraction module; Input the text information features into the element marking module, mark the element types of the text information sequence through the element marking module, and output the text element sequence of the text information sequence according to the element type marking result; Calculate the information weight of the text element sequence through the key degree module, and output the key degree vector of the text element sequence according to the information weight; Input all the text information features and the key degree vector into the information classification module in sequence for information classification, and output the information classification result of the text information through the information classification module.

3. The method according to claim 2, characterized in that, The steps for extracting the text information features of the text information sequence through the feature extraction module are as follows: Convert the text information sequence into an information vector sequence through the feature extraction module; Use the feature extraction module to complete the context semantic analysis of the information vector sequence, and extract the text semantic features of all character vectors in the information vector sequence according to the context semantic analysis result; Analyze the semantic correlation degree between all character vectors based on the feature extraction module, and assign weights to all text semantic features according to the semantic correlation degree; Weight and fuse all the text semantic features with weights assigned through the feature extraction module, and output the text information features.

4. The method according to claim 2, characterized in that The steps for calculating the information weight of the text element sequence through the key degree module and outputting the key degree vector of the text element sequence according to the information weight are as follows: Calculate the sequence self-information of all text elements in the text element sequence through the key degree module and using the self-information formula, and calculate the information entropy weight of all text elements based on the sequence self-information; Calculate the sequence mutual information between all text elements through the key degree module and using the mutual information formula, and calculate the mutual information weight of all text elements based on the sequence mutual information; Integrate all information entropy weights and mutual information weights to obtain the information weight of the text element sequence; Construct the text key degree matrix of the text element sequence based on the information weight; Vectorize the text key degree matrix and output the key degree vector of the text element sequence through the key degree module.

5. The method according to claim 1, characterized in that, The non-text information other than text information in the medical area medical information is vectorized, and the vectorized non-text information is compressed and encrypted to obtain non-text encrypted information, including the following steps: Vectorize the non-text information to obtain a non-text information vector; Decompose the non-text information vector based on the empirical mode decomposition algorithm to obtain multiple non-text sub-informations; For any non-text sub-information, perform lossless compression on the non-text sub-information based on the coding technology to obtain non-text compressed information; Use the chaotic system to encrypt the non-text compressed information chaotically to obtain non-text encrypted sub-information; Integrate all non-text encrypted sub-informations to obtain non-text encrypted information.

6. The method according to claim 5, characterized in that, The non-text information vector is decomposed based on the empirical mode decomposition algorithm to obtain multiple non-text sub-informations, including the following steps: Extract the extreme value interval of the non-text information vector based on the empirical mode decomposition algorithm; Calculate the local average value of the non-text information vector according to the extreme value interval and through the integral average formula, and generate the information mean curve of the non-text information vector based on the local average value; Iteratively extract all information components of the non-text information vector based on the information mean curve; After reconstructing all information components, obtain multiple vector information matrices; Construct multiple non-text sub-informations based on all vector information matrices.

7. The method according to claim 5, characterized in that, The lossless compression of the non-text sub-information is completed based on the coding technology to obtain non-text compressed information, as follows: For any information element in the non-text sub-information, generate the predicted value of the information element according to multiple adjacent elements of the information element and through the linear prediction algorithm; Calculate the absolute value of the difference between the element value of the information element and the predicted value to obtain the predicted element error; Integrate all predicted element errors to obtain the predicted error non-text; Decompose the predicted error non-text based on the bit plane decomposition technology to obtain multiple error bit planes; For any error bit plane, divide the error bit plane into multiple non-overlapping error blocks; Traverse all non-overlapping error blocks to extract elements, and synthesize element planes according to the element extraction results; Use the coding technology to complete the double coding of all element planes, and complete the element compression of all element planes based on the double coding results to obtain non-text compressed information.

8. The method according to claim 5, characterized in that, The non-text compressed information is encrypted chaotically using the chaotic system to obtain non-text encrypted sub-information, including the following steps: Drive the chaotic system through a preset initial key to complete the chaotic mapping of the non-text compressed information and generate a chaotic non-text sequence; Quantize the chaotic non-text sequence and generate an integer index sequence according to the quantization result; Construct a Latin square based on the integer index sequence and using the dynamic filling rule; Scramble the non-text compressed information based on the Latin square to obtain scrambled non-text; Perform modulo addition operation on the scrambled non-text and the diffusion key, and obtain intermediate non-text, where the diffusion key is generated based on the Latin square; Perform diffusion operation on the intermediate non-text according to the initial key to obtain diffused non-text; Diffuse the non-text as new non-text compressed information, and repeat the above scrambling and diffusion steps until the preset maximum encryption times are reached, and output the non-text encrypted sub-information.

9. The method according to claim 1, characterized in that, The method further includes the following steps: Receive the medical encrypted information, and extract the pre-stored reference hash information in the medical encrypted information; Generate a hash path according to the medical encrypted information, and calculate the root hash of the medical encrypted information according to the hash path; Integrate the hash path and the root hash to obtain the verification hash information; Verify whether the reference hash information is consistent with the verification hash information. If the verification hash information is consistent with the verification hash information, it is determined that the medical encrypted information passes the verification; If the verification hash information is inconsistent with the verification hash information, it is determined that the medical encrypted information fails the verification.

10. A regional medical information sharing system based on artificial intelligence, characterized in that, Including: A memory configured to store instructions; And A processor configured to call instructions from the memory and capable of implementing a method for sharing regional medical information based on artificial intelligence according to any one of claims 1 to 9 when executing the instructions.

Citation Information

Patent Citations

  • Text feature selection method based on inter-class information entropy

    CN110580286A

  • System and method for notifying work information of medical informatization system

    CN115691771A

  • Medical data desensitization method and system

    CN115859372A

  • Digital information processing method and system based on Hash encryption algorithm

    CN116484341A

  • Medical data encryption system for multi-modal large language model

    CN117951746A

Cited By

  • Diffusion-scrambling-diffusion image encryption method and system

    CN121217875A

  • Heat stroke medical record text classification method based on reinforcement learning

    CN121278096A

  • Training method of clinical auxiliary diagnosis model

    CN121747910A

  • Supply chain information sharing method based on cloud platform

    CN122053063A