Content security management method and system in medical information system

CN122800288APending Publication Date: 2026-09-22HANGZHOU YUANGUI NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611020653.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-09
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0002]医院屏幕终端广泛用于展示健康教育视频、医院公告等医疗资讯,现有技术中,内容安全审核主要依赖服务器端完成,通常采用单一的关键词过滤或简单规则匹配进行内容检测,难以识别复杂的医学语义关系,导致医疗信息准确性和安全性无法得到有效保障

Benefits of technology

[0015]本发明通过构建双通道语义分析网络,将医学专业知识与通用语义理解能力有机结合,显著提高了医学实体识别的准确性,使系统能够精准理解复杂医疗内容语义。关键区与非关键区的智能划分方法实现了对医疗内容的差异化处理,提高系统处理效率和精准度。在屏幕终端本地实现实时内容校验,有效解决了服务器端审核延迟问题,及时拦截不合规内容。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122800288A_ABST
    Figure CN122800288A_ABST
Patent Text Reader

Abstract

The application provides a content security management method and system in a medical information system, relates to the technical field of medical information, and identifies medical entities by constructing a dual-channel semantic analysis network comprising a medical professional channel and a general semantic channel, according to which the content is divided into a key area and a non-key area and verification is respectively performed, a semantic consistency network is constructed to calculate the overall reliability, meanwhile, the content is calculated for a sensitivity score, and a compliance score is calculated by using a medical regulation knowledge graph; a security evaluation vector is constructed based on the reliability, the sensitivity and the compliance, evaluation is performed in combination with terminal scene attributes, and then the content and evaluation data are written into a display buffer; when the content is updated, the security evaluation process is re-executed, and it is ensured that the medical information is safely and reliably displayed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to medical information technology, and more particularly to a method and system for content security management in medical information systems. Background Technology

[0002] Hospital screen terminals are widely used to display health education videos, hospital announcements, and other medical information. Currently, content security verification primarily relies on the server side, typically employing single keyword filtering or simple rule matching for content detection. This fails to identify complex medical semantic relationships, resulting in inadequate assurance of the accuracy and security of medical information. With the diversification of information dissemination, content downloaded to screen terminals may face the risk of tampering or mistransmission. Existing content management methods lack a real-time local verification mechanism on the terminal, making it difficult to quickly locate problematic content without increasing storage burden.

[0003] Existing medical content security assessments mostly rely on static rule bases, lacking dynamic adaptability and struggling to cope with the rapid updates to medical knowledge and the display needs of different scenarios. Furthermore, traditional systems often conduct sensitivity and compliance assessments in a fragmented manner, lacking systematicity and coherence, leading to biased security assessment results. Moreover, the display needs of medical information vary significantly across different terminal devices, scenarios, and user permissions, lacking targeted adjustment mechanisms. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a content security management method and system for medical information systems, which can solve the problems in existing technologies.

[0005] A first aspect of the present invention provides a content security management method for a medical information system, comprising: A dual-channel semantic analysis network is constructed, comprising a medical professional channel based on a medical ontology database and a general semantic channel based on a general language model; the content to be processed is input into the dual-channel semantic analysis network, and medical entities are identified through fusion via a convolutional attention mechanism; Based on the identification results of the medical entities, the content to be processed is divided into key areas and non-key areas; professional verification and integrity verification are performed on the key areas, and standardization checks are performed on the non-key areas; a medical content semantic consistency network is constructed based on the processing results of different areas, the coherence of information between areas is analyzed, and the overall credibility of the content is calculated. Calculate a sensitivity score for the content to be processed; Construct a medical regulations knowledge graph, semantically match the medical entities with the medical regulations knowledge graph, and calculate a compliance score; A security assessment vector is constructed based on the overall credibility, the sensitivity score, and the compliance score. The security assessment vector is evaluated according to the security threshold and the scenario attributes of the display terminal. When the security assessment is passed, the content to be processed and its security assessment data are written into the terminal display buffer. When the content of the terminal display buffer needs to be updated, the security assessment process is re-executed on the updated content.

[0006] Optionally, The steps for constructing a dual-channel semantic analysis network, comprising a medical specialty channel based on a medical ontology database and a general semantic channel based on a general language model, include: A hierarchical relationship tree is established based on medical concepts in the Unified Medical Language System. The hierarchical relationship tree includes hierarchical relationships and partial whole relationships between medical concepts. The semantic types of the medical concepts are extracted based on the hierarchical relationship tree, and a set of semantic types of the medical concepts is constructed through a semantic type mapping function. Based on the set of semantic types, the medical concepts are converted into medical feature vectors through transpose operation, and a medical entity embedding matrix is ​​constructed. The distance matrix between medical concepts is calculated based on the hierarchical relationship tree. A hierarchical reachability mask is generated based on the distance matrix. The distance matrix and the hierarchical reachability mask are combined to obtain a concept-level attention matrix. A multi-layer stacked attention encoder is constructed, each layer of which includes a multi-head self-attention module and a feedforward neural network. In the multi-head self-attention module, the input features are divided into multiple attention heads, and each attention head calculates an attention score through a query matrix, a key matrix, and a value matrix. A position encoding matrix is ​​generated, and the position encoding matrix calculates text position information through a sinusoidal position encoding function.

[0007] Optionally, The steps of inputting the content to be processed into the dual-channel semantic analysis network and identifying medical entities in the content through convolutional attention mechanism include: The content to be processed is segmented and standardized to obtain a standardized text sequence; The standardized text sequence and the medical entity embedding matrix are used to perform feature calculations to obtain medical channel features. The medical channel features are then interacted with the concept-level attention matrix to obtain the medical channel hidden layer representation. The standardized text sequence is input into the attention encoder, and the output features are superimposed with the position encoding matrix to obtain general semantic channel features; Feature extraction with different convolution kernel sizes is performed on the medical channel hidden layer representation and the general semantic channel features to obtain multi-scale convolution features; inter-channel attention scores are calculated based on the multi-scale convolution features; and the medical channel hidden layer representation and the general semantic channel features are weighted and fused according to the inter-channel attention scores to obtain fused features. The number of medical concepts and the number of medical concept relationships are identified from the content to be processed, and the text length of the content to be processed is calculated; a professionalism score is calculated based on the number of medical concepts, the number of medical concept relationships, and the text length; the professionalism score is combined with the feature entropy of the fused features to generate boundary optimization weights. The desired boundary is calculated based on the preset internal structure rules of the medical entity; the boundary sensitivity of the initially identified medical entity boundary and the desired boundary is calculated to obtain the structural loss; the medical entity boundary is iteratively optimized based on the structural loss and the boundary optimization weight to obtain the medical entity recognition result.

[0008] Optionally, The steps include: dividing the content to be processed into critical and non-critical areas based on the identification results of the medical entities; performing professional and completeness verification on the critical areas and standardization checks on the non-critical areas; constructing a semantic consistency network for medical content based on the processing results of different areas; analyzing the coherence of information between areas; and calculating the overall credibility of the content. Importance information of the medical entities is obtained from the medical knowledge base, and regional importance scores are calculated based on the distribution of the medical entities. Based on the regional importance scores, regional boundaries are determined by minimizing the variance between regions, and the content to be processed is divided into key regions and non-key regions. For the critical area, a set of medical feature points is extracted, and the semantic similarity and contextual consistency of each feature point in the set with the standard concept set are calculated to obtain a professionalism verification matrix. Based on the professionalism verification matrix, medical statements are extracted, and the medical statements are matched with clinical guidelines to calculate the guideline compliance. The coverage of medical elements in the medical statements is also calculated to obtain a completeness score. The guideline compliance and the completeness score are combined to obtain a safety score for the critical area. For the non-critical areas, the terminology standardization score and contextual semantic similarity are calculated to generate a normative score for the non-critical areas; Feature vector transformation is performed on the key regions and the non-key regions to calculate the semantic similarity, entity overlap, and information flow intensity between regions, thus constructing a semantic consistency network. In the semantic consistency network, an initial state matrix is ​​constructed based on the statement credibility of the region nodes, and the initial state matrix is ​​updated through iterative propagation to obtain a consistency propagation matrix. The degree of conflict between regions is calculated based on the consistency propagation matrix. The security score, standardization score, and conflict level are fused together, and the overall credibility of medical content is determined by an adaptive threshold.

[0009] Optionally, The steps for performing professional verification and integrity verification on the critical areas include: The medical feature point set is grouped according to feature type to construct a feature group index table; a compressed standard concept set is created based on the feature group index table, and a concept fast matching index is established using the locality-sensitive hashing method; The key area is scanned using a sliding window method, and candidate medical statements are extracted based on a preset semantic template. The candidate medical statements are matched with the standard concept set to obtain an initial matching result. The semantic similarity and contextual consistency of the initial matching result are combined incrementally to generate a professional verification matrix. Based on the verification results of the professional verification matrix, valid medical statements are selected from the candidate medical statements; the valid medical statements are matched with clinical guideline rule templates, and the rule coverage rate is calculated to obtain the guideline compliance. A hierarchical medical element structure tree is constructed, and the effective medical statements are mapped to the medical element structure tree. The coverage ratio of the target nodes is calculated to obtain an integrity score.

[0010] Optionally, The steps of constructing a medical regulatory knowledge graph, semantically matching the medical entities with the medical regulatory knowledge graph, and calculating compliance scores include: A multi-level concept tree is constructed based on medical regulatory texts to extract regulatory entities and relationships. The regulatory entities are used as nodes and the regulatory relationships are used as edges to establish the basic structure of a medical regulatory knowledge graph. The attribute information of the regulatory entities is expanded in a recursive manner to establish reasoning rules between entities. Based on the reasoning rules, a constraint propagation network is constructed to realize dynamic reasoning of regulatory knowledge. The medical entities are mapped to the medical regulatory knowledge graph, and the path similarity and semantic similarity between entities are calculated. An entity alignment matrix is ​​constructed based on the path similarity and semantic similarity. The entity alignment matrix is ​​used to perform reasoning in the constraint propagation network to identify potential violating entities. The aforementioned entities that violated regulations were statistically weighted, and a compliance score was calculated based on the severity of the violation.

[0011] Optionally, A security assessment vector is constructed based on the overall credibility, the sensitivity score, and the compliance score; the security assessment vector is evaluated according to the security threshold and the scenario attributes of the display terminal; when the security assessment is passed, the step of writing the content to be processed and its security assessment data into the terminal display buffer includes: The overall credibility, sensitivity score, and compliance score are normalized and weighted to construct a security assessment vector. Based on the device type, usage scenario, and user permission information of the display terminal, terminal relevance, scenario relevance, and permission relevance are calculated respectively; the terminal relevance, scenario relevance, and permission relevance are weighted and combined to obtain the scenario adaptation coefficient; the preset security benchmark threshold is multiplied by the scenario adaptation coefficient to obtain the dynamic security threshold; The security assessment vector is compared with the dynamic security threshold to generate a security level identifier; a content tracing index is established based on the security level identifier; the content tracing index is bound to the operator's identity information and operation time information; and a digital watermark containing the content tracing index is embedded in the content to be processed. Based on the security level identifier, the content to be processed is divided into general level, guidance level and restricted level; the location attribute information and target audience attribute information of the display terminal are obtained, the content display level is determined, and the content to be processed is filtered or replaced according to the content display level to generate the content to be displayed; The content to be displayed is subjected to network security testing. After the test is passed, the content to be displayed is combined with scene information and time information to generate an anti-tampering verification value and written into the terminal display buffer.

[0012] Optionally, The steps of obtaining location attribute information and target audience attribute information of the display terminal, determining the content display level, filtering or replacing the content to be processed according to the content display level, and generating the content to be displayed include: The spatial location information and time information of the display terminal are obtained, and the location sensitivity is calculated; the time change rate and spatial second derivative of the location sensitivity are calculated, and the spatial dynamic characteristics are obtained by combining the time change rate and the spatial second derivative. Obtain the proportion of professionals and non-professionals in the area where the display terminal is located, and construct a crowd proportion vector; establish a group interaction matrix, and multiply the crowd proportion vector with the group interaction matrix to obtain an audience feature vector; combine the audience feature vector with the spatial dynamic features to generate scene features; The content display level of the display terminal is determined based on the scene characteristics; when the content display level is lower than the content classification level, the level difference is calculated; when the level difference exceeds a preset filtering threshold, the corresponding content is filtered; when the level difference does not exceed the filtering threshold, the content to be processed is segmented, and the local sensitivity of each segment and the correlation between content segments are calculated; the local sensitivity and the correlation are combined to obtain the global sensitivity; the content replacement intensity is determined based on the scene characteristics and the global sensitivity, and the content to be processed is processed according to the content replacement intensity to obtain the replaced content; The replaced content is subjected to multi-objective optimization to obtain the optimal replacement scheme. The multi-objective optimization includes a comprehensive evaluation of information entropy, sensitive information loss degree and context bias. The content to be displayed is generated according to the optimal replacement scheme.

[0013] Regarding content security management in the second medical information system, a system is provided, including: The first unit is used to construct a dual-channel semantic analysis network, which includes a medical professional channel based on a medical ontology library and a general semantic channel based on a general language model. The content to be processed is input into the dual-channel semantic analysis network and fused through a convolutional attention mechanism to identify medical entities. The second unit is used to divide the content to be processed into key areas and non-key areas based on the identification results of the medical entities; to perform professional verification and integrity verification on the key areas, and to perform standardization checks on the non-key areas; to construct a medical content semantic consistency network based on the processing results of different areas, to analyze the coherence of information between areas, and to calculate the overall credibility of the content. The third unit is used to calculate a sensitivity score for the content to be processed. The fourth unit is used to construct a medical regulatory knowledge graph, semantically match the medical entities with the medical regulatory knowledge graph, and calculate a compliance score. The fifth unit is used to construct a security assessment vector based on the overall credibility, the sensitivity score, and the compliance score; to evaluate the security assessment vector according to the security threshold and the scene attributes of the display terminal; when the security assessment is passed, the content to be processed and its security assessment data are written into the terminal display buffer; when the content of the terminal display buffer needs to be updated, the security assessment process is re-executed on the updated content.

[0014] Thirdly, a computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0015] This invention constructs a dual-channel semantic analysis network, organically combining medical expertise with general semantic understanding capabilities, significantly improving the accuracy of medical entity recognition and enabling the system to accurately understand the semantics of complex medical content. The intelligent segmentation method of key and non-key areas allows for differentiated processing of medical content, improving system processing efficiency and accuracy. Real-time content verification is implemented locally on the screen terminal, effectively solving the server-side review delay problem and promptly intercepting non-compliant content.

[0016] This invention innovatively introduces a medical regulatory knowledge graph and constraint propagation network to achieve dynamic compliance assessment of medical content, accurately identify potential violators, and provide a tracing path. Through the construction of security assessment vectors and the application of dynamic security thresholds, the system can adaptively adjust security standards according to different scenarios, balancing information security and usability. The embedded digital watermarking method, combined with content tracing indexes, enables problem content to be quickly traced back to the operator, significantly improving problem-solving efficiency.

[0017] The scene-aware technology of this invention can dynamically adjust the content display level according to the terminal location, time, and audience characteristics, thereby achieving precise delivery of medical information. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating the content security management method in a medical information system according to an embodiment of the present invention. Figure 2 A chart comparing the accuracy of compliance scoring using different methods. Detailed Implementation

[0019] The technical solutions of the present invention will be described below with reference to the accompanying drawings. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0020] Figure 1 This is a flowchart illustrating the content security management method in the medical information system of the present invention, such as... Figure 1 As shown, the method includes: A dual-channel semantic analysis network is constructed, comprising a medical professional channel based on a medical ontology database and a general semantic channel based on a general language model; the content to be processed is input into the dual-channel semantic analysis network, and medical entities are identified through fusion via a convolutional attention mechanism; Based on the identification results of the medical entities, the content to be processed is divided into key areas and non-key areas; professional verification and integrity verification are performed on the key areas, and standardization checks are performed on the non-key areas; a medical content semantic consistency network is constructed based on the processing results of different areas, the coherence of information between areas is analyzed, and the overall credibility of the content is calculated. Calculate a sensitivity score for the content to be processed; Construct a medical regulations knowledge graph, semantically match the medical entities with the medical regulations knowledge graph, and calculate a compliance score; A security assessment vector is constructed based on the overall credibility, the sensitivity score, and the compliance score. The security assessment vector is evaluated according to the security threshold and the scenario attributes of the display terminal. When the security assessment is passed, the content to be processed and its security assessment data are written into the terminal display buffer. When the content of the terminal display buffer needs to be updated, the security assessment process is re-executed on the updated content.

[0021] Optionally, The steps to construct a dual-channel semantic analysis network include: A hierarchical relationship tree is established based on medical concepts in the Unified Medical Language System. The hierarchical relationship tree includes hierarchical relationships and partial whole relationships between medical concepts. The semantic types of the medical concepts are extracted based on the hierarchical relationship tree, and a set of semantic types of the medical concepts is constructed through a semantic type mapping function. Based on the set of semantic types, the medical concepts are converted into medical feature vectors through transpose operation, and a medical entity embedding matrix is ​​constructed. The distance matrix between medical concepts is calculated based on the hierarchical relationship tree. A hierarchical reachability mask is generated based on the distance matrix. The distance matrix and the hierarchical reachability mask are combined to obtain a concept-level attention matrix. A multi-layer stacked attention encoder is constructed, each layer of which includes a multi-head self-attention module and a feedforward neural network. In the multi-head self-attention module, the input features are divided into multiple attention heads, and each attention head calculates an attention score through a query matrix, a key matrix, and a value matrix. A position encoding matrix is ​​generated, and the position encoding matrix calculates text position information through a sinusoidal position encoding function.

[0022] In this embodiment, over 2 million medical concepts are extracted from the Unified Medical Language System (UMLS) database, covering multiple medical fields such as diseases, symptoms, drugs, surgeries, and anatomical structures. The hierarchical relationship tree contains the hierarchical relationships and partial-whole relationships between medical concepts. For example, "heart disease" is a subordinate concept of "cardiovascular disease," and "ventricle" is a partial concept of "heart." An adjacency list storage structure is used to represent the hierarchical relationship tree, with each node storing a unique identifier for the concept, the concept name, and pointers to its parent and child nodes.

[0023] UMLS contains approximately 135 semantic types, such as "disease or syndrome," "anatomical structure," and "drug." A semantic type mapping function is used to construct a set of semantic types for medical concepts. This function is implemented using a hash table, where the key is the concept identifier and the value is a set of semantic type identifiers. For example, the concept "aspirin" maps to the semantic types "drug" and "organic chemical substance." In practical applications, a medical concept may correspond to multiple semantic types, with an average of 2.3 semantic types associated with each concept.

[0024] A semantic type-concept matrix is ​​constructed, where rows represent semantic types and columns represent medical concepts. Matrix elements have values ​​of 0 or 1, indicating whether a concept belongs to that semantic type. This matrix is ​​then transposed to obtain a concept-semantic type matrix, where each row represents a feature vector of a medical concept. Finally, the feature vectors are normalized so that the magnitude of each vector is 1. Thus, medical concepts are represented as feature vectors of dimension 135. The feature vectors of all medical concepts are combined to form a medical entity embedding matrix, with a dimension of the number of concepts × 135.

[0025] In a hierarchical relationship tree, the distance between two concepts is defined as the length of the shortest path between them. A distance matrix is ​​formed by calculating the distance between any two concepts using a breadth-first search algorithm. To improve computational efficiency, a block-based computation strategy is adopted, dividing the concept set into multiple subsets and calculating the distance between concepts within each subset. For concept pairs whose distance exceeds a preset threshold (set to 8 in this example), the distance value is set to infinity. The hierarchical reachability mask is a binary matrix; when the distance between concept i and concept j is less than the preset threshold, the matrix element value is 1; otherwise, it is 0. The preset threshold is set to 4, indicating that concepts with a distance less than or equal to 4 in the hierarchical tree are considered semantically related.

[0026] The distance matrix and the hierarchical reachability mask are combined as follows: for elements with a mask value of 1, the attention value is equal to 1 minus the normalized distance value; for elements with a mask value of 0, the attention value is set to 0. Normalization is achieved by dividing the distance value by the maximum distance value (4 in this example). Thus, concepts that are closer in distance correspond to larger attention values, with a maximum of 1 (the concept itself) and a minimum of 0 (unrelated concepts).

[0027] The attention encoder consists of 6 layers, each composed of a multi-head self-attention module and a feedforward neural network. The multi-head self-attention module has an input dimension of 768 and 12 attention heads, each with a dimension of 64. The feedforward neural network contains two fully connected layers, with the intermediate layer having a dimension of 3072. Normalization and residual connections are added after each attention encoder layer to improve the model's training stability and expressive power.

[0028] In the multi-head self-attention module, the input features are divided into multiple attention heads. For each attention head, the input features are transformed into query vectors, key vectors, and value vectors using query matrices, key matrices, and value matrices, respectively. These matrices are trainable parameters, initially randomly generated, with a dimension of 768×64. The dot product of the query vector and the key vector is calculated, divided by a scaling factor (8 in this example, the square root of the head dimension), and then the softmax function is applied to obtain the attention weights. The attention weights are then weighted and summed with the value vectors to obtain the output of that attention head. Finally, the outputs of all attention heads are concatenated, and the final output of the multi-head attention is obtained through the output projection matrix (768×768 dimension).

[0029] A positional encoding matrix is ​​generated. Positional codes provide location information for the input sequence and are calculated using a sinusoidal positional encoding function. For the pos-th position in the sequence, the positional code value of dimension i is calculated as follows: when i is even, a sine function is used; when i is odd, a cosine function is used. The function period increases exponentially with the dimension to ensure that different positions have unique codes. The dimension of the positional encoding matrix is ​​the maximum sequence length × 768, and the maximum sequence length is set to 512.

[0030] For the model training, a standard pre-training-fine-tuning paradigm was adopted. In the pre-training phase, a masked language model task was used, randomly masking 15% of the words in a medical literature corpus (including PubMed articles, clinical guidelines, etc.), and training the model to predict these masked words. The Adam optimizer was used with a learning rate of 0.00002, a batch size of 32, and 8 training epochs. In the fine-tuning phase, for the medical entity recognition task, a manually annotated medical text dataset (containing 50,000 annotated medical records) was used, employing the cross-entropy loss function with a learning rate of 0.00005, and 4 training epochs.

[0031] The dual-channel semantic analysis network constructed in this invention can simultaneously process medical expertise and general semantic information, significantly improving the accuracy and comprehensiveness of medical content analysis, ensuring that the system can accurately identify medical entities, and laying the foundation for subsequent security assessment and hierarchical display.

[0032] Optionally, The steps of inputting the content to be processed into the dual-channel semantic analysis network and identifying medical entities in the content through convolutional attention mechanism include: The content to be processed is segmented and standardized to obtain a standardized text sequence; The standardized text sequence and the medical entity embedding matrix are used to perform feature calculations to obtain medical channel features. The medical channel features are then interacted with the concept-level attention matrix to obtain the medical channel hidden layer representation. The standardized text sequence is input into the attention encoder, and the output features are superimposed with the position encoding matrix to obtain general semantic channel features; Feature extraction with different convolution kernel sizes is performed on the medical channel hidden layer representation and the general semantic channel features to obtain multi-scale convolution features; inter-channel attention scores are calculated based on the multi-scale convolution features; and the medical channel hidden layer representation and the general semantic channel features are weighted and fused according to the inter-channel attention scores to obtain fused features. The number of medical concepts and the number of medical concept relationships are identified from the content to be processed, and the text length of the content to be processed is calculated; a professionalism score is calculated based on the number of medical concepts, the number of medical concept relationships, and the text length; the professionalism score is combined with the feature entropy of the fused features to generate boundary optimization weights. The desired boundary is calculated based on the preset internal structure rules of the medical entity; the boundary sensitivity of the initially identified medical entity boundary and the desired boundary is calculated to obtain the structural loss; the medical entity boundary is iteratively optimized based on the structural loss and the boundary optimization weight to obtain the medical entity recognition result.

[0033] In this embodiment, the content to be processed undergoes word segmentation and standardization. The word segmentation process employs a medical-specific word segmenter that integrates general word segmentation rules and a medical terminology dictionary. For Chinese medical text, a conditional random field-based word segmentation algorithm is used, combined with a professional dictionary of over 500,000 medical terms to accurately identify the boundaries of medical terms. For example, for the input text "The patient's blood pressure is high, and the doctor recommends taking nifedipine sustained-release tablets," the word segmentation result is "patient / blood pressure / high / , / doctor / recommend / take / nifedipine sustained-release tablets". Standardization processing includes removing punctuation marks, converting to lowercase, and replacing synonyms. Medical terminology standardization uses UMLS's preferred terminology replacement mechanism to map variant terms to standard terms. For example, "hypertension" and "elevated blood pressure" are both standardized to "hypertension". The processed standardized text sequence is stored in the form of a word list, with each word associated with its original position information.

[0034] Feature computation is performed on the standardized text sequence and the medical entity embedding matrix. The feature computation uses a search-and-match approach: for each word in the standardized text sequence, the corresponding feature vector is searched in the medical entity embedding matrix. If a direct match fails, a fuzzy matching algorithm is used to calculate the similarity between the word and the medical concept, selecting the feature vector of the concept with the highest similarity exceeding a threshold (set to 0.85). If no match is still found, zero vectors are used for padding. After matching, a medical channel feature matrix with dimensions of sequence length × 135 is obtained.

[0035] The medical channel features are interacted with a concept-level attention matrix. This interaction process is essentially feature enhancement, adjusting feature weights based on the hierarchical relationships between concepts. Specifically, for each feature vector in the medical channel feature matrix, the concept-level attention matrix is ​​used to calculate its attention score with all concepts. These scores are then used as weights to perform a weighted summation on the medical entity embedding matrix, resulting in the enhanced feature vector. This step integrates the hierarchical relationship information of concepts into the feature representation, allowing semantically similar concept features to mutually enhance each other. The resulting hidden layer representation of the medical channel has a dimension of sequence length × 256.

[0036] Simultaneously, the standardized text sequence is input into an attention encoder, and the output features are superimposed with the positional encoding matrix to obtain general semantic channel features. The standardized text sequence is first converted into an initial feature vector through a word embedding layer with a dimension of 768. Then, the positional encoding and word embedding vectors are superimposed and input into a 6-layer stacked attention encoder. The attention encoder performs self-attention computation and feature transformation on the input sequence, capturing the semantic information and contextual relationships of the text. The final output is a feature matrix with a dimension of sequence length × 768, i.e., general semantic channel features.

[0037] To extract text features at different granularities, four types of convolutional kernels—1×1, 3×3, 5×5, and 7×7—were used to perform convolution operations on the features of two channels. The output channels of each convolutional layer were set to 128. For the hidden layer representation of the medical channel, four feature matrices of size 128 (sequence length × 128) were obtained after the convolution operation; similarly, four feature matrices of size 128 (sequence length × 128) were obtained for the general semantic channel features. These eight feature matrices constitute the multi-scale convolutional features.

[0038] Inter-channel attention scores are calculated based on multi-scale convolutional features. The convolutional features of the same size for two channels are multiplied by a dot product, and then normalized using the softmax function to obtain the inter-channel attention matrix. This matrix represents the correlation between features of two channels, with dimensions of sequence length × sequence length. By summing and normalizing the row directions of the attention matrix, the channel attention score at each position is obtained, representing the relative importance of the medical channel and the general semantic channel at that position.

[0039] The hidden representation of the medical channel and the features of the general semantic channel are weighted and fused based on the inter-channel attention scores. The weighted fusion adopts a gating mechanism, with the channel attention score as the weight of the medical channel and 1 minus the attention score as the weight of the general semantic channel. The two weighted channel features are concatenated and then reduced in dimensionality by a fully connected layer to obtain a fused feature with a dimension of sequence length × 512.

[0040] Preliminary medical entity boundary identification is performed based on fusion features. A sequence labeling method is employed, using a Bidirectional Long Short-Term Memory (BiLSTM) network combined with a Conditional Random Field (CRF) to process the fusion features, assigning a label to each term in the sequence. The labeling scheme adopts the BIO tagging scheme: B represents the start of an entity, I represents the interior of an entity, and O represents a non-entity part. The BiLSTM layer contains 256 hidden units to capture the contextual dependencies of the sequence; the CRF layer considers the transition probability between labels to ensure the validity of the label sequence. For example, for the text "Patient's blood pressure is high, doctor recommends taking nifedipine sustained-release tablets", the preliminary identification result is "Patient (O) blood pressure (B) high (I), doctor (O) recommends (O) taking (O) nifedipine sustained-release tablets (B)", identifying two medical entities, "high blood pressure" and "nifedipine sustained-release tablets," and their boundaries.

[0041] The number of medical concepts is obtained by matching against a medical terminology dictionary, and the number of medical concept relationships is identified using a preset relationship template. For example, for the input text "Long-term use of metformin by diabetic patients may lead to vitamin B12 deficiency," three medical concepts—"diabetes," "metformin," and "vitamin B12 deficiency"—are identified, along with two medical concept relationships: "use" and "lead to." The text length is 20 characters. The professionalism score is calculated using a weighted method: the number of medical concepts has a weight of 0.5, the number of medical concept relationships has a weight of 0.3, and the inverse of the text length has a weight of 0.2. These three factors are normalized and then weighted and summed to obtain a professionalism score between 0 and 1. For example, the professionalism score of the above example text is 0.78.

[0042] Feature entropy reflects the uncertainty of features and is obtained by calculating the entropy value of the fused features. Boundary optimization weights are calculated by weighting the expertise score and feature entropy, with weights of 0.7 and 0.3 respectively. These weights are used in the subsequent boundary optimization process; the higher the expertise score and the lower the feature entropy, the larger the boundary optimization weight.

[0043] The pre-defined internal structural rules for medical entities include prefix rules, suffix rules, and combination rules. For example, drug names typically consist of the generic name, dosage form, and strength, such as "nifedipine sustained-release tablets 10mg". Based on these rules, the expected boundary positions of the entities are calculated. Boundary sensitivity is calculated by comparing the initially identified medical entity boundaries with the expected boundaries, yielding the structural loss. Boundary sensitivity is defined as the degree of deviation between the actual boundary and the expected boundary, calculated through a weighted sum of positional differences. The structural loss reflects the accuracy of boundary identification; a smaller value indicates a more accurate boundary.

[0044] The boundary of the medical entity is iteratively optimized based on structural loss and boundary optimization weights. The iterative optimization process uses a boundary adjustment algorithm, which calculates and adjusts the step size according to the structural loss and boundary optimization weights, gradually moving the entity boundary position until the structural loss is less than a threshold or the maximum number of iterations is reached. Finally, the optimized boundary of the medical entity is obtained, completing the medical entity recognition.

[0045] The medical entity recognition method implemented in this invention exhibits high accuracy and robustness, accurately identifying medical terminology and entity boundaries in medical texts. The dual-channel fusion mechanism fully leverages medical expertise and general semantic understanding capabilities, while the convolutional attention mechanism effectively captures text features at different granularities. The boundary optimization algorithm significantly improves the accuracy of entity boundary recognition. This method provides a solid technical foundation for content security management in medical information systems, ensuring that the system can accurately understand the semantics of medical content.

[0046] Optionally, The steps include: dividing the content to be processed into critical and non-critical areas based on the identification results of the medical entities; performing professional and completeness verification on the critical areas and standardization checks on the non-critical areas; constructing a semantic consistency network for medical content based on the processing results of different areas; analyzing the coherence of information between areas; and calculating the overall credibility of the content. Importance information of the medical entities is obtained from the medical knowledge base, and regional importance scores are calculated based on the distribution of the medical entities. Based on the regional importance scores, regional boundaries are determined by minimizing the variance between regions, and the content to be processed is divided into key regions and non-key regions. For the critical area, a set of medical feature points is extracted, and the semantic similarity and contextual consistency of each feature point in the set with the standard concept set are calculated to obtain a professionalism verification matrix. Based on the professionalism verification matrix, medical statements are extracted, and the medical statements are matched with clinical guidelines to calculate the guideline compliance. The coverage of medical elements in the medical statements is also calculated to obtain a completeness score. The guideline compliance and the completeness score are combined to obtain a safety score for the critical area. For the non-critical areas, the terminology standardization score and contextual semantic similarity are calculated to generate a normative score for the non-critical areas; Feature vector transformation is performed on the key regions and the non-key regions to calculate the semantic similarity, entity overlap, and information flow intensity between regions, thus constructing a semantic consistency network. In the semantic consistency network, an initial state matrix is ​​constructed based on the statement credibility of the region nodes, and the initial state matrix is ​​updated through iterative propagation to obtain a consistency propagation matrix. The degree of conflict between regions is calculated based on the consistency propagation matrix. The security score, standardization score, and conflict level are fused together, and the overall credibility of medical content is determined by an adaptive threshold.

[0047] In this embodiment, the importance information of medical entities comes from professional medical knowledge bases, such as UMLS and MeSH. The importance information includes two parts: entity type weight and entity frequency weight. The entity type weight is predefined based on the semantic type of the medical entity; for example, the weight for disease type is 0.9, for symptom type is 0.8, for drug type is 0.85, and for examination type is 0.75. The entity frequency weight is calculated based on the frequency of the entity's occurrence in medical literature, using logarithmic normalization. For the content to be processed, it is first initially segmented according to natural paragraph or sentence boundaries; each segment is called a candidate region. For each candidate region, the number and type of medical entities it contains are counted, and the region importance score is calculated. The calculation method is: the importance values ​​of all medical entities in the region are weighted and summed, then divided by the region text length for normalization. For example, for a region containing "type 2 diabetes" (disease type, importance 0.92) and "insulin resistance" (symptom type, importance 0.83), the region importance score is (0.92 + 0.83) / number of characters in the region.

[0048] An adaptive thresholding method is used to rank the importance scores of regions, calculate the difference between adjacent scores, and select the position with the largest difference as the dividing point. Alternatively, a clustering method, such as K-means clustering (K=2), is used to divide the candidate regions into two categories. Regions with high importance scores are marked as key regions, and regions with low scores are marked as non-key regions. To avoid excessive fragmentation, adjacent and similar small regions are merged. The merging criteria are: the difference in importance scores between adjacent regions is less than a preset threshold (e.g., 0.15), and both belong to either key or non-key regions. For example, in the medical article "Symptoms and Treatment of Diabetes," after analysis, paragraphs with dense medical entities, such as "Diabetes can lead to various complications, including cardiovascular disease, kidney disease, and retinopathy," are classified as key regions, while paragraphs with sparse medical entities, such as "This article will introduce the basic knowledge of diabetes," are classified as non-key regions.

[0049] Medical feature points are combinations of medical terms, relational terms, and descriptive terms within the key area. Extraction employs a hybrid rule-based and statistical approach, including term matching, syntactic analysis, and keyword extraction. The standard concept set is derived from authoritative medical knowledge bases, such as the International Classification of Diseases (ICD) and medical practice guidelines. For each feature point, its semantic similarity to concepts in the standard concept set is calculated. Semantic similarity is calculated using cosine similarity of word vectors, generated using a pre-trained model in the medical field. Contextual consistency is calculated by matching the context of the feature point with the context of the standard concept. Semantic similarity and contextual consistency are weighted and combined (with weights of 0.6 and 0.4, respectively) to form a professional verification matrix. Rows in the matrix represent feature points, columns represent standard concepts, and element values ​​are combined similarities. For example, the feature point "oral metformin controls blood sugar" has a semantic similarity of 0.87, a contextual consistency of 0.92, and a combined similarity of 0.89 with the standard concept "metformin is used to treat type 2 diabetes." Medical statements are sentences or phrases within the key area that express a complete medical viewpoint. Medical statements are constructed by selecting feature points with similarity scores higher than a threshold (e.g., 0.75) using a professional verification matrix. Guideline compliance is calculated by matching the medical statement with clinical guideline rules. Clinical guideline rules are stored in "if-then" format, such as "If a patient has type 2 diabetes, then metformin is recommended as a first-line treatment." The matching process combines template matching and semantic matching to calculate the degree of conformity between the statement and the rule. Medical element coverage is calculated by checking whether the statement contains necessary medical information elements (such as disease, treatment, dosage, indications, etc.). Different sets of necessary elements are predefined according to the statement type. A completeness score is obtained by weighted combination of guideline compliance (weight 0.7) and medical element coverage (weight 0.3). The guideline compliance and completeness scores are further weighted combination (weights of 0.65 and 0.35, respectively) to obtain the safety score for the critical area.

[0050] The terminology standardization score reflects the degree of standardization in the use of medical terminology in non-critical areas, calculated by checking whether the terminology uses standard names. For example, "high blood sugar" should be standardized as "hyperglycemia". The degree of standardization is calculated by the matching rate between the terminology and the standard dictionary. Contextual semantic similarity is calculated by the semantic relevance between the non-critical area and adjacent critical areas, using the cosine similarity of the text embedding vectors. The terminology standardization score (weight 0.4) and contextual semantic similarity (weight 0.6) are weighted and combined to obtain the standardization score of the non-critical area.

[0051] Feature vector transformation is performed on both critical and non-critical regions using text embedding technology, converting each region into a 512-dimensional feature vector. Semantic similarity between regions is calculated using the cosine similarity of the feature vectors. Entity overlap is calculated by dividing the number of shared medical entities between two regions by the total number of medical entities in those regions. Information flow intensity reflects the degree to which information is transmitted from one region to another, calculated through inter-regional citation relationships and topic coherence. A semantic consistency network is constructed based on these three indicators, where network nodes represent regions, edges represent inter-regional relationships, and edge weights are the weighted sum of the three indicators (weights are 0.4, 0.3, and 0.3, respectively).

[0052] In a semantic consistency network, statement credibility is derived from the security score of critical regions and the prescriptive score of non-critical regions. The initial state matrix is ​​a diagonal matrix, with diagonal elements representing the statement credibility of each region. Iterative propagation uses a random walk algorithm to calculate how information from one region is transmitted to other regions through the network. The iterative formula is: the product of the current state and the transition probability matrix plus the product of the initial state and the damping factor. The transition probability matrix is ​​obtained by normalizing the network edge weights, and the damping factor is set to 0.15. Iteration continues until convergence or the maximum number of iterations (e.g., 20) is reached, yielding the consistency propagation matrix.

[0053] The degree of conflict between regions is calculated based on the consistency propagation matrix. The degree of conflict reflects the inconsistency of information between regions and is calculated using the variance of the element values ​​in the consistency propagation matrix. A larger variance indicates a higher degree of conflict between regions. For example, if regions A and B have different descriptions of the same medical concept, their consistency propagation values ​​will differ significantly, leading to a high degree of conflict.

[0054] The security score, standardization score, and conflict level are fused using a weighted average method. The security score for critical areas has a weight of 0.5, the standardization score for non-critical areas has a weight of 0.3, and the inverse value of the conflict level has a weight of 0.2. The resulting overall credibility score ranges from 0 to 1. An adaptive threshold is dynamically adjusted based on content type and application scenario; for example, the threshold is set to 0.8 for professional medical papers and 0.7 for health science articles. Content is considered credible when the overall credibility exceeds the threshold; otherwise, it is marked as requiring further review.

[0055] This method precisely divides medical content into critical and non-critical areas, performs targeted verification on different regions, and constructs a semantic consistency network to analyze the information coherence between regions, thus achieving refined security management of medical content. Compared with traditional methods, this solution can effectively identify professional defects, integrity issues, and inter-regional conflicts in medical content, significantly improving the accuracy and reliability of content security assessment. By distinguishing between critical and non-critical areas, it avoids the inefficiency of uniform processing of the entire text, allowing the system to concentrate resources on important content and improving processing efficiency. This semantic consistency network-based approach provides a new approach to the overall credibility assessment of medical content and offers strong technical support for content security management in medical information systems.

[0056] Optionally, The steps for performing professional verification and integrity verification on the critical areas include: The medical feature point set is grouped according to feature type to construct a feature group index table; a compressed standard concept set is created based on the feature group index table, and a concept fast matching index is established using the locality-sensitive hashing method; The key area is scanned using a sliding window method, and candidate medical statements are extracted based on a preset semantic template. The candidate medical statements are matched with the standard concept set to obtain an initial matching result. The semantic similarity and contextual consistency of the initial matching result are combined incrementally to generate a professional verification matrix. Based on the verification results of the professional verification matrix, valid medical statements are selected from the candidate medical statements; the valid medical statements are matched with clinical guideline rule templates, and the rule coverage rate is calculated to obtain the guideline compliance. A hierarchical medical element structure tree is constructed, and the effective medical statements are mapped to the medical element structure tree. The coverage ratio of the target nodes is calculated to obtain an integrity score.

[0057] In this embodiment, the medical feature point set is grouped according to feature type to construct a feature group index table. Medical feature points are medically meaningful terms, phrases, or statements within the key area. Feature types include multiple dimensions such as disease, symptom, drug, treatment, and examination. The grouping process uses a predefined feature type dictionary and classification rules. For example, "hypertension" belongs to the disease category, "headache" belongs to the symptom category, and "aspirin" belongs to the drug category. The feature group index table is implemented using a hash table, where the key is the feature type identifier and the value is the set of all feature points under that type. The index table structure is designed to support fast querying and dynamic updates.

[0058] A compressed standard concept set is created based on a feature group index table. This standard concept set is derived from authoritative medical knowledge bases, such as UMLS. The compression process, based on the feature group index table, selects a subset of concepts related to the current key region from the complete knowledge base, reducing the computational load of subsequent matching. For example, if the feature group mainly contains cardiovascular disease-related terms, standard concepts in the cardiovascular field are prioritized. The selection rules are based on the mapping relationship between feature types and concept categories, and consider the semantic association between superior and subordinate concepts. For each standard concept, its key attributes are extracted to form a feature vector, typically with dimensions between 256 and 512. Locality Sensitive Hashing (LSH) is used to construct a fast matching index. The basic principle is to divide the high-dimensional feature space into multiple buckets, mapping similar feature vectors to the same or adjacent buckets. In implementation, 20 hash functions are selected, each mapping a feature vector to a hash value, which are then combined to form a hash signature. The similarity threshold is set to 0.8, meaning concepts with a cosine similarity greater than 0.8 are considered matches. Through LSH indexing, the time complexity of concept matching is reduced from O(n) to O(1), significantly improving retrieval efficiency.

[0059] Key regions are scanned using a sliding window approach. The sliding window size is set to a variable length, with a minimum of 5 words and a maximum of 20 words, and a step size of 2 words. The window size selection is based on the average length of medical statements, covering more than 90% of common statement lengths. Semantic templates are predefined sentence structure patterns used to identify statements with medical significance. Templates include more than ten core templates such as "disease-symptom relationship template," "drug-indication template," and "treatment-effect template." Each template contains slots and constraints, such as "<disease> may cause <symptom>" and "<drug> is used to treat <disease>." Template matching uses part-of-speech tagging and dependency parsing to identify fragments in the text that conform to the template structure. For the input text "Long-term use of metformin by patients with type 2 diabetes can reduce the risk of cardiovascular complications," after sliding window scanning, "metformin can reduce the risk of cardiovascular complications" is extracted as a candidate medical statement based on the "drug-effect template."

[0060] Candidate medical statements are matched against a set of standard concepts. The matching process utilizes the previously constructed LSH index to convert candidate statements into feature vectors, and then searches the index for standard concepts with high similarity. Feature vector conversion employs a pre-trained language model in the medical field, such as BioBERT or MedBERT, to extract the semantic representation of the statements. For each candidate statement, the top 5 standard concepts with the highest similarity are returned as the initial matching results. For example, the statement "Metformin can reduce the risk of cardiovascular complications" might match the standard concepts "Metformin reduces cardiovascular events" (similarity 0.92) and "Diabetes drugs reduce the risk of heart disease" (similarity 0.85), etc.

[0061] An incremental approach is used to combine the semantic similarity and contextual consistency of the initial matching results to generate a professional validation matrix. The incremental approach means that the validation matrix is ​​calculated and updated incrementally for each candidate statement, rather than processing all statements at once. This method reduces memory usage and improves processing efficiency. Semantic similarity has already been obtained in the previous step, and contextual consistency is calculated by comparing the context of the statement with the standard context of the standard concept. The context includes the content of the three sentences before and after the statement, represented by a TF-IDF weighted bag-of-words model. The standard context comes from typical usage scenarios of the concept in medical literature. The combination method uses a weighted average, with a semantic similarity weight of 0.7 and a contextual consistency weight of 0.3. The rows of the validation matrix represent candidate statements, the columns represent standard concepts, and the element values ​​are the combined similarity scores, ranging from 0 to 1. For example, for the candidate statement "Aspirin can prevent heart disease," the semantic similarity with the standard concept "Aspirin is used to prevent myocardial infarction" is 0.88, the contextual consistency is 0.76, and the combined score is 0.844.

[0062] Based on the verification results of the professionalism verification matrix, valid medical statements are selected from candidate medical statements. The selection criteria are based on a threshold set according to the combination score, typically between 0.75 and 0.85. If the combination score of a candidate statement with any standard concept exceeds the threshold, the statement is considered valid. In addition, a multi-concept consistency check is introduced, requiring the statement to have a high similarity with multiple related standard concepts to avoid misjudgments caused by single-point matching. In specific implementation, the average score of the statement with the top 3 most similar standard concepts is taken, requiring the average score to be no less than 0.7. Through these two screenings, it is ensured that valid medical statements have sufficient professional accuracy. For the input text "Long-term smoking increases the risk of lung cancer, especially for people with a family history," after verification, it was confirmed as a valid medical statement. Its combination score with the standard concept "Smoking is a risk factor for lung cancer" is 0.91, and its average score with related concepts is 0.82.

[0063] Clinical guideline rule templates are derived from authoritative medical guidelines, such as WHO treatment guidelines, and are stored in a structured format. Each rule template includes a condition section and a recommendation section, such as "IF <patient condition> THEN <recommended action>". The matching process first decomposes valid medical statements into three parts: subject (e.g., "patients with type 2 diabetes"), predicate (e.g., "should use"), and object (e.g., "metformin"), and then performs semantic matching with the condition and recommendation sections of the rule. For example, the statement "adult patients with hypertension should take thiazide diuretics" has a match score of 0.85 with the rule "IF adult patients with essential hypertension THEN recommend thiazide diuretics as initial treatment". Rule coverage is calculated by dividing the number of rule clauses covered by the statement by the total number of applicable rule clauses. Guideline compliance is calculated using a weighted average rule coverage rate, with weights related to rule importance and update time.

[0064] A medical element structure tree is a hierarchical knowledge representation structure. The root node is the statement topic, and the child nodes are different types of medical elements, such as disease characteristics, treatment plans, and medication information. The depth of the structure tree is typically 3 to 5 levels, with the number of nodes at each level varying depending on the domain complexity. For example, the drug treatment branch might include child nodes such as drug name, dosage, administration method, administration time, and contraindications. The structure tree is built based on a medical knowledge graph, with different target node sets for different types of medical statements. The mapping process uses entity recognition and relation extraction techniques to map the information elements in the statement to the corresponding nodes in the structure tree. The coverage ratio is calculated by dividing the number of target nodes contained in the statement by the total number of expected target nodes. For example, for the drug treatment statement "Adult patients with hypertension take amlodipine 5mg orally daily, at night," the expected target nodes include disease (hypertension), population (adult), drug (amlodipine), dosage (5mg), administration method (oral), and administration time (at night), a total of 6 nodes. This statement covers all 6 nodes, with a coverage ratio of 1.0 and an integrity score of 100%. The statement "hypertensive patients take amlodipine" only covers 3 points, with a completeness score of 50%.

[0065] This invention combines medical knowledge bases, clinical guidelines, and structured analysis techniques to achieve refined verification of medical statements. The combination of sliding windows and semantic templates improves the accuracy of medical statement extraction, the professional verification matrix ensures consistency between the content and standard medical knowledge, and the medical element structure tree guarantees the integrity of medical information.

[0066] Optionally, The steps of constructing a medical regulatory knowledge graph, semantically matching the medical entities with the medical regulatory knowledge graph, and calculating compliance scores include: A multi-level concept tree is constructed based on medical regulatory texts to extract regulatory entities and relationships. The regulatory entities are used as nodes and the regulatory relationships are used as edges to establish the basic structure of a medical regulatory knowledge graph. The attribute information of the regulatory entities is expanded in a recursive manner to establish reasoning rules between entities. Based on the reasoning rules, a constraint propagation network is constructed to realize dynamic reasoning of regulatory knowledge. The medical entities are mapped to the medical regulatory knowledge graph, and the path similarity and semantic similarity between entities are calculated. An entity alignment matrix is ​​constructed based on the path similarity and semantic similarity. The entity alignment matrix is ​​used to perform reasoning in the constraint propagation network to identify potential violating entities. The aforementioned entities that violated regulations were statistically weighted, and a compliance score was calculated based on the severity of the violation.

[0067] In this embodiment, a multi-level concept tree is constructed based on medical regulatory texts. The sources of these texts include national laws and regulations such as the *Drug Administration Law*, the *Regulations on the Supervision and Administration of Medical Devices*, and the *Measures for the Administration of Medical Advertising*, as well as industry standards and guiding principles. These texts undergo preprocessing, including segmentation, removal of redundant information, and standardization of formatting. Natural language processing (NLP) technology is used to perform syntactic analysis and semantic understanding of the regulatory texts, identifying their chapter structure and hierarchical relationships. The construction of the multi-level concept tree follows the "general-specific-detailed" principle. The top level represents the major regulatory categories, such as "Drug Administration" and "Medical Advertising"; the middle level contains specific regulatory entries; and the bottom level contains detailed regulations and explanations. For example, under the category of "Medical Advertising," there are middle-level concepts such as "prohibited content" and "access conditions," while "prohibited content" is further subdivided into specific regulations such as "exaggerating efficacy" and "using patient images." The concept tree is stored using a tree-like data structure, with each node containing fields such as concept ID, concept name, parent node ID, and concept description.

[0068] This study extracts legal entities and relationships, using legal entities as nodes and legal relationships as edges to establish the basic structure of a medical regulatory knowledge graph. Legal entities include legal subjects (e.g., "medical institutions," "pharmaceutical manufacturers"), regulated objects (e.g., "prescription drugs," "medical advertisements"), violations (e.g., "false advertising"), and penalties (e.g., "fines," "license revocation"). Legal relationships represent semantic connections between entities, such as "prohibition relationships," "regulatory relationships," and "penalty relationships." Entity and relationship extraction employs named entity recognition and relationship extraction techniques, combined with rule templates and deep learning methods. For example, for the regulatory text "Medical institutions shall not publish false medical advertisements; violators shall be fined," the extracted entities include "medical institutions," "false medical advertisements," and "fines," and the extracted relationships include "medical institutions - prohibition - false medical advertisements" and "false medical advertisements - penalties - fines." Entities and relationships are stored in a graph database, supporting efficient graph query and traversal operations.

[0069] The attribute information of regulatory entities is expanded recursively. This attribute information includes the entity's definition scope, applicable conditions, and exceptions. Recursive expansion refers to starting from an initial entity and continuously expanding related entities and attributes through association relationships. For example, for the entity "prescription drug," recursively expanding its attributes to include "requires a prescription" and "cannot be advertised," with related entities including "pharmacist" and "medical institution." Inference rules are established based on conditional judgments and logical relationships in the regulations, using a "premise-conclusion" structure, such as "IF the drug is a prescription AND advertising in public media THEN violates advertising laws." Inference rules are expressed using rule languages, such as SWRL (Semantic Web Rule Language), supporting automatic inference and rule matching.

[0070] A constraint propagation network is constructed based on inference rules to achieve dynamic reasoning of regulatory knowledge. The constraint propagation network is a graph-based reasoning framework where network nodes represent regulatory entities and attributes, and edges represent inference rules. The construction process first converts all inference rules into constraints, and then establishes propagation paths between nodes. For example, the rule "prescription drugs must not be advertised" is converted into constraint edges from the "prescription drug" node to the "advertising prohibited" node. The network supports forward reasoning (from condition to conclusion) and backward reasoning (from conclusion to condition), enabling multi-step chained reasoning. The constraint propagation algorithm uses a message passing mechanism; when the state of a node changes, the impact is propagated to connected nodes through edges, triggering new inference. For example, when a drug is identified as a prescription drug, compliance restrictions on its advertising behavior can be deduced through constraint propagation. The network implementation adopts a rule engine-based architecture, such as Drools, supporting efficient processing of complex rule sets.

[0071] Medical entities are mapped to a medical regulatory knowledge graph, and path similarity and semantic similarity between entities are calculated. Medical entities originate from the aforementioned medical entity identification results, including disease names, drug names, and treatment methods. The mapping process first standardizes medical entities to match the entity format in the knowledge graph, and then performs entity matching. Matching methods include exact matching and fuzzy matching. Exact matching directly finds entities with identical names; fuzzy matching calculates the edit distance of entity names or uses the cosine similarity of word vectors, with a threshold set to 0.85. Path similarity measures the similarity of the structural relationships between two entities in the graph, calculated by comparing the shortest path lengths from them to common entities. For example, "aspirin" and "ibuprofen" are both directly connected to the "over-the-counter drug" entity, resulting in high path similarity. Semantic similarity measures the proximity of entities in the semantic space, calculated using entity embedding techniques. Entity embedding uses methods such as TransE to represent entities as low-dimensional vectors (typically 100-dimensional), and then calculates the cosine similarity between vectors.

[0072] An entity alignment matrix is ​​constructed based on path similarity and semantic similarity. This matrix maps medical entities to regulatory entities, with behaviors representing medical entities, columns representing regulatory entities, and element values ​​representing alignment scores. The alignment score is calculated as a weighted average of path similarity (weight 0.4) and semantic similarity (weight 0.6). For example, the medical entity "ibuprofen tablets" has a path similarity of 0.9 and a semantic similarity of 0.85 with the regulatory entity "over-the-counter drug," resulting in an alignment score of 0.87. The matrix is ​​constructed using sparse matrix storage, only storing elements with alignment scores exceeding a threshold (e.g., 0.7) to improve storage and computational efficiency.

[0073] Reasoning is performed within a constraint propagation network using an entity alignment matrix. The reasoning process, based on a constraint propagation algorithm, maps medical entities to regulatory entities via the alignment matrix. Then, corresponding nodes are activated in the constraint propagation network, propagating the impact along constraint edges to check for violations of regulations. For example, if the medical entity "atorvastatin" is mapped to "prescription drug" and advertising is detected in the content, the constraint rule "prescription drugs must not be advertised" is triggered, marking it as a potential violation. The reasoning result includes information such as the type of violation, the basis for the violation, and the severity of the violation.

[0074] A weighted statistical analysis is performed on non-compliant entities, and a compliance score is calculated based on the severity of the violation. The weighted analysis considers the number, type, and contextual importance of the non-compliant entities, with different types of violations having different weights. The severity of the violation is based on the legal definitions of penalties and social impact, and is categorized into three levels: mild (1 point), moderate (2 points), and severe (3 points). For example, "prescription drug advertising violation" is at the severe level (3 points), and "exaggerated medical effects" is at the moderate level (2 points). The compliance score uses a full-score deduction system, with an initial score of 100 points. Each non-compliant entity has points deducted according to its severity and weight. The calculation formula is: Compliance Score = 100 - Sum of all violation deductions. If the final score is less than 0, it is set to 0. For example, if one prescription drug advertising violation (15 points deducted) and two exaggerated effect descriptions (8 points each deducted) are detected, the compliance score is 100 - 15 - 16 = 69 points.

[0075] Figure 2The graph compares the accuracy of different compliance scoring methods, with the horizontal axis representing dataset size (in thousands of medical records) and the vertical axis representing the percentage of accuracy. The method of this invention (marked with black circles) demonstrates significant performance advantages. As the dataset size increases from 1,000 to 25,000 records, its accuracy steadily improves from 87.5% to 96.1%, exhibiting good scalability and learning ability. Particularly on large-scale datasets, the accuracy improvement remains stable, indicating that this method can effectively utilize incremental data to continuously optimize performance. The rule-matching method (marked with dark gray squares) has relatively low accuracy, improving only from 75.2% to 77.8%, and showing a slight downward trend after the data volume exceeds 15,000 records. This reflects the limitations of rule-matching methods in handling complex and varied medical content, making it difficult to extract additional benefits from more data. The keyword filtering method (marked with light gray triangles) performs the worst, with its accuracy consistently hovering around 70%, indicating that simple keyword matching strategies cannot effectively address the complex requirements of medical regulatory compliance assessments. This invention, combining medical entity recognition technology and semantic matching algorithms, effectively identifies potential violations in content and provides reasonable compliance scores. Compared to traditional keyword filtering or rule matching methods, this solution possesses stronger semantic understanding and reasoning capabilities, enabling it to handle complex regulatory clauses and implicit compliance requirements. The knowledge graph construction supports structured representation and flexible querying of regulatory knowledge, the constraint propagation network enables dynamic reasoning and chain judgments, and entity alignment technology solves the mapping problem between medical terminology and regulatory concepts. This multi-technology integration significantly improves the accuracy, interpretability, and reliability of medical content compliance assessment, providing strong protection for content security management in medical information systems.

[0076] Optionally, A security assessment vector is constructed based on the overall credibility, the sensitivity score, and the compliance score; the security assessment vector is evaluated according to the security threshold and the scenario attributes of the display terminal; when the security assessment is passed, the step of writing the content to be processed and its security assessment data into the terminal display buffer includes: The overall credibility, sensitivity score, and compliance score are normalized and weighted to construct a security assessment vector. Based on the device type, usage scenario, and user permission information of the display terminal, terminal relevance, scenario relevance, and permission relevance are calculated respectively; the terminal relevance, scenario relevance, and permission relevance are weighted and combined to obtain the scenario adaptation coefficient; the preset security benchmark threshold is multiplied by the scenario adaptation coefficient to obtain the dynamic security threshold; The security assessment vector is compared with the dynamic security threshold to generate a security level identifier; a content tracing index is established based on the security level identifier; the content tracing index is bound to the operator's identity information and operation time information; and a digital watermark containing the content tracing index is embedded in the content to be processed. Based on the security level identifier, the content to be processed is divided into general level, guidance level and restricted level; the location attribute information and target audience attribute information of the display terminal are obtained, the content display level is determined, and the content to be processed is filtered or replaced according to the content display level to generate the content to be displayed; The content to be displayed is subjected to network security testing. After the test is passed, the content to be displayed is combined with scene information and time information to generate an anti-tampering verification value and written into the terminal display buffer.

[0077] In this embodiment, the overall credibility, sensitivity score, and compliance score are normalized and weighted to construct a security assessment vector. The specific implementation method for calculating the sensitivity score of the content to be processed is as follows: A conventional hierarchical filtering technique is used to calculate the sensitivity score of the content to be processed. The system maintains a medical sensitive word database, including patient privacy terms, medical risk terms, and specific disease names, categorized into high, medium, and low sensitivity levels. The content to be processed is segmented into words to identify the contained sensitive words and calculate their density. Simultaneously, rule matching is used to identify sensitive information in specific formats, such as ID card numbers, phone numbers, and other personal identifiers. The system accumulates the original sensitivity scores of the identified sensitive elements according to preset weights (high: 3 points, medium: 2 points, low: 1 point), and then normalizes the scores by combining this with the content length, ultimately obtaining a sensitivity score in the range of 0-1. For example, a patient's surgical plan might score 0.75 after sensitive word matching, suitable for display at a medical workstation but not for public display on a waiting area screen.

[0078] Normalization uses a min-max normalization method to map each score to a range of 0 to 1. Specifically, the calculation is as follows: subtract the minimum value for that dimension from the original score, then divide by the difference between the maximum and minimum values. For example, if the overall credibility score is 85 out of 100, the normalized score is 0.85; if the sensitivity score is 60 out of 100 (a higher score indicates lower sensitivity), the normalized score is 0.6; and if the compliance score is 90 out of 100, the normalized score is 0.9. Weighted combination uses a linear weighting method, with weights adjusted according to the application scenario: the overall credibility score has a weight of 0.4, the sensitivity score has a weight of 0.3, and the compliance score has a weight of 0.3. The three normalized scores are multiplied by their respective weights and then summed to obtain the security assessment value. For example, if the normalized overall credibility, sensitivity score, and compliance score are 0.85, 0.6, and 0.9 respectively, then the security assessment value is 0.85×0.4+0.6×0.3+0.9×0.3=0.79. The security assessment vector consists of these three normalized scores and their weighted combination value, in the format of a quadruple (0.85,0.6,0.9,0.79).

[0079] Based on the device type, usage scenario, and user permission information of the display terminal, terminal relevance, scenario relevance, and permission relevance are calculated respectively. Terminal relevance reflects the degree of adaptation between content and the display device type, calculated using a lookup table matching score. Device types include fixed large screens in hospitals, mobile trolley screens, doctor workstations, and patient terminals. Different types of content correspond to different device type adaptation scores. For example, professional medical image content has a terminal relevance of 0.95 on a doctor's workstation and a relevance of 0.4 on a patient terminal. Scenario relevance indicates the degree of matching between content and the usage scenario, including consultation rooms, waiting areas, wards, and public areas. The calculation method is also a lookup table matching. For example, medication instructions have a scenario relevance of 0.9 in a consultation room scenario and a relevance of 0.6 in a public area. Permission relevance measures the degree of matching between content and user permissions, which are divided into levels such as doctor, nurse, patient, and visitor. Higher permissions allow access to a wider range of content. For example, the correlation between surgical details and the doctor's permissions was 0.95, while the correlation with the patient's permissions was 0.3.

[0080] The scenario adaptability coefficient is obtained by weighting and combining terminal relevance, scenario relevance, and permission relevance. The weighting also uses a linear weighting method, with the weights allocated as follows: terminal relevance 0.3, scenario relevance 0.4, and permission relevance 0.3. For example, if the terminal relevance of certain content is 0.8, the scenario relevance is 0.7, and the permission relevance is 0.9, then the scenario adaptability coefficient is 0.8 × 0.3 + 0.7 × 0.4 + 0.9 × 0.3 = 0.79. The dynamic security threshold is obtained by multiplying the preset security baseline threshold by the scenario adaptability coefficient. The security baseline threshold is the minimum security assessment value required by the system, usually set to 0.7. The dynamic security threshold is adjusted by multiplying the security baseline threshold by the scenario adaptability coefficient, allowing it to change flexibly according to specific scenarios. For example, if the security baseline threshold is 0.7 and the scenario adaptability coefficient is 0.79, then the dynamic security threshold is 0.7 × 0.79 = 0.553.

[0081] The security assessment vector is compared with the dynamic security threshold to generate a security level label. The comparison method is as follows: first, it is determined whether the weighted combination value is greater than or equal to the dynamic security threshold. If so, it is further checked whether the scores of each sub-item exceed their respective minimum requirements. The minimum requirement for overall credibility is 0.6, the minimum requirement for sensitivity score is 0.5, and the minimum requirement for compliance score is 0.7. Based on the comparison results, the content security level is divided into three levels: high security (Level A), medium security (Level B), and low security (Level C). If all indicators meet the requirements and the weighted combination value exceeds the dynamic security threshold by more than 20%, it is Level A; if all indicators meet the requirements and the weighted combination value is between the dynamic security threshold and 120% of the dynamic security threshold, it is Level B; if only the weighted combination value meets the requirements, or a certain indicator is slightly below the requirement but not more than 10%, it is Level C; otherwise, it fails and no security level label is generated.

[0082] A content traceability index is established based on security level identifiers. The content traceability index is a unique identifier composed of a security level code, content type code, timestamp, and random string, in the format "security level-content type-timestamp-random string". For example, "A-MED-20250605123045-8F7D9A" represents medical content with level A security, generated at 12:30:45 PM on June 5, 2025, followed by the random string. The content traceability index is bound to operator identity information and operation time information. Operator identity information includes user ID, role type, department, etc., and operation time information includes creation time, modification time, and review time. The bound information is stored in a blockchain or secure database to ensure immutability. A digital watermark containing the content traceability index is embedded in the content to be processed. The digital watermark embedding adopts frequency domain watermarking technology. For text content, text interval adjustment or invisible character insertion methods are used; for image content, DCT coefficient modification methods are used; and for video content, inter-frame difference coding methods are used. The watermark strength is adjusted according to the content type to ensure that the watermark does not affect the content quality while having sufficient robustness.

[0083] Based on security level identifiers, content to be processed is categorized into general, guidance, and restricted levels. Level A security content corresponds to the general level and is suitable for widespread dissemination; Level B security content corresponds to the guidance level and should be used under the guidance of professionals; Level C security content corresponds to the restricted level and is only accessible to specific groups of people. The system obtains the location attribute information and target audience attribute information of the display terminal, filters or replaces the content to be processed according to the content display level, and generates content to be displayed.

[0084] The system performs environmental security checks and access permission verification on the content to be displayed. Environmental security checks include network security checks and physical environment checks. Network security checks verify whether the communication between the terminal and the server is encrypted and whether the connection is secure. Physical environment checks analyze the surrounding environment through the terminal's camera to confirm the presence of unsuitable users. Access permission verification checks whether the current user has permission to view the content, using a role-based access control model combined with biometric authentication (such as fingerprints and facial recognition) to enhance security. After successful verification, the content to be displayed is combined with scene information and time information to generate an anti-tampering verification value. The anti-tampering verification value is calculated using the HMAC algorithm, which combines the content hash value with scene information, time information, and the system key. Finally, the content to be displayed, content traceability index, security level identifier, and anti-tampering verification value are written to the terminal display buffer. The writing process uses atomic operations to ensure data consistency. Before the terminal display engine reads the buffer data for rendering, it first verifies the anti-tampering verification value to ensure that the content has not been illegally modified.

[0085] This invention constructs a multi-dimensional security assessment vector, enabling the system to comprehensively measure the credibility, sensitivity, and compliance of content. Through dynamic security thresholds, it achieves adaptive matching between content security standards and display scenarios. Content traceability indexing and digital watermarking technologies ensure content traceability and tamper-proof capabilities. A scenario-based content grading and adjustment mechanism ensures the secure and adaptable display of medical information in different environments. Compared to traditional static content review methods, this solution offers higher intelligence and adaptability, effectively balancing the security and usability of medical information, providing strong security guarantees for content display on hospital screen terminals, and significantly improving the content security management level of medical information systems.

[0086] Optionally, The steps of obtaining location attribute information and target audience attribute information of the display terminal, determining the content display level, filtering or replacing the content to be processed according to the content display level, and generating the content to be displayed include: The spatial location information and time information of the display terminal are obtained, and the location sensitivity is calculated; the time change rate and spatial second derivative of the location sensitivity are calculated, and the spatial dynamic characteristics are obtained by combining the time change rate and the spatial second derivative. Obtain the proportion of professionals and non-professionals in the area where the display terminal is located, and construct a crowd proportion vector; establish a group interaction matrix, and multiply the crowd proportion vector with the group interaction matrix to obtain an audience feature vector; combine the audience feature vector with the spatial dynamic features to generate scene features; The content display level of the display terminal is determined based on the scene characteristics; when the content display level is lower than the content classification level, the level difference is calculated; when the level difference exceeds a preset filtering threshold, the corresponding content is filtered; when the level difference does not exceed the filtering threshold, the content to be processed is segmented, and the local sensitivity of each segment and the correlation between content segments are calculated; the local sensitivity and the correlation are combined to obtain the global sensitivity; the content replacement intensity is determined based on the scene characteristics and the global sensitivity, and the content to be processed is processed according to the content replacement intensity to obtain the replaced content; The replaced content is subjected to multi-objective optimization to obtain the optimal replacement scheme. The multi-objective optimization includes a comprehensive evaluation of information entropy, sensitive information loss degree and context bias. The content to be displayed is generated according to the optimal replacement scheme.

[0087] In this embodiment, spatial location information is obtained through the hospital's positioning system, including the building, floor, area, and specific coordinates of the terminal. For example, a terminal is located in the internal medicine clinic area on the 2nd floor of the outpatient building, with coordinates (X=125.5, Y=78.3). Time information includes the current time, date type (weekday / holiday), and time period category (outpatient / non-outpatient). Location sensitivity is calculated using a preset area sensitivity mapping table, with different areas having different sensitivity values ​​ranging from 0 to 1. For example, the sensitivity of a general waiting area is 0.3, that of a specialist clinic is 0.7, and that of the intensive care unit is 0.9. Time factors adjust location sensitivity; peak outpatient hours increase sensitivity in public areas, while non-outpatient hours decrease sensitivity. The adjustment method involves adding or subtracting a time adjustment factor from the base sensitivity. This factor is determined based on the flow of people during the time period, ranging from -0.2 to 0.2. For example, the adjustment factor during peak outpatient hours is 0.15, and the adjustment factor at night is -0.1. If a terminal is located in a regular waiting area (base sensitivity 0.3), and the current time is 9:00 AM on a weekday (peak outpatient period, adjustment factor 0.15), then the location sensitivity is 0.45.

[0088] The system calculates the temporal rate of change and the spatial second derivative of location sensitivity. The temporal rate of change reflects the trend of location sensitivity over time and is calculated by comparing the sensitivity at the current moment with that at the previous moment. The system calculates location sensitivity every 15 minutes, recording the values ​​at the most recent 8 time points to form a time series. The temporal rate of change is calculated by subtracting the previous value from the current value and then dividing by the time interval (15 minutes). For example, if the current sensitivity is 0.45 and the previous value was 0.42, then the temporal rate of change is 0.2 (i.e., an increase of 0.12 per hour). The spatial second derivative reflects the degree of change in location sensitivity in space and is calculated by comparing the differences in sensitivity between adjacent regions. The system constructs a region adjacency graph, recording the sensitivity difference between each region and its neighboring regions. For each region, the standard deviation of its sensitivity differences with all neighboring regions is calculated to obtain the spatial first derivative. The spatial second derivative is calculated by comparing the differences in the spatial first derivatives between adjacent regions. For example, the spatial second derivative at the boundary between the waiting area and the consultation room area is relatively high (0.6), indicating drastic changes in sensitivity; while the spatial second derivative at different locations within the same waiting area is relatively low (0.1), indicating gradual changes in sensitivity. The spatial dynamic characteristics are obtained by weighted combination of the time rate of change (weight 0.4) and the spatial second derivative (weight 0.6), reflecting the dynamic changes in location sensitivity.

[0089] The proportion of professionals and non-professionals in the area where the display terminal is located is obtained to construct a crowd proportion vector. Crowd data is obtained through a video analytics system or a hospital access control system. The video analytics system uses deep learning algorithms to classify crowds in surveillance videos, dividing them into medical professionals (such as doctors and nurses) and non-professionals based on clothing characteristics (such as white coats, work uniforms, and nurses' uniforms) and identification features (such as name tags). The hospital access control system records card swipes or registration information for different types of personnel, distinguishing between professionals holding work permits and other personnel. The crowd proportion vector is a two-dimensional vector, representing the proportion of professionals (doctors, nurses, etc.) and non-professionals (patients, visitors, etc.), with the sum of the two being 1. For example, the crowd proportion vector of an internal medicine clinic area is (0.3, 0.7), indicating that professionals account for 30% and non-professionals account for 70%.

[0090] The group interaction matrix is ​​a 2×2 matrix describing the interaction strength between different groups. The rows and columns represent professionals and non-professionals, respectively, and the element values ​​represent the interaction strength between the corresponding groups, ranging from 0 to 1. For example, the interaction strength between professionals is 0.5, the interaction strength between professionals and non-professionals is 0.7, and the interaction strength between non-professionals is 0.4. The group interaction matrix is ​​obtained through a crowd behavior analysis system that tracks crowd aggregation distribution and interaction frequency, and statistically analyzes the interaction between different groups. Multiplying the crowd proportion vector by the group interaction matrix yields a two-dimensional audience feature vector, representing the actual influence distribution after considering group interaction. For example, multiplying the crowd proportion vector (0.3, 0.7) by the group interaction matrix yields the audience feature vector (0.34, 0.66), indicating that the actual influence of professionals is 34% and that of non-professionals is 66%. The audience feature vector is combined with spatial dynamic features to generate scene features. The combination method is to construct a three-dimensional vector, with the first two dimensions being the audience feature vector and the third dimension being the spatial dynamic features. For example, the scene features are (0.34, 0.66, 0.35) when the audience feature vector (0.34, 0.66, 0.35) is combined with the spatial dynamic feature 0.35.

[0091] The content display level of the display terminal is determined based on scene characteristics. There are three levels: General, Guidance, and Restricted, corresponding to security level identifiers. The determination of the display level uses a combination of rule-based and machine learning methods. The rule component includes: if the influence of professionals is below 20% and the influence of non-professionals is above 80%, the highest display level is General; if the spatial dynamic feature is above 0.6, the highest display level is Guidance, etc. The machine learning component uses a random forest classifier, taking scene feature vectors as input and outputting the content display level. The classifier is trained on historical data containing over 5000 labeled scene-level correspondence records. The final content display level is obtained through a combination of rule-based and classifier-based judgments. For example, scene features (0.34, 0.66, 0.35) are classified as Guidance.

[0092] When the content display level is lower than the content classification level, a level difference is calculated. The content classification level is the sensitivity level of the content itself, also divided into three levels: general, guidance, and restricted, corresponding to safety levels A, B, and C. The level difference is the difference between the content classification level and the display level. For example, if the content classification is restricted and the current display level is guidance, the level difference is 1. When the level difference exceeds a preset filtering threshold, the corresponding content is filtered. The filtering threshold is usually set to 2, meaning that if the difference exceeds two levels, the content is directly filtered. Filtering includes two methods: complete removal and partial filtering. Complete removal is suitable for content with high overall sensitivity that cannot be divided; partial filtering is suitable for divisible content, removing only the highly sensitive parts. For example, if drug adverse reaction details (restricted level) are to be displayed on the waiting area display screen (general level), and the level difference is 2, reaching the filtering threshold, the system will completely remove the content.

[0093] When the level difference does not exceed the filtering threshold, the content to be processed is segmented, and the local sensitivity of each segment and the degree of correlation between content segments are calculated. Segmentation is based on the natural boundaries (such as paragraphs, chapters) or semantic boundaries of the content, with each segment typically consisting of 3-5 sentences. Local sensitivity is calculated through a combination of sensitive word identification and semantic analysis. Sensitive word identification uses a medical sensitive dictionary containing tens of thousands of sensitive terms and their weights; semantic analysis uses the BERT model to determine the implicit sensitive information in the content. The combined score ranges from 0 to 1, with higher scores indicating higher sensitivity. The degree of correlation between content segments represents the semantic coherence and information dependence between paragraphs. This is calculated by measuring the semantic similarity and referential strength between paragraphs, ranging from 0 to 1, with higher scores indicating stronger correlation. For example, after segmenting the contents of the drug instructions, the local sensitivity of the "Indications" section is 0.3, the sensitivity of the "Dosage and Administration" section is 0.5, and the sensitivity of the "Adverse Reactions" section is 0.8. The correlation between "Indications" and "Dosage and Administration" is 0.7, and the correlation between "Dosage and Administration" and "Adverse Reactions" is 0.4.

[0094] The global sensitivity is obtained by combining local sensitivity and relevance. The combination method uses a graph structure, with content segments as nodes, relevance as edge weights, and local sensitivity as the initial node value. Through an iterative propagation algorithm, nodes influence each other, ultimately yielding a global sensitivity that considers relevance. The global sensitivity remains a value between 0 and 1. The content replacement intensity is determined based on scene features and the global sensitivity. The replacement intensity is a value between 0 and 1, representing the degree of content replacement. The calculation method is as follows: find matching scene templates based on scene features, determine the basic replacement intensity by combining it with the global sensitivity, and then adjust it according to the level difference. For example, in a typical waiting area scene (few professionals, many non-professionals), the global sensitivity is 0.7, the level difference is 1, and the replacement intensity is 0.6. The content to be processed is then processed according to the content replacement intensity to obtain the replaced content. Replacement processing includes: terminology replacement (replacing technical terms with common expressions), detail omission (omitting sensitive details), abstract generalization (replacing specific descriptions with general descriptions), and reorganization (adjusting the content structure). The system maintains a thesaurus of medical terms, containing professional terms and their colloquial expressions at different levels. For example, "myocardial infarction" can be replaced with "heart attack" (mild replacement) or "heart disease" (severe replacement). Different replacement strategy combinations are selected based on the replacement intensity to generate the replaced content.

[0095] The optimal replacement scheme is obtained through multi-objective optimization of the replaced content. Multi-objective optimization considers three objectives: information entropy, loss of sensitive information, and context bias. Information entropy represents the amount of information retained in the content, calculated as the information ratio before and after replacement; loss of sensitive information represents the degree to which sensitive information is filtered, calculated as the proportion of sensitive content removed; context bias represents the semantic deviation between the replaced content and the original content, calculated through semantic similarity. A genetic algorithm is used for multi-objective optimization, with a population size of 100 and an evolutionary generation of 30. The replacement scheme is iteratively optimized to find the optimal balance between the three objectives. Finally, the optimal solution on the Pareto front is selected as the final replacement scheme. Content to be displayed is generated based on the optimal replacement scheme, including the content itself, display format, and layout information.

[0096] This embodiment constructs a comprehensive scene perception model by integrating spatial location, temporal changes, and crowd characteristics, enabling precise perception of dynamic changes in the display environment. By introducing the concepts of level difference and global sensitivity, it achieves refined adjustments to content processing strategies, providing strong support for the intelligent and scenario-based display of hospital information screen content. This not only ensures the safe dissemination of medical information but also enhances the relevance and effectiveness of information display, which is of great significance for improving the hospital's informatization level and patient satisfaction.

[0097] Optionally, the step of performing multimodal scanning to calculate the sensitivity score of the content to be processed includes: separating the content to be processed into text, image, and table modes; constructing character-level feature sequences and word-level feature sequences for the text content, and extracting sensitive text features using a bidirectional scanning method; segmenting the image content, extracting visual and text features of the image blocks, and identifying sensitive image regions based on an attention mechanism; extracting cell relationship features and numerical distribution features for the table content, and identifying sensitive data combinations; matching the sensitive text features, the sensitive image regions, and the sensitive data combinations with preset sensitive feature templates to obtain a mode-level sensitivity score; constructing a cross-modal feature association graph, and analyzing the correlation strength of sensitive information between different modes based on the cross-modal feature association graph; and weighting and combining the mode-level sensitivity score with the correlation strength of sensitive information to obtain the overall sensitivity score of the content to be processed.

[0098] In this embodiment, the content to be processed is modally separated into text, images, and tables. Modal separation is based on the content's format features and tag information, and is performed recursively. For HTML or XML content, it is distinguished according to tag type; for Word, PDF, and other formats, different modal content is extracted by parsing the document object model. Image content includes independent images and embedded graphics; table content includes HTML tables, Word tables, and table structures within images; the remaining parts are classified as text content.

[0099] Character-level and word-level feature sequences are constructed from the text content, and a bidirectional scanning method is used to extract sensitive text features. Character-level feature sequences are constructed using character embedding, with each character mapped to a fixed-dimensional (typically 64-dimensional) vector. Word-level feature sequences are constructed using word embedding after word segmentation, employing a pre-trained word vector model in the medical field with a vector dimension of 200. Bidirectional scanning includes both left-to-right and right-to-left directions, each processed using a Bidirectional Long Short-Term Memory (BiLSTM) network. The hidden layer size of the BiLSTM is set to 128, and the output sequence passes through a self-attention layer to obtain the weight distribution. Sensitive text feature extraction consists of three levels: basic sensitive word identification, semantic sensitivity analysis, and contextual sensitivity assessment. Basic sensitive word identification uses a medical sensitivity dictionary, including multiple categories such as patient privacy, medical disputes, and adverse reactions, totaling approximately 8000 entries. Semantic sensitivity analysis is based on pre-trained models such as BERT to determine the implicit sensitive information in the text. Contextual sensitivity assessment analyzes the sensitivity of words in a specific context. For example, the sensitivity of the word "tumor" varies in different contexts: it has a sensitivity of 0.3 in "tumor is a common disease," while it reaches 0.85 in "a patient was diagnosed with a malignant tumor." Text sensitivity features are represented as triples (position, type, score), where position identifies the start and end points of the sensitive text, type represents the sensitivity category, and the score ranges from 0 to 1.

[0100] Image content is segmented into blocks, and visual and textual features of each block are extracted. Sensitive image regions are identified based on an attention mechanism. Image segmentation employs an adaptive grid method, dividing the image into irregular blocks, with block size dynamically adjusted according to the complexity of the image content. For medical images, blocks are smaller in dense areas and larger in sparse areas. Visual feature extraction uses convolutional neural networks (such as ResNet50 or EfficientNet), with the final convolutional feature map serving as the visual feature, having a dimension of 7×7×2048. Text in the image is extracted using optical character recognition (OCR) and then processed using text feature extraction methods. The attention mechanism combines channel attention and spatial attention, calculating the attention weight for each image block. Sensitive image region identification is based on several aspects: patient identification information (e.g., face, ID number), medically sensitive areas, and abnormal medical findings (e.g., tumors, lesions). The system uses instance-based segmentation to accurately locate the boundaries of sensitive regions. For each identified sensitive region, a sensitivity score is calculated, ranging from 0 to 1. For example, for a chest X-ray, the system may identify nodules in the upper right lung region (sensitivity 0.78) and blurred areas in the lower left lung region (sensitivity 0.65), while also marking patient information at the image edges (sensitivity 0.95).

[0101] The system extracts cell relationship features and numerical distribution features from table content to identify sensitive data combinations. The table structure is parsed using cell boundary detection and merged cell recognition to construct a logical tree structure. Cell relationship features describe the row and column relationships and semantic associations between cells, calculated through the correspondence between table headers and data cells. Numerical distribution features analyze the statistical distribution characteristics of numerical cells, including mean, variance, and outlier ratio. Sensitive data combinations refer to sets of cells that may reveal privacy or sensitive information. Identification methods combine rule matching and machine learning. Rule matching targets known sensitive data patterns, such as personal ID numbers and contact information; machine learning methods are used to identify complex sensitive combinations, such as combinations of abnormal test results and diagnoses. The system constructs a sensitive data pattern library containing over 100 common medical sensitive data patterns. For each identified sensitive data combination, its location, type, and sensitivity score are recorded. For example, in a blood test report table, the system might identify a combination of HIV test results (sensitivity 0.9) and abnormal liver function indicators (sensitivity 0.7).

[0102] Sensitive text features, sensitive image regions, and sensitive data combinations are matched against predefined sensitive feature templates to obtain modality-level sensitivity scores. Sensitive feature templates are predefined sets of sensitive features, categorized according to different application scenarios and sensitivity levels. For example, in a medical teaching scenario, the patient privacy information template has a "high" sensitivity level, while the professional terminology template has a "low" sensitivity level; whereas in a public relations scenario, the professional terminology template has a "medium" sensitivity level. The matching process combines similarity calculation and rule-based judgment. Similarity calculation uses cosine similarity or the Jaccard coefficient, with a threshold of 0.75. For each modality, a sensitivity score is calculated based on the degree of matching between its sensitive features and the template. The calculation method is: the score of the sensitive feature is multiplied by the weight of the corresponding template, and then normalized. For example, the text modality identified 5 sensitive features and matched 3 templates, resulting in a sensitivity score of 0.72; the image modality identified 3 sensitive regions and matched 2 templates, resulting in a score of 0.85; and the table modality identified 2 sensitive data combinations and matched 1 template, resulting in a score of 0.63.

[0103] A cross-modal feature association graph is constructed, and the strength of the association between sensitive information in different modalities is analyzed based on the graph. The cross-modal feature association graph is an undirected weighted graph where nodes represent sensitive features in each modality, and edges represent the association relationships between features. Association relationships are established through the following methods: spatial proximity (features in adjacent locations may be related), semantic similarity (features describing the same topic), and referencing relationships (one modality refers to content from another modality). For example, the text mentions "such as..." Figure 2 The tumor region shown is associated with the tumor region marked in the image. The association strength is calculated by comprehensively considering association type, distance, and semantic similarity, ranging from 0 to 1. For example, the association strength between the text description "CT image shows a 2.3cm × 1.8cm nodule in the upper lobe of the right lung" and the corresponding nodule region in the image is 0.92. For the constructed association graph, graph analysis algorithms are used to calculate the centrality and community structure of the nodes. Sensitive features with high centrality have a greater impact on the overall sensitivity; closely associated features form sensitive information communities, representing related sensitive topics.

[0104] The overall sensitivity score of the content to be processed is obtained by weighting and combining the modality-level sensitivity scores with the association strength of sensitive information. The weighting combination considers the following factors: the importance weight of each modality, the association structure of sensitive features, and the sensitivity distribution. The importance weight is set according to the application scenario. For example, in medical image reports, the image modality has a high weight (0.5), followed by the text modality (0.3), and the table modality has the lowest weight (0.2). The association structure influences the calculation of graph features, with higher weights for sensitive features with high association. The sensitivity distribution adopts a non-linear mapping, giving higher weights to the high sensitivity range (0.8-1.0) to reflect the risk aversion characteristic of the sensitivity score. The combination method is as follows: first calculate the association-adjusted modality scores, then sum them according to the weights, and finally map them to the range of 0 to 1 through an S-shaped function. For example, the text modality score is 0.72 (weight 0.3), the image modality score is 0.85 (weight 0.5), and the table modality score is 0.63 (weight 0.2). After considering the text-image association strength of 0.92 and the text-table association strength of 0.75, the final overall sensitivity score is 0.79.

[0105] This embodiment enables the system to comprehensively capture various forms of sensitive information through collaborative analysis of text, images, and table content; it improves the accuracy of sensitive feature extraction by employing bidirectional scanning and attention mechanisms; and it reveals the intrinsic connections between sensitive information in different modalities based on the analysis of cross-modal feature association graphs, effectively solving complex sensitive scenarios that cannot be identified by single-modal analysis.

[0106] Secondly, it provides a content security management system for medical information systems, including: The first unit is used to construct a dual-channel semantic analysis network, which includes a medical professional channel based on a medical ontology library and a general semantic channel based on a general language model. The content to be processed is input into the dual-channel semantic analysis network and fused through a convolutional attention mechanism to identify medical entities. The second unit is used to divide the content to be processed into key areas and non-key areas based on the identification results of the medical entities; to perform professional verification and integrity verification on the key areas, and to perform standardization checks on the non-key areas; to construct a medical content semantic consistency network based on the processing results of different areas, to analyze the coherence of information between areas, and to calculate the overall credibility of the content. The third unit is used to calculate a sensitivity score for the content to be processed. The fourth unit is used to construct a medical regulatory knowledge graph, semantically match the medical entities with the medical regulatory knowledge graph, and calculate a compliance score. The fifth unit is used to construct a security assessment vector based on the overall credibility, the sensitivity score, and the compliance score; to evaluate the security assessment vector according to the security threshold and the scene attributes of the display terminal; when the security assessment is passed, the content to be processed and its security assessment data are written into the terminal display buffer; when the content of the terminal display buffer needs to be updated, the security assessment process is re-executed on the updated content.

[0107] Thirdly, a computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

Claims

1. A content security management method in a medical information system, characterized in that, include: A dual-channel semantic analysis network is constructed, which includes a medical professional channel and a general semantic channel; The content to be processed is input into the dual-channel semantic analysis network, and medical entities are identified through fusion via a convolutional attention mechanism. Based on the identification results of the medical entities, the content to be processed is divided into key areas and non-key areas; professional verification and integrity verification are performed on the key areas, and standardization checks are performed on the non-key areas; a medical content semantic consistency network is constructed based on the processing results of different areas, the coherence of information between areas is analyzed, and the overall credibility of the content is calculated. Calculate a sensitivity score for the content to be processed; Construct a medical regulations knowledge graph, semantically match the medical entities with the medical regulations knowledge graph, and calculate a compliance score; A security assessment vector is constructed based on the overall credibility, the sensitivity score, and the compliance score. The security assessment vector is evaluated according to the security threshold and the scenario attributes of the display terminal. When the security assessment is passed, the content to be processed and its security assessment data are written into the terminal display buffer. When the content of the terminal display buffer needs to be updated, the security assessment process is re-executed on the updated content.

2. The method according to claim 1, characterized in that, The steps to construct a dual-channel semantic analysis network include: A hierarchical relationship tree is established based on medical concepts in the Unified Medical Language System. The hierarchical relationship tree includes hierarchical relationships and partial whole relationships between medical concepts. The semantic types of the medical concepts are extracted based on the hierarchical relationship tree, and a set of semantic types of the medical concepts is constructed through a semantic type mapping function. Based on the set of semantic types, the medical concepts are converted into medical feature vectors through transpose operation, and a medical entity embedding matrix is ​​constructed. The distance matrix between medical concepts is calculated based on the hierarchical relationship tree. A hierarchical reachability mask is generated based on the distance matrix. The distance matrix and the hierarchical reachability mask are combined to obtain a concept-level attention matrix. A multi-layer stacked attention encoder is constructed, each layer of which includes a multi-head self-attention module and a feedforward neural network. In the multi-head self-attention module, the input features are divided into multiple attention heads, and each attention head calculates an attention score through a query matrix, a key matrix, and a value matrix. A position encoding matrix is ​​generated, and the position encoding matrix calculates text position information through a sinusoidal position encoding function.

3. The method according to claim 2, characterized in that, The steps of inputting the content to be processed into the dual-channel semantic analysis network and identifying medical entities in the content through convolutional attention mechanism include: The content to be processed is segmented and standardized to obtain a standardized text sequence; The standardized text sequence and the medical entity embedding matrix are used to perform feature calculations to obtain medical channel features. The medical channel features are then interacted with the concept-level attention matrix to obtain the medical channel hidden layer representation. The standardized text sequence is input into the attention encoder, and the output features are superimposed with the position encoding matrix to obtain general semantic channel features; Feature extraction with different convolution kernel sizes is performed on the medical channel hidden layer representation and the general semantic channel features to obtain multi-scale convolution features; inter-channel attention scores are calculated based on the multi-scale convolution features; and the medical channel hidden layer representation and the general semantic channel features are weighted and fused according to the inter-channel attention scores to obtain fused features. The number of medical concepts and the number of medical concept relationships are identified from the content to be processed, and the text length of the content to be processed is calculated; a professionalism score is calculated based on the number of medical concepts, the number of medical concept relationships, and the text length; the professionalism score is combined with the feature entropy of the fused features to generate boundary optimization weights. The desired boundary is calculated based on the preset internal structure rules of the medical entity; the boundary sensitivity of the initially identified medical entity boundary and the desired boundary is calculated to obtain the structural loss; the medical entity boundary is iteratively optimized based on the structural loss and the boundary optimization weight to obtain the medical entity recognition result.

4. The method according to claim 1, characterized in that, Based on the identification results of the medical entities, the content to be processed is divided into critical areas and non-critical areas; professional verification and integrity verification are performed on the critical areas, and standardization checks are performed on the non-critical areas. The steps involved in constructing a semantic consistency network for medical content based on processing results from different regions, analyzing the coherence of information between regions, and calculating the overall credibility of the content include: Importance information of the medical entities is obtained from the medical knowledge base, and regional importance scores are calculated based on the distribution of the medical entities. Based on the regional importance scores, regional boundaries are determined by minimizing the variance between regions, and the content to be processed is divided into key regions and non-key regions. For the critical area, a set of medical feature points is extracted, and the semantic similarity and contextual consistency of each feature point in the set with the standard concept set are calculated to obtain a professionalism verification matrix. Based on the professionalism verification matrix, medical statements are extracted, and the medical statements are matched with clinical guidelines to calculate the guideline compliance. The coverage of medical elements in the medical statements is also calculated to obtain a completeness score. The guideline compliance and the completeness score are combined to obtain a safety score for the critical area. For the non-critical areas, the terminology standardization score and contextual semantic similarity are calculated to generate a normative score for the non-critical areas; Feature vector transformation is performed on the key regions and the non-key regions to calculate the semantic similarity, entity overlap, and information flow intensity between regions, thus constructing a semantic consistency network. In the semantic consistency network, an initial state matrix is ​​constructed based on the statement credibility of the region nodes, and the initial state matrix is ​​updated through iterative propagation to obtain a consistency propagation matrix. The degree of conflict between regions is calculated based on the consistency propagation matrix. The security score, standardization score, and conflict level are fused together, and the overall credibility of medical content is determined by an adaptive threshold.

5. The method according to claim 4, characterized in that, The steps for performing professional verification and integrity verification on the critical areas include: The medical feature point set is grouped according to feature type to construct a feature group index table; a compressed standard concept set is created based on the feature group index table, and a concept fast matching index is established using the locality-sensitive hashing method; The key area is scanned using a sliding window method, and candidate medical statements are extracted based on a preset semantic template. The candidate medical statements are matched with the standard concept set to obtain an initial matching result. The semantic similarity and contextual consistency of the initial matching result are combined incrementally to generate a professional verification matrix. Based on the verification results of the professional verification matrix, valid medical statements are selected from the candidate medical statements; the valid medical statements are matched with clinical guideline rule templates, and the rule coverage rate is calculated to obtain the guideline compliance. A hierarchical medical element structure tree is constructed, and the effective medical statements are mapped to the medical element structure tree. The coverage ratio of the target nodes is calculated to obtain an integrity score.

6. The method according to claim 1, characterized in that, The steps of constructing a medical regulatory knowledge graph, semantically matching the medical entities with the medical regulatory knowledge graph, and calculating compliance scores include: A multi-level concept tree is constructed based on medical regulatory texts to extract regulatory entities and relationships. The regulatory entities are used as nodes and the regulatory relationships are used as edges to establish the basic structure of a medical regulatory knowledge graph. The attribute information of the regulatory entities is expanded in a recursive manner to establish reasoning rules between entities. Based on the reasoning rules, a constraint propagation network is constructed to realize dynamic reasoning of regulatory knowledge. The medical entities are mapped to the medical regulatory knowledge graph, and the path similarity and semantic similarity between entities are calculated. An entity alignment matrix is ​​constructed based on the path similarity and semantic similarity. The entity alignment matrix is ​​used to perform reasoning in the constraint propagation network to identify potential violating entities. The aforementioned entities that violated regulations were statistically weighted, and a compliance score was calculated based on the severity of the violation.

7. The method according to claim 1, characterized in that, A security assessment vector is constructed based on the overall credibility, the sensitivity score, and the compliance score; the security assessment vector is evaluated according to the security threshold and the scenario attributes of the display terminal; when the security assessment is passed, the step of writing the content to be processed and its security assessment data into the terminal display buffer includes: The overall credibility, sensitivity score, and compliance score are normalized and weighted to construct a security assessment vector. Based on the device type, usage scenario, and user permission information of the display terminal, terminal relevance, scenario relevance, and permission relevance are calculated respectively; the terminal relevance, scenario relevance, and permission relevance are weighted and combined to obtain the scenario adaptation coefficient; the preset security benchmark threshold is multiplied by the scenario adaptation coefficient to obtain the dynamic security threshold; The security assessment vector is compared with the dynamic security threshold to generate a security level identifier; a content tracing index is established based on the security level identifier; the content tracing index is bound to the operator's identity information and operation time information; and a digital watermark containing the content tracing index is embedded in the content to be processed. Based on the security level identifier, the content to be processed is divided into general level, guidance level and restricted level; the location attribute information and target audience attribute information of the display terminal are obtained, the content display level is determined, and the content to be processed is filtered or replaced according to the content display level to generate the content to be displayed; The content to be displayed is subjected to network security testing. After the test is passed, the content to be displayed is combined with scene information and time information to generate an anti-tampering verification value and written into the terminal display buffer.

8. The method according to claim 7, characterized in that, The steps of obtaining location attribute information and target audience attribute information of the display terminal, determining the content display level, filtering or replacing the content to be processed according to the content display level, and generating the content to be displayed include: The spatial location information and time information of the display terminal are obtained, and the location sensitivity is calculated; the time change rate and spatial second derivative of the location sensitivity are calculated, and the spatial dynamic characteristics are obtained by combining the time change rate and the spatial second derivative. Obtain the proportion of professionals and non-professionals in the area where the display terminal is located, and construct a crowd proportion vector; establish a group interaction matrix, and multiply the crowd proportion vector with the group interaction matrix to obtain an audience feature vector; combine the audience feature vector with the spatial dynamic features to generate scene features; The content display level of the display terminal is determined based on the scene characteristics; when the content display level is lower than the content classification level, the level difference is calculated; when the level difference exceeds a preset filtering threshold, the corresponding content is filtered; when the level difference does not exceed the filtering threshold, the content to be processed is segmented, and the local sensitivity of each segment and the correlation between content segments are calculated; the local sensitivity and the correlation are combined to obtain the global sensitivity; the content replacement intensity is determined based on the scene characteristics and the global sensitivity, and the content to be processed is processed according to the content replacement intensity to obtain the replaced content; The replaced content is subjected to multi-objective optimization to obtain the optimal replacement scheme. The multi-objective optimization includes a comprehensive evaluation of information entropy, sensitive information loss degree and context bias. The content to be displayed is generated according to the optimal replacement scheme.

9. A content security management system in a medical information system, used to implement the method of any one of claims 1-8, characterized in that, include: The first unit is used to construct a dual-channel semantic analysis network, which includes a medical professional channel based on a medical ontology database and a general semantic channel based on a general language model. The content to be processed is input into the dual-channel semantic analysis network, and medical entities are identified through fusion via a convolutional attention mechanism. The second unit is used to divide the content to be processed into key areas and non-key areas based on the identification results of the medical entities; to perform professional verification and integrity verification on the key areas, and to perform standardization checks on the non-key areas; to construct a medical content semantic consistency network based on the processing results of different areas, to analyze the coherence of information between areas, and to calculate the overall credibility of the content. The third unit is used to calculate a sensitivity score for the content to be processed. The fourth unit is used to construct a medical regulatory knowledge graph, semantically match the medical entities with the medical regulatory knowledge graph, and calculate a compliance score. The fifth unit is used to construct a security assessment vector based on the overall credibility, the sensitivity score, and the compliance score; to evaluate the security assessment vector according to the security threshold and the scene attributes of the display terminal; when the security assessment is passed, the content to be processed and its security assessment data are written into the terminal display buffer; when the content of the terminal display buffer needs to be updated, the security assessment process is re-executed on the updated content.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 8.