Internal threat identification and early warning method based on behavior analysis

By structurally decomposing and clustering medical staff query logs, and combining time series data with a permission database, the monitoring strategy is dynamically adjusted. This solves the problem of accuracy in identifying and warning of medical staff query behavior, enables real-time monitoring and risk assessment of internal threats, and protects medical data security.

CN120930180APending Publication Date: 2025-11-11GUANGZHOU FENGCHUAN NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511031779.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing technologies struggle to dynamically adjust analysis strategies and accurately assess the relevance of healthcare workers' query behavior to their professional fields. This results in low accuracy in identifying internal threats and an inability to respond promptly to changes in behavioral patterns, increasing the risk of sensitive information leakage.

Method used

By structurally analyzing medical staff's query behavior logs, and employing cluster analysis and time series analysis, combined with a domain responsibility and permission database, monitoring rules are dynamically adjusted to identify and issue warnings for abnormal behaviors.

Benefits of technology

It enables real-time monitoring and risk assessment of medical staff's query behavior, effectively preventing unauthorized access to and misuse of medical data, and protecting patient privacy and medical information security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930180A_ABST
    Figure CN120930180A_ABST
Patent Text Reader

Abstract

The invention provides an internal threat identification and early warning method based on behavior analysis, and the method comprises the steps: analyzing the matching degree of the record of a query object unassociated patient and the retrieval range exceeding the responsibility authority for a classified behavior feature subset, and carrying out the comparison through a pre-established field responsibility authority library, if the query time is concentrated in a non-working period or the retrieval time presents a periodic rule, judging that the behavior deviates from an expected range, and determining a preliminary abnormal behavior mark; and for the deep behavior analysis result, analyzing the similarity between the rare disease category related to the query content and the specific medical history focused by the retrieval content, and if the retrieval range exceeds the responsibility authority and the query behavior lacks the business context, determining the classification label of the illegal behavior to obtain the illegal behavior judgment conclusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information technology, and in particular to a method for identifying and warning of internal threats based on behavioral analysis. Background Technology

[0002] Research in the field of medical data security is crucial in today's information age, directly impacting patient privacy protection and the credibility of medical institutions. With the widespread application of medical information systems, the protection of sensitive data has become central to ensuring the quality of medical services and public trust. However, the identification and prevention of insider threats, especially inappropriate querying behavior by healthcare professionals, remains a critical area requiring significant breakthroughs, its importance self-evident. Currently, while many solutions attempt to restrict inappropriate querying behavior through rule settings or access control, these methods often overlook the complexity and dynamism of healthcare professionals' behavioral patterns. Existing methods largely rely on pre-set static rules, making it difficult to adapt to the diversity of individual behavioral habits and effectively capture adjustments in query needs caused by changes in work environment or responsibilities. This limitation often leads to the overlooking of potential insider threats, increasing the risk of sensitive information leakage. Against this backdrop, the main challenges in research focus on how to accurately identify the relevance of querying behavior to professional fields and how to address new problems arising from changes in behavioral patterns. Determining the relevance of querying behavior to professional fields is the primary challenge, as healthcare professionals' work may involve cross-disciplinary collaboration, and simply classifying the rationality of queries based on professional labels can easily lead to misjudgments. When behavioral patterns change, such as frequent queries for disease information unrelated to one's job duties or concentrated searches for data on specific groups, it becomes difficult to distinguish between normal work requirements and potential threatening behaviors if the analysis strategy cannot be dynamically adjusted. These two factors are closely linked: the former determines the accuracy of identification, while the latter further tests the system's adaptability, together constituting the technical bottleneck of insider threat identification. Therefore, how to design a mechanism that can dynamically adjust the analysis strategy to accurately assess the relevance of query behavior to the professional field and promptly identify potential threats when behavioral patterns change has become a key issue that this research urgently needs to address. Summary of the Invention

[0003] This invention provides a method for internal threat identification and early warning based on behavioral analysis, mainly including:

[0004] The system retrieves query behavior logs from medical information systems, structurally decomposes query operations by time, object, and content, forming an initial behavioral feature set. Based on this initial behavioral feature set, and combined with records of specific age groups and medical histories identified by the search objects, cluster analysis is used to group query behaviors, distinguishing behavioral patterns outside of responsibilities and permissions, resulting in a categorized subset of behavioral features. Using this categorized subset, the matching degree between query objects with unrelated patient records and search scopes exceeding responsibilities and permissions is analyzed. Combined with comparison to a domain responsibility and permission database, preliminary abnormal behavior markers are identified for behaviors where query times are concentrated during non-working hours or where search times exhibit periodic patterns. Based on these preliminary abnormal behavior markers, time series analysis is used to detect changes in query frequency and analyze query behavior deficiencies. The lack of business context and subsequent operation records are used to obtain a dynamic behavior change description. Based on this description, for queries with no subsequent medical actions or involving sensitive data fields, contextual information about unauthorized terminals and abnormal terminal location changes from the query source device is extracted to obtain deep behavior analysis results. According to these results, the similarity of query content involving rare diseases and the concentration of medical history is analyzed to determine the classification tags for violations. Using these classification tags, combined with abnormal terminal locations and unauthorized device access characteristics, an updated real-time monitoring scheme is generated. This updated scheme is used to analyze newly collected query behavior logs to detect whether behavior patterns deviate from the expected range, thus obtaining a violation risk assessment result.

[0005] Furthermore, the process involves obtaining query behavior logs from the medical information system, structurally breaking down the query operations into their time, object, and content to form an initial set of behavioral characteristics, including:

[0006] The system retrieves timestamps, patient identifiers, and disease codes from login records and operation logs of medical staff using database queries. It then extracts query time periods, operation frequencies, and disease type information using field parsing methods, resulting in structured query behavior data records. For the disease codes in these records, the system verifies the matching relationship between disease codes and professional field codes by combining them with a medical staff professional qualification table, marking query records that deviate from their professional fields. By statistically analyzing the frequency of identical medical staff identifiers in the query behavior data records, the system determines whether the operation frequency exceeds the workload threshold. Based on the timestamps in the query behavior data records, the system uses a time interval division method to identify whether the query time falls within non-working hours. Combined with a patient table, the system verifies the correlation of patient identifiers, constructing an initial set of behavioral features.

[0007] Furthermore, based on the initial set of behavioral features, combined with records of specific age groups and medical histories identified by the search object, cluster analysis is used to group the query behaviors, distinguishing behavioral patterns outside of responsibilities and authority, resulting in a categorized subset of behavioral features, including:

[0008] By matching disease codes in the initial behavioral feature set with the disease classification directory, rare disease information involved in the query content is identified, and the age group characteristics of the search object are obtained, resulting in a comprehensive feature record containing professional deviation, rare disease identification, age group code, and medical history content. For the medical history content and age group code in the comprehensive feature record, it is verified whether the medical history content involves a specific disease type and whether the age group code is concentrated in a specific range, marking specific objects for focused queries. The comprehensive feature record is grouped using the K-means clustering method to obtain cluster group identifiers for query behavior. Combined with the scope of medical staff's responsibilities and authority, it is verified whether the cluster group identifiers exceed preset permissions, marking behavioral patterns outside of responsibilities and authority, and generating a categorized subset of behavioral features.

[0009] Furthermore, by analyzing the matching degree of unrelated patient records and retrieval scopes exceeding the scope of responsibilities and permissions through the categorized behavioral feature subset, and combining this with comparison with the domain responsibility and permission database, preliminary abnormal behavior markers are determined for behaviors where query times are concentrated during non-working hours or where retrieval times exhibit periodic patterns, including:

[0010] The matching relationship between the query object identifiers in the categorized behavioral feature subset and the patient records managed by medical staff is verified through database association queries. The number of unrelated patient records and the degree of permission deviation are calculated. For the number of unrelated patient records and the degree of permission deviation, a comparison is made with the domain responsibility and permission database to mark permission violations. The query time information in the categorized behavioral feature subset is analyzed to identify non-working hours, calculate and detect periodic repetitive patterns, and determine time violation identifiers. Based on the permission violations and time violation identifiers, query behaviors that deviate from the expected range are marked.

[0011] Furthermore, based on the initial abnormal behavior markers, time series analysis is used to detect changes in query frequency, analyze the characteristics of query behavior lacking business context and subsequent operation records, and obtain a description of dynamic behavior changes, including:

[0012] The system acquires historical query records and current environment information from medical staff, extracts the historical query frequency baseline and the current query frequency value; for the current query frequency value, it verifies whether it exceeds the historical frequency baseline threshold and exhibits a short-term surge characteristic, marking the frequency abnormal query; it verifies the business rationality of the frequency abnormal query by comparing the operation sequence; and it generates a dynamic behavior change description based on the frequency abnormal query and business rationality verification results.

[0013] Furthermore, the description of dynamic behavior changes involves extracting contextual information about unauthorized terminals and abnormal location changes of the query source device for queries that lack subsequent diagnostic or treatment actions or involve sensitive data fields, thereby obtaining deep behavior analysis results, including:

[0014] The system identifies sensitive information types in the description of dynamic behavior changes, verifies whether query records generate medical activity records, and marks high-risk query behaviors that lack subsequent medical actions and have a sensitive information access frequency exceeding a threshold. It verifies the source device of the high-risk query behavior by comparing it with the device authorization list, identifying unauthorized terminal access records. It verifies the terminal location of the detected high-risk query behavior by comparing its location, analyzing the degree of location change anomalies. Based on the high-risk query behavior, unauthorized terminal access records, and the degree of location change anomalies, it determines the severity level of potential violations and generates in-depth behavior analysis results. Further, based on the in-depth behavior analysis results, it analyzes the similarity of the query content involving rare diseases and the concentration of medical history, determining the classification tags of the violations, including:

[0015] The system identifies rare disease types in the deep behavior analysis results, analyzes the concentration of medical history using content similarity comparison, and generates query content analysis data. Based on this data, and considering the scope of responsibilities and permissions, it verifies whether the search scope exceeds the permission boundaries and whether the rare disease types and the concentration of medical history deviate from the normal range, thus marking unauthorized query behavior. Through medical business relevance verification, it determines the business background of the unauthorized query behavior, identifies the category of violation, and generates classification tags for the violation.

[0016] Furthermore, the updated real-time monitoring scheme is generated by combining the classification tags of violations with the characteristics of abnormal terminal locations and unauthorized device access, including:

[0017] Obtain the violation type identifier, abnormal terminal location, and unauthorized device access records from the classification tags of the violations, and analyze the severity and frequency distribution of the violations; adjust the trigger threshold parameters of the behavior monitoring rules based on the severity and frequency distribution, and identify the time periods and device types where violations are concentrated; construct special monitoring rules for querying and periodic retrieval behaviors during non-working hours, and generate targeted monitoring conditions; integrate the trigger threshold parameters and special monitoring rules into the monitoring configuration, update the monitoring rule priority, and generate an updated real-time monitoring scheme.

[0018] Furthermore, the updated real-time monitoring scheme analyzes the newly collected query behavior logs to detect whether the behavior pattern deviates from the expected range, and obtains the risk assessment result of the violation, including:

[0019] The updated real-time monitoring scheme analyzes newly collected query behavior logs to obtain query time, operation frequency, user identifier, and device information. For the operation frequency, it verifies whether it exceeds the normal range threshold or exhibits short-term surge characteristics, determines the degree of deviation from the behavior pattern, and identifies abnormal behavior detection markers. A classification and analysis process based on the abnormal behavior detection markers is initiated to extract new violation behavior feature information. The new violation behavior feature information is fed back to the monitoring rule adjustment mechanism to generate a violation risk assessment result that includes risk level evaluation and early warning notification content.

[0020] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:

[0021] This invention discloses an internal threat identification and early warning method based on behavioral analysis. It extracts features such as deviation from professional domains, abnormal frequency, and abnormal timing by performing structured analysis on the query logs of medical staff, forming an initial set of behavioral features. Subsequently, cluster analysis is used to group query behaviors, identifying behavioral patterns outside of responsibilities and authority. Time series analysis is combined to detect changes in behavior frequency and extract dynamic behavioral change features. Further, deep behavioral analysis is used to analyze the probability of potential violations, ultimately leading to a conclusion on whether a violation has occurred. This invention can also dynamically adjust monitoring rules based on the characteristics of violations, enabling real-time monitoring and risk assessment of new query behaviors, effectively preventing unauthorized access and misuse of medical data, and protecting patient privacy and medical information security. Attached Figure Description

[0022] Figure 1 This is a flowchart of an internal threat identification and early warning method based on behavior analysis according to the present invention. Detailed Implementation

[0023] The technical solutions of the embodiments of the present invention will be clearly and thoroughly described below with reference to the accompanying drawings. The described embodiments are merely some embodiments of the present invention.

[0024] like Figure 1 This embodiment of an internal threat identification and early warning method based on behavior analysis may specifically include:

[0025] Step S101: Obtain the query behavior logs of medical staff from the medical information system, and structurally decompose the time, object and content of each query operation. Extract the original behavior dataset containing features such as querying disease types that deviate from the professional field, query frequency that exceeds the normal range, query time concentrated in non-working hours and query objects that have no related patient records, to form an initial behavior feature set.

[0026] Based on the login records and operation logs of medical staff in the medical information system, database query statements are used to obtain the timestamp, patient identifier, and disease code for each query operation. Query time period, operation frequency, and disease type information are extracted using field parsing methods to obtain structured query behavior data records. For the disease code in the query behavior data records, if the disease code does not match the professional field code in the medical staff's professional qualification table, the query record is marked as a professional deviation query. The frequency of occurrence of the same medical staff identifier in the query behavior data records is used to determine whether the operation frequency exceeds the normal working load threshold. Based on the professional deviation query identifier and the operation frequency records exceeding the threshold, a time interval division method is used to identify whether the query time falls within the non-working period. The matching relationship of patient identifiers is verified by querying the medical staff's responsible patient table to determine the patient association identifier. For the professional deviation query identifier, threshold exceeding identifier, non-working period identifier, and patient association identifier, a behavioral feature vector is constructed containing characteristics of disease type deviation, abnormal query frequency, abnormal query time, and no association of query objects, resulting in an initial behavioral feature set of abnormal query behaviors of medical staff.

[0027] Specifically, extracting query behavior data from medical information systems is a fundamental step in identifying abnormal operations.

[0028] In one possible implementation, the original record containing a timestamp field, a user identifier field, and an operation type field is obtained by accessing the user operation log table in the database.

[0029] Specifically, the timestamp field records the precise time the query occurred, such as 14:30:15 on July 22, 2025; the user identifier field stores the unique number of the medical staff; and the operation type field identifies the specific query action. The database query statement filters records within a specific time range using WHERE conditions, and the field parsing method breaks down composite fields into independent data items. This structured processing transforms previously scattered operational information into a standardized data format that facilitates subsequent analysis. Based on this structured data, professional matching verification becomes a crucial step in determining the compliance of the query.

[0030] In one embodiment, a professional qualification table for healthcare personnel stores the professional field code for each healthcare worker; cardiologists are assigned code 001, and neurosurgeons are assigned code 002. When a cardiologist queries a neurosurgical disease, the disease code does not match the professional field code, and the system automatically identifies the query as a professional deviation. Simultaneously, operation frequency statistics are calculated by counting the number of queries made by the same healthcare worker per unit of time. Under normal circumstances, healthcare workers should not query patient information more than 20 times per hour; exceeding this threshold is marked as an abnormal frequency. This dual verification mechanism effectively identifies cross-professional queries and excessive querying behavior. Based on professional deviation and abnormal frequency analysis, time-dimensional analysis further refines the accuracy of abnormal behavior identification.

[0031] It should be noted that the time interval division method divides a 24-hour day into two intervals: working hours and non-working hours. Working hours are typically from 8:00 to 18:00, with the remaining time considered non-working hours. When a query timestamp falls within non-working hours, such as 22:30 or 2:00 AM, the system identifies it as an abnormal time query. Furthermore, patient association verification is achieved by comparing the queried patient identifier with records in the patient responsibility table for medical staff. If a nurse queries information about patients outside their assigned ward, it is marked as an unrelated query. This dual constraint of time and association can identify potential unauthorized access behavior. Combining the aforementioned identification results, the behavioral feature vector construction integrates multi-dimensional abnormal information into a unified data structure.

[0032] Specifically, the feature vector contains four Boolean fields: professional deviation identifier, frequency anomaly identifier, time anomaly identifier, and association anomaly identifier.

[0033] For example, a query with a feature vector of [1, 0, 1, 0] indicates a deviation from professional scope and an anomaly in timing, but with normal frequency and correlation. By converting all query records into such feature vectors, a complete dataset of abnormal query behavior features is formed.

[0034] Step S102: Based on the initial set of behavioral features, targeting the characteristics of querying disease types that deviate from the professional field and query content that involves rare diseases, and combining the search object to lock in a specific age group and the search content to focus on records of specific medical history, a cluster analysis method is used to group the query behavior, distinguish the behavioral patterns outside of responsibilities and authority, and obtain a subset of classified behavioral features.

[0035] Based on the professional deviation identifier and disease coding fields in the medical staff behavior characteristic dataset, rare disease information involved in the query content is matched through the disease classification directory to obtain the age group characteristics of the search object, resulting in a comprehensive feature record containing professional deviation degree, rare disease identifier, age group code, and medical history content. For the medical history content and age group code in the comprehensive feature record, if the medical history content involves a specific disease type and the age group code is concentrated within a specific range, the query record is marked as a specific object focused query, and the similarity value between query records is obtained through feature matching degree calculation. Based on the specific object focused query marker and similarity value, the query records are grouped using the K-means clustering method. Clustering operations are used to obtain cluster group identifiers for different query behaviors, determining the behavior pattern group to which each query record belongs. A comparison and verification between the behavior pattern group and the scope of medical staff responsibilities and authority is performed. If the query behavior in a group exceeds the preset scope of medical staff responsibilities and authority, the group is marked as an out-of-responsibility behavior pattern, resulting in a categorized subset of behavior features distinguishing between inside and outside of responsibilities and authority.

[0036] Specifically, disease classification catalog matching is the core mechanism for identifying rare disease queries.

[0037] In one possible implementation, the disease classification directory is based on the International Classification of Diseases (ICD) and includes a mapping between common disease codes and rare disease codes. When healthcare professionals search for a disease, they match the queried disease code with the rare disease codes in the directory. If the code falls under the rare disease category, it is marked as a rare disease query.

[0038] For example, pulmonary hypertension, Gaucher disease, and Fabry disease fall under the rare disease coding range, while the common cold and hypertension are classified as common diseases. This classification and matching mechanism can accurately identify healthcare professionals' queries for specific disease types, providing foundational data support for subsequent anomaly detection. The age range extraction method further refines the characteristics of the query objects based on rare disease identification.

[0039] Specifically, age range extraction calculates the current age by parsing the birth date field in the patient's information and then groups them according to preset age ranges, such as 0-5 years old for infants, 6-18 years old for adolescents, and 19-65 years old for adults. When medical staff search for patients in a specific age group, the system can recognize this focused query pattern.

[0040] For example, a pediatrician frequently accessing adult patient information, or an adult specialist extensively accessing infant and toddler case studies, demonstrates the advantages of multi-dimensional anomaly detection when combined with rare disease identifiers, based on these cross-age-group queries and comprehensive feature records. The labeling process for object-specific focused queries, based on these comprehensive feature records, showcases the advantages of multi-dimensional anomaly detection.

[0041] In one embodiment, when a query record simultaneously meets the conditions of a rare disease identifier and a specific age group concentration, the system automatically marks it as a focused query. This dual-condition constraint filters out normal medical query behavior, focusing on identifying query patterns that may pose a privacy risk. Feature matching degree calculation quantifies the degree of correlation between query behaviors by comparing the similarity of different query records in dimensions such as disease type, age group, and query time. The higher the similarity degree value, the more likely the query behavior belongs to the same abnormal pattern, which provides a quantitative basis for subsequent clustering grouping. The application of the K-means clustering method transforms the scattered query records into ordered behavioral pattern groups.

[0042] It's important to note that K-means clustering groups similar query behaviors into the same cluster by calculating the Euclidean distance between query records. During the clustering process, the algorithm automatically determines the cluster centroids and assigns each query record to the nearest cluster based on the principle of minimum distance. This unsupervised learning method can automatically discover behavioral patterns hidden in large amounts of query data without requiring manually preset classification criteria. The generation of cluster group labels provides a clear classification for each query record, transforming the originally chaotic query data into a structured set of behavioral patterns. Verification of the scope of responsibilities and permissions is the final step in distinguishing between compliant and non-compliant query behaviors.

[0043] Specifically, the scope of authority for medical staff is predefined based on their position, department, and job responsibilities, including restrictions such as the types of patients that can be queried, the range of diseases, and the level of data access. When a query within a cluster exceeds the authority of the corresponding medical staff, that cluster is identified as an out-of-authority behavior pattern. This clear delineation of authority boundaries effectively identifies violations such as unauthorized queries, cross-department queries, and access to data beyond the permitted scope. The resulting subset of behavioral characteristics provides a precise risk identification and early warning mechanism for medical information security management.

[0044] Step S103: For the categorized subset of behavioral features, analyze the matching degree of unrelated patient records and retrieval scopes that exceed the scope of duties and permissions of the query object. Compare with the pre-established domain duty and permission database. If the query time is concentrated in non-working hours or the retrieval time shows a periodic pattern, it is determined that the behavior deviates from the expected range and a preliminary abnormal behavior marker is determined.

[0045] Based on the query object identifier and patient association field in the categorized behavioral feature subset, a database association query method is used to verify the matching relationship between the query object and the patient records under the responsibility of medical staff. The number of unrelated patient records and the degree of deviation of the search scope from the scope of responsibilities and permissions are obtained through association degree calculation, resulting in permission matching data containing the association verification results and the degree of permission deviation. This permission matching data is compared and verified with a pre-established domain responsibility and permission database. If the degree of permission deviation exceeds a preset threshold and the number of unrelated patient records exceeds the normal range, the query behavior is marked as a permission violation. The specific time information of the query behavior is obtained through time field parsing. Based on the permission violation mark and the specific time information, it is identified whether the query time is outside of working hours. The search time is calculated to detect whether it exhibits a periodic repetition pattern, determining the time violation mark and the periodic anomaly mark. For the permission violation mark, time violation mark, and periodic anomaly mark, if the query behavior simultaneously meets the conditions of permission violation and time anomaly, it is determined that the behavior deviates from the expected range, resulting in query data records with preliminary abnormal behavior marks.

[0046] Specifically, database association query methods are the core technical means to verify the compliance of query objects.

[0047] In one possible implementation, relational queries identify unrelated queries by comparing the patient identifier of the query object with records in the table of patients under the care of healthcare professionals.

[0048] Specifically, when a nurse queries patient information from a ward outside their assigned area, and the patient identifier has no matching record in their assigned patient table, the number of unrelated patient records increases. The correlation degree is calculated based on the ratio of the number of successfully matched records to the total number of queried records; a ratio below 0.7 indicates a large number of unrelated queries. This correlation verification mechanism can accurately identify unauthorized access to patient information by medical staff, providing a quantitative basis for subsequent permission deviation judgments. The calculation of the degree of permission deviation further quantifies the severity of the violation based on the correlation verification.

[0049] It should be noted that the search scope exceeding the scope of the responsibilities and permissions is achieved by comparing the data type of the query with the scope of permissions of medical staff. When the data type involved in the query exceeds the preset permissions, the deviation value increases accordingly.

[0050] For example, when ordinary nurses query physician-specific information such as surgical records and medication regimens, the degree of permission deviation increases significantly. The generation of permission matching data integrates the correlation verification results and deviation levels into a unified data structure, including query compliance scores and severity indicators of permission violations. This data integration provides a comprehensive assessment basis for subsequent identification of violations. The comparison and verification with the domain responsibility permission database demonstrates the advantages of rule-driven anomaly detection.

[0051] In one embodiment, a domain responsibility and permission library predefines data access permissions for different positions, including detailed specifications such as queried patient types, data fields, and operation permissions. The comparison and verification process determines whether the query complies with the position requirements by matching the query behavior with the rule entries in the permission library. When a cardiologist queries neurosurgery-specific brain imaging data, the comparison result shows a permission mismatch, triggering a permission violation flag. Time field parsing is performed simultaneously to extract the specific time information of the query, preparing data for anomaly detection in the time dimension. Based on the permission violation flag, a time range judgment method identifies the compliance of the query time.

[0052] Specifically, non-working hours are typically defined as 10:00 PM to 6:00 AM the following day, as well as weekends and public holidays. Queries occurring within these periods trigger a time violation flag. Time interval calculation identifies periodic patterns by analyzing the time differences between consecutive queries.

[0053] For example, a healthcare worker regularly checks specific patient information every Wednesday at 2:00 AM, with a 7-day cyclical pattern. This unusual pattern suggests potentially premeditated privacy theft. The generation of periodic anomaly markers provides crucial clues for identifying organized violations. A comprehensive judgment mechanism integrates these multi-dimensional anomaly markers into the final behavioral assessment result.

[0054] In one embodiment, the credibility of abnormal behavior is significantly improved when the query behavior has both permission violation and time abnormality characteristics.

[0055] For example, if a nurse queries sensitive medical information of a patient who is not under her care during the early morning hours, triggering a triple alert of permission violation, time violation, and periodic anomaly, such behavior is judged as seriously deviating from the expected scope.

[0056] Step S104: By initially marking abnormal behavior, time series analysis is used to detect changes in behavior frequency, targeting the characteristics of query frequency exceeding the normal range and retrieval frequency surging in a short period of time. By comparing the difference between the standard operation frequency of medical staff querying patient medical records in their daily routine and the time interval of the current retrieval behavior, the specific manifestations of query behavior lacking business context and retrieval results without subsequent operation records are analyzed to obtain a description of dynamic behavior changes.

[0057] Based on query data records with preliminary abnormal behavior markers, historical query records and current work environment information of medical staff are obtained. Historical query frequency baselines and current query frequency values ​​are extracted through frequency statistics calculations, resulting in frequency comparison data containing historical frequency baselines, current query frequencies, and shift schedule information. Regarding the query frequency changes in the frequency comparison data, if the current query frequency exceeds a preset threshold of the historical frequency baseline and shows a surge within a short period, it is marked as an abnormal query. The time interval difference value of the query behavior is obtained to determine the severity of the frequency anomaly. Based on the severity of the frequency anomaly and the standard operating procedure for medical staff to query patient medical records daily, the business rationality of the query behavior is verified. Subsequent operation records are checked to determine whether the search results generate corresponding medical activity records, obtaining business relevance verification results. Based on the severity of the frequency anomaly and the business relevance verification results, if the query time deviates from the normal shift time range of medical staff and the search content does not match the type of patients currently treated in the department, an abnormal feature of lacking corresponding medical records after the query is determined, resulting in a dynamic behavioral change description including time deviation, department mismatch, and missing medical records.

[0058] Specifically, database join query methods are a key technical component in constructing baselines of healthcare worker behavior.

[0059] In one possible implementation, the join query associates the current query record table with the historical query record table through a JOIN operation, establishing a data relationship based on the unique identifier of the medical staff. Frequency statistics are calculated by statistically analyzing the average number of queries performed by the same medical staff over the past 30 days to establish a personalized historical frequency baseline.

[0060] For example, a nurse who queried patient information an average of 15 times per shift over a 30-day period saw a 300% increase in frequency when she queried 45 times during her current shift. Extracting work environment information includes key data such as shift schedules, departmental assignments, and patient lists for the shift. This information provides the necessary business context and judgment basis for subsequent abnormal behavior identification. Anomaly detection based on frequency comparison data demonstrates the technical advantages of dynamic threshold monitoring.

[0061] Specifically, the preset threshold is typically set at 200% to 300% of the historical frequency baseline. An anomaly alert is triggered when the current query frequency exceeds this range. Short-term surges are identified by analyzing changes in query density within a continuous time window; for example, a sudden increase in queries from 5 to 25 within one hour. This surge pattern clearly deviates from the normal working rhythm. The time interval calculation method identifies abnormal query behavior rhythms by measuring the time difference between two consecutive queries. Normal medical query intervals are typically 10-30 minutes, while abnormal queries may exhibit dense intervals of 2-5 minutes. This time compression pattern suggests that the query behavior may have non-medical purposes. Based on the identification of frequency anomalies, the operation sequence comparison method further verifies the medical rationale for the query behavior.

[0062] It should be noted that the standard operating procedure (SOP) defines the typical process of normal medical inquiries, including a series of steps such as reviewing basic patient information, examination and test results, reviewing medical history records, and developing a treatment plan. Operation sequence comparison identifies abnormal patterns that deviate from normal medical logic by analyzing whether the inquiry behavior follows this standard process.

[0063] For example, a healthcare worker might only view patient personal information and contact details, skipping over medical-related data. This selective query pattern exposes a potential intent to steal privacy. Subsequent operation record checks verify whether the query generates medical documents such as nursing records, medication records, and treatment plans to determine the medical necessity of the query. Further analysis of business relevance verification reveals the degree of alignment between the query behavior and medical practice.

[0064] In one embodiment, shift time range verification identifies unusual access outside of working hours by comparing query times with the medical staff's shift schedule. When a nurse makes numerous queries during leave or off-duty hours, this time misalignment suggests suspicious query motives. Departmental patient type matching identifies unusual cross-specialty queries by comparing the queried patient's disease type with the current department's specialty scope.

[0065] For example, cardiology nurses frequently querying pediatric patient information; this departmental mismatch violates the basic principle of medical specialization. The missing medical record check verifies the medical consequences of the query by tracing the medical activity records following the query. When the query does not generate any medical documents or treatment activities, it indicates that the query purpose deviates from normal medical needs. The resulting dynamic behavioral changes provide multi-dimensional risk identification and precise abnormal behavior profiling for medical information security management.

[0066] Step S105: Generate abnormal features based on the description of dynamic behavior changes. For the frequency of detected query records without subsequent diagnosis and treatment actions or retrieval behaviors involving sensitive data fields, trigger deep behavior analysis, extract context information of abnormal changes in the location of unauthorized terminals of query source devices and retrieval terminals, analyze the possibility of potential violations, and obtain deep behavior analysis results.

[0067] Based on the behavioral anomaly descriptions including time deviation, department mismatch, and missing medical records, the types of sensitive information involved in the query records are identified. Medical operation record queries are used to verify whether corresponding medical activity records were generated after the query. If the query record lacks subsequent medical actions and the access frequency of sensitive information exceeds a preset threshold, it is marked as a high-risk query behavior, resulting in a high-risk query behavior identifier. For the high-risk query behavior identifier and the device identifier information of the query source, the device from which the query originated is verified against the device authorization list to determine whether it belongs to the pre-authorized terminal device range of the medical institution. Device compliance checks are used to obtain unauthorized terminal access records, determining the compliance status identifier of the device access. Based on the compliance status identifier of the device access and the network location information of the query terminal, a location comparison verification method is used to detect whether the location of the query terminal deviates from the medical staff's regular work location. Location change records are analyzed to determine the degree of anomaly in the terminal's geographical location, obtaining location anomaly assessment data. For high-risk query behavior identifiers, device access compliance status identifiers, and location anomaly assessment data, if the query behavior simultaneously exhibits characteristics of sensitive information access, unauthorized device access, and location anomalies, the severity level of potential violations is determined through comprehensive risk assessment, resulting in in-depth behavior analysis results that include severity level and violation characteristic descriptions.

[0068] Specifically, data field classification standards are the core technological foundation for identifying access to sensitive information.

[0069] In one possible implementation, data field classification divides medical information into three sensitivity levels: basic information, clinical data, and personal privacy. The personal privacy level includes highly sensitive fields such as patient ID numbers, home addresses, and contact information. When healthcare professionals query these highly sensitive fields, the access is automatically flagged as sensitive information access. Medical procedure record queries verify the medical necessity of the query by checking for medical documents such as nursing records, medication records, and examination requests generated within 30 minutes of the query time.

[0070] For example, if a nurse queries a patient's detailed personal information but fails to generate any medical records within a specified time, such a query lacks medical business support and is identified as suspicious behavior. This dual verification mechanism can effectively distinguish between normal medical queries and potential privacy theft. Based on the results of sensitive information access identification, device authorization management further strengthens the technical barriers to access control.

[0071] Specifically, the authorized device list contains key information such as hardware identifiers, network addresses, and device types for all compliant terminal devices within the medical institution. Device compliance checks identify access behavior from unauthorized devices by comparing the identifier information of the source device with the authorized list records. When the query originates from a non-medical terminal such as a personal mobile phone or home computer, an unauthorized device alert is immediately triggered. This device-level access control fills the security gaps of traditional account verification, ensuring that even if medical staff use legitimate accounts, access via unauthorized devices can still be detected and blocked in a timely manner. Building upon device compliance verification, a location comparison verification method enhances the accuracy of abnormal behavior detection from a geographical perspective.

[0072] It should be noted that healthcare workers' regular work locations are typically confined to specific medical buildings. Location comparison analyzes the GPS coordinates or network location information of the query terminal to determine whether the accessed location deviates from the preset work area. When queries originate from locations outside the hospital, especially from residential areas, entertainment venues, or other locations unrelated to medical work, the degree of location anomaly increases significantly. Location change record analysis identifies abnormal geographical location jump patterns by tracking the changes in the query location of the same healthcare worker at different points in time.

[0073] For example, if a medical professional makes inquiries from different cities within a short period, this pattern of location changes clearly violates normal work patterns, indicating that the account may have been compromised or that other security risks exist. Comprehensive risk assessment integrates multi-dimensional anomaly information into a unified risk judgment standard.

[0074] In one embodiment, the risk assessment employs a weighted scoring mechanism, with access to sensitive information accounting for 40% of the weight, unauthorized device access for 35%, and abnormal location for 25%. When a query triggers anomaly indicators in all three dimensions, the overall risk score reaches the highest level and is deemed a serious violation. Severity levels are categorized into low, medium, and high risk, each corresponding to different response measures. High-risk behaviors trigger immediate alerts and temporary account locking, medium-risk behaviors are added to a key monitoring list, and low-risk behaviors are recorded for future analysis. The generation of in-depth behavioral analysis results provides a precise risk identification and tiered response mechanism for medical information security management, enabling a shift from passive monitoring to proactive protection and improving the technical level and management efficiency of medical data privacy protection.

[0075] Step S106: Based on the deep behavior analysis results, analyze the similarity between query content involving rare diseases and search content focusing on specific medical histories. If the search scope exceeds the scope of responsibilities and authority and the query behavior lacks business context, determine the classification label of the violation and obtain the violation judgment conclusion.

[0076] Based on the deep behavioral analysis results, which include severity levels and descriptions of violation characteristics, rare disease types involved in the query content are identified through disease coding classification queries. Content similarity comparison methods are used to analyze the concentration of search content in specific medical histories, resulting in query content analysis data containing rare disease types and medical history concentration indicators. Regarding the query content analysis data and the scope of medical staff's responsibilities and authority, if the search scope exceeds the boundaries of responsibilities and authority and both rare disease types and medical history concentration indicators deviate from the normal range, it is marked as an over-authority query. Medical business relevance verification is used to determine whether the query behavior has a corresponding business background. Based on the over-authority query behavior marking and medical business relevance verification results, if the query behavior lacks a corresponding medical business background and the over-authority characteristics are obvious, the specific violation category is determined through violation type identification rules, resulting in query behavior identifier data labeled with violation categories. For query behavior identifier data labeled with violation categories, if the query behavior simultaneously meets the judgment conditions of exceeding authority, sensitive content, and lack of business background, the violation identification level is determined, resulting in a violation judgment conclusion including the identification level and violation category.

[0077] Specifically, disease code classification query is the core technical step in identifying abnormalities in query content.

[0078] In one possible implementation, disease coding classification establishes a hierarchical directory based on the International Classification of Diseases (ICD-1), categorizing diseases into three main types: common diseases, chronic diseases, and rare diseases. The rare disease category includes diseases with an incidence rate of less than one in ten thousand, such as Huntington's disease, Marfan syndrome, and Wilson's disease. When healthcare professionals query these rare diseases, the disease code is automatically matched to the rare disease category, triggering an anomaly content flag. The content similarity comparison method analyzes medical history keywords in query records to calculate the degree of concentration of query content across specific disease types, symptom descriptions, and treatment plans.

[0079] For example, if a healthcare worker repeatedly searches for multiple patients with the same rare genetic disease, the concentration index of medical history increases significantly, indicating that the search behavior has a clear target orientation. Based on the results of anomaly identification of search content, the verification of responsibility and authority boundaries further determines the compliance of the search behavior.

[0080] Specifically, the scope of duties and authority for medical staff is predefined based on their job responsibilities, departmental division of labor, and professional qualifications, including specific restrictions such as the types of diseases that can be queried, the range of patients, and data fields. When a regular nurse queries complex surgical records or rare disease gene testing results, the search scope clearly exceeds the boundaries of their duties and authority. The quantitative judgment of authority boundaries is achieved by calculating the matching degree between the query content and the scope of authority; an over-authority alert is triggered when the matching degree is below 70%. Medical business relevance verification verifies the medical necessity and business rationality of the query behavior by checking whether the query behavior is directly related to the medical staff's current work tasks, patient care plans, and departmental admission status. Based on authority verification, violation type identification rules further refine abnormal query behavior into specific violation categories.

[0081] It should be noted that the violation type identification rules predefine characteristic patterns for different violations, including categories such as privacy theft, data breach, unauthorized access, and malicious querying. A typical characteristic of privacy theft violations is the concentrated querying of sensitive personal information of specific patients without a medical business context, while data breach violations manifest as the downloading or exporting of large amounts of medical data without any corresponding work requirement. Automatic violation type identification is achieved by matching query behavior characteristics with preset rule patterns. When a query behavior simultaneously meets multiple violation characteristics, it is categorized into the corresponding violation category. This refined violation classification provides a clear basis and guidance for subsequent handling measures. The comprehensive judgment mechanism integrates multi-dimensional violation information into a violation determination conclusion.

[0082] In one embodiment, the determination of a violation is based on a comprehensive assessment of three core elements: the degree of exceeding authority, the level of content sensitivity, and the degree of lack of business context. The degree of exceeding authority is quantified by calculating the deviation between the query scope and the user's authorized duties; the level of content sensitivity is determined based on the data type and privacy level of the query; and the degree of lack of business context is assessed by verifying the relevance of the query behavior to medical work. When the scores of all three elements exceed preset thresholds, the query behavior is considered a definitive violation. The determination is categorized into three levels: suspected violation, minor violation, and serious violation, with different handling procedures and disciplinary measures corresponding to each level.

[0083] Step S107: Based on the violation type identified in the violation judgment conclusion and the abnormal terminal location and unauthorized device access characteristics found in the deep behavior analysis results, the threshold of the behavior monitoring rules in the system is dynamically adjusted in combination with the severity and frequency distribution of the violation. A special monitoring strategy is generated for query times concentrated in non-working hours and retrieval times showing periodic patterns, resulting in an updated monitoring plan.

[0084] Based on the violation judgment conclusions including the identification level and violation category, violation type identifiers, abnormal terminal location information, and unauthorized device access records are obtained. The severity distribution and frequency patterns of the violations are analyzed to obtain comprehensive violation data containing the severity and frequency distribution characteristics of the violations. For the severity and frequency distribution characteristics in the comprehensive violation data, if the severity and frequency of the violations exceed the detection capability of the current monitoring threshold, the trigger threshold parameters of the existing behavior monitoring rules are adjusted through threshold calculation. Simultaneously, the time periods and device types where violations are concentrated are identified to determine the key monitoring time range and device range. Based on the key monitoring time range and device range, specific monitoring rules for non-working-hour query and periodic retrieval behaviors are constructed. Targeted monitoring conditions are generated through time pattern matching and device type filtering, resulting in a set of specific monitoring rules including time monitoring conditions and device monitoring conditions. For the adjusted trigger threshold parameters and specific monitoring rule set, the new threshold parameters and specific rules are integrated into the existing monitoring configuration through a configuration merging operation. The execution priority and triggering mechanism of the monitoring rules are updated, resulting in an updated monitoring scheme integrating dynamic threshold adjustment and specific monitoring rules.

[0085] Specifically, data extraction is a key technical step in building the foundation for violation analysis.

[0086] In one possible implementation, data extraction automatically identifies key information such as violation type identifiers, abnormal terminal location coordinates, and unauthorized device MAC addresses by parsing structured fields in the violation judgment conclusion. Violation type identifiers include predefined categories such as privacy theft, unauthorized access, and data leakage. Abnormal terminal location information records the geographic coordinates and network location where the query occurred. Unauthorized device access records store device fingerprints and access timestamps. Statistical calculation methods analyze the distribution characteristics of this extracted data to calculate quantitative indicators such as the frequency of occurrence, severity distribution range, and temporal concentration of different violation types.

[0087] For example, when a certain type of violation occurs more than 50 times within 30 days and is concentrated outside of working hours, the frequency distribution shows a clear abnormal clustering pattern, providing a quantitative basis for subsequent threshold adjustments. Based on the results of statistical analysis of violation data, the threshold calculation and adjustment reflect a dynamic response monitoring and optimization mechanism.

[0088] Specifically, the current monitoring thresholds are set based on historical normal behavior patterns, but as violation patterns change, fixed thresholds may become ineffective or overly sensitive. Threshold calculations redetermine the critical values ​​for triggering alarms by analyzing the severity distribution of violations.

[0089] For example, the original query frequency threshold was set at 20 times per hour, but newly discovered violation patterns showed that abnormal query frequencies were concentrated in the range of 35-50 times per hour. Therefore, the threshold was adjusted to 30 times per hour to improve detection accuracy. The identification of time periods and device types was achieved by statistically analyzing the time distribution and device source distribution of violations, identifying the 22:00-06:00 time period and mobile device types as high-risk factors. This identification provides a clear target scope for specialized monitoring. Based on threshold optimization, the rule generator further refined the technical implementation of targeted monitoring strategies.

[0090] It should be noted that the specialized monitoring rules are designed for specific abnormal behavior patterns and complement the general monitoring rules. Time pattern matching identifies periodic abnormal behaviors by defining time windows and frequency patterns.

[0091] For example, a healthcare worker performs a large number of queries every Wednesday from 2:00 AM to 3:00 AM. Time pattern matching rules automatically capture this periodic anomaly. Device type filtering distinguishes between authorized medical terminals and suspicious devices by establishing whitelists and blacklists. When a query originates from an unregistered personal mobile device, the device type filtering rules immediately trigger an alert. The generation of specialized monitoring rule sets combines time and device conditions into composite monitoring logic, forming a precise detection mechanism for specific violation patterns. Configuration merging integrates scattered monitoring rules into a unified monitoring system architecture.

[0092] In one embodiment, configuration merging involves complex processing such as rule prioritization, conflict detection, and execution order optimization. New threshold parameters need to be compatible with existing monitoring configurations to avoid false positives or false negatives.

[0093] For example, the adjusted query frequency threshold needs to be coordinated with device type filtering rules to ensure that normal high-frequency queries to authorized devices do not trigger false alarms. Execution priority is set based on the severity and urgency of the violation, with monitoring rules for serious violations having higher execution priority. Updates to the triggering mechanism include key elements such as alarm level settings, response time requirements, and subsequent processing procedures. The generation of updated monitoring schemes represents a shift from static rules to dynamic adaptation, significantly improving the response speed and detection accuracy of medical information security monitoring, and providing medical institutions with more intelligent and precise security protection capabilities.

[0094] Step S108: By updating the monitoring scheme, the newly collected query behavior logs are analyzed. If the query frequency exceeds the normal range or the search frequency surges in a short period of time, the behavior pattern deviates from the expected range, the classification and analysis process is triggered. At the same time, the newly discovered violation behavior characteristics are fed back to the monitoring rule adjustment mechanism to obtain the latest violation risk assessment results and issue an early warning.

[0095] Based on the updated monitoring scheme integrating dynamic threshold adjustment and specific monitoring rules, newly collected query behavior logs are parsed. Through data field separation and structured transformation, query time, operation frequency, user identifier, and device information are obtained, resulting in standardized behavioral data containing query frequency indicators and operation pattern characteristics. For the query frequency indicators and operation pattern characteristics in the standardized behavioral data, if the query frequency exceeds the normal range threshold set in the updated monitoring scheme or the retrieval frequency shows a surge in a short period, deviation metric calculation is used to determine the degree to which the behavioral pattern deviates from the expected range, identifying an abnormal behavior detection flag. Based on the abnormal behavior detection flag, the behavior classification and deep analysis processing flow is initiated. Simultaneously, newly identified violation behavior feature information is obtained through abnormal feature extraction methods. This violation behavior feature information is fed back to the monitoring rule adjustment mechanism for parameter optimization, resulting in rule update data containing the new features. Based on the processing results of behavior classification and deep analysis, as well as the rule update data, the severity and potential impact of the violation are comprehensively evaluated and analyzed. Based on the evaluation results, a corresponding level of security warning notification is generated, resulting in the latest violation risk assessment result containing risk level assessment and warning notification content.

[0096] Specifically, real-time data processing methods are the core technological foundation for building a closed-loop monitoring system.

[0097] In one possible implementation, real-time processing uses streaming data processing technology to instantly parse continuously generated query behavior logs, avoiding the time delay issues of traditional batch processing. Data field separation uses predefined field mapping rules to extract information such as timestamps, user IDs, query content, and device identifiers from the original logs into independent data fields. Structured transformation converts unstructured log text into a standardized data record format, including quantitative indicators such as query frequency, time intervals, and operation types.

[0098] For example, the original log "2025-07-22 14:30:15 User 001 queried patient Zhang's medical records" was converted into a structured record containing standardized data items such as time, user, operation type, and target object fields. This real-time processing capability ensures timely detection and rapid response to abnormal behavior. Anomaly detection based on standardized data demonstrates the technical advantages of dynamic threshold monitoring.

[0099] Specifically, the query frequency metric is calculated by counting the number of queries per unit of time. Operational pattern characteristics are identified by analyzing the time distribution of query behavior, target selection, and content type. The normal range threshold is established based on historical normal behavior data; comparing the current query frequency with the historical baseline provides a quantitative basis for anomaly detection. Surge characteristics are identified by comparing frequency changes within a continuous time window; an anomaly alert is triggered when the query frequency increases significantly within a short period. Deviation metric calculation quantifies the severity of abnormal behavior by measuring the degree of difference between the current behavioral pattern and the expected pattern.

[0100] For example, a medical staff member historically queried patient information an average of 12 times per hour, but this hour they queried it 38 times, a frequency deviation of 216%, far exceeding the normal range and thus identified as abnormal behavior. Based on anomaly detection, the automatic scheduling mechanism achieves intelligent and systematic monitoring and response.

[0101] It's important to note that automated scheduling avoids the delays and inconsistencies of manual intervention and subjective judgment. Through pre-set triggering conditions and processing procedures, it ensures timely and standardized handling of abnormal behavior. The behavior classification and deep analysis process comprises multiple consecutive analysis stages, from basic behavioral feature extraction to complex violation pattern recognition, forming a complete abnormal behavior analysis chain. The abnormal feature extraction method identifies new violation features and behavioral patterns by comparing newly discovered abnormal behaviors with known violation patterns. The feedback mechanism for these new feature information enables monitoring rules to continuously learn and adapt to new threat patterns. Parameter optimization is achieved by adjusting detection thresholds, expanding the monitoring scope, and refining rule conditions. Comprehensive evaluation calculations integrate multi-dimensional information into a unified risk assessment standard.

[0102] In one embodiment, a comprehensive assessment considers multiple factors, including the type, frequency, scope of impact, and potential consequences of violations, and calculates the overall risk level using a weighted scoring mechanism. Severity assessment quantifies the potential impact of violations on patient privacy, medical data security, and institutional reputation. Scope of impact analysis is determined by assessing dimensions such as the number of patients involved, data types, and time span. Security alerts are generated using a tiered response mechanism based on risk level: high-risk behaviors trigger immediate alerts and emergency procedures, medium-risk behaviors enter key monitoring and investigation processes, and low-risk behaviors are recorded for trend analysis. The generation of the latest violation risk assessment results represents a shift from passive monitoring to proactive protection. Through continuous learning and optimization mechanisms, the intelligence level and response efficiency of medical information security protection are significantly improved, establishing a complete closed-loop security monitoring system for medical institutions.

[0103] Specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. Unless otherwise specified, the embodiments and features described in this application can be arbitrarily combined with each other.

Claims

1. A method for identifying and issuing early warning of internal threats based on behavioral analysis, characterized in that, The method includes: The system retrieves query behavior logs from medical information systems, structurally decomposes query operations by time, object, and content, forming an initial behavioral feature set. Based on this initial behavioral feature set, and combined with records of specific age groups and medical histories identified by the search objects, cluster analysis is used to group query behaviors, distinguishing behavioral patterns outside of responsibilities and permissions, resulting in a categorized subset of behavioral features. Using this categorized subset, the matching degree between query objects with unrelated patient records and search scopes exceeding responsibilities and permissions is analyzed. Combined with comparison to a domain responsibility and permission database, preliminary abnormal behavior markers are identified for behaviors where query times are concentrated during non-working hours or where search times exhibit periodic patterns. Based on these preliminary abnormal behavior markers, time series analysis is used to detect changes in query frequency and analyze query behavior deficiencies. The lack of business context and subsequent operation records are used to obtain a dynamic behavior change description. Based on this description, for queries with no subsequent medical actions or involving sensitive data fields, contextual information about unauthorized terminals and abnormal terminal location changes from the query source device is extracted to obtain deep behavior analysis results. According to these results, the similarity of query content involving rare diseases and the concentration of medical history is analyzed to determine the classification tags for violations. Using these classification tags, combined with abnormal terminal locations and unauthorized device access characteristics, an updated real-time monitoring scheme is generated. This updated scheme is used to analyze newly collected query behavior logs to detect whether behavior patterns deviate from the expected range, thus obtaining a violation risk assessment result.

2. The method for internal threat identification and early warning based on behavioral analysis according to claim 1, characterized in that, The process involves obtaining query behavior logs from medical staff within the medical information system, structurally breaking down the time, object, and content of query operations to form an initial set of behavioral characteristics, including: The system retrieves timestamps, patient identifiers, and disease codes from login records and operation logs of medical staff using database queries. It then extracts query time periods, operation frequencies, and disease type information using field parsing methods, resulting in structured query behavior data records. For the disease codes in these records, the system verifies the matching relationship between disease codes and professional field codes by combining them with a medical staff professional qualification table, marking query records that deviate from their professional fields. By statistically analyzing the frequency of identical medical staff identifiers in the query behavior data records, the system determines whether the operation frequency exceeds the workload threshold. Based on the timestamps in the query behavior data records, the system uses a time interval division method to identify whether the query time falls within non-working hours. Combined with a patient table, the system verifies the correlation of patient identifiers, constructing an initial set of behavioral features.

3. The method for internal threat identification and early warning based on behavioral analysis according to claim 1, characterized in that, The process involves using an initial set of behavioral characteristics, combined with records of specific age groups and medical histories identified by the search object, to group query behaviors using cluster analysis, distinguishing behavioral patterns outside of authorized duties, and obtaining a categorized subset of behavioral characteristics, including: By matching disease codes in the initial behavioral feature set with the disease classification directory, rare disease information involved in the query content is identified, and the age group characteristics of the search object are obtained, resulting in a comprehensive feature record containing professional deviation, rare disease identification, age group code, and medical history content. For the medical history content and age group code in the comprehensive feature record, it is verified whether the medical history content involves a specific disease type and whether the age group code is concentrated in a specific range, marking specific objects for focused queries. The comprehensive feature record is grouped using the K-means clustering method to obtain cluster group identifiers for query behavior. Combined with the scope of medical staff's responsibilities and authority, it is verified whether the cluster group identifiers exceed preset permissions, marking behavioral patterns outside of responsibilities and authority, and generating a categorized subset of behavioral features.

4. The method for internal threat identification and early warning based on behavioral analysis according to claim 1, characterized in that, The process involves analyzing the matching degree of unrelated patient records and searches exceeding the scope of responsibilities and permissions using a subset of categorized behavioral features. This is combined with comparison against the domain responsibility and permission database. Preliminary abnormal behavior markers are identified for behaviors where queries are concentrated during non-working hours or exhibit periodic patterns during search times. These markers include: The matching relationship between the query object identifiers in the categorized behavioral feature subset and the patient records managed by medical staff is verified through database association queries. The number of unrelated patient records and the degree of permission deviation are calculated. For the number of unrelated patient records and the degree of permission deviation, a comparison is made with the domain responsibility and permission database to mark permission violations. The query time information in the categorized behavioral feature subset is analyzed to identify non-working hours, calculate and detect periodic repetitive patterns, and determine time violation identifiers. Based on the permission violations and time violation identifiers, query behaviors that deviate from the expected range are marked.

5. The method for internal threat identification and early warning based on behavioral analysis according to claim 1, characterized in that, Based on the initial abnormal behavior markers, time series analysis is used to detect changes in query frequency. This analysis identifies the characteristics of query behavior lacking business context and subsequent operation records, resulting in a description of dynamic behavior changes, including: The system acquires historical query records and current environment information from medical staff, extracts the historical query frequency baseline and the current query frequency value; for the current query frequency value, it verifies whether it exceeds the historical frequency baseline threshold and exhibits a short-term surge characteristic, marking the frequency abnormal query; it verifies the business rationality of the frequency abnormal query by comparing the operation sequence; and it generates a dynamic behavior change description based on the frequency abnormal query and business rationality verification results.

6. The method for internal threat identification and early warning based on behavioral analysis according to claim 1, characterized in that, The description of dynamic behavior changes involves extracting contextual information about unauthorized terminals and abnormal location changes of the query source device for queries that lack subsequent diagnostic or treatment actions or involve sensitive data fields. This yields deep behavior analysis results, including: The system identifies sensitive information types in the description of dynamic behavior changes, verifies whether query records generate medical activity records, and marks high-risk query behaviors that lack subsequent medical actions and have a sensitive information access frequency exceeding a threshold. It verifies the source device of the high-risk query behavior by comparing it with the device authorization list, and identifies unauthorized terminal access records. It verifies the terminal location of the detected high-risk query behavior by comparing its location, and analyzes the degree of location change anomalies. Based on the high-risk query behavior, unauthorized terminal access records, and the degree of location change anomalies, it determines the severity level of potential violations and generates in-depth behavior analysis results.

7. The method for internal threat identification and early warning based on behavioral analysis according to claim 1, characterized in that, Based on the deep behavioral analysis results, the similarity of query content involving rare diseases and the concentration of medical history is analyzed to determine the classification tags of the violations, including: The system identifies rare disease types in the deep behavior analysis results, analyzes the concentration of medical history using content similarity comparison, and generates query content analysis data. Based on this data, and considering the scope of responsibilities and permissions, it verifies whether the search scope exceeds the permission boundaries and whether the rare disease types and the concentration of medical history deviate from the normal range, thus marking unauthorized query behavior. Through medical business relevance verification, it determines the business background of the unauthorized query behavior, identifies the category of violation, and generates classification tags for the violation.

8. The method for internal threat identification and early warning based on behavior analysis according to claim 1, characterized in that, The updated real-time monitoring scheme is generated by combining the classification tags of violations with the characteristics of abnormal terminal locations and unauthorized device access, including: Obtain the violation type identifier, abnormal terminal location, and unauthorized device access records from the classification tags of the violations, and analyze the severity and frequency distribution of the violations; adjust the trigger threshold parameters of the behavior monitoring rules based on the severity and frequency distribution, and identify the time periods and device types where violations are concentrated; construct special monitoring rules for querying and periodic retrieval behaviors during non-working hours, and generate targeted monitoring conditions; integrate the trigger threshold parameters and special monitoring rules into the monitoring configuration, update the monitoring rule priority, and generate an updated real-time monitoring scheme.

9. The method for internal threat identification and early warning based on behavioral analysis according to claim 1, characterized in that, The updated real-time monitoring scheme analyzes newly collected query behavior logs to detect whether behavior patterns deviate from the expected range, and obtains a risk assessment result for violations, including: The newly collected query behavior logs in the updated real-time monitoring scheme are analyzed to obtain query time, operation frequency, user ID and device information; for the operation frequency, it is verified whether it exceeds the normal range threshold or has short-term surge characteristics, the degree of deviation of the behavior pattern is judged, and the abnormal behavior detection mark is determined. Initiate the classification and parsing process based on the abnormal behavior detection identifier to extract new violation behavior feature information; The new violation characteristics are fed back to the monitoring rule adjustment mechanism to generate a violation risk assessment result that includes risk level evaluation and early warning notification.