Intelligent supervision method and system for medical institutions based on mobile terminal and video cloud supervision
By introducing mobile terminals and video cloud supervision into the medical institution supervision scenario, and using topological state calculation and finite state machine for trajectory extrapolation, the problems of monitoring blind spots and data spatiotemporal asynchrony are solved, realizing the continuity of cross-regional supervision and the automated generation of evidence chains, thereby improving supervision efficiency and the reliability of evidence.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 杨亚龙
- Filing Date
- 2026-03-18
- Publication Date
- 2026-07-31
AI Technical Summary
In existing medical institution supervision scenarios, single video surveillance has physical space blind spots, which makes target tracking prone to interruption. Data collected by multi-source heterogeneous devices is not synchronized in time and space and it is difficult to form a complete chain of evidence. The determination of violations in complex cross-regional business logic lacks multimodal data verification and automated scheduling response mechanisms.
A smart supervision method based on mobile terminals and video cloud supervision is adopted. The target trajectory is deduced in a directed acyclic graph and finite state machine by using the topology state calculation module on the cloud server. Combined with the topology anchor tags and optical sensors of the mobile law enforcement terminal, human-machine collaboration is achieved to generate joint evidence files, realizing the automation of cross-regional supervision and the integrity of the evidence chain.
It enables continuous cross-regional regulatory tracking, automatically generates a complete chain of evidence, reduces evidence collection costs and improves regulatory efficiency, and provides real-time risk warnings and visual management.
Smart Images

Figure CN122494142A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of smart healthcare and digital supervision technology, specifically to a smart supervision method and system for medical institutions based on mobile terminals and video cloud supervision. Background Technology
[0002] As medical institutions expand their operations, regulatory requirements for core aspects such as infection control, medical waste disposal, and personnel qualifications are becoming increasingly stringent. Traditional medical supervision relies heavily on manual on-site inspections and post-event manual review of records. This approach not only consumes significant human resources but also struggles to achieve continuous, real-time, and cross-regional monitoring of medical procedures. In recent years, some medical institutions have begun to introduce video surveillance and basic intelligent analytics technologies, attempting to capture violations through edge computing nodes. However, in complex real-world business scenarios, existing regulatory technologies still exhibit significant limitations.
[0003] Existing video intelligent analysis technologies are insufficient for continuous target tracking in complex physical spaces. Medical institutions have complex building structures, resulting in numerous video surveillance blind spots, such as enclosed medical waste storage rooms, stairwell corners without surveillance coverage, or private passageways. When a monitored target crosses these physical blind spots, single visual feature extraction algorithms are prone to failure due to the target's temporary disappearance from the field of view, drastic changes in ambient lighting, or physical occlusion, leading to visual tracking interruptions. Traditional trajectory deduction logic typically cannot handle this target re-identification problem across blind spots, causing what should be a continuous workflow to be fragmented into isolated video segments within the monitoring system. The system cannot automatically determine whether these separated segments belong to the same monitored event.
[0004] Furthermore, due to the lack of a deep fusion mechanism for multi-source sensing data, existing systems cannot spatiotemporally stitch together the sensing data from online video gateways with the on-site verification trajectories of offline law enforcement personnel when tracking gaps occur. Clock drift is common among multi-source heterogeneous devices, and existing time synchronization methods are rather crude, often leading to temporal logic conflicts during cross-domain data fusion. In the evidence consolidation stage, discrete video clips captured by existing systems, business metadata, and on-site verification records are often stored independently, lacking automated joint splicing and tamper-proof consolidation methods. This results in administrative law enforcement personnel still needing to spend a significant amount of time manually editing, comparing, and filling out documents to generate the final case file, making it difficult to form a complete closed-loop evidence chain with rigorous legal effect, severely restricting the execution efficiency of smart supervision. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a smart supervision method and system for medical institutions based on mobile terminals and video cloud supervision. The technical problems it solves are that in existing medical institution supervision scenarios, single video surveillance has physical space blind spots that make target tracking prone to interruption, data collected by multi-source heterogeneous devices are not synchronized in time and space and it is difficult to form a complete chain of evidence, and there is a lack of multimodal data verification and automated scheduling response mechanism for violation judgment of complex cross-regional business logic.
[0006] To address the above problems, the present invention provides the following technical solution: The first aspect of this invention provides a smart supervision method for medical institutions based on mobile devices and video cloud monitoring, comprising the following steps: The cloud server receives daily supervision tasks dispatched by the management backend or alarm information generated by the video cloud supervision platform through analysis of the original video stream, and generates verification tasks. The cloud server utilizes the topology state calculation module to perform target trajectory deduction based on video visual tracking within the constructed directed acyclic graph and finite state machine; When visual tracking is determined to be broken, the physical node corresponding to the broken link location is recorded, the finite state machine enters the suspended state, and the cloud server sends the corresponding alarm video clip to the mobile law enforcement terminal. Mobile law enforcement terminals scan the topological anchor point labels of the corresponding area, extract physical coordinates and business attributes as data to be reported to the cloud server; The cloud server generates dynamic virtual nodes in the directed acyclic graph based on the data entered, and calculates the confidence of human-machine collaboration stitching between the dynamic virtual nodes and the broken physical nodes. When the preset conditions are met, the trajectory stitching is completed, and the finite state machine is awakened to determine the violation. After the finite state machine outputs the violation determination result, the cloud server extracts the corresponding alarm video stream segment and the data filled in by the mobile law enforcement terminal to generate a joint evidence file.
[0007] Furthermore, in the aforementioned trajectory extrapolation and visual tracking stages, the cloud server sets a dynamic time window based on the spatial distance and average movement speed of adjacent physical nodes. Under the constraints of the time window, the visual similarity of adjacent high-dimensional feature vectors is obtained through dot product operation and norm calculation, and a smoothing constant is introduced to prevent algebraic singularity. When the maximum visual similarity of the candidate targets in the entire network is continuously lower than the confidence threshold, it is determined that a visual tracking chain break has occurred.
[0008] Furthermore, the mobile law enforcement terminal establishes near-field interaction with the topological anchor tag through a radio frequency antenna and optical sensor; extracts Bluetooth signal strength to estimate physical spatial distance and introduces radio frequency attenuation compensation for smoothing; comprehensively extracts acceleration variance features and decodes encrypted QR codes to determine anti-counterfeiting; after verification and the estimated distance falls within the safe range, physical coordinates and business attributes are extracted.
[0009] Furthermore, the calculation process of the human-machine collaborative stitching confidence score includes: calculating the spatial confidence score component based on the spatial connectivity distance between the broken physical node and the dynamic virtual node using a Gaussian kernel function; calculating the dynamic time deviation by combining the absolute timestamp, movement speed, and connectivity distance, and calculating the time confidence score component using an exponential decay function; extracting business attributes and calculating attribute confidence score components through Boolean logic and vector similarity; and linearly weighting and summing the above components to obtain the confidence score.
[0010] Furthermore, in the closed-loop drive of the finite state machine, the feature sequence of consecutive associated nodes with timestamp alignment is extracted and input into the long short-term memory network model to predict the business anomaly probability of the output sequence; the nonlinear amplified delay timeout penalty component and the shortest physical connectivity path deviation penalty component are calculated; the above values are normalized and weighted to obtain the comprehensive violation probability, and when the probability exceeds the preset threshold, the state machine is triggered to jump to the violation alarm node.
[0011] Furthermore, the process of generating the joint evidence file includes: calculating a cropping window to extract video segments using timestamps calibrated by the Network Time Protocol, and filtering effective slices using a three-dimensional convolutional network; inserting a sequence of transition frames containing virtual fast-forward magnification and text markers between adjacent video segments with physical blind spots; concatenating the underlying binary data of the video, the illegal metadata, and the execution order of the electronic signature, and generating and encapsulating a tamper-proof evidence hash value using a secure hash algorithm.
[0012] Furthermore, the time synchronization mechanism of the system is as follows: the terminal obtains the network round-trip delay and basic time offset from the cloud, and after smoothing with a sliding window, it combines the dynamic drift rate calculated by the hardware clock based on the first-order difference principle, and superimposes the linear compensation increment on the local timestamp to establish an absolute benchmark for spatiotemporal extrapolation.
[0013] Furthermore, this method also includes automated scheduling and early warning mechanisms: The time markers of all pending tasks are synchronized and calibrated. The dynamic comprehensive priority is derived by combining the queuing time of static risk and nonlinear mapping, and then the tasks are rearranged in descending order. Tasks are pushed based on the distance, historical latency, and signal strength of idle terminals; Multidimensional operational data with weights assigned based on information entropy is used to output instantaneous risk using a long short-term memory network and to calculate a dynamic risk index by integrating historical cumulative risk. This index is then input into a four-color early warning mapping model to trigger an emergency response.
[0014] A second aspect of the present invention provides a smart monitoring system for medical institutions based on mobile terminals and video cloud supervision, for implementing the above method, comprising: The video cloud monitoring platform is used to access front-end video streams for feature extraction and event recognition, and to generate alarm information. The management backend is used to assign daily monitoring tasks and display data. The cloud server is used to receive tasks, generate verification tasks, and perform trajectory deduction in a directed acyclic graph and finite state machine. An alarm video is issued when a chain break is detected and the state machine is suspended. Based on the reported data, dynamic virtual nodes are generated to calculate the stitching confidence level to wake up the state machine, and a joint evidence dossier is generated after the violation is determined. The mobile law enforcement terminal is used to receive video, scan topological anchor point tags to extract physical coordinates and business attributes, and then report them.
[0015] This invention provides a method and system for intelligent supervision of medical institutions based on mobile devices and video cloud monitoring. It has the following beneficial effects: 1. This invention innovatively introduces a virtual node mapping mechanism and multi-dimensional confidence calculation by combining intelligent video analysis and graph models. When encountering physical monitoring blind spots or visual disconnections, it can automatically complete spatiotemporal trajectory stitching in human-machine collaboration, effectively overcoming environmental interference and ensuring the continuity of cross-regional monitoring and tracking.
[0016] 2. This invention establishes an automatic mechanism for splicing joint evidence files, performing precise temporal alignment and blind spot transition processing on multi-source videos. The system integrates on-site verification data to generate hash digests with digital signatures, achieving full automation from violation alerts to electronic document generation, significantly reducing evidence collection costs and ensuring the validity of evidence.
[0017] 3. This invention constructs a dynamic risk calculation model and a macro-data dashboard, utilizing deep learning networks to analyze multi-dimensional operational data to output early warning levels. Combining 3D visualization and visual encoding technologies, it transforms abstract risk data into intuitive visual features, assisting managers in real-time monitoring of the overall risk evolution and quickly locating abnormal nodes. Attached Figure Description
[0018] Figure 1 This is a diagram illustrating the architecture of a smart medical institution monitoring system according to an embodiment of the present invention. Figure 2 This is a flowchart of a smart supervision method for medical institutions according to an embodiment of the present invention; Figure 3 This is a bar chart comparing the system performance of the present invention. Detailed Implementation
[0019] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0020] See attached document Figure 1 This invention provides a smart supervision system for medical institutions based on mobile terminals and video cloud supervision, including: mobile law enforcement terminals, video cloud supervision platforms, cloud servers, and management backends.
[0021] The video cloud supervision platform is deployed at network edge nodes or servers, and includes: a video access gateway for acquiring real-time video streams, a video cloud storage module for storing data, and a video intelligent analysis engine. This engine employs a ResNet-50 architecture with a fused Feature Pyramid Network (FPN), taking normalized RGB video frame streams as input and outputting the physical bounding box coordinate sequence of targets and high-dimensional feature vectors representing their appearance. It then identifies violations and generates alarms according to rules. Its model training uses a cross-entropy loss function to optimize the classification branch and introduces a triplet loss function to constrain the feature space, thereby enhancing the intra-class compactness and inter-class separability of the high-dimensional feature vectors.
[0022] The cloud server communicates with the aforementioned platform and includes: a task scheduling module (generating verification tasks based on alarms or instructions), a data storage module (storing archive records), a risk warning module (calculating risk indices based on historical data and alarm frequency), and a topology state calculation module. The topology state calculation module constructs a directed acyclic graph and a finite state machine to perform trajectory stitching and compliance judgment; it expands the static directed acyclic graph representing the physical space by injecting dynamic virtual nodes and a time dimension, and dynamically compensates for communication latency of each data stream based on the embedded Network Time Protocol (NTP), ensuring that all event nodes accessing the graph are strictly monotonically increasing in the time dimension.
[0023] The mobile law enforcement terminal communicates with a cloud server and includes modules for task reception, file retrieval, intelligent inspection, remote verification, offline caching, and document generation. This terminal is used to scan topological anchor point tags at the scene to obtain physical coordinates and generate reporting data containing timestamps and business attributes.
[0024] The management backend connects to the cloud server and includes: a resource configuration module (configuring form templates and analysis rules), a task management module (dispatching and tracking tasks), and a data analysis module (extracting cloud data to generate visual charts).
[0025] See attached document Figure 2 This invention provides a smart monitoring method based on the above system, comprising the following steps: S10: The management backend dispatches daily supervision tasks, or the video cloud supervision platform analyzes the accessed video streams to generate alarm information; the cloud server receives task instructions or alarm information to generate verification tasks.
[0026] S20, the cloud server utilizes the topology state calculation module to perform target trajectory deduction within the constructed directed acyclic graph and finite state machine. Based on the aforementioned graph theory deduction principle, the system needs to quantify the consistency of visual-physical features of targets across different view domains. The cloud server extracts adjacent high-dimensional feature vectors from the time series. and ,in, Defined as the target at a historical moment The high-dimensional feature vector representing the physical state of the appearance is extracted by the video intelligent analysis engine; Defined as a candidate target at a subsequent time. The extracted high-dimensional feature vectors representing the physical state of the object are used to evaluate their matching degree using cosine similarity. Considering that in real-world conditions, occlusion or drastic changes in lighting can cause the feature extraction network to output a minimum or even zero vector, leading to the denominator in division operations approaching zero... To address the singularity problem, this embodiment introduces a visual feature smoothing constant. Smooth the denominator. Visual similarity. The calculation formula is expressed as: .in, These represent the corresponding eigenvectors. Norm, The value range is set to 10. −6 Up to 10 −5 When the system finds the maximum value among all candidate targets in the entire network. When the value remains below the preset confidence threshold (which is usually set to 0.75 to 0.85 based on the false alarm rate tolerance of historical samples), a visual tracking break is determined, the finite state machine enters a suspended state, and the cloud server sends the corresponding alarm video clip to the mobile law enforcement terminal for verification.
[0027] S30: The mobile law enforcement terminal scans the topological anchor point tags of the corresponding area, extracts the physical coordinates and business attributes, and reports them to the cloud server. The cloud server generates dynamic virtual nodes in the directed acyclic graph based on the reported data. On this basis, the system further calculates the confidence level of human-machine collaboration stitching between the dynamic virtual nodes and the disconnected physical nodes. The technical purpose of this calculation step is to overcome the problem of failure of a single visual feature through multi-dimensional data verification. The output judgment of the human-machine collaboration stitching confidence level does not rely on a single extreme value index, but comprehensively introduces multiple dimensions of correlation factors such as spatial physical distance attenuation, time difference attenuation, and logical matching of business attributes for weighted summation calculation. When the human-machine collaboration stitching confidence level meets the preset conditions, trajectory stitching is completed, and the finite state machine is activated to determine violations.
[0028] After the S40 finite state machine outputs the violation determination result, the cloud server extracts the video stream segments and the data filled in by the mobile law enforcement terminal to generate a joint evidence file; the mobile law enforcement terminal generates law enforcement documents based on the joint evidence file through the document generation module and obtains the electronic signature data stream.
[0029] S50, the mobile law enforcement terminal uploads the law enforcement documents and joint evidence files containing electronic signature data streams to the data storage module. The risk warning module extracts violation event parameters to calculate the risk index of the target medical institution. In this embodiment, the selection of specific input parameters is based on the business causal logic of administrative supervision. The system extracts basic file deduction items, the frequency of violations automatically identified by AI, and the severity coefficient confirmed by manual review as multiple input variables through the file query module. The calculation result of the risk index is presented as a four-level mapping of red, orange, yellow, and green to update the risk level of the medical institution.
[0030] S60, the management backend summarizes and verifies records and risk level data, and generates visualized data reports containing regulatory statistical indicators and risk distribution information through the data analysis module.
[0031] Based on the system architecture, the video cloud monitoring platform is deployed on edge computing nodes or dedicated server rooms. The video access gateway acquires real-time video streams from the image acquisition devices at the front end of the medical institution based on standard communication protocols. For the data packet encapsulation and parsing mechanisms of standard communication protocols (such as GB / T 28181 protocol and real-time streaming protocol), those skilled in the art can refer to relevant network communication standards and specifications for implementation. The underlying interaction principles are well-known technologies in this field and will not be elaborated here.
[0032] A multidimensional edge-aware processing method is used to perform parallel analysis and structured data extraction on video streams at edge nodes. This method includes the following steps: S11, the video access gateway decodes and performs fixed-frequency frame extraction on the acquired real-time video stream, and synchronously stores the video stream data to the video cloud storage module. The system scales the decoded keyframe images to a fixed resolution and performs mean subtraction and pixel value normalization preprocessing. The processed image matrix is converted into multi-dimensional tensor data, which serves as the input source for subsequent deep learning algorithm models.
[0033] S12, the video intelligent analysis engine, receives preprocessed tensor data and runs multiple parallel task analysis models in memory. The task analysis models employ a target detection neural network architecture based on region proposal or single-stage regression. Their underlying shared feature extraction backbone network contains cascaded two-dimensional convolutional kernels and spatial pyramid pooling modules. During model training, a medical institution monitoring dataset is extracted as samples, combined with the intersection-union (IU) loss function for predicting bounding boxes. Focus loss function with class confidence Perform end-to-end backpropagation optimization.
[0034] In practice, the system loads a medical waste management model, an infection control model, and a personnel qualification verification model. The medical waste management model detects the bounding box coordinates of the medical waste storage bins and the coordinates of the center point of the soles of moving personnel. The system uses a ray-crossing algorithm to calculate the spatial inclusion relationship between the center point coordinates and the preset medical waste storage area boundary. When the number of consecutive frames in which the center point falls within this area exceeds a dwell time threshold, it is determined that there is an intrusion into the area. At the same time, the system extracts local feature data of the storage bins to determine whether they are uncovered or overflowing.
[0035] The infection control model extracts a feature map of the head region after detecting a person, inputs it into the attribute classification network branch, and outputs a binary probability value indicating whether the target is wearing a medical mask or work cap correctly. The system introduces a smoothing weighting logic based on a time sliding window to extract the probability sequence of the target over multiple consecutive frames. When the weighted average probability is greater than the classification threshold, a violation judgment is triggered.
[0036] The personnel qualification verification model uses a keypoint localization algorithm to capture facial region images, extract facial feature vectors, and compares them with the personnel qualification feature database in the management backend for similarity. The calculation logic for facial feature similarity is as follows: In the formula, This represents the dot product operation between the currently extracted vector and the reference library vector. Its physical purpose is to characterize the degree of overlap of the projections of the two sets of high-dimensional data in the feature space. These represent the corresponding eigenvectors. Norms are used to eliminate the influence of absolute values on directional measurements. This is because in real-world scenarios, extreme lighting or abnormal output minima or even zero vectors from feature extraction networks can cause the denominator in division operations to approach zero. To address the singularity problem, this embodiment introduces a facial feature regularization constant. Smooth the denominator. Wherein, The value range is set to 10. −6 Up to 10 −5 In addition, the system performs a multi-dimensional matrix weighted comparison of facial feature similarity and gait feature similarity extracted based on overall appearance. If the comprehensive similarity value calculated after traversal fails to reach the matching threshold for identity determination (this threshold is usually dynamically set to 0.65 to 0.75 depending on the security level), the system determines that the person in the current field of view is not on the institution's list of practitioners, triggering an abnormal personnel qualification determination.
[0037] In step S13, the video intelligent analysis engine summarizes the analysis results from the parallel processing described above and generates a structured metadata package. This metadata package is based on the absolute timestamp of the current video frame, injected using the edge gateway's local hardware-level clock, and includes the node device's unique identifier, the violation event type code, and the extracted target high-dimensional feature vector. The video access gateway then reports the metadata package to the cloud server via an encrypted channel to drive cross-domain tracing and business scheduling processes.
[0038] Based on the aforementioned system architecture and edge-aware network, this embodiment provides a dual-track task generation and scheduling method, including the following steps: S14, in this embodiment, the management backend responds to the operation command through the task management module and generates the first track of manually dispatched tasks. The system extracts the target organization identifier, the standardized inspection form template identifier configured by the resource configuration module, and the completion time limit parameters from the command, encapsulates them into a manual task data package, and sends it to the task scheduling module of the cloud server. For the serialization and hypertext transfer protocol encapsulation mechanism of the front-end web page form data, those skilled in the art can refer to relevant network communication specifications for implementation; it is a well-known technology in the field and will not be elaborated upon here.
[0039] S15, the cloud server receives the structured metadata package reported by the video access gateway and executes the second-track automated task triggering logic. The task scheduling module parses the violation event type code in the metadata package. When the event type code indicates a static single-point violation, the system instantiates a single-point alarm verification task and binds the abnormal video frame as evidence to the task record. When the event type code indicates a cross-regional business anomaly, the system extracts the target high-dimensional feature vector and the current node's physical identifier from the metadata, driving the topology state calculation module to instantiate a cross-domain tracking task.
[0040] S16, the task scheduling module merges manually generated task data packets with automatically triggered task instances and injects them into a unified global pending task queue in memory. As a preferred approach, considering the heterogeneity between manually dispatched instructions and the hardware time of edge device nodes, the system performs cross-source data time alignment logic before merging tasks. The system extracts the network time protocol reference time from the cloud server as an absolute reference axis and synchronizes and calibrates the initial time identifiers of tasks from various sources entering the queue. To ensure rapid response to high-risk business events and prevent scheduling starvation caused by low-priority tasks, based on dynamic queuing theory and anti-starvation scheduling principles, the system performs dynamic priority calculation on the tasks to be assigned in the queue. Let the number of tasks in the queue be... The dynamic overall priority of each task is The system extracts the static baseline risk of the task and performs a weighted calculation based on the queuing time. The dynamic overall priority of each task is The calculation formula is as follows: In the formula, Indicates the system is the first The static baseline risk weight assigned to each task is obtained by mapping the value from the preset level configuration matrix based on the trigger source and business type. This represents the cumulative time the task has resided in the global pending task queue since its timestamp was generated. It addresses the calculation difference caused by minute clock skews in distributed nodes during actual operation. In case of negative values, the system introduces a maximum value function. Perform zero-boundary truncation; This represents the time decay coefficient, used to control the saturation rate of priority increase over time. Its value is usually set to the range of 0.01 to 0.1. These represent the normalized weight parameters for the baseline risk component and the time compensation component, respectively. To ensure a consistent weighting calculation scale, the system imposes mandatory constraints. The physical purpose of introducing the exponential term containing the base of the natural constant is to perform a nonlinear bounded mapping on the waiting time. This algorithm structure allows the compensation priority of low-priority tasks to gradually increase and approach the asymptotic upper bound when they have not been processed for a long time, forcibly changing the queue order to prevent system scheduling deadlock.
[0041] S17, The task scheduling module updates dynamic comprehensive priorities based on computation. The global pending task queue is reordered in descending order. The system maintains a heartbeat detection mechanism to obtain the online status and geographic coordinates of mobile law enforcement terminals. Based on the principle of multi-objective optimization decision-making, the system constructs a multivariate weighted evaluation function by comprehensively considering the absolute linear distance between idle mobile law enforcement terminals and target nodes, the average time taken for historical task completion by the terminals, and the current network signal strength. The system selects the mobile law enforcement terminal with the highest comprehensive score and pushes the high-priority tasks at the top of the queue to that terminal. The mobile law enforcement terminal receives and parses the task data stream through the task receiving module and renders the business verification interface on the front end.
[0042] This embodiment provides a physical space graph theory abstraction method based on static directed acyclic graphs and finite state machines, including the following steps: S21, In this embodiment, the cloud server abstracts the physical building floor plan of the medical institution into a static directed acyclic graph. .in, It represents a static set of physical nodes, whose specific features are implemented as a field of view area with fixed cameras or a forcibly set physical entrance / exit. This represents a static set of directed edges, used to characterize walkable connected paths between adjacent physical nodes. The system construction dimension is... ( The adjacency matrix (where the total number of static physical nodes is the number of nodes) When node With nodes When there are one-way or two-way travel paths, matrix elements The value is assigned to the straight-line distance between the two or the actual walking distance along the corridor; if there is a rigid barrier preventing direct access, then... The value is assigned to infinity. For the contiguous memory storage of matrix data and the dimensionality reduction access mechanism using adjacency lists, those skilled in the art can refer to standard data structure specifications for implementation; these are well-known technologies in the field and will not be elaborated upon here.
[0043] S22, the system is in a static directed acyclic graph. Based on this, a time constraint dimension is injected to calculate the time window matrix for cross-node transitions. Based on the principles of temporal and spatial kinematic constraints of multi-source data, the technical purpose of this calculation step is to provide a dynamic and reasonable time determination boundary for subsequent multi-channel video cross-domain tracking, filtering out false alarms caused by spatiotemporal logical paradoxes. The system extracts a set of static directed edges. The system calculates the baseline cross-node transfer time by combining the physical distance data of each side with the average statistical speed of personnel movement under specific business scenarios. As a preferred approach, considering the fluctuations in time consumption caused by personnel's brief stays or conversations in actual working conditions, the system does not rely on a single extreme value to determine time consistency, but instead constructs a dynamic time window that includes an upper and lower bound. The formulas for calculating the upper and lower bounds of the time window are as follows: ; In the formula, Nodes extracted from the adjacency matrix and The physical connectivity distance between them; and These represent the average maximum and minimum movement speeds of personnel, derived from historical multi-source working condition statistics. To prevent the denominator from approaching zero in algebraic operations... This triggers a system crash singularity. and All are strictly constrained to be strictly greater than a preset non-zero constant (e.g., 0.1 m / s). Meanwhile, to ensure the self-consistency of the physical logic, The value of is constrained to be within the maximum speed threshold for normal people to run (usually set to 3 to 5 meters per second). The time compensation offset set for the system is used to absorb the small residual errors that exist in the process of multi-source node clock synchronization. Its value range is usually set to 1 to 3 seconds. This is the maximum tolerable dwell time parameter set in the business flow rules. For time alignment logic in multi-source data fusion, the system uses the network time protocol. The absolute reference clock of the server is extracted and differentially calibrated to eliminate timestamp drift at heterogeneous edge nodes. Furthermore, to prevent singularities in the graph Laplacian matrix caused by physically isolated dead-end nodes during global shortest path calculation, which could lead to matrix inversion failure, the system performs differential calibration on all diagonal elements. Add a minimal constant damping term (This damping term) The value is usually set to 10. −5 That is, to construct a regularized Laplace matrix, thereby ensuring the full rank and invertibility of the matrix.
[0044] S23, the cloud server is based on a static directed acyclic graph. A finite state machine model is instantiated for specific regulatory business, and this model is defined as a quintuple. .in, A set of discrete physical states representing the business lifecycle (e.g., in the medical waste scenario, this includes the states of generation and collection, regional transfer, temporary storage and handover upon discharge). The alphabet representing the external events that trigger state transitions, with input sources including structured metadata packages and time window timeout alarm commands; Indicates the initial activation status of the service; This represents the set of final business states, including both compliant closed loops and non-compliant / abnormal final states. Represents the state transition function, with the mapping relationship as follows: The system will use a static physical node set. With state transition function Perform forced memory mapping binding. This applies when entities are in a directed acyclic graph. When cross-node displacement occurs and metadata is generated, the finite state machine extracts the node's physical identifier from the metadata and drives the internal state to perform synchronous transitions along directed edges. The system comprehensively considers the similarity of features across multiple consecutive frames of edge nodes, the fit of physical time windows, and multi-dimensional data from business flow specifications for weighted joint judgment, and outputs the state transition result. This mechanism ensures high-dimensional consistency between displacement in the physical space and the compliant state in the digital logical space, effectively avoiding discrete state machine jump errors caused by instantaneous physical occlusion or visual misdetection.
[0045] Based on the above physical space topology and state machine mapping model, this embodiment provides a cross-view feature extraction and visual tracking chain break determination method, including the following steps: S24, In this embodiment, the system constructs and deploys a deep neural network model for target identity re-identification. As a preferred approach, this model employs a fused feature pyramid network. The system utilizes the ResNet-50 architecture. After preprocessing the target local image data captured from edge nodes through interpolation, scaling, and pixel value normalization, it is adjusted into fixed-dimensional tensor data (tensor dimension set to...). ,in For batch size, and The input sources were set to 256 and 128 pixels respectively. The network then... The module integrates multi-scale features, ultimately outputting a 512-dimensional high-dimensional feature vector representing the physical state of the target's appearance. During model training, trajectory videos from specific medical institution scenarios are extracted as samples, with labels consisting of unique target identification codes across the entire network. Training employs a multi-task joint optimization strategy, with its loss function being a weighted combination of classification cross-entropy loss and triplet loss from batch hard sample mining, thereby enhancing the intra-class compactness and inter-class separability of the feature vector. For the underlying convolutional calculations and forward propagation mechanisms of the deep neural network model, those skilled in the art can refer to standard deep learning frameworks for implementation; these are well-known technologies in the field and will not be elaborated upon here.
[0046] S25, the cloud server receives the target high-dimensional feature vector reported by the edge nodes and performs cross-view feature matching calculation in memory. Based on the principles of metric learning and high-dimensional spatial distance representation, the system needs to quantify the visual-physical feature consistency of cross-view targets. Continuing the definition in the previous embodiments, the system extracts adjacent high-dimensional feature vectors in the time series. and .in, Defined as the target at a historical moment Leave the source node The high-dimensional feature vector extracted during the field of view; Defined as a candidate target at a subsequent time. Enter node The high-dimensional feature vector is extracted from the field of view. To filter out impossible correlations in physical space, the system forcibly extracts the aforementioned constructed dynamic time window. Determine the time difference Whether it strictly falls within a valid time interval. To address potential time difference calculation anomalies caused by clock drift at distributed edge nodes, the system uniformly calls the network time protocol before performing subtraction operations. Reference clock pair and Synchronous calibration and alignment are performed. For candidate feature pairs that meet spatiotemporal constraints, the system evaluates their matching degree by calculating cosine similarity. Visual similarity. The calculation formula is as follows: In the formula, The dot product operation represents the dot product of two sets of high-dimensional vectors. Its technical purpose is to characterize the degree of overlap of the projections of the two sets of data in the feature space. These represent the corresponding eigenvectors. The norm is used to eliminate the influence of the absolute value of a vector on the angular dimension. To prevent computational overflow caused by extracting zero features from visual blind spots, this embodiment introduces a fixed visual feature smoothing constant into the denominator. Perform denominator smoothing. This visual feature smoothing constant... The value range is strictly set to 10. −6 Up to 10 −5 Order of magnitude.
[0047] S26, the system summarizes the matching calculation results of candidate targets and performs a tracking chain break determination based on multi-dimensional data. Based on multi-source data fusion and probabilistic decision-making theory, and considering that changes in clinical ambient lighting can easily cause the failure of a single visual feature, the determination of the output result no longer solely relies on the aforementioned visual similarity. The system constructs a comprehensive stitching confidence evaluation model that includes visual features, spatiotemporal fit, and business attribute constraints. The selection of multidimensional input parameters is based on the continuity causal constraints of the same physical entity in spatiotemporal displacement. The system extracts edge-structured attribute labels carried by the target when entering the field of view (such as whether it carries a medical waste storage bin of a specific color, or whether personnel are wearing specific uniforms), and calculates attribute matching scores using Boolean logic. (Assign a value when the attributes are completely identical) When a conflict exists, the value assigned is To avoid symbolic semantic conflicts with the weight parameters in the aforementioned task priority scheduling model, this step constructs an independent confidence weight factor, which is then used to synthesize and stitch the confidence scores. The calculation formula is as follows: In the formula, Represents a node To node The historical average transfer time between them; The smoothing coefficient for adjusting the time decay rate is physically designed to control the sensitivity of the spatiotemporal deviation penalty. Its value is typically set between 10 and 30 seconds, and the system enforces verification and constraints on this. To prevent algebraic division by zero errors; and These represent the normalized confidence weights for visual similarity, temporal fit, and attribute matching, respectively. To ensure consistency in the weighted calculation scale, the system enforces strict constraints. This represents an exponential function. The physical purpose of using an exponential function, which includes both the absolute value and the base of the natural constant, to handle time differences is to non-linearly penalize lag behavior that deviates from the mean; that is, the further the target's appearance time deviates from the historical statistical mean, the more drastically the confidence score for classifying it as the same entity decreases. The system traverses the target nodes. Extract the maximum comprehensive suture confidence from all candidate targets within the field of view. When this maximum value remains below the preset association matching threshold for an extended period exceeding the set business tolerance window, the system determines that the current target has experienced a visual tracking break. The aforementioned association matching threshold is typically set between 0.65 and 0.75, based on the crossover point of the false positive and false negative rates of historical samples. After triggering the break determination, the cloud server drives a finite state machine. The internal state of the corresponding instance is forcibly switched from the normal transition state to the abnormal suspension state. At the same time, the alarm video clip of the preceding node of the broken link is captured and sent to the mobile law enforcement terminal to make up for the perception gap in the physical space monitoring blind spot through manual intervention.
[0048] To address the issue that video intelligent analysis engines struggle to cover blind spots in the physical spaces within medical institutions (such as the interior of enclosed medical waste storage rooms and stairwell corners without monitoring), this embodiment provides a topological anchor point label and mobile law enforcement terminal interaction dimensionality reduction mapping mechanism. This mechanism injects the physical trajectories of on-site law enforcement personnel into a static topological structure as dynamic virtual nodes, including the following steps: S31, In this embodiment, the system pre-deploys topology anchor tags in specific physical blind spots. The topology anchor tag is specifically implemented as a composite hardware device combining a low-power Bluetooth beacon and an encrypted QR code. Each tag internally stores a unique spatial physical identifier code and a code derived from the Building Information Model (BIM). Or high-precision three-dimensional spatial coordinates calibrated by on-site laser ranging .
[0049] S32, when the mobile law enforcement terminal is performing on-site verification tasks, it establishes a near-field interactive connection with the topological anchor tag through its built-in radio frequency antenna and optical sensor to extract the spatial physical identification code. To prevent location spoofing caused by personnel scanning tag photos or cloning radio frequency signals in different locations, the system enforces physical proximity verification logic based on received signal strength indication. Based on the general principle of logarithmic distance path loss for radio signals propagating in space, the mobile law enforcement terminal extracts Bluetooth signal strength from multiple consecutive sampling periods, performs smoothing filtering, and calculates the estimated physical spatial distance between the terminal and the tag. The system constructs a distance estimation model to calculate the estimated distance between the terminal and the tag. The calculation formula is as follows: In the formula, This parameter represents the reference received signal strength at a distance of 1 meter from the tag. Its value depends on the tag hardware factory calibration (usually in the range of -50dBm to -60dBm). This represents the time-smoothed signal strength actually measured by the mobile law enforcement terminal; The environmental path loss index (ARBI) is used to characterize the attenuation effect of specific indoor building materials on electromagnetic wave signals. Its value is typically set between 2.0 and 4.0 based on the density of building obstructions. To prevent the path loss index from failing to read under extreme abnormal conditions, causing the denominator to approach zero... To address the algebraic collapse problem, this embodiment introduces a small amount of radio frequency attenuation compensation in the denominator term. Perform computational smoothing. The value is set to 10. −5 .
[0050] Considering that a single radio frequency signal is highly susceptible to multipath effects or interference from people moving around in complex indoor electromagnetic environments, which can cause abrupt changes in ranging, the determination of the output result no longer relies solely on the aforementioned estimated distance. The single extreme value. The system comprehensively incorporates the inertial measurement unit built into the mobile law enforcement terminal. The system extracts acceleration variance features during scanning and decodes the encrypted QR code using the optical sensor. This data is used to construct a multi-dimensional weighted anti-counterfeiting judgment logic that incorporates radio frequency distance, device motion stability, and optical interaction timeliness. Specifically, the system normalizes the extracted acceleration variance, decoding time, and estimated radio frequency distance, and then performs a linear weighted sum based on a preset confidence weight matrix to calculate a comprehensive anti-counterfeiting judgment score. The score is determined when the comprehensive judgment score meets a preset anti-counterfeiting threshold, and the estimated distance... Only when the mobile law enforcement terminal falls within the physical adjacency safety zone (usually set to 0 to 2.0 meters, or less than or equal to 2.0 meters) will the system determine that the spatial location of the current mobile law enforcement terminal is legal and valid, and allow subsequent interaction logic to proceed.
[0051] S33, after the proximity verification is passed, the mobile law enforcement terminal extracts the underlying hardware timestamp in real time. Mobile law enforcement terminals use the Network Time Protocol (NTP). Extract the base clock from the cloud server and analyze the underlying hardware timestamp. Perform phase offset compensation to obtain the calibrated timestamp. This ensures that the timestamp and the timeline of the video intelligent analysis engine are in the same spatiotemporal reference frame. The mobile law enforcement terminal will use the spatial physical identifier code and the calibrated timestamp. The terminal device hardware identification code and the structured business status data verified and entered on-site through the intelligent inspection module are packaged into an interactive mapping data packet. For the asymmetric encryption encapsulation of the data packet and the network communication reporting mechanism based on the transmission control protocol, those skilled in the art can refer to well-known information security encryption specifications and underlying network transmission standards for implementation; these are well-known technologies in the field and will not be elaborated upon here.
[0052] S34: The cloud server receives the interactive mapping data packet and performs a dimensionality reduction mapping calculation from the continuous physical coordinate system to the discrete graph theory matrix system in memory. Based on the reported spatial physical identifier code, the system queries the system's database to resolve the corresponding three-dimensional spatial coordinates. Subsequently, in the original static directed acyclic graph Instantiate a dynamic virtual node with time attributes. This dimensionality reduction mapping operation mathematically transforms the irregular spatial patrol behavior occasionally observed by human supervisors into standard node elements in graph theory space capable of performing matrix operations. Based on the principle of three-dimensional spatial geometric distance measurement, the system further drives the topology state calculation module to update the adjacency matrix of the graph structure and calculate the newly generated dynamic virtual nodes. With each pre-existing static physical node The relative positional parameters between them. Spatial connectivity distance. The calculation formula is as follows: In the formula, Represents static physical nodes The three-dimensional coordinate position in the global coordinate system; This represents the passageway polyline compensation coefficient, which physically corrects the nonlinear error between the Euclidean straight-line distance in three-dimensional space and the actual walking distance around the corridor within the building. The value range is preset to 1.2 to 1.5 based on the complexity of the medical institution's internal structure. The system traverses and calculates dynamic virtual nodes. With all static physical nodes The distance parameter is used to calculate the effective distance vector, which is then added as new row and column elements to the edge of the original adjacency matrix, thus dynamically expanding the matrix data dimension. To prevent the Laplacian matrix singularity problem in the graph theory model caused by dynamic node injection, which could lead to the failure of matrix inversion calculation during global trajectory optimization, the system adds the adjacent elements of the dynamic virtual node itself to the diagonal of the expanded adjacency matrix. Forced injection of a minimal normal number damping term (the value is usually set to 10). −5 ).
[0053] To establish an absolute spatiotemporal reference for multi-source heterogeneous devices in spatiotemporal stitching simulation, this embodiment provides a cross-modal dynamic clock calibration method, which specifically includes the following steps: S35, in this embodiment, the cloud server is set as a globally unified level-zero reference time source. The mobile law enforcement terminal and each edge video node act as time synchronization clients, periodically initiating clock synchronization requests to the cloud server. The time synchronization client extracts the sending timestamp from the local operating system and encapsulates it into a clock synchronization probe data packet, which is then sent to the cloud server. The cloud server records the arrival timestamp of this data packet and, combined with the processing time, generates a response data packet with the cloud sending timestamp, returning it to the client. For the underlying network socket communication and data packet encapsulation and serialization mechanism, those skilled in the art can refer to the standard Transmission Control Protocol specification for implementation; this is well-known technology in the field and will not be elaborated upon here.
[0054] S36, the time synchronization client receives the response data packet and extracts the local received timestamp, calculating the network round-trip delay and base time offset for a single communication. Considering the sudden network congestion and routing jitter in the internal wireless LAN environment of medical institutions, the time offset calculated in a single instance is highly susceptible to interference from high-frequency delay spikes. As a preferred approach, the system does not rely on the extreme values of a single probe data packet for forced clock overlay, but instead constructs a delay-weighted historical sliding window filtering model. The system extracts the base time offset and network round-trip delay for multiple consecutive sampling periods, calculating the smoothed baseline time offset. The calculation formula is as follows: In the formula, This represents the reference time offset after smoothing filtering; This represents the total number of valid samples within the sliding window, and its value is usually set to 5 to 15 based on the balance between memory consumption and real-time computation. Indicates the first The base time offset calculated from each sample; Represented as the first Each sample is assigned a confidence weight component. This is to prevent the weights from accumulating and approaching a certain level due to insufficient valid samples during system initialization. The algebraic collapse problem that this leads to is addressed by introducing a minimal positive constant damping term into the denominator in this embodiment. (Its value is usually set to 10) −6 This weight component The physical purpose is to suppress the negative impact of high-latency network packets on global time calibration. The system determines the weight values based on inverse proportional mapping logic. The calculation formula is as follows: In the formula, Indicates the first The absolute network round-trip delay for each sample. To prevent extremely short, occasional delays in the physical link from causing the denominator to approach zero. To address the algebraic overflow anomaly, this embodiment introduces a minimal constant smoothing term into the denominator. This constant is usually set to 10. −4 Order of magnitude.
[0055] S37, the system constructs a linear frequency drift compensation model for the underlying hardware oscillator based on the smoothed reference time offset. The system extracts the reference time offset between the current synchronization cycle and the previous synchronization cycle, as well as the local system's absolute timestamp, and calculates the dynamic drift rate of the hardware clock based on the first-order difference principle. Dynamic drift rate. The calculation formula is as follows: In the formula, Indicates the dynamic drift rate; and These represent the reference time offset between the current calculation cycle and the previous calculation cycle, respectively. and These represent the local system absolute timestamps when the two computation operations occurred; The aforementioned minimal constant smoothing term is reused to prevent division-to-zero crashes.
[0056] The system injects the dynamic drift rate into the kernel timer of the local operating system. When the mobile law enforcement terminal generates business data, the system extracts the local uncalibrated timestamp, uses the static time offset of the last successful synchronization as the calibration baseline, and superimposes a linear compensation increment calculated based on the elapsed time and the dynamic drift rate to output a compensated timestamp. The calculation formula is as follows: In the formula, Indicates the compensation timestamp; This represents the local, uncalibrated timestamp when business data was generated.
[0057] S38, the system performs a comprehensive reliability assessment of the compensated timestamps to address the issue of linear compensation model failure caused by long-term network blind spots. The system constructs confidence evaluation logic that includes the dimensions of network latency variance and time drift span. Time compensation confidence. The calculation formula is as follows: In the formula, Indicates the time-compensated confidence level; The statistical variance of network round-trip delay within the sliding window represents the stability of recent communication links; It represents the absolute time difference that has elapsed since the last valid clock synchronization to the generation of the current business data; and These represent the delay sensitivity coefficient and the time sensitivity coefficient, respectively (the system constraint requires them to be strictly greater than 0). and Normalized weighting factors (constraints) representing network stability and time freshness, respectively. ); This represents the natural exponential function, used to perform a nonlinear penalty mapping on operating conditions with extremely long delays or prolonged periods of nonsynchronization.
[0058] When the calculated time-compensated confidence level When the preset time reliability threshold is met (usually set in the range of 0.80 to 0.90), the cloud server directly accepts the compensated timestamp. When the confidence level is below this threshold, the system determines that the current data source is in a low-precision clock decay state. The cloud server extracts the above time difference and network variance, and dynamically reverses the upper and lower bounds of the time determination window between nodes in the directed acyclic graph (i.e., the aforementioned embodiment). and This is achieved by relaxing the spatiotemporal topology constraints to force the absorption of residual clock uncertainties in the underlying hardware.
[0059] Based on the constructed static directed acyclic graph and dynamic virtual node expansion mechanism, this embodiment provides a human-machine collaborative spatiotemporal stitching and confidence calculation method when video tracking breaks down and causes the finite state machine to suspend, specifically including the following steps: In S39, the cloud server responds to the suspend command of the finite state machine, intercepts the historical running context of the preceding node of the visual tracking chain break, and extracts the node index identifier, absolute timestamp, and structured business feature tags of the broken physical node. When the mobile law enforcement terminal caches and scans the topology anchor tags and uploads data in the physical blind zone through the offline caching module, the system parses the interactive mapping data packet, extracts the spatial coordinates of the dynamic virtual node, the clock-calibrated terminal timestamp, and the business attribute parameters entered on-site, to establish a unified spatiotemporal measurement benchmark for discrete human-machine collaborative data.
[0060] S310: The system extracts the spatial topology parameters of broken-link physical nodes and dynamic virtual nodes, and performs spatial physical distance attenuation determination. The system calls the calculated connectivity distance to construct a spatial consistency evaluation model based on the Gaussian kernel function and calculates the spatial confidence components. The calculation formula is as follows: In the formula, This represents the spatial confidence component, whose values are mapped and constrained to the interval between 0 and 1; This represents the spatial connectivity distance between the broken-chain physical nodes and the dynamic virtual nodes extracted from the aforementioned expanded adjacency matrix; This represents the standard deviation constant for spatial diffusion, and its physical meaning lies in defining the threshold of a reasonable range of activity under a specific physical environment. This constant is typically set to 15 to 25 meters based on the average span of the corridors in medical facility buildings. To prevent algebraic overflow anomalies caused by the denominator approaching zero during system initialization or under extremely overlapping point deployment conditions, this embodiment introduces a spatially minimal positive constant damping term into the denominator. Its value is usually set to 10. −5 .
[0061] S311, the system performs time drift attenuation determination between cross-modal nodes. The system extracts the timestamp parameters of the disconnected physical node and the dynamic virtual node, calculates the dynamic time deviation by combining the physical connectivity distance, and outputs the time confidence component. The calculation formula is as follows: In the formula, Indicates the time confidence component; This represents the absolute timestamp at which the target was last captured by video at the physical node where the link was broken; This represents the absolute timestamp when the mobile law enforcement terminal generates the dynamic virtual node. This indicates the average moving speed of personnel, set to 1.0 to 1.5 meters per second; The velocity smoothing constant is set to 10. −3 ; This indicates the time decay tolerance factor, set to 60 to 120 seconds; The time deviation smoothing constant is set to 10. −4 The absolute value calculation of the above formula is used to characterize the absolute span of the residual between the actual elapsed time and the theoretically expected elapsed time.
[0062] S312, the system performs logical matching calculations for structured business attributes. The system extracts a set of feature attributes, constructs a hybrid decision model combining Boolean logic and multi-dimensional vector weighted summation, and calculates the confidence components of the business attributes. The calculation formula is as follows: In the formula, Indicates the total number of feature dimensions; and These represent the disconnected physical node and the dynamic virtual node at the [number]th [time]. The attribute feature values of each feature dimension; This represents a feature matching function that performs character comparisons for discrete enumerated data and calculates cosine similarity for continuous vector data. Represented as the first The system assigns normalized weight coefficients to each feature dimension and constrains the sum of the weights to 1. Features with absolute exclusion relationships are assigned the highest weights.
[0063] S313, the system integrates multi-dimensional evaluation components to calculate the confidence level of human-machine collaboration stitching to drive state machine transitions. The system constructs a multivariate weighted decision model, and the confidence level of human-machine collaboration stitching... The calculation formula is as follows: In the formula, Indicates the confidence level of human-machine collaborative suturing; and The normalized weighting factors correspond to the spatial, temporal, and business attribute components, respectively, and the sum of the three is strictly 1. The allocation ratio is dynamically adjusted according to the specific task focus. For the memory allocation and multi-threaded concurrent read protection mechanism of the above weight matrix, those skilled in the art can refer to the standard operating system synchronization lock specification for implementation, which is a well-known technology in this field and will not be elaborated here.
[0064] When the confidence level of human-machine collaborative stitching When the preset spatiotemporal stitching wake-up threshold (dynamically set between 0.75 and 0.85) is met, the system determines that the spatiotemporal stitching is successful, drives the finite state machine to exit the abnormal suspension state, and injects the dynamic virtual node as a valid predecessor state node into the path deduction link. If the threshold is not met within multiple consecutive inspection windows, the system triggers a final state abnormality alarm, defines the target as lost, and generates a forced manual intervention command.
[0065] Based on the above-mentioned topological state calculation and multimodal trajectory stitching mechanism, in order to achieve closed-loop management of regulatory operations, this embodiment provides a method for state transition and violation event triggering based on a finite state machine, specifically including the following steps: S41, In this embodiment, the cloud server drives a finite state machine to perform state transition determination based on the extracted sequence of continuously related nodes. As a preferred approach, the system is constructed based on a long short-term memory network. The trajectory compliance assessment model. To eliminate out-of-order interference caused by network latency, the system calls a cloud-based reference clock to perform monotonically increasing alignment and resampling on the timestamps of each node. The aligned input data is constructed into a dimension... tensor ( For batch size, For the time step of circulation, The node feature dimension includes specific features such as the target's absolute dwell time, business attribute vector, and personnel identity verification code, and zero-mean normalization is performed. Internally, the model consists of a fully connected data input layer and stacked layers. Circular unit module and multilayer perceptron Output layer. The model extracts business flow dependency features from the sequence data and outputs a predicted probability value of the current business flow sequence deviating from standard operating procedures. During the model training phase, historical compliant and non-compliant transaction records are extracted as samples, and end-to-end gradient backpropagation optimization is performed using a binary cross-entropy loss function.
[0066] S42, the system extracts the high-dimensional sequence features from the hidden layer output of the aforementioned network and, combined with pre-set rigid administrative regulatory rules, performs a deterministic hybrid compliance measurement calculation. Based on the principle of information completeness review, the determination of the output result no longer relies solely on the black-box prediction of the neural network's extreme value, but comprehensively introduces a multi-dimensional and explicit physical attenuation penalty logic. The system calculates the detention timeout penalty component. For regulatory tasks with strict time constraints, such as the temporary storage of medical waste and the transfer of special drugs, their actual flow must meet the time window boundary constraints in the graph theory model. The physical causality of choosing absolute residence time as the input parameter lies in the fact that the risk of biosafety pollution and the duration of personnel's environmental exposure are positively correlated and accumulate exponentially. Detention timeout penalty component The calculation formula is as follows: In the formula, This represents the calculated penalty component for excessive detention time. This represents the difference in absolute timestamps of the actual cumulative residence of a target within a specific regulated physical area; This indicates the upper limit threshold for the legal stay time in this area preset by the management backend (usually based on the equivalent benchmark of 24 to 48 hours set by health administrative regulations in seconds). This represents the timeout tolerance decay coefficient, whose physical purpose is to control the steepness of the divergence of the penalty curve after it crosses the threshold. It is usually dynamically set to 3600 to 7200 seconds depending on the type of regulatory object with different risk levels. The smoothing term is a very small constant, and its value is set to 10. −4 This is used to prevent algebraic overflow exceptions caused when the penalty coefficient configuration approaches zero or the input parameters are abnormal, ensuring that the denominator of the division operation does not approach 0. The physical purpose of the maximum value function nested outside the above formula is to enforce zero penalty truncation on compliant behaviors that have not timed out, while imposing a nonlinear exponential amplification penalty that conforms to the natural physical diffusion law on delayed behaviors that have indeed timed out.
[0067] S43, the system synchronously performs deviation assessment of the spatial flow path. Based on the principle of spatial topological connectivity physical constraints, certain high-risk flow tasks (such as the inter-ward transfer of infectious medical waste) must strictly follow a pre-set dedicated physical isolation route. The system extracts the actual sequence of on-site nodes traversed by the target and calculates its topological difference from the standard compliant path in the adjacency matrix of a static directed acyclic graph. Path deviation penalty component. The calculation formula is as follows: In the formula, This indicates the path deviation penalty component used for violation determination; This represents the shortest physical connectivity deviation between the actual flow node and the standard-specified path node in the spatial topology; This parameter represents the standard deviation of the route drift tolerance. Its value is determined based on the passage width of the physical building corridor of the medical institution and the reasonable avoidance buffer zone, and is usually in the range of 5 to 10 meters. This is a spatially minimal constant damping term, typically set to 10. −5 The scale is designed to prevent potential breakage due to division-by-zero singularities caused by the failure of the standard deviation setting in extremely coincident coordinate systems. This calculation mechanism effectively eliminates minor coordinate jitter interference caused by local multipath effects of radio frequency signals, ensuring that the system only penalizes actions involving crossing substantially prohibited areas or illegal detours.
[0068] S44, the system integrates neural network temporal prediction with the aforementioned multi-dimensional deterministic penalty features to calculate the comprehensive violation trigger probability, thus completing the closed-loop drive of the finite state machine. Based on the multi-source evidence chain cross-verification logic, the comprehensive violation judgment system can effectively avoid systemic false alarms caused by sudden failures of a single edge sensing hardware. Comprehensive violation probability. The calculation formula is as follows: In the formula, This represents the overall probability of triggering a violation that controls the transitions and jumps of the finite state machine. The foregoing The probability value of the predicted anomaly in the sequence of business outputs by the model; and These correspond to the normalized weighted adjustment coefficients for the sequence prediction component, timeout penalty component, and path deviation component, respectively. To ensure the consistency of the joint calculation dimensions, the system enforces verification and constrains the sum of the above three coefficients to be strictly equal to 1. The selection of specific input weighting parameters is based on the physical hazard mapping of various regulated objects. For example, for medical waste transportation events with a high risk of contact transmission, the system's underlying scheduling logic will automatically increase the weight of the spatial path deviation component. The ratio of [value] is used to strengthen the monitoring tolerance control of physical isolation in space.
[0069] When the comprehensive violation probability is calculated in real time When the preset business blocking threshold is exceeded (this threshold is typically set within a decision range of 0.75 to 0.85 based on the effective recall rate distribution curve of historical law enforcement cases), the finite state machine forcibly terminates its current internal state (including normal path deduction flow state or suspended waiting state triggered by visual tracking chain breakage) and directly jumps to the final state violation alarm node. The cloud server then truncates the current flow context of the target in memory, driving the data storage module to generate an immutable snapshot of violation event evidence carrying an absolute network time protocol signature. The system pushes a high-priority abnormal alarm data packet containing key video feature slices, device underlying hardware identifiers, and on-site topology anchor coordinates to the management backend and associated mobile law enforcement terminals, and forces administrative law enforcement personnel to perform on-site intervention and document solidification through the remote verification module. For the protocol encapsulation of the underlying binary format of the alarm data packet and the distribution and subscription mechanism of the network message queue, those skilled in the art can refer to the standard data middleware interface specification for implementation, which is a well-known technology in the field and will not be elaborated here.
[0070] Based on the violation trigger probability and state transition results output by the finite state machine, this embodiment provides an automatic splicing and alignment mechanism for cross-domain joint evidence files of violations in order to integrate discrete sensing data distributed across multiple physical nodes into legally valid solidified evidence. The mechanism includes the following steps: S45, in this embodiment, the cloud server receives the violation alarm command and the target's unique identification code sent by the finite state machine, and extracts the node sequence of the target flowing in the directed acyclic graph. Based on the physical identifier and absolute timestamp in the node sequence, the system sends an asynchronous data retrieval request to the data storage module. The first-stage source data set extracted includes original surveillance video clips and topological anchor point scanning records (photos and multimedia data collected on-site will be added and merged in subsequent law enforcement confirmation stages). For the indexing and cross-table join query mechanism of relational databases, those skilled in the art can refer to the standard Structured Query Language specification for implementation; this is a well-known technology in the field and will not be elaborated upon here.
[0071] S46, the system performs cropping and alignment on the retrieved multi-source heterogeneous video segments based on a unified timeline. Considering the differences in frame rate configurations and video file segmentation and packaging strategies among different camera devices in a distributed network environment, simply relying on file creation time cannot achieve precise frame-level alignment. The system utilizes the absolute reference timestamp calibrated by the network time protocol in the aforementioned embodiment to map the local playback time of each video segment to a globally unified physical timeline. To achieve accurate evidence extraction, assume the target flow path includes... For the valid video capture node, for the _ ... For each video capture node, the system extracts the absolute timestamps of the target entering and leaving the node's field of view, and calculates a cropping time window for extracting valid video evidence. Based on this cropping time window, the system starts from the [node name missing]. Extract the corresponding first video from the original video stream uploaded by the first video capture node. A video clip. The formula for calculating the time-cropping window is as follows: In the formula, Indicates that for the first The time clipping window range of each video capture node; and These represent the target entering and leaving the first stage, respectively. The absolute timestamp of the field of view of each video capture node; This indicates the preset buffer time offset, set to 3 to 5 seconds, to preserve complete environmental background features.
[0072] To filter out invalid video frames, the system introduces a video quality assessment network based on a 3D convolutional architecture (3D-CNN). This model internally consists of an input layer, three cascaded spatiotemporal 3D convolutional and max-pooling layers, a global average pooling layer, and a fully connected output layer. The input data dimension is set to [missing information]. The fully connected layer outputs a continuous scalar representing the probability of feature clarity. The model is trained using a binary cross-entropy loss function. The system performs a secondary screening of the cropped video slices based on the output probability state combined with a set quality threshold (set to 0.75).
[0073] S47, based on the principles of non-linear video editing and temporal rendering, performs temporal stitching and rendering splicing on multiple cropped video segments to address spatial blind spots. When the target crosses two physical nodes without surveillance coverage, there is an objective discontinuity in the video evidence chain on the timeline. To accurately reflect the physical displacement time of the target in the joint evidence file, while ensuring the continuity of playback, the system in the... The video clip and the first Artificially generated transition frame sequences are inserted between video clips. The lower-level features of each transition frame are specifically implemented as a sequence of solid-color background image frames containing the physical coordinates of the start and end points, a time elapsed counter, and path text for the nodes of a directed acyclic graph. The system calculates the duration of this transition frame sequence. The formula for calculating the transition playback duration is as follows: In the formula, This indicates the actual playback duration of the inserted transition frame sequence in the spliced video; Indicates the target leaving the preceding sequence The absolute timestamp of each video capture node; This indicates that the target has entered the subsequent stage. The absolute timestamp of each video capture node; This indicates the maximum allowed black screen duration. The purpose of this parameter is to prevent excessive expansion of the generated spliced video file size due to the target remaining in the physical blind spot for an extended period. Its value is typically set to 5 to 10 seconds based on the buffering characteristics of conventional media players. A minimum value function is introduced at the outer layer of the formula. The logic is that when the actual physical elapsed time is less than the upper limit, the transition frame is rendered according to the 1:1 ratio of the actual time; when the actual elapsed time exceeds the upper limit, the rendering time is forcibly truncated.
[0074] To quantify the playback rate distortion caused by spatiotemporal compression in the blind zone, the system further calculates the virtual fast-forward ratio during the transition phase, as shown in the following formula: In the formula, The numerator represents the time-lapse rate of the video after temporal compression; the numerator represents the actual time spent in physical space; and the denominator represents the actual rendering time allocated to the video file. The rendering time smoothing constant set for the system (with a constant value range of 10). −5 (Orders of magnitude) to handle division-by-zero exceptions when rendering time approaches zero. The system applies this scaling factor to transition frame images. Added text markers to indicate the passage of time during fast-forwarding, in order to maintain the compactness of the video file and the rigor of the evidence.
[0075] S48, based on the principles of multimodal evidence cross-verification and cryptographic solidification, extracts structured on-site attribute data uploaded by mobile law enforcement terminals and performs cross-modal data evidence solidification and tamper-proof hash signing. The system uses multimodal weighted logic for comprehensive reliability determination, calculated using the following formula: In the formula, The overall quality score of the joint evidence files; and These represent the validity of the spliced video, the completeness of the metadata, and the legality of the signature verification, respectively. It is the corresponding normalized weight coefficient, and its sum is constrained to 1.
[0076] When the overall quality score After the preset acceptance threshold is met, the system uses a cryptographic digest algorithm to calculate the comprehensive evidence hash value, and the calculation formula is as follows: In the formula, This represents the calculated fixed-length aggregated evidence hash value; This indicates that the system uses a 256-bit secure hash algorithm. This represents the underlying binary data source of the cross-domain complete video stream generated in the preceding steps; The extracted structured metadata of the violation event includes the target identity code, the text of the pre-defined regulatory rule clauses that triggered it, and the absolute timestamp sequence of each node. This represents the electronic signature data stream generated by the mobile law enforcement terminal through its embedded security encryption unit; symbol This indicates that a sequential concatenation operation is performed on the aforementioned multimodal underlying byte stream. For the specific compression mechanism of the secure hash algorithm, those skilled in the art can refer to well-known cryptographic standards and specifications, which are well-known technologies in the field and will not be elaborated upon here.
[0077] S49, the system encapsulates multimodal data and signature hash values to generate standardized joint evidence case files. These joint evidence case files consist of spliced video files, structured data description documents, and a preliminary draft data package of violations to assist on-site law enforcement. The system pushes the file to the data storage module for archiving and sends the file's absolute access path (e.g., / mnt / evidence_storage / 2026 / 03 / case_WF20260311.zip) and retrieval instructions to the management backend. Upon receiving the instructions, the management backend renders a case details tree diagram on its display interface for supervisory personnel to perform subsequent operations.
[0078] Based on prior compliance assessment and the establishment of a video evidence chain, this embodiment triggers on-site verification and document generation processes on the mobile terminal, realizing a closed-loop management system from online monitoring to offline handling. Specific steps include: Task Issuance and Parsing: The cloud server encapsulates digital evidence files containing the target's unique identifier, a summary of the violation, video evidence slices, and the target's last known spatial coordinates, along with on-site verification instructions, into a task data package, which is then issued via an encrypted link. The mobile law enforcement terminal parses this data package within its local security sandbox and displays the details of the event to be verified.
[0079] On-site spatial location comparison: The mobile law enforcement terminal acquires the current spatial location coordinates and verifies the distance with the last known coordinates of the target. In outdoor environments, the distance to the ground is calculated based on the spherical cosine theorem using GPS coordinates; in indoor environments, the three-dimensional spatial coordinates are extracted by scanning nearby topological anchor tags, and the spatial deviation is calculated based on the three-dimensional Euclidean distance. The system determines whether the actual spatial deviation is less than or equal to a preset effective radius threshold for on-site verification (e.g., 10 to 50 meters).
[0080] On-site evidence collection and image evaluation: After unlocking permissions, the terminal acquires supplementary multimedia evidence from the scene. The acquired raw images, after scaling and normalization preprocessing, are input into a built-in image quality and relevance evaluation model based on a deep residual convolutional network (internally containing an input layer, a cascaded residual convolutional module, a global average pooling layer, and a multi-label classification output head). The model outputs a confidence vector representing image sharpness and scene consistency. This model is generated through backpropagation training using a historical image sample set with cross-entropy as the loss function.
[0081] Data-level fusion and evidence solidification: After the image evaluation is qualified, the system executes time alignment logic to ensure that the deviation between the supplementary evidence generation time and the standard network timestamp of the trusted timestamp server is within a tolerance window (e.g., 500 milliseconds). Subsequently, using a one-way hashing and cascading authentication mechanism, the initial hash digest value of the original digital evidence file, the hash digest value of the supplementary evidence data sequence collected on site, and the standard network timestamp are concatenated and spliced into a data byte stream, and the final joint evidence hash digest is generated uniformly through a secure hashing algorithm.
[0082] Automated document generation: The terminal retrieves local law enforcement document templates based on the joint evidence summary, parses data placeholders, and maps and fills the corresponding data nodes with the elements of the draft facts of violations, time and location attributes, and anti-counterfeiting verification codes of solidified evidence in the task data package, thereby generating structured electronic documents.
[0083] Digital Signature and Archiving: Personnel use their personal digital certificate private key stored in a secure chip to encrypt the global data feature hash value of the electronic document, generating a digital signature data segment. The system embeds this digital signature data segment, public key certificate data, and digital timestamp object together at the end of the electronic document, generating a closed-loop electronic document, and then sends it back to the business supervision system for unified archiving.
[0084] In this embodiment, the dynamic risk index calculation and four-color early warning mapping model for medical institutions is used to convert multi-dimensional operational characteristics of medical institutions into quantitative risk values and determine the safety alarm level. The specific execution process includes the following steps: S51 collects a multi-dimensional operational data sequence from medical institutions within a preset time window. This data sequence includes administrative regulatory violation characteristic data extracted in the previous steps (such as basic file deductions and AI violation frequency), and simultaneously integrates medical resource load status data, patient visit flow data, and medical equipment operation status data. The system introduces timestamp alignment and interpolation completion mechanisms to ensure the logical consistency of multi-source heterogeneous data at the same time point. The system performs data cleaning, outlier removal, and extreme value normalization on the acquired multi-dimensional operational data sequence. A non-zero bias micro-value (such as 10) is introduced during the normalization process. −5 To avoid division by zero anomalies, a standardized data matrix is output.
[0085] S52 calculates the dynamic weight coefficients corresponding to each data feature based on a standardized data matrix. The system assigns weights to data features based on the principle of information entropy variation, identifies and amplifies sensitive indicators with drastic fluctuations, and introduces a time decay factor to adjust the ratio of historical data to current data in the weight allocation. The specific calculation process for determining the basic weights of each evaluation indicator using the entropy weight method or principal component analysis can be derived by those skilled in the art based on the degree of indicator variation; the weight allocation logic is well-known in the field and will not be elaborated here.
[0086] S53, based on a standardized data matrix and dynamic weighting coefficients, uses a risk calculation model to solve for the dynamic risk index of medical institutions. This risk calculation model integrates the instantaneous risk status and historical risk accumulation trends of medical institutions, and its calculation formula is as follows: In the formula, This indicates that medical institutions are at all times The dynamic risk index ranges from 0 to 100. The weighting constant representing the impact of instantaneous risk is set to between 0.4 and 0.6. Indicates the total number of evaluation indicators; Indicates the first Each evaluation indicator at time Dynamic weighting coefficients; Indicates the first Each evaluation indicator at time Standardized values; This represents the weighting constant for the historical cumulative risk impact, and satisfies... ; This indicates the preset historical backtracking time window length, set to 4 to 12 hours from the past. This represents the time decay constant, with a value ranging from 0.01 to 0.1. Representing historical moments The instantaneous risk status assessment value.
[0087] To obtain the instantaneous risk status assessment value The system uses a long short-term memory network. The model performs temporal feature extraction and state evaluation. The input data dimension of the model is defined as follows: The model uses a standardized feature vector and its corresponding timestamp sequence, and extracts historical segments by setting a sliding window (e.g., a 15-minute step size). Internally, the model consists of an input layer, two layers with 128 hidden units each, and so on. The model consists of stacked layers, fully connected layers, and a sigmoid output layer, ultimately mapping and outputting a one-dimensional risk probability value to characterize the probability of resource overload in healthcare institutions. Model training uses a historical operational database as samples, labeling periods of healthcare overload as positive samples and stable operational periods as negative samples. The training process employs binary cross-entropy as the loss function, uses the Adam optimizer for parameter updates, and includes an early stopping mechanism.
[0088] S54. The dynamic risk index is input into the four-color early warning mapping model, and the risk warning level is determined by comparing it with the preset risk threshold range. The model has low-risk, medium-risk, and high-risk thresholds, which are obtained based on the intersection points of historical operational characteristic distribution curves. The mapping logic is as follows: when the dynamic risk index is less than the low-risk threshold, it is mapped to a green warning; when it is greater than or equal to the low-risk threshold and less than the medium-risk threshold, it is mapped to a yellow warning; when it is greater than or equal to the medium-risk threshold and less than the high-risk threshold, it is mapped to an orange warning; and when it is greater than or equal to the high-risk threshold, it is mapped to a red warning. For the dynamic division of the warning thresholds, those skilled in the art can use clustering algorithms such as K-Means to cluster historical risk index samples to determine boundary values. The clustering process is a well-known technique in this field and will not be described in detail here.
[0089] S55 triggers corresponding emergency response actions based on the warning color level output by the four-color warning mapping model. When a green warning is triggered, the normal medical resource allocation mechanism is maintained; when a yellow warning is triggered, a risk warning message is sent and the backup medical supply warehouse is pre-checked; when an orange warning is triggered, the number of on-duty medical staff is increased, backup treatment wards are opened, and waiting patients are guided to related medical institutions; when a red warning is triggered, the highest level of emergency response is activated, the admission of non-critical patients is suspended, a distress signal is sent to the higher-level health administration department, and the emergency center is coordinated to dispatch ambulances for inter-hospital transfer and diversion.
[0090] In this embodiment, the macro-data dashboard is used to visually present the system's operating status, dynamic risk quantification results, and evolution trends, providing decision support for control personnel. The method specifically includes the following steps: S61, Obtain the dynamic risk quantification data set. As a preferred method, the control node receives indicator parameters output by the risk warning module via the data bus, specifically including the frequency of violations in each physical area, the probability of medical resource overload, and the overall dynamic risk index of the institution. The system adds a unified timestamp to each parameter based on the Network Time Protocol (NTP) and establishes a synchronous cache queue with fixed time windows. The system's underlying layer uses a WebSocket bidirectional communication protocol based on Transmission Control Protocol (TCP) to establish a data stream channel, periodically pulling risk quantification data in the form of data frames.
[0091] S62 performs visual feature mapping of multidimensional data. The acquired risk quantification data has multidimensional attributes. Based on the general principles of visual encoding in information visualization, the system constructs a visual mapping model to convert abstract data dimensions into geometric attribute parameters recognizable by the rendering engine. Since risk values often exhibit a continuous distribution in business scenarios, in specific implementation, the system maps the real-time risk value of a node to a color vector and scaling factor for the rendered primitive. The color mapping process is implemented based on a linear interpolation algorithm, and its mapping formula is: ,in, Represents a node The corresponding target color vector, whose value range in the RGB color space is usually defined as a three-dimensional array of [0, 255] for each channel; Represents a node The current risk assessment value; and These represent the lower and upper risk thresholds set by the system, respectively. The determination of these two thresholds is based on the statistical extreme values under the historical operating conditions of the system or the safe operating boundary of the physical equipment itself. This represents the starting color vector corresponding to the lower risk limit; This represents the termination color vector corresponding to the risk ceiling. In the above calculation steps, to avoid the denominator being incorrect due to the same upper and lower limit values... When the value approaches 0, triggering a division-by-zero exception in the underlying operation, the system has built-in protective capture logic: when At that time, interpolation operations are no longer performed, and the value is directly set to... The technical purpose of this mapping operation is to enable controllers to intuitively perceive the evolution of local risks in the system in real time from the smooth gradation of visual colors.
[0092] S63 invokes the graphics rendering engine to generate the view layer of the macro data cockpit. The front-end system constructs a 3D scene tree, converting the underlying network topology into mesh nodes and connected objects in the scene tree. The system employs instanced rendering technology, merging node primitives with the same geometric topology but different spatial locations and color attributes into a single draw call to reduce the overhead of graphics instruction submission to the central processing unit. For the processing of the 3D graphics rendering pipeline based on WebGL technology, those skilled in the art can use programmable shaders to implement the spatial matrix transformation of vertex coordinates and pixel fragment shading. Its 3D projection transformation and depth testing are well-known technologies in the field and will not be elaborated here.
[0093] S64 generates and overlays a decision-aid rendering layer. On top of the basic 3D topology view, the system overlays decision-aid information based on preset business logic. Considering that instantaneous extreme value deviations of a single node may cause false alarms due to sensor noise or network jitter, the triggering of the early warning rendering mechanism no longer relies solely on a single risk feature value, but introduces multi-dimensional weighted logic. The system calculates the risk deviation degree of the current operating state to determine the alarm level and generate corresponding visual feedback instructions. The risk deviation degree calculation formula is: ,in, Represents a node Risk deviation; This indicates the safety threshold of the region where the node is located; The smallest positive real number preset by the system (e.g., 10) −5 This is used to ensure that the formula maintains strict mathematical completeness when the safety threshold is 0; Represents a node The importance weight in the current network topology is calculated based on a combination of node degree centrality and current service capacity, with a normalized value range of [0,1]. The physical meaning of this calculation model is to accurately quantify the relative severity of the risk spillover of the current node to the system's tolerance benchmark, and to amplify it by combining its criticality in the global network topology.
[0094] Based on the above multi-dimensional logical judgment, the system follows... Within the specified numerical range, the underlying particle system is invoked to render dynamic flickering effects or outward-spreading radiating ripple effects at the corresponding node's three-dimensional coordinates. To further reduce the risk of biased judgment, when... If the risk level exceeds the set high-risk blocking threshold for an extended period of time, the system will automatically generate a two-dimensional floating window layer containing the physical area number involved, the type of violation, and the emergency response strategy. Simultaneously, it will be rendered on the top layer of the screen interface using orthogonal projection, guiding operators to make manual interventions through an unobstructed visual interaction layer.
[0095] Specific application examples: This example uses the infectious medical waste transfer operation of a certain tertiary general hospital on a certain day of a certain month of a certain year as the background.
[0096] The system deployed the aforementioned intelligent monitoring platform on edge computing nodes and cloud servers. Medical waste transport personnel (identification code: WF-20260311) are responsible for transporting infectious medical waste from the temporary storage room on the third floor of the inpatient department (physical node). The data was transferred to the centralized collection station (physical node) at the bottom of the hospital ground floor. During the transfer process, the system triggered a cross-view tracking and human-machine collaborative integration mechanism because the route passed through a freight elevator area without surveillance coverage.
[0097] Specific application process and core calculations: Visual tracking and chain break detection: The transfer personnel entered the third-floor corridor of the inpatient department (node) When monitoring the field of view, the video intelligent analysis engine extracts its high-dimensional feature vector. After entering the freight elevator and leaving the field of vision, record the time. When a candidate target is at the bottom-level freight elevator exit (node) When a feature vector appears, the system extracts a high-dimensional feature vector. and record the time. .
[0098] Visual similarity calculation: Let's assume the vector dot product after extraction and normalization. ,correspond norm Smoothing constant The system calculates visual similarity based on a formula. ; Comprehensive suture confidence calculation: node To node Historical average transfer time seconds, smoothing coefficient Seconds. The transport personnel carried yellow medical waste bins and wore prescribed uniforms; attribute matching score. Preset normalized weight parameters: The system extracts the time difference after network time protocol alignment. Seconds. Calculate the overall suture confidence level according to the formula. ; Substitute the values:
[0099] The calculated result of 0.642596 is lower than the system's preset association matching threshold of 0.75. The system determines that a visual tracking chain has been broken and the finite state machine enters a suspended state.
[0100] Seamless integration of mobile terminal interaction and human-computer collaboration: The mobile enforcement terminal scans topological anchor point tags in the freight elevator exit area. Environmental path loss index during scanning. Smooth signal strength Reference parameters RF attenuation compensation .
[0101] Distance estimation calculation: rice; If the estimated distance falls within the legally adjacent range, the system allows the interaction. Law enforcement officers confirm the business attributes and upload the data. The cloud server constructs a dynamic virtual node and performs human-machine collaborative stitching of confidence levels. Calculate (assuming the weighted result of the spatial, temporal, and attribute confidence components here satisfies the 0.80 threshold), and wake up the finite state machine.
[0102] Violation probability assessment and document generation: Due to excessive processing time, the system introduces penalty logic. (Stay-at-home difference) Seconds, legal limit Seconds, tolerance coefficient Second, .
[0103] Calculation of penalty for overstaying: ; ; Assume the model predicts the probability of violation. Path deviation component Weight parameters .
[0104] Overall probability of violation calculation: ; ; The state machine records the violation index. If the cumulative violation index from multiple nodes exceeds the threshold, the system retrieves the template via the mobile law enforcement terminal, generates a joint evidence file with an electronic signature, and saves the complete ZIP package to / mnt / evidence_storage / 2026 / 03 / case_WF20260311.zip.
[0105] Experimental verification and effect comparison: To verify the actual operational effectiveness of the intelligent monitoring system of this invention, 30 days of operational data (a total of 12,500 workflow tasks) were extracted from the hospital and compared with a traditional video analysis system that did not introduce multimodal spatiotemporal alignment and penalty components.
[0106] Combined with appendix Figure 3 As can be seen, the system of this invention has achieved a significant generational improvement in core performance indicators. In this figure, the light gray bars represent traditional video analysis systems, and the dark gray bars represent the system of this invention. Specifically, this invention, by introducing multimodal fusion computing, significantly improves the automatic link recovery rate from the traditional 42.5% to 89.4%; in terms of cross-domain feature matching, the matching accuracy jumps from 78.1% to 96.5%; at the same time, through the dual constraints of spatiotemporal dynamic confidence and penalty model, the false alarm rate of the system is effectively suppressed, sharply reduced from 15.2% to 3.8%.
[0107]
Claims
1. A smart supervision method for medical institutions based on mobile terminals and video cloud monitoring, characterized in that: Includes the following steps: The cloud server receives daily supervision tasks dispatched by the management backend or alarm information generated by the video cloud supervision platform through analysis of the original video stream, and generates verification tasks. The cloud server uses the topology state calculation module to perform target trajectory deduction based on video visual tracking in the constructed directed acyclic graph and finite state machine; When visual tracking is determined to be interrupted, the physical node corresponding to the interruption location is recorded, the finite state machine enters a suspended state, and the cloud server sends the corresponding alarm video clip to the mobile law enforcement terminal. The mobile law enforcement terminal scans the topological anchor point labels of the corresponding area, extracts the physical coordinates and business attributes as data to be reported to the cloud server. The cloud server generates dynamic virtual nodes in the directed acyclic graph based on the data entered, and calculates the confidence level of human-machine collaborative stitching between the dynamic virtual nodes and the broken physical nodes. When the preset conditions are met, the trajectory stitching is completed, and the finite state machine is awakened to determine the violation. After the finite state machine outputs the violation determination result, the cloud server extracts the corresponding alarm video stream segment and generates a joint evidence file with the data filled in by the mobile law enforcement terminal.
2. The intelligent supervision method for medical institutions based on mobile terminal and video cloud supervision according to claim 1, characterized in that, The steps of performing target trajectory deduction in the constructed directed acyclic graph and finite state machine, and determining the occurrence of visual tracking chain breakage, include: The cloud server calculates a dynamic time window with upper and lower bounds based on the spatial straight-line distance between adjacent physical nodes in the directed acyclic graph and the average speed of personnel movement. High-dimensional feature vectors adjacent to each other in the time series are extracted. Under the premise of satisfying the dynamic time window constraint, visual similarity is obtained by dot product operation and norm calculation. A visual feature smoothing constant is introduced in the calculation to smooth the denominator and prevent singularity. When the maximum visual similarity of all candidate targets in the network is consistently lower than the preset confidence threshold, the visual tracking chain is determined to have broken.
3. The intelligent supervision method for medical institutions based on mobile terminal and video cloud supervision according to claim 1, characterized in that, The mobile law enforcement terminal scans the topological anchor tags of the corresponding area, including: The mobile law enforcement terminal establishes a near-field interactive connection with the topological anchor tag through a radio frequency antenna and an optical sensor. The physical space distance between the terminal and the tag is estimated by extracting the Bluetooth signal strength of continuous sampling periods, and a small amount of radio frequency attenuation compensation is introduced to smooth the denominator. The anti-counterfeiting comprehensive judgment is made by combining the estimated distance in physical space, the acceleration variance features extracted by the inertial measurement unit, and the decoding time index of the encrypted QR code read by the optical sensor. When the verification is successful and the estimated distance in physical space falls within the physical adjacency safe range, the extraction of the physical coordinates and business attributes is allowed.
4. The intelligent supervision method for medical institutions based on mobile terminal and video cloud supervision according to claim 1, characterized in that, The steps for calculating the confidence level of human-machine collaborative stitching between the dynamic virtual node and the disconnected physical node include: Based on the spatial connectivity distance between broken physical nodes and dynamic virtual nodes extracted from the extended adjacency matrix, the spatial confidence component is calculated using the Gaussian kernel function. The dynamic time deviation is calculated by extracting the absolute timestamp parameter and combining it with the average movement speed of personnel and the spatial connectivity distance. The time confidence component is calculated by using an exponential decay function that includes the absolute value and the base of the natural constant. Structured business attributes are extracted, and confidence components of business attributes are calculated through a hybrid matching of Boolean logic and vector similarity. The spatial confidence component, temporal confidence component, and business attribute confidence component are assigned weight factors and linearly weighted and summed to obtain the human-machine collaboration stitching confidence score used to drive the state machine flow.
5. The intelligent supervision method for medical institutions based on mobile terminal and video cloud supervision according to claim 1, characterized in that, The steps for waking up the finite state machine to determine violations include: Alignment processing is performed on the timestamps of consecutive related node sequences, and feature sequences are extracted and input into the long short-term memory network model to output the probability value of business anomaly prediction for the sequence. Based on the difference between the actual cumulative dwell time of the target within the preset monitored physical area and the upper limit threshold of the legal dwell time, the overstay penalty component is calculated using a nonlinear amplification algorithm with truncation properties. The path deviation penalty component is calculated based on the shortest physical connectivity deviation distance between the actual flow node and the standard compliant path node in the spatial topology. The abnormal prediction probability value of the sequence service, the delay timeout penalty component and the path deviation penalty component are normalized and weighted to obtain the comprehensive violation probability. When the comprehensive violation probability exceeds the preset service blocking threshold, the finite state machine is triggered to jump to the final state violation alarm node.
6. The intelligent supervision method for medical institutions based on mobile terminal and video cloud supervision according to claim 1, characterized in that, The step of extracting video stream segments and generating a joint evidence file from the data entered by the mobile law enforcement terminal includes: The cropping time window is calculated using an absolute reference timestamp calibrated by the Network Time Protocol, video segments are extracted from the original video stream, and a video quality evaluation network with a 3D convolutional architecture is used to filter effective video slices. Artificially generated transition frame sequences are inserted between adjacent video segments with physical blind spots. The virtual fast-forward ratio is calculated based on the actual physical time consumption and the maximum allowed black screen occupation time limit, and text marks are added to generate a spliced video stream. The underlying binary data source of the spliced video stream, the structured metadata of the violation event extracted based on the filled data, and the electronic signature data stream generated by the mobile law enforcement terminal are sequentially spliced together. A tamper-proof comprehensive evidence hash value is generated through a secure hash algorithm and packaged into the joint evidence file.
7. The intelligent supervision method for medical institutions based on mobile terminal and video cloud supervision according to claim 1, characterized in that, Before generating the dynamic virtual node, the mobile law enforcement terminal performs a cross-modal dynamic clock calibration step: Periodically initiate synchronization requests to the cloud server to obtain network round-trip latency and base time offset; Based on the inverse proportional mapping logic, the network round-trip delay is used as the base time offset to allocate confidence weight components, and the smoothed base time offset is calculated within a sliding window. The dynamic drift rate of the hardware clock is calculated based on the first-order difference principle. The local uncalibrated timestamp is superimposed with a linear compensation increment using the reference time offset and the dynamic drift rate. The calibrated timestamp is then output for packaging interactive mapping data packets.
8. The intelligent supervision method for medical institutions based on mobile terminal and video cloud supervision according to claim 1, characterized in that, The method also includes an early warning step based on the operational characteristics of medical institutions: Collect and standardize multidimensional operational data sequences of medical institutions that include characteristics of administrative regulatory violations and patient flow, and calculate dynamic weight coefficients for various evaluation indicators based on the principle of information entropy variation. A long short-term memory network model is used to evaluate the extracted historical operational feature segments and output instantaneous risk status assessment values; The dynamic risk index is calculated by integrating the instantaneous risk impact component calculated at the current moment with the historical cumulative risk component of the instantaneous risk state assessment value adjusted by an exponential decay function. The dynamic risk index is input into the four-color early warning mapping model, compared with the preset risk threshold range to determine the corresponding early warning color level and trigger the corresponding emergency response action.
9. The intelligent supervision method for medical institutions based on mobile terminal and video cloud supervision according to claim 1, characterized in that, After the cloud server generates the verification task, the process also includes a queue allocation step executed by the task scheduling module: The manual task data packets containing management backend instructions are merged with the task instances automatically triggered by video alarms and uniformly injected into the global pending task queue. The network time protocol reference time is extracted as the absolute reference axis, and the initial time identifiers of each source task entering the global pending task queue are synchronized and calibrated. For the tasks to be assigned in the global task queue, the static baseline risk and queuing time of the tasks to be assigned are extracted, and an exponential term containing the natural constant base is introduced to perform a nonlinear bounded mapping on the queuing time to obtain a time compensation component. The static baseline risk and the time compensation component are weighted and calculated to obtain the dynamic comprehensive priority of the tasks to be assigned. Based on the calculated dynamic comprehensive priority, the global pending task queue is reordered in descending order. The absolute spatial straight-line distance of idle mobile law enforcement terminals, the average time spent on historically completed tasks, and the current network signal strength are comprehensively evaluated to filter target mobile law enforcement terminals. High-priority tasks at the top of the global pending task queue are then pushed to the target mobile law enforcement terminal.
10. A smart monitoring system for medical institutions based on mobile terminals and video cloud supervision, characterized in that: The method for implementing the intelligent supervision of medical institutions based on mobile terminal and video cloud supervision as described in any one of claims 1 to 9 includes: The video cloud monitoring platform is deployed at network edge nodes or servers to access video streams from front-end devices, perform feature extraction and event recognition, and generate alarm information. The management backend is used to assign daily monitoring tasks and display data. The cloud server communicates with the video cloud supervision platform and management backend, and is used to receive task or alarm information to generate verification tasks. The topology state calculation module is used to deduce the target trajectory in the constructed directed acyclic graph and finite state machine. When visual tracking is determined to be disconnected and the finite state machine enters a suspended state, an alarm video clip is issued; based on the reported data, dynamic virtual nodes are generated in the directed acyclic graph and the confidence of human-machine collaboration stitching is calculated to wake up the finite state machine; And after determining a violation, a joint evidence file is generated; The mobile law enforcement terminal is connected to the cloud server and is used to receive the alarm video clips, scan the topological anchor point tags at the scene, extract physical coordinates and business attributes, and report them to the cloud server.