A machine vision-based dispensing monitoring method and system
By using a multi-view polarized red-green-blue depth camera and dynamic neural radiation field technology, the problems of high light reflection and transparent object occlusion were solved, enabling high-precision drug dispensing monitoring and safety detection, and improving the interpretability and robustness of the drug dispensing process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHAANXI BAOFANG TECHNOLOGY CO LTD
- Filing Date
- 2026-04-30
- Publication Date
- 2026-06-26
AI Technical Summary
Existing technologies struggle to effectively address issues such as target loss, measurement inaccuracies, and poor interpretability caused by high light reflection, transparent objects, and dynamic occlusion in complex environments, especially in centralized intravenous medication preparation scenarios, which can compromise medication safety.
Multi-view polarized red-green-blue depth cameras are used to simultaneously acquire images. Through physical-level de-highlighting and transparent material area recognition, combined with dynamic neural radiation fields, the visual characteristics of occluded areas are restored, realizing a virtual monitoring view with de-occlusion and de-highlighting, and performing 3D reconstruction and security detection.
It enables effective perception of highly reflective and transparent objects, reduces false alarms and recognition interruptions caused by occlusion, improves the accuracy and operational interpretability of three-dimensional dose measurement, and enhances human-machine trust and safety.
Smart Images

Figure CN122290214A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of computer vision and medical information technology, and in particular to a drug dispensing monitoring method and system based on machine vision. Background Technology
[0002] In scenarios such as centralized dispensing of intravenous medications, automated monitoring based on machine vision is crucial for ensuring patient medication safety. Existing technologies can verify medications and dosages through methods such as object detection and text recognition, but they still face significant challenges in complex environments.
[0003] First, the drug preparation environment contains numerous transparent containers and reflective surfaces. Stainless steel countertops, glass ampoules, and syringes generate strong highlights, causing the target outline to break and making it difficult to detect liquid level boundaries. Traditional de-reflection methods, while removing highlights, often lose crucial dosage information beneath the transparent objects.
[0004] Secondly, dynamic hand occlusion is a long-standing problem. Frequent hand movements by dispensing personnel can obstruct critical targets such as medication labels and syringe markings. Most existing methods can only detect the presence of occlusion and issue an alarm, but cannot recover the visual characteristics of the obscured target, leading to frequent interruptions in the recognition process and loss of dosage readings.
[0005] Furthermore, the lack of depth information in two-dimensional images prevents the system from detecting safety indicators in three-dimensional space, such as whether the needle has touched the bottom of the ampoule. Ordinary multi-view 3D reconstruction techniques produce severe reconstruction artifacts when dealing with transparent materials.
[0006] Finally, when an alarm occurs, the existing system's visual monitoring has poor continuity and interpretability, failing to provide the reviewing pharmacist with a clear, unobstructed view of the erroneous operation, making manual review difficult and reducing human-machine trust. Summary of the Invention
[0007] The technical problem to be solved by the present invention is to provide a drug dispensing monitoring method and system based on machine vision to address the shortcomings of the prior art, thereby solving the problems of target loss, inaccurate measurement and poor interpretability caused by high light reflection, transparent objects and dynamic occlusion in the prior art.
[0008] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a drug dispensing monitoring method and system based on machine vision.
[0009] This invention proposes a machine vision-based method for monitoring medication dispensing, comprising: Acquire multi-view polarization image sequences and corresponding depth image sequences of the drug dispensing operation area; Physical-level de-highlighting is performed on the multi-view polarized image, and transparent material regions are identified to generate multimodal fusion data with transparency prior. Real-time detection of the operator's hand posture generates a dynamic occlusion mask; A dynamic neural radiation field is constructed and trained. The training process introduces temporal de-occlusion prior loss to utilize the unoccluded spatiotemporal information to recover the visual features of the area marked by the dynamic occlusion mask. Based on the trained dynamic neural radiation field, the hand area is removed, and a virtual monitoring view with de-occlusion and de-highlighting is generated.
[0010] Furthermore, the acquisition of the multi-view polarization image sequence and the corresponding depth image sequence of the drug dispensing operation area specifically includes: At least three polarized red-green-blue depth cameras are deployed in a hemispherical shape around the drug dispensing area to simultaneously acquire RGB images, polarization degree images, polarization angle images, and depth maps.
[0011] Furthermore, the physical-level de-highlighting processing of the multi-view polarized image and the identification of transparent material regions therein to generate multimodal fusion data with transparency prior specifically includes: Based on the Fresnel reflection model, the diffuse reflection component and specular reflection component are separated using the polarization angle image to generate a diffuse reflection map without specular highlights. Adaptive threshold segmentation is performed on the polarization image, and the transparent material region is identified and assigned an initial transmittance weight by combining the edge information of the depth map. The diffuse reflection map without specular highlights, the initial transmittance weights, and the depth map are fused at the channel level to form the multimodal fused data with transparency prior.
[0012] Furthermore, the real-time detection of the operator's hand posture and the generation of a dynamic occlusion mask specifically includes: A spatiotemporal graph convolutional network for hand key points is used to estimate the key points of the operator's hands in real time and project them onto the imaging planes of each camera to generate an initial occlusion mask. Using the bidirectional optical flow method, the initial occlusion mask is propagated between consecutive frames to compensate for the mask boundary offset caused by rapid hand movements, thereby generating a time-consistent dynamic occlusion mask.
[0013] Furthermore, the construction and training of a dynamic neural radiation field includes: The dynamic neural radiation field is constructed by taking the three-dimensional coordinates of the spatial sampling point, the observation direction, the timestamp, the polarization feature vector, and a transparent material indicator variable as inputs, and taking the volume density of the point, the polarization-sensing color related to the observation direction, and the transparency coefficient as outputs.
[0014] Furthermore, the temporal deocclusion prior loss is specifically as follows: For pixels in the current frame that are marked as occluded by the dynamic occlusion mask, obtain the observation information of their corresponding spatial points at adjacent unoccluded times or adjacent unoccluded viewpoints; By using stereo matching, an approximate realistic texture of the occluded pixel is generated using the observation information; Using the approximate realistic texture as a supervision signal, the loss is calculated, driving the dynamic neural radiation field to learn and complete the geometric and appearance information of the occluded area.
[0015] Furthermore, the training process also introduces a refraction correction loss for transparent material regions. This refraction correction loss is based on a physical refraction model and constrains the propagation path of light at the interface of transparent objects in order to accurately reconstruct the three-dimensional position of the internal structure of the transparent object.
[0016] Furthermore, the step of removing the hand region and rendering a de-occluded and de-highlighted virtual monitoring view based on the trained dynamic neural radiation field specifically includes: To obtain real hand posture information, during the spatial sampling process of the dynamic neural radiation field, the density field area corresponding to the hand is directly covered. Set up a virtual camera view for medication dispensing monitoring; The light rays along the virtual camera's viewpoint are sampled, and the density field after the hand is covered is used for volume rendering to generate the de-occluded and de-highlighted virtual monitoring view. The volume rendering is accelerated using multi-resolution hash encoding and model distillation techniques.
[0017] Furthermore, it also includes: Based on the depth map corresponding to the virtual monitoring view, the syringe is segmented into instances to identify its outer wall, piston and liquid column area. Combined with the known syringe geometric dimensions, the liquid column area is reconstructed in three dimensions, and the actual liquid volume is calculated by volume integration. In addition, the depth difference between the syringe needle tip and the bottom of the medicine bottle is calculated based on the depth map, and a bottoming risk alarm is issued when the depth difference is less than a safety threshold; And / or, based on the identified sequence of key points of the operator's hand bones and the reconstructed three-dimensional model of the drug in the virtual monitoring view, a spatiotemporal motion comparison is performed to identify the compliance of the drug dispensing operation steps.
[0018] Based on the same inventive concept, this invention also proposes a machine vision-based medication monitoring system, comprising: The multi-view polarization-depth acquisition module is used to simultaneously acquire multi-view polarization image sequences and corresponding depth image sequences of the drug dispensing operation area; The multimodal preprocessing module is used to perform physical-level de-highlighting on the multi-view polarization image, identify transparent material regions, and generate multimodal fusion data with transparency prior. The hand posture sensing module is used to detect the operator's hand posture in real time and generate a dynamic occlusion mask; The neural radiation field training and reconstruction module is used to construct and train a dynamic neural radiation field. Its training process introduces temporal de-occlusion prior loss and refraction correction loss to recover the visual features of the occluded area using the unoccluded spatiotemporal information and accurately reconstruct the internal structure of transparent objects. The virtual view rendering and monitoring module is used to remove the hand area based on the trained dynamic neural radiation field, render and generate a virtual monitoring view that is de-occluded and de-highlighted, and perform drug dispensing verification and risk alarm based on this view.
[0019] The present invention provides a drug dispensing monitoring method and system based on machine vision, which has the following advantages compared with the prior art: At the information acquisition level, this invention achieves effective perception of highly reflective and transparent objects in a medication dispensing scenario through simultaneous acquisition by multi-view polarized red-green-blue depth cameras and physical-level polarization processing. Specifically, based on the Fresnel reflection model, the diffuse reflection component and specular reflection component are physically separated using polarization angle images. This completely removes strong reflections from stainless steel surfaces and glassware surfaces while fully preserving key texture information such as drug label text and syringe markings that are obscured by light spots. Through adaptive threshold segmentation of the polarization degree image and combined with edge consistency constraints of the depth map, the invention achieves for the first time accurate calibration and transmittance weight assignment of transparent material areas such as glass ampoule walls and syringe barrels, making these objects, which are almost "invisible" in traditional vision systems, structured entities that can be processed by subsequent algorithms. Finally, the diffuse reflection image without high light, the transmittance weight map, and the depth map are fused at the channel level to form multimodal fusion data that combines texture, physical properties, and spatial structure information, providing the neural radiation field with initialization conditions far superior to those of a single modality input.
[0020] At the feature recovery level, this invention achieves a paradigm shift from "passive occlusion detection" to "active information recovery" through an innovative "temporal de-occlusion prior loss" mechanism and a temporally consistent dynamic occlusion mask generation method. For key targets such as drug labels and syringe markings obscured by the user's hand, the system no longer simply classifies them as "invisible" and triggers an alarm. Instead, it proactively utilizes multi-view spatiotemporal redundancy information—including unoccluded observations from different viewpoints at the same time, and unoccluded observations from adjacent times within the same viewpoint—to generate an approximate realistic texture of the occluded area through stereo matching. This texture is then used as a supervisory signal to drive dynamic neural radiation field learning to complete the geometric and appearance information of the occluded area. Simultaneously, by employing a spatiotemporal graph convolutional network of hand keypoints combined with bidirectional optical flow correction, the generated dynamic occlusion mask has clear boundaries and temporal consistency, effectively avoiding mask jitter caused by rapid hand movements. Thanks to this, in the entire medication dispensing process where the probability of the operator's hands covering the target exceeds 60%, the system can maintain continuous identification of key targets for more than 90% of the operation time, significantly reducing false alarms and identification interruptions caused by obstruction.
[0021] This invention, at the 3D reconstruction level, achieves sub-millimeter-level precise reconstruction of the internal structure of transparent objects by introducing refractive correction loss specifically designed for transparent materials into a dynamic neural radiation field. This supports clinical-grade 3D dose measurement and spatial safety monitoring. Traditional neural radiation fields suffer severe geometric distortion due to light refraction when encountering transparent objects. This invention utilizes a physical refraction model to explicitly constrain the propagation path of light at interfaces such as air-glass-liquid-glass-air, enabling unprecedentedly accurate reconstruction of the 3D positions of internal structures such as the liquid surface inside the ampoule and the drug column inside the syringe. Based on this precise 3D model, the system can automatically segment the outer wall, piston, and liquid column areas of the syringe. Combining this with known syringe geometric parameters, it directly performs volume density integration calculations on the liquid column segment, improving the measurement accuracy of drug volume to within 0.05 mL, meeting the stringent verification standards for high-risk drug preparation. Furthermore, by calculating the depth difference between the needle tip and the bottom of the ampoule based on the depth reconstruction results, a "high risk of bottom contact" warning can be issued in a timely manner to avoid the risk of insoluble particles being drawn into the syringe during medication administration. This three-dimensional spatial safety indicator is completely imperceptible to traditional two-dimensional monitoring systems.
[0022] At the application interaction level, this invention provides unprecedented interpretability, privacy protection, and end-to-end traceability for medication dispensing monitoring by rendering unobstructed and specular-free views from any specified virtual perspective. When the system detects an abnormal event triggering an alarm, it can simultaneously generate a clear, unobstructed, and specular-free virtual perspective image. Pharmacists can intuitively review erroneous images and make quick decisions without having to go to the dispensing room or repeatedly replay obstructed recordings, greatly improving the efficiency of manual review and the trust level of human-machine collaboration. Based on a trained dynamic neural radiation field, virtual camera parameters can be freely set during the inference stage to render views from any perspective, such as a top-down view or a frontal view, as if a "virtual shadowless camera" has been deployed above the medication dispensing area, allowing for comprehensive examination of operational details. In terms of privacy protection, the system can be deployed on edge computing nodes, encrypting and uploading only keyframes and alarm events; privacy data such as the operator's face in the original video stream does not need to leave the local machine. Simultaneously, the system records hand skeletal movement sequences, changes in the 3D drug model, and timestamps of key steps, forming a complete 3D operation trajectory, which can be used for post-event auditing and traceability, as well as standardized training for newly hired pharmacists. In summary, this invention, through synergistic innovation across these four levels, upgrades the traditional "recording-detection" model of visual drug dispensing monitoring to a proactive prevention closed loop of "understanding-reconstruction-verification," achieving a qualitative breakthrough in robustness, accuracy, interpretability, and safety. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating a machine vision-based drug dispensing monitoring method provided by the present invention.
[0024] Figure 2 This is a schematic diagram of the structural connection of a machine vision-based drug dispensing monitoring system provided by the present invention. Detailed Implementation
[0025] like Figure 1 As shown in the figure, the drug dispensing monitoring method based on machine vision provided by this invention can be deployed inside the biosafety cabinet or horizontal laminar flow table of a intravenous medication preparation center (PIVAS). The following description, using the monitoring of the extraction process of a glass ampoule in a PIVAS as an example, details each step of the method.
[0026] Step 1: Obtain the multi-view polarization image sequence and the corresponding depth image sequence of the drug dispensing operation area. This step is the foundation for data acquisition in the entire method, corresponding to the step "acquiring the multi-view polarization image sequence and the corresponding depth image sequence of the drug preparation operation area".
[0027] In terms of hardware deployment, at least three polarization-RGB-D depth cameras are deployed in a hemispherical shape around the drug dispensing area. Each camera has a rotatable and adjustable polarizer mounted in front of its optical lens and integrates a time-of-flight (ToF) or structured light depth sensor. Each camera covers the drug dispensing area from different angles (e.g., directly above, 45 degrees to the left front, and 45 degrees to the right front), forming a multi-view redundant observation. All cameras are synchronously calibrated via external hardware trigger lines to ensure that each viewpoint triggers exposure at the same millisecond level. Precise geometric registration is performed using methods such as Zhang Zhengyou's calibration method to establish the intrinsic parameter matrix and extrinsic parameter pose transformation relationship between each camera.
[0028] After the invention begins, each camera synchronously and continuously acquires and outputs a multimodal video stream. Each frame of image data contains four modalities: an RGB image (recording color and texture information), a polarization degree image (recording the linear polarization degree of each pixel, reflecting the degree of polarization of the reflected light at that point), a polarization angle image (recording the linear polarization angle of each pixel, reflecting the polarization direction of the reflected light at that point), and a depth map (recording the distance from each pixel's corresponding spatial point to the camera). This multi-view, multimodal synchronous acquisition method provides a rich and complementary raw data foundation for subsequent specular highlight removal, transparent object detection, and occlusion recovery.
[0029] Step 2: Perform physical-level specular de-highlighting on the multi-view polarized image and identify transparent material regions within it to generate multimodal fusion data with transparency prior. This step performs multimodal preprocessing based on the raw data obtained in step 1, corresponding to the step "perform physical-level de-highlighting on the multi-view polarization image and identify transparent material regions to generate multimodal fusion data with transparency priors". This step can be further subdivided into the following three sub-steps.
[0030] Sub-step 2.1: Physical specular removal based on Fresnel reflection model This sub-step corresponds to the step "Based on the Fresnel reflection model, use the polarization angle image to separate the diffuse reflection component and the specular reflection component to generate a diffuse reflection map without specular highlights".
[0031] The physical principle is as follows: when light shines on the surface of an object, the reflected light consists of diffuse reflection and specular reflection components. Specular reflection maintains a high degree of polarization during reflection, while diffuse reflection is essentially depolarized after multiple scatterings within the object. By analyzing the variation pattern of the polarization angle image within the pixel's neighborhood, the two can be effectively distinguished. Specifically, this invention iterates through each pixel, constructs a polarization feature vector using the polarization angle values of that pixel and its neighborhood, and solves for the intensity of the diffuse and specular reflection components of that pixel using Fresnel equations. When the specular reflection component of a pixel exceeds a threshold, it is determined to be a specular pixel, and only its diffuse reflection component is retained as the true texture color of that pixel, thus generating a specular-free diffuse reflection image.
[0032] Unlike traditional specular removal methods based on image inpainting or deep learning, this physical-level separation method does not rely on training data or blur missing areas. Instead, it precisely strips away specular reflection components at the physical level. Therefore, while removing strong reflective spots from stainless steel tabletops and glassware surfaces, the fine textures such as the strokes of the text on medicine labels and the markings on syringes, which were originally completely obscured by the spots, are completely preserved without any loss of texture or blurring of edges.
[0033] Sub-step 2.2: Calibration of transparent material regions based on polarization degree and depth edge This sub-step corresponds to the step "Perform adaptive threshold segmentation on the polarization image, and combine it with the edge information of the depth map to mark the transparent material region and assign an initial transmittance weight".
[0034] The principle behind this invention lies in the significant difference in polarization between transparent and opaque objects. When light passes through transparent media such as glass and liquids, its polarization state undergoes a regular change, resulting in a clear statistical separation between the polarization distribution of these regions and that of opaque diffuse reflective objects. This invention first uses the Otsu method or other adaptive thresholding algorithms to perform preliminary segmentation of the polarization image, extracting candidate regions whose polarization characteristics match those of transparent materials.
[0035] However, polarization information alone is insufficient to distinguish transparent objects from some smooth, non-transparent objects (such as polished metal). Therefore, a depth map is introduced for joint edge consistency determination: for the aforementioned candidate regions, the system extracts their visible light edges in the RGB image and their depth edges in the depth map. If a region has a clear edge in the RGB image (such as the outline of an ampoule wall), but there is no actual depth jump at the corresponding position in its depth map (because the interior of a transparent object is a single connected medium with no physical reflective surface), then the region is identified as a transparent material region. In this way, transparent objects such as glass ampoule walls, syringe barrel shells, and transparent films of infusion bags are accurately identified. The present invention assigns an initial transmittance weight τ to each pixel in these regions, with a value ranging from [0,1], representing the degree of attenuation of light passing through that point; non-transparent regions are set to 1.
[0036] Sub-step 2.3: Multimodal data channel-level fusion This sub-step corresponds to the step "to perform channel-level fusion of the diffuse reflection map without specular highlights, the initial transmittance weights, and the depth map to form the multimodal fusion data with transparency prior".
[0037] In practical terms, this invention splices the specular-free diffuse reflection map (RGB three channels) generated in sub-step 2.1, the transmittance weight map (single channel) generated in sub-step 2.2, and the depth map (single channel) obtained in step 1 along the channel dimension after spatial registration, forming a five-channel multimodal fusion data tensor. Each spatial location of this tensor not only contains the true color texture after specular removal, but also includes the physical properties of transmittance and three-dimensional depth geometry information at that point. This data structure provides a much better initialization condition than a single RGB input for the subsequent construction of the dynamic neural radiation field, enabling the network to possess prior knowledge from the outset to distinguish between transparent and non-transparent objects and perceive three-dimensional spatial structures.
[0038] Step 3: Real-time detection of the operator's hand posture to generate a dynamic occlusion mask. This step corresponds to the step "real-time detection of the operator's hand posture and generation of dynamic occlusion mask", and is a pre-processing step for subsequent intelligent occlusion removal. It can be further divided into two sub-steps.
[0039] Sub-step 3.1: Hand keypoint estimation based on spatiotemporal graph convolutional network The corresponding step in this sub-step is "using a spatiotemporal graph convolutional network of hand key points to estimate the key points of the operator's hands in real time and project them onto the imaging planes of each camera to generate an initial occlusion mask".
[0040] This invention employs a pre-trained lightweight spatiotemporal graph convolutional network (ST-GCN) with 21 keypoints in the hand as the hand pose estimator. In the graph structure of this network, each node corresponds to a skeletal keypoint of the hand (such as fingertip, knuckle, wrist, etc., 21 nodes per hand). Spatial edges connect adjacent physical joints within the same frame, and temporal edges connect the correspondence of the same node between adjacent frames. By stacking multiple layers of spatiotemporal graph convolutional operations, the network can simultaneously capture the spatial skeletal configuration of the hand within a single frame and the dynamic motion information across frames. The network takes a sequence of consecutive RGB image frames as input and outputs the three-dimensional spatial coordinates of 42 keypoints for both hands in real time (frame rate ≥ 30fps).
[0041] After obtaining the 3D coordinates of the key points, this invention utilizes the known intrinsic and extrinsic pose parameters of each camera to project the 3D skeleton points onto the 2D imaging plane of each camera, forming a sparse 2D point set. Convex hull operations are performed on these point sets, and a certain pixel margin is extended outwards to completely cover the hand contour, generating an initial hand occlusion mask for that viewpoint.
[0042] Sub-step 3.2: Mask timing propagation and correction based on bidirectional optical flow This sub-step corresponds to the step "using bidirectional optical flow to propagate the initial occlusion mask between consecutive frames, compensate for the mask boundary offset caused by rapid hand movement, and generate a time-consistent dynamic occlusion mask".
[0043] During medication preparation, the operator's hand movements are rapid, and masks detected independently frame by frame are prone to edge jitter and inter-frame inconsistencies. To address this, this invention introduces a bidirectional optical flow correction mechanism. First, utilizing the dense optical flow field between adjacent frames, the mask from the previous frame is propagated forward to the current frame via motion vectors, while the mask from the next frame is propagated backward. Then, the three masks (the mask directly generated in the current frame, the forward propagation mask, and the backward propagation mask) are weighted and fused at the pixel level, with the weights dynamically adjusted based on the confidence level of the corresponding region for each mask. The resulting dynamic occlusion mask has smooth boundaries and a consistent temporal sequence, accurately marking the pixel regions occluded by the hand from each viewpoint, providing a reliable supervised region indication for the de-occlusion training in step 4.
[0044] Step 4: Construct and train a dynamic neural radiation field. The training process introduces a temporal de-occlusion prior loss to utilize the unoccluded spatiotemporal information to recover the visual features of the area marked by the dynamic occlusion mask. This step is the core of the entire method, corresponding to the step "Construct and train a dynamic neural radiation field, the training process of which introduces temporal demasking prior loss...".
[0045] Sub-step 4.1: Network construction of dynamic neural radiation field This sub-step corresponds to the step "taking the three-dimensional coordinates of the spatial sampling point, the observation direction, the timestamp, the polarization feature vector, and a transparent material indicator variable as inputs, and the volume density of the point, the polarization-sensing color related to the observation direction, and the transparency coefficient as outputs to construct the dynamic neural radiation field".
[0046] Before the actual medication preparation operation begins, this invention first performs an "empty scene scan": the operator temporarily removes all items from the worktable, and each camera captures a 10-second multimodal video of the empty scene, which is used to pre-train a static scene neural radiation field F_static. F_static takes spatial coordinates and the viewing direction as input and outputs volume density and color. After training convergence, its network weights implicitly encode and store the static geometric and appearance information of the medication preparation worktable, background environment, etc.
[0047] Subsequently, this invention builds a dynamic version, F_dynamic, based on F_static. F_dynamic extends both its input and output. The inputs include: the 3D coordinates x of the spatial sampling point for spatial location; the viewing direction d for modeling viewpoint-related reflection effects; a timestamp t for capturing dynamic changes in the scene over time; a polarization feature vector (extracted from the polarization degree map and polarization angle image at the corresponding time using a small convolutional encoder) for injecting polarization physical properties; and a binary transparent material indicator variable flag, indicating whether the sampling point is located within the transparent material region defined in step 2. The outputs include: the volume density σ of the point, representing the light attenuation capability of the spatial point; the polarization-sensing color c=(r,g,b) related to the viewing direction, determined by both the polarization feature and the viewing direction; and a transparency coefficient α for adjusting the light transmittance of the transparent region.
[0048] Sub-step 4.2: Design and mechanism of temporal de-occlusion prior loss This sub-step corresponds to the step "For a pixel in the current frame that is marked as occluded by the dynamic occlusion mask, obtain the observation information of its corresponding spatial point at the adjacent unoccluded time or the adjacent unoccluded viewpoint; generate an approximate real texture of the occluded pixel using the observation information through stereo matching; use the approximate real texture as a supervision signal to calculate the loss and drive the dynamic neural radiation field to learn and complete the geometric and appearance information of the occluded region."
[0049] Standard neural radiation field training employs color photometric loss, which minimizes the difference between the colors of pixels rendered by the network and those captured by the real camera. However, if pixels occluded by a hand are directly included in this loss calculation, the network will be incorrectly guided to learn the color of the hand instead of the occluded target. The traditional approach is to discard these pixels, but this results in the permanent loss of information about the occluded area.
[0050] This invention's innovative "temporal de-occlusion prior loss" completely circumvents the aforementioned difficulties. Its workflow is as follows: For a pixel p marked as occluded by a hand mask in the current frame t and the current camera C1's viewpoint, the corresponding real-world spatial point X can be determined based on multi-view geometry principles. This invention automatically retrieves instances of this spatial point X that are not occluded in all available spatiotemporal observations—possibly from camera C2 at the same time t but with a different viewpoint (where X is not occluded by the hand due to parallax), or from observations by the same camera C1 at adjacent times t-1 or t+1 (when the hand has not yet moved to that position or has moved away). Once an available observation is retrieved, this invention utilizes the calibrated multi-camera pose relationships or temporal scene flow information to perform stereo matching and projection transformation on the texture of the spatial point X, generating an "approximately real texture," i.e., a pseudo-label, for that pixel in the current frame. Then, using this pseudo-label as a monitoring signal, the loss is calculated on the color rendered by F_dynamic at pixel p.
[0051] The core of this mechanism lies in its ability to force the neural network not to ignore occluded pixels as information gaps, but to actively "reason" and "complete" the geometry and texture of the hidden region from the spatiotemporal redundancy information from multiple perspectives and times. Figuratively speaking, the network learns to "imagine what the object behind the hand looks like" through training.
[0052] Sub-step 4.3: Introduction and mechanism of refractive correction loss This sub-step corresponds to the step "The training process also introduces a refraction correction loss for transparent material regions. The refraction correction loss is based on a physical refraction model and constrains the propagation path of light at the interface of transparent objects in order to accurately reconstruct the three-dimensional position of the internal structure of the transparent object".
[0053] Traditional neural radiation fields assume that light travels in a straight line in a scene, but this assumption fails when encountering transparent objects. When light passes through the interface of media such as air-glass or glass-liquid, it undergoes significant refraction and bending, resulting in severe geometric distortions in the reconstruction of the internal structure of transparent objects (such as the liquid surface inside an ampoule or the liquid column inside a syringe) by standard NeRF.
[0054] This invention overcomes this problem by introducing refractive correction loss into the training loss. Its physical basis is Snell's law of refraction. This invention predetermines the refractive index parameters of the main media in the drug preparation scenario (air ≈ 1.00, borosilicate glass ≈ 1.47, water / medicine solution ≈ 1.33). During volume rendering sampling, when light passes through an area marked as a transparent medium by a transparent material indicator variable, the light sampling direction no longer extends in a straight line. Instead, according to Snell's law, it undergoes real-time deflection correction at the medium interface, forming a physically correct bent sampling path. The refractive correction loss constrains the pixel color and depth values obtained from volume rendering along this bent path to be consistent with real camera observations; simultaneously, it constrains the geometry of the sampling path at the interface to conform to the physical refraction model. Through this strong physical constraint, F_dynamic can accurately reconstruct the true thickness of the ampoule glass wall, the absolute three-dimensional height of the internal liquid surface, and other internal geometric information, achieving sub-millimeter-level reconstruction accuracy, laying the physical foundation for subsequent accurate volume measurement.
[0055] Step 5: Based on the trained dynamic neural radiation field, remove the hand area and render a de-occluded and de-highlighted virtual monitoring view. This step corresponds to the step "Based on the trained dynamic neural radiation field, remove the hand area and render a virtual monitoring view that is de-occluded and de-highlighted," which is the core operation of the method in the inference application stage.
[0056] Sub-step 5.1: Density field removal of the hand region This sub-step corresponds to the step "obtaining real hand posture information and directly covering the density field area corresponding to the hand during the spatial sampling process of the dynamic neural radiation field".
[0057] During online monitoring, when a critical medication verification is required (e.g., after the operator completes the medication aspiration action), this invention first acquires the key skeletal points of the operator's hands, estimated in real-time during step 3. Based on the three-dimensional coordinates of these key points, this invention fits a coarse three-dimensional envelope of the hands and forces the volume density σ of all sampling points in the region occupied by this envelope to zero in the density field space of F_dynamic. The effect of this operation is to directly "remove" the operator's hands from the three-dimensional scene representation, ensuring they do not contribute any light attenuation or color contribution in subsequent rendering.
[0058] Sub-step 5.2: Setting up the virtual viewpoint and executing volume rendering This sub-step corresponds to the step "Setting a virtual camera viewpoint for medication dispensing monitoring; sampling the light along the virtual camera viewpoint, using the density field after covering the hand for volume rendering, and generating the de-occluded and de-highlighted virtual monitoring view".
[0059] This invention generates a virtual camera whose viewing angle parameters can be freely set by the monitoring pharmacist as needed, such as a "God's-eye view" (vertically looking down from directly above) or a "frontal view" (horizontally observing from directly in front). For each pixel of this virtual camera, this invention emits a ray of light that passes through the scene density field (with hands removed). The ray samples spatial points along its path at certain steps, and each sampled point is fed into an F_dynamic network for forward inference to obtain the volume density σ and color c of that point. Finally, using a volume rendering integral formula, the color and density of each sampled point are accumulated along the ray direction to generate the final color value of the pixel.
[0060] Since the input data has undergone polarization despecculation in step 2, and the hands area has been removed during rendering, the generated virtual view exhibits the following effect: the operator's hands are completely invisible, the stainless steel tabletop and glassware surfaces have no reflective spots, and the text on the medicine label and the syringe markings are clear and sharp. The entire scene appears as if it were filmed under a "virtual shadowless lamp." Simultaneously, the corresponding depth map is also rendered and output, containing precise three-dimensional spatial distance information.
[0061] Sub-step 5.3: Rendering acceleration methods This sub-step corresponds to the step "The volume rendering is accelerated by multi-resolution hash coding and model distillation techniques".
[0062] Volume rendering itself is computationally intensive. To meet the real-time monitoring requirement of ≥25fps, this method introduces two acceleration techniques. First, multi-resolution hashing is used to efficiently map spatial coordinates into features, quickly encoding 3D coordinates into sparse, learnable feature vectors, significantly reducing the computational load of network queries. Second, after the complete F_dynamic network training converges, knowledge distillation is used to train a simplified, lightweight student network to learn and approximate the output distribution of the teacher network. During inference, only the student network is deployed, compressing the rendering time of a single frame of virtual view to less than 20 milliseconds without significantly sacrificing rendering quality.
[0063] Step 6: Auxiliary Functions – Three-Dimensional Dosage Measurement and Safety Detection This step is a high-precision analysis function performed on the high-quality virtual view and depth map generated in step 5, corresponding to the various auxiliary operations in the step.
[0064] Sub-step 6.1: Precise measurement of three-dimensional dose volume The sub-step involves "segmenting the syringe based on the depth map corresponding to the virtual monitoring view, identifying its outer wall, piston, and liquid column area, reconstructing the liquid column area in three dimensions based on the known syringe geometric parameters, and calculating the actual liquid volume through volume integration."
[0065] This invention utilizes instance segmentation networks (such as Mask R-CNN) to precisely segment syringes in a virtual view, identifying and delineating the syringe's outer wall contour, the piston rubber tip front boundary, and the medication column region. Since clinically used syringes are standard industrial products, their actual physical inner diameter (e.g., 14.5 mm for a 10 mL syringe) is loaded as a known prior parameter. Using a sub-millimeter precision depth map corresponding to the virtual view, this invention can obtain the three-dimensional spatial coordinates of each pixel on the liquid column surface. Combined with the known inner diameter parameter, it directly performs three-dimensional reconstruction and volume integration calculation of the irregular spatial volume of the liquid column segment to obtain the precise volume of the actual aspirated medication. The measurement error is controlled within 0.05 mL, and is automatically compared with the electronic prescription dosage transmitted from the hospital. If the error exceeds the tolerance threshold, a dosage error alarm is triggered. This accuracy far exceeds traditional measurement methods based on two-dimensional image pixel height estimation, meeting the stringent verification standards for high-risk drug preparation.
[0066] Sub-step 6.2: Risk detection of needle tip touching the bottom The step in this document is to "calculate the depth difference between the syringe tip and the bottom of the vial based on the depth map, and issue a bottoming risk alarm when the depth difference is less than a safety threshold".
[0067] During medication extraction, if the syringe needle tip touches the bottom of the ampoule, it may scratch the glass, generating insoluble particles, or aspirate sediment from the bottom, posing a potential medication safety hazard. This invention utilizes the high-precision depth map rendered in step 5 to pinpoint the three-dimensional coordinates of the syringe needle tip and the lowest point on the inner wall of the ampoule, calculating the Euclidean distance between the two points. When this distance is less than a preset safety threshold (e.g., 1 mm), this invention determines there is a risk of the needle tip touching the bottom and immediately issues a "high risk of bottom contact" alarm, prompting the operator to adjust the needle insertion depth. This three-dimensional spatial safety indicator is completely imperceptible to this invention using traditional two-dimensional monitoring.
[0068] Sub-step 6.3: Automatic check of compliance of medication preparation operation. The sub-step in this document involves "comparing the identified key point sequence of the operator's hand bones with the reconstructed 3D model of the drug in the virtual monitoring view to identify the compliance of the drug dispensing operation steps through spatiotemporal motion comparison."
[0069] This invention analyzes the motion trajectories of key points in the skeletal structure of both hands over a continuous time series and matches them with 3D models of medicines (ampoules, sterile swabs, infusion bags, etc.) reconstructed in the same spatial coordinate system for spatiotemporal relationship matching. A state machine model of standard dispensing procedures is predefined, with each step including specific contact relationships, relative positions, and motion patterns between the hand and the object. For example, the "sterilization" step is defined as the distance between the key points of the hand bones and the sterile swab being less than a threshold, and the swab contacting the neck area of the ampoule; the "cutting" step is defined as the specific relative position and motion trajectory between the key points of the hand bones and the cutting tool. This invention performs frame-by-frame matching of each step. If a step is missed (e.g., entering the cutting step without sterilization) or the order is reversed, it is immediately marked as an operational violation, and a timestamp is recorded for alarm and post-event traceability.
[0070] like Figure 2 As shown, corresponding to the above method, the present invention also provides a drug dispensing monitoring system based on machine vision. This system includes: a multi-view polarization-depth acquisition module, consisting of at least three time-stamped and geometrically calibrated polarized RGB-D cameras and trigger circuits; a multi-modal preprocessing module, deployed on the GPU of an edge computing node, used to perform physical de-highlighting and transparent material calibration in step 2; a hand pose perception module, running a lightweight ST-GCN network, used to perform hand pose estimation and mask generation in step 3; a neural radiation field training and reconstruction module, as the core computing unit, used to perform network construction and training in step 4, with its loss function embedding temporal de-occlusion prior loss and refraction correction loss; and a virtual view rendering and monitoring module, including an interactive terminal and an analysis unit, used to perform virtual view rendering in step 5 and three-dimensional dose measurement, safety depth detection, and action compliance detection in step 6. All alarm information is automatically linked to the hospital information system for encrypted uploading and archiving.
[0071] Through the coordinated operation of the above steps and modules, this invention completely upgrades the traditional visual monitoring of medication dispensing from a passive recording mode of "capture-detection" to a closed-loop proactive prevention system of "understanding-reconstruction-verification", significantly improving the safety, accuracy and traceability of intravenous medication dispensing.
[0072] The above description is merely a preferred embodiment of the present invention and does not constitute any limitation on the present invention. Any simple modifications, alterations, or equivalent structural changes made to the above embodiments based on the technical essence of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A machine vision-based dispensing monitoring method, characterized by, include: Acquire multi-view polarization image sequences and corresponding depth image sequences of the drug dispensing operation area; Physical-level de-highlighting is performed on the multi-view polarized image, and transparent material regions are identified to generate multimodal fusion data with transparency prior. Real-time detection of the operator's hand posture generates a dynamic occlusion mask; A dynamic neural radiation field is constructed and trained. The training process introduces temporal de-occlusion prior loss to utilize the unoccluded spatiotemporal information to recover the visual features of the area marked by the dynamic occlusion mask. Based on the trained dynamic neural radiation field, the hand area is removed, and a virtual monitoring view with de-occlusion and de-highlighting is generated.
2. The method of claim 1, wherein, The acquisition of the multi-view polarization image sequence and the corresponding depth image sequence of the drug dispensing operation area specifically involves: At least three polarized red-green-blue depth cameras are deployed in a hemispherical shape around the drug dispensing area to simultaneously acquire RGB images, polarization degree images, polarization angle images, and depth maps.
3. The method of claim 2, wherein, The process of performing physical-level specular de-illumination on the multi-view polarization image and identifying transparent material regions within it to generate multimodal fusion data with transparency priors specifically includes: Based on the Fresnel reflection model, the diffuse reflection component and specular reflection component are separated using the polarization angle image to generate a diffuse reflection map without specular highlights. Adaptive threshold segmentation is performed on the polarization image, and the transparent material region is identified and assigned an initial transmittance weight by combining the edge information of the depth map. The diffuse reflection map without specular highlights, the initial transmittance weights, and the depth map are fused at the channel level to form the multimodal fused data with transparency prior.
4. The method of claim 3, wherein, The real-time detection of the operator's hand posture and the generation of a dynamic occlusion mask specifically includes: A spatiotemporal graph convolutional network for hand key points is used to estimate the key points of the operator's hands in real time and project them onto the imaging planes of each camera to generate an initial occlusion mask. Using the bidirectional optical flow method, the initial occlusion mask is propagated between consecutive frames to compensate for the mask boundary offset caused by rapid hand movements, thereby generating a time-consistent dynamic occlusion mask.
5. The method of claim 4, wherein, The construction and training of a dynamic neural radiation field includes: The dynamic neural radiation field is constructed by taking the three-dimensional coordinates of the spatial sampling point, the observation direction, the timestamp, the polarization feature vector, and a transparent material indicator variable as inputs, and taking the volume density of the point, the polarization-sensing color related to the observation direction, and the transparency coefficient as outputs.
6. The method of claim 5, wherein, The temporal de-occlusion prior loss is specifically as follows: For pixels in the current frame that are marked as occluded by the dynamic occlusion mask, obtain the observation information of their corresponding spatial points at adjacent unoccluded times or adjacent unoccluded viewpoints; By using stereo matching, an approximate realistic texture of the occluded pixel is generated using the observation information; Using the approximate realistic texture as a supervision signal, the loss is calculated, driving the dynamic neural radiation field to learn and complete the geometric and appearance information of the occluded area.
7. The method of claim 6, wherein, The training process also introduces a refraction correction loss for transparent material regions. This refraction correction loss is based on a physical refraction model and constrains the propagation path of light at the interface of transparent objects in order to accurately reconstruct the three-dimensional position of the internal structure of the transparent object.
8. The method of claim 1, wherein, The process of removing the hand region and rendering a de-occluded and de-highlighted virtual monitoring view based on the trained dynamic neural radiation field specifically includes: To obtain real hand posture information, during the spatial sampling process of the dynamic neural radiation field, the density field area corresponding to the hand is directly covered. Set up a virtual camera view for medication dispensing monitoring; The light rays along the virtual camera's viewpoint are sampled, and the density field after the hand is covered is used for volume rendering to generate the de-occluded and de-highlighted virtual monitoring view. The volume rendering is accelerated using multi-resolution hash encoding and model distillation techniques.
9. The method according to claim 1 or 8, characterized in that, Also includes: Based on the depth map corresponding to the virtual monitoring view, the syringe is segmented into instances to identify its outer wall, piston and liquid column area. Combined with the known syringe geometric dimensions, the liquid column area is reconstructed in three dimensions, and the actual liquid volume is calculated by volume integration. In addition, the depth difference between the syringe needle tip and the bottom of the medicine bottle is calculated based on the depth map, and a bottoming risk alarm is issued when the depth difference is less than a safety threshold; And / or, based on the identified sequence of key points of the operator's hand bones and the reconstructed three-dimensional model of the drug in the virtual monitoring view, a spatiotemporal motion comparison is performed to identify the compliance of the drug dispensing operation steps.
10. A machine vision based dispensing monitoring system, characterized in that include: The multi-view polarization-depth acquisition module is used to simultaneously acquire multi-view polarization image sequences and corresponding depth image sequences of the drug dispensing operation area; The multimodal preprocessing module is used to perform physical-level de-highlighting on the multi-view polarization image, identify transparent material regions, and generate multimodal fusion data with transparency prior. The hand posture sensing module is used to detect the operator's hand posture in real time and generate a dynamic occlusion mask; The neural radiation field training and reconstruction module is used to construct and train a dynamic neural radiation field. Its training process introduces temporal de-occlusion prior loss and refraction correction loss to recover the visual features of the occluded area using the unoccluded spatiotemporal information and accurately reconstruct the internal structure of transparent objects. The virtual view rendering and monitoring module is used to remove the hand area based on the trained dynamic neural radiation field, render and generate a virtual monitoring view that is de-occluded and de-highlighted, and perform drug dispensing verification and risk alarm based on this view.