Deep Learning-Based Monitoring System for the Production and Filling of Skin and Mucous Membrane Disinfectants
Patent Information
- Application Number
- CN202511916751.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-18
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2045-12-18
AI Technical Summary
传统灌装监控系统图像获取视角单一,难以全面捕捉液滴在空间中的运动轨迹,特别是在液滴偏移或喷溅等异常状态发生时易出现识别盲区,影响检测准确性;现有图像识别算法对液滴边缘识别能力有限,在存在复杂背景光干扰或液滴透明度变化的场景下识别精度下降,无法稳定输出边界框用于后续分析;环境因素如温湿度和气压对液滴行为具有显著影响,但现有方案普遍未能将多源环境数据纳入统一分析流程,缺乏有效的多维融合机制,导致对异常风险的判别结果波动大且响应滞后
本发明通过构建区域势能分布图与改进Deformable DETR模型,针对皮肤黏膜消毒制剂灌装过程中液滴状态监测受视角局限、边缘识别能力差及环境扰动影响大的问题,采用多视角高帧率图像获取与区域势能场聚焦机制相结合,提升液滴边界识别的空间聚焦能力与检测稳定性;在液滴状态判别环节,融合液滴边界框时序变化构建液滴运动轨迹,并设定位移与角度变化规则实现对正常下落、散射喷溅与偏移逸散状态的精确分类;针对环境因素对液滴行为的潜在干扰影响,构建多源环境传感数据与图像数据的时间同步多维监测数据集,结合设定安全阈值进行环境风险量化建模;在残留风险评估中,引入液滴状态、液滴运动轨迹与环境参数的分项风险值加权机制,实现总风险评分值的量化输出,并通过设定分级风险阈值将结果划分为低、中、高三等级;最终,在响应控制环节根据不同风险等级联动执行继续灌装、灌装针头清洁或紧急中断报警等操作,构建形成从识别、评估到响应的智能闭环控制系统,有效提升消毒制剂灌装过程的安全性、稳定性与自动化水平。
Smart Images

Figure CN121686367B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent monitoring technology for filling production, and in particular to a deep learning-based monitoring system for the production and filling of skin and mucous membrane disinfectant preparations. Background Technology
[0002] With the pharmaceutical industry's ever-increasing standards for filling quality and residue control, monitoring and intelligent management technologies for the filling process of products with high safety requirements, such as skin and mucous membrane disinfectants, have received widespread attention. Existing filling process monitoring mainly relies on manual inspections or droplet detection systems based on single-view image recognition, but these systems generally suffer from the following problems in actual production: Traditional filling monitoring systems rely on a single image acquisition perspective, making it difficult to comprehensively capture the movement trajectory of droplets in space. This is especially true when abnormal states such as droplet deviation or splashing occur, which can lead to blind spots and affect detection accuracy. Existing image recognition algorithms have limited ability to identify droplet edges, and their recognition accuracy decreases in scenarios with complex background light interference or changes in droplet transparency, making it impossible to stably output bounding boxes for subsequent analysis. Environmental factors such as temperature, humidity, and air pressure have a significant impact on droplet behavior, but existing solutions generally fail to incorporate multi-source environmental data into a unified analysis process and lack an effective multi-dimensional fusion mechanism, resulting in large fluctuations and delayed responses in the judgment of abnormal risks.
[0003] Therefore, how to provide a deep learning-based monitoring system for the production and filling of skin and mucous membrane disinfectants is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0004] One objective of this invention is to propose a deep learning-based monitoring system for the production and filling of skin and mucous membrane disinfectant preparations. This invention integrates regional potential energy distribution maps with an improved Deformable DETR model, and combines multi-angle image acquisition with multi-dimensional monitoring data analysis to achieve accurate identification and dynamic tracking of droplet states during the filling process. By constructing a fused image sequence and introducing a regional potential energy field focusing mechanism, the accuracy of target detection is effectively improved. At the same time, risk assessment is performed by combining multi-source environmental sensor data to achieve residual droplet risk level discrimination and graded response control. It has significant advantages such as high identification accuracy, timely response control, and strong adaptability to complex working conditions.
[0005] The deep learning-based monitoring system for the production and filling of skin and mucous membrane disinfectant preparations according to an embodiment of the present invention includes the following modules: The multi-angle image acquisition module is used to acquire image information of the filling needle during the filling cycle and obtain the original image sequence; The image preprocessing module is used to preprocess the original image sequence to obtain a preprocessed image sequence; The regional potential energy construction module is used to construct a regional potential energy distribution map and fuse the regional potential energy distribution map with a preprocessed image sequence to generate a fused image sequence. The droplet detection module is used to input the fused image sequence into the improved Deformable DETR model for droplet target detection. The improved Deformable DETR model introduces a regional potential energy field focusing mechanism and outputs droplet bounding boxes. The state discrimination module is used to extract the droplet motion trajectory based on the position change of the droplet bounding box in the fused image sequence, and to determine the droplet state according to a preset discrimination rule; The multi-source data synchronization module is used to collect multi-source environmental sensor data during each filling process and synchronize it with droplet trajectory and droplet state information in time to build a multi-dimensional monitoring dataset. The risk assessment module is used to calculate the total risk score based on the multidimensional monitoring dataset and obtain the residual risk level. The response control module is used to execute the corresponding response control strategy according to the residual risk level.
[0006] The deep learning-based monitoring method for the production and filling of skin and mucous membrane disinfectant preparations according to embodiments of the present invention includes the following steps: Step 1: Set up industrial cameras at multiple angles at the filling station for producing skin and mucous membrane disinfectant to collect image information of the filling needle during the filling cycle and obtain the original image sequence; Step 2: Preprocess the original image sequence to obtain a preprocessed image sequence; Step 3: Construct a regional potential energy distribution map, and fuse the regional potential energy distribution map with the preprocessed image sequence to generate a fused image sequence; Step 4: Input the fused image sequence into the improved Deformable DETR model for droplet target detection. The improved Deformable DETR model introduces a regional potential energy field focusing mechanism and outputs droplet bounding boxes. Step 5: Based on the positional changes of the droplet bounding box in the fused image sequence, extract the droplet trajectory and combine it with historical droplet trajectory data to determine the droplet state; Step 6: Collect corresponding multi-source environmental sensor data during each filling process, synchronize the multi-source environmental sensor data with the droplet motion trajectory and droplet state in time, and construct a multi-dimensional monitoring dataset; Step 7: Calculate the total risk score based on the multidimensional monitoring dataset to obtain the residual risk level, including low risk level, medium risk level and high risk level; Step 8: Implement the corresponding response control strategy based on the residual risk level.
[0007] Optionally, step one specifically includes: Industrial cameras are arranged at multiple angles relative to the filling needle at the filling station. The industrial cameras include a main camera located directly in front of the filling needle, and two auxiliary cameras respectively located on the left and right sides in front of the filling needle, forming an angle of 30° to 60° with the main camera in the horizontal direction. The installation height of the main camera and the auxiliary cameras is the same as the height of the filling needle nozzle. The industrial camera has an image resolution greater than 3840×2160 pixels, an image acquisition frame rate greater than 120 frames per second, and the acquired images are time-synchronized by an external trigger signal. The acquisition cycle corresponds to the filling cycle. An LED ring cold light source with an oblique incidence angle of 30° to 60° is set within the field of view of the industrial camera. The LED ring cold light source is installed above the filling needle to form a reflection enhancement area by oblique illumination. A background light shield and a linear backlight are set on the back of the filling needle. The surface reflectivity of the background light shield material is less than 15%, and the illuminance of the linear backlight ranges from 15,000 to 25,000 Lux. The acquired multi-angle image frames are arranged in chronological order to form an original image sequence.
[0008] Optionally, step two specifically involves: Each frame of the original image sequence is subjected to two-dimensional convolution processing using a 3×3 Laplacian convolution operator to obtain an edge response image; The edge response images are reduced to 1 / 2 and 1 / 4 of their original size, respectively, to obtain two downsampled images, which together with the edge response images form a three-layer image pyramid; The two downsampled images are upsampled to the same size as the edge response image using bilinear interpolation. The gray values of each pixel in the three images are then weighted and averaged to output a fused scale image. The pixel grayscale values of the fused scale image are divided into 8 equal-width grayscale intervals within the range of 0 to 255. RGB three-channel values are set for each interval. Pixels in each grayscale interval are assigned preset combinations of red, green, and blue channel values to generate a pseudo-color image. All pseudo-color images are arranged in chronological order of image acquisition time to obtain a preprocessed image sequence.
[0009] Optionally, step three specifically includes: Using the pixel coordinates of the lower end of the filling needle in the image as the center point, establish a two-dimensional matrix with the same size as the image, and calculate the Euclidean distance from each pixel in the matrix to the center point. Divide the Euclidean distance by the length of the image diagonal to obtain a normalized distance value, and subtract the normalized distance value from 1 to obtain the regional potential energy value of each pixel, thereby generating a regional potential energy distribution map. The gray value of each pixel in the corresponding frame image in the preprocessed image sequence is multiplied by the regional potential energy value at the same position in the regional potential energy distribution map to obtain the fused pixel value, which constitutes the fused image frame. All fused image frames are arranged in chronological order of acquisition time to form a fused image sequence.
[0010] Optionally, the improved Deformable DETR model specifically includes a backbone feature extraction network, an attention encoder, a region potential guided decoder, and a bounding box regression head; The backbone feature extraction network adopts the ResNet-50 structure, receives the fused image sequence and extracts image features through four stages of convolutional layers. The first stage outputs a shallow feature map with a size of 1 / 4 of the input image, the second stage outputs a mid-shallow feature map with a size of 1 / 8, the third stage outputs a mid-deep feature map with a size of 1 / 16, and the fourth stage outputs a deep feature map with a size of 1 / 32. The attention encoder is composed of a multi-layer Transformer encoder structure stacked together. Each layer includes a multi-head self-attention module and a feedforward neural network module, which encode the feature maps output from the four stages respectively to generate four sets of encoded feature maps. The regional potential energy guided decoder has several preset initial target query vectors. The initial target query vectors are generated by selecting several pixel positions with the highest regional potential energy values from the regional potential energy distribution map, extracting two-dimensional coordinates and performing linear position encoding, and serving as the initial input of the query vector. The regional potential energy guided decoder introduces a regional potential energy field focusing mechanism, which performs element-wise multiplication of the initial target query vector with the regional potential energy value at the corresponding position in the regional potential energy distribution map to obtain a weighted target query vector; the weighted target query vector interacts with the encoded feature map through cross-attention in the multi-layer decoder, and iteratively updates the output target representation vector layer by layer. The bounding box regression head consists of two cascaded fully connected networks. The first fully connected network maps the target representation vector to an intermediate feature vector of length 128, and the second fully connected network maps the intermediate feature vector to four-dimensional bounding box parameters, representing the center x-coordinate, center y-coordinate, width, and height, respectively, and outputs the droplet bounding box.
[0011] Optionally, step five specifically includes: Based on the positional changes of the droplet bounding box in the fused image sequence, the coordinates of the center point of the droplet bounding box in consecutive frames are extracted, and the droplet motion trajectory is constructed in chronological order. Calculate the lateral and longitudinal displacements between points on the continuous droplet trajectory, and generate a velocity vector sequence based on the changing direction of the droplet bounding box center point between adjacent frames; The state of the droplet is determined according to a preset discrimination rule, which specifically includes: If the lateral displacement does not exceed the set displacement threshold in all consecutive frames, the vertical displacement continues to be downward, and the velocity vector change angle is less than or equal to the set amplitude threshold, the droplet is determined to be in a normal falling state. If the lateral displacement exceeds the set displacement threshold for three or more consecutive frames, and the velocity vector change angle is greater than the set amplitude threshold, the droplet state is determined to be a scattering and splashing state. If the droplet's trajectory is biased towards one side of the horizontal plane of the image, and the lateral displacement direction remains consistent across multiple consecutive frames while the cumulative displacement exceeds a set boundary distance, and the velocity vector direction deviates from the vertical direction by more than a set amplitude threshold across multiple consecutive frames, the droplet is determined to be in an offset and dissipation state.
[0012] Optionally, the multi-source environmental sensing data specifically includes ambient temperature data, ambient humidity data, and ambient air pressure data.
[0013] Optionally, step seven specifically includes: The droplet states in the multidimensional monitoring dataset are assigned different values according to their state types: 0 for normal falling state, 1 for scattering and splashing state, and 2 for offset and escaping state, to obtain the droplet state risk value. The lateral cumulative offset and the average change angle of the velocity vector of the droplet motion trajectory in the multidimensional monitoring dataset are normalized and then added together to obtain the trajectory risk value. The ambient temperature, ambient humidity, and ambient air pressure from the multi-source environmental sensor data are compared with their respective set safety thresholds. For the portion that exceeds the set safety threshold range, the deviation is multiplied by a preset proportional coefficient and then added to obtain the environmental risk value. The total risk score is calculated by adding the droplet state risk value, trajectory risk value, and environmental risk value. The residual risk level is obtained based on the total risk score. If the total risk score is less than a set first risk threshold, it is determined to be a low risk level; if the total risk score is greater than or equal to the set first risk threshold and less than or equal to a set second risk threshold, it is determined to be a medium risk level; if the total risk score is greater than the set second risk threshold, it is determined to be a high risk level.
[0014] Optionally, step eight specifically includes: When the residual risk level is low risk, the current filling process remains unchanged and the filling operation continues. When the residual risk level is medium risk level, the filling needle is automatically controlled to perform a cleaning action, including spraying cleaning fluid onto the outer wall of the filling needle and drying it with airflow, while briefly stopping the one-time filling operation, and resuming filling after cleaning is completed; When the residual risk level is high, the filling process should be immediately interrupted, and an audible and visual alarm should be triggered to alert the operator to conduct on-site investigation and handling.
[0015] The beneficial effects of this invention are: This invention addresses the challenges of limited viewing angles, poor edge recognition, and significant environmental disturbances in droplet state monitoring during the filling of skin and mucous membrane disinfectant preparations by constructing a regional potential energy distribution map and an improved Deformable DETR model. It combines multi-view high-frame-rate image acquisition with a regional potential energy field focusing mechanism to enhance the spatial focusing capability and detection stability of droplet boundary recognition. In the droplet state discrimination stage, it integrates the temporal changes of droplet boundary boxes to construct droplet trajectories and sets rules for displacement and angle changes to accurately classify normal falling, scattering / splashing, and drifting states. To address the potential interference of environmental factors on droplet behavior, it constructs a multi-source environmental sensor data and image data system. A time-synchronized multidimensional monitoring dataset is used to quantitatively model environmental risks by combining it with set safety thresholds. In residual risk assessment, a weighted mechanism for the risk values of droplet state, droplet trajectory, and environmental parameters is introduced to achieve quantitative output of the total risk score. The results are divided into low, medium, and high levels by setting graded risk thresholds. Finally, in the response control stage, operations such as continuing filling, cleaning filling needles, or emergency interruption alarms are executed in conjunction with different risk levels to build an intelligent closed-loop control system from identification and assessment to response, effectively improving the safety, stability, and automation level of the disinfectant filling process. Attached Figure Description
[0016] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a schematic diagram of the deep learning-based monitoring system for the production and filling of skin and mucous membrane disinfectant preparations proposed in this invention. Figure 2 This is an overall flowchart of the deep learning-based monitoring method for the production and filling of skin and mucous membrane disinfectant preparations proposed in this invention. Detailed Implementation
[0017] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0018] refer to Figure 1 A deep learning-based monitoring system for the production and filling of skin and mucous membrane disinfectant preparations includes the following modules: The multi-angle image acquisition module is used to acquire image information of the filling needle during the filling cycle and obtain the original image sequence; The image preprocessing module is used to preprocess the original image sequence to obtain a preprocessed image sequence; The regional potential energy construction module is used to construct a regional potential energy distribution map and fuse the regional potential energy distribution map with a preprocessed image sequence to generate a fused image sequence. The droplet detection module is used to input the fused image sequence into the improved Deformable DETR model for droplet target detection. The improved Deformable DETR model introduces a regional potential energy field focusing mechanism and outputs droplet bounding boxes. The state discrimination module is used to extract the droplet motion trajectory based on the position change of the droplet bounding box in the fused image sequence, and to determine the droplet state according to a preset discrimination rule; The multi-source data synchronization module is used to collect multi-source environmental sensor data during each filling process and synchronize it with droplet trajectory and droplet state information in time to build a multi-dimensional monitoring dataset. The risk assessment module is used to calculate the total risk score based on the multidimensional monitoring dataset and obtain the residual risk level. The response control module is used to execute the corresponding response control strategy according to the residual risk level.
[0019] refer to Figure 2 A deep learning-based method for monitoring the production and filling of skin and mucous membrane disinfectant preparations includes the following steps: Step 1: Set up industrial cameras at multiple angles at the filling station for producing skin and mucous membrane disinfectant to collect image information of the filling needle during the filling cycle and obtain the original image sequence; Step 2: Preprocess the original image sequence to obtain a preprocessed image sequence; Step 3: Construct a regional potential energy distribution map, and fuse the regional potential energy distribution map with the preprocessed image sequence to generate a fused image sequence; Step 4: Input the fused image sequence into the improved Deformable DETR model for droplet target detection. The improved Deformable DETR model introduces a regional potential energy field focusing mechanism and outputs droplet bounding boxes. Step 5: Based on the positional changes of the droplet bounding box in the fused image sequence, extract the droplet trajectory and combine it with historical droplet trajectory data to determine the droplet state; Step 6: Collect corresponding multi-source environmental sensor data during each filling process, synchronize the multi-source environmental sensor data with the droplet motion trajectory and droplet state in time, and construct a multi-dimensional monitoring dataset; Step 7: Calculate the total risk score based on the multidimensional monitoring dataset to obtain the residual risk level, including low risk level, medium risk level and high risk level; Step 8: Implement the corresponding response control strategy based on the residual risk level.
[0020] In this embodiment, step one specifically includes: Industrial cameras are arranged at multiple angles relative to the filling needle at the filling station. The industrial cameras include a main camera located directly in front of the filling needle, and two auxiliary cameras respectively located on the left and right sides in front of the filling needle, forming an angle of 30° to 60° with the main camera in the horizontal direction. The installation height of the main camera and the auxiliary cameras is the same as the height of the filling needle nozzle. The industrial camera has an image resolution greater than 3840×2160 pixels, an image acquisition frame rate greater than 120 frames per second, and the acquired images are time-synchronized by an external trigger signal. The acquisition cycle corresponds to the filling cycle. An LED ring cold light source with an oblique incidence angle of 30° to 60° is set within the field of view of the industrial camera. The LED ring cold light source is installed above the filling needle to form a reflection enhancement area by oblique illumination. A background light shield and a linear backlight are set on the back of the filling needle. The surface reflectivity of the background light shield material is less than 15%, and the illuminance of the linear backlight ranges from 15,000 to 25,000 Lux. The acquired multi-angle image frames are arranged in chronological order to form an original image sequence.
[0021] In this embodiment, step two specifically includes: Each frame of the original image sequence is subjected to two-dimensional convolution processing using a 3×3 Laplacian convolution operator to obtain an edge response image; The edge response images are reduced to 1 / 2 and 1 / 4 of their original size, respectively, to obtain two downsampled images, which together with the edge response images form a three-layer image pyramid; The two downsampled images are upsampled to the same size as the edge response image using bilinear interpolation. The gray values of each pixel in the three images are then weighted and averaged to output a fused scale image. The pixel grayscale values of the fused scale image are divided into 8 equal-width grayscale intervals within the range of 0 to 255. RGB three-channel values are set for each interval. Pixels in each grayscale interval are assigned preset combinations of red, green, and blue channel values to generate a pseudo-color image. All pseudo-color images are arranged in chronological order of image acquisition time to obtain a preprocessed image sequence; In this invention, a multi-stage image preprocessing procedure is performed on the original image sequence to improve the robustness and accuracy of droplet detection and recognition. A 3×3 Laplacian convolution operator is applied to each frame of the image for two-dimensional convolution, highlighting the edge structure in the image and enhancing the droplet boundary features. The edge response images are reduced to 1 / 2 and 1 / 4 of their original size, constructing a three-layer image pyramid to represent the image's features at different spatial scales. Finally, bilinear interpolation is used to restore the downsampled image to a size similar to the original. Figure 1 To achieve consistent image size, a fused scale image is generated by weighted averaging of corresponding pixel grayscale values, thereby enhancing the unified expression of multi-scale information in the image. The fused scale image is then divided and mapped into grayscale ranges, and different grayscale segments are encoded according to preset red, green, and blue channel values to achieve pseudo-color enhancement display, which helps to improve the visual and computational dimensions of droplet features. All processed images are organized in chronological order to form a complete preprocessed image sequence, providing a unified standard and data support for regional potential construction and target detection.
[0022] In this embodiment, step three specifically includes: Using the pixel coordinates of the lower end of the filling needle in the image as the center point, establish a two-dimensional matrix with the same size as the image, and calculate the Euclidean distance from each pixel in the matrix to the center point. Divide the Euclidean distance by the length of the image diagonal to obtain a normalized distance value, and subtract the normalized distance value from 1 to obtain the regional potential energy value of each pixel, thereby generating a regional potential energy distribution map. The gray value of each pixel in the corresponding frame image in the preprocessed image sequence is multiplied by the regional potential energy value at the same position in the regional potential energy distribution map to obtain the fused pixel value, which constitutes the fused image frame. All fused image frames are arranged in chronological order of acquisition time to form a fused image sequence; In this invention, a regional potential energy distribution map is used to highlight image features in the area below the filling needle, enhancing the model's ability to focus on key droplet regions. A two-dimensional matrix of the same size as the image is established, centered on the pixel coordinates of the lower end of the filling needle in the image. The Euclidean distance from each pixel to the center point is calculated. The Euclidean distance is divided by the diagonal length of the image to obtain a normalized distance value. Subtracting this normalized value from 1 assigns a corresponding regional potential energy value to each pixel; the closer the value is to the center point, the higher the potential energy. The regional potential energy distribution map is multiplied pixel-by-pixel with the preprocessed image to generate a fused image frame, highlighting key region features, increasing the weight of key region information in the image, and simultaneously improving droplet detection accuracy.
[0023] In this embodiment, the improved Deformable DETR model specifically includes a backbone feature extraction network, an attention encoder, a region potential guided decoder, and a bounding box regression head; The backbone feature extraction network adopts the ResNet-50 structure, receives the fused image sequence and extracts image features through four stages of convolutional layers. The first stage outputs a shallow feature map with a size of 1 / 4 of the input image, the second stage outputs a mid-shallow feature map with a size of 1 / 8, the third stage outputs a mid-deep feature map with a size of 1 / 16, and the fourth stage outputs a deep feature map with a size of 1 / 32. The attention encoder consists of a stacked multi-layer Transformer encoder structure. Each layer includes a multi-head self-attention module and a feedforward neural network module, which encode the feature maps output from the four stages to generate four sets of encoded feature maps. Specifically, the multi-head self-attention module receives four sets of multi-scale feature maps output from the backbone feature extraction network and performs self-attention calculation on each set of feature maps. Each feature map is first transformed linearly to obtain multiple query, key, and value vector subspaces, and the dot product similarity between the query vector and the key vector is calculated. Combined with the Softmax operation and the weighted summation of the value vectors, the expression of each position in the global context is obtained, and the same-dimensional attention feature map is output.
[0024] The feedforward neural network module receives the output of the multi-head attention module and independently performs two-layer linear transformation and LeakyReLU activation operations on the feature vector at each location, further enhancing feature representation and nonlinear modeling capabilities. Each layer of the encoder repeats the above structure, updating four sets of feature maps layer by layer, which enhances cross-region information fusion while maintaining the original spatial structure. Finally, it outputs four sets of encoded feature maps, corresponding to 1 / 4, 1 / 8, 1 / 16, and 1 / 32 of the original input image size, respectively, which are used to guide the cross-attention mechanism in the region potential guide decoder. The entire process strengthens the modeling of semantic relationships between different scales and image regions, improving the feature representation effect of the droplet target. The regional potential energy guided decoder has several preset initial target query vectors. The initial target query vectors are generated by selecting several pixel positions with the highest regional potential energy values from the regional potential energy distribution map, extracting two-dimensional coordinates and performing linear position encoding, and serving as the initial input of the query vector. The region potential energy guided decoder introduces a region potential energy field focusing mechanism, performing element-wise multiplication of the initial target query vector with the corresponding region potential energy value in the region potential energy distribution map to obtain a weighted target query vector. This weighted target query vector interacts with the encoded feature map through cross-attention in a multi-layer decoder, iteratively updating the output target representation vector layer by layer. Specifically, in the region potential energy guided decoder, the weighted target query vector is sequentially input into multiple decoding layers, each containing a cross-attention calculation unit and a feedforward update unit. In the cross-attention calculation unit, the weighted target query vector is linearly mapped to generate query vector subspaces, and the feature vectors in the encoded feature map are linearly mapped to generate key vector and value vector subspaces. Centered on each weighted target query vector, the dot product similarity between the query vector and the key vectors at each position in the encoded feature map is calculated, and the similarity is scaled and Softmax normalized to obtain the attention weight distribution for the encoded feature map. This attention weight is then weighted and summed with the corresponding value vector to form the cross-attention output feature of the current decoding layer.
[0025] The cross-attention output features enter the feedforward update unit. Each target query vector is transformed independently through a two-layer fully connected structure. The first fully connected layer increases the feature dimension of the vector, and the second fully connected layer maps it back to the original dimension. ReLU activation is added between the two layers to form a complete feedforward update structure. The updated target query vector is added to the weighted target query vector through a residual connection and then normalized by the layer before being used as the input of the next decoding layer. The above processing is repeated in each decoding layer. The target query vector updated by the previous layer is passed to the next layer and cross-attention calculation and feedforward update are performed again with the encoded feature map. After iterative updates by multiple layers of decoders, the final output target representation vector contains the spatial structure information of the droplet target in the fused image sequence, providing feature basis for the bounding box regression head to generate the droplet bounding box; The bounding box regression head consists of two cascaded fully connected networks. The first fully connected network maps the target representation vector to an intermediate feature vector of length 128. The second fully connected network maps the intermediate feature vector to four-dimensional bounding box parameters, representing the center x-coordinate, center y-coordinate, width, and height, respectively, outputting the droplet bounding box. The intermediate feature vector is a fixed-length one-dimensional vector, and its dimensions are composed of high-dimensional semantic features extracted progressively by the first fully connected network and activation units. The second fully connected network contains a set of weight vectors of length 4 and corresponding bias terms. Each component in the intermediate feature vector is multiplied term by term with the corresponding component of the weight vector. All products are then summed and the bias is added to obtain the first output value representing the center x-coordinate. The same linear operation is used with the other three sets of weights and biases to generate three output values corresponding to the center y-coordinate, width, and height, respectively. Finally, the four output values are combined in sequence to form the bounding box parameters, which are used to describe the detection position of the droplet in the fused image sequence.
[0026] This invention improves the traditional Deformable DETR model in several ways to meet the visual monitoring requirements of small-scale, unstable-shaped droplets and strong background interference in the production and filling scenarios of skin and mucous membrane disinfectant preparations. A four-stage ResNet-50 structure is adopted in the backbone feature extraction part to obtain multi-scale features from shallow edge contours to deep semantic information, providing a hierarchical input basis for droplet localization. A multi-layer Transformer encoding structure is introduced into the attention encoder. Through sequential processing by a multi-head self-attention module and a feedforward network, a stable balance is achieved between long-range dependency modeling and local feature preservation in the four-stage feature maps, forming four sets of encoded feature maps that provide multi-scale correlation information for the decoder.
[0027] In the decoding stage, this invention introduces a regional potential energy field focusing mechanism as a core improvement. By extracting the positions of several pixels with the highest potential energy values from the regional potential energy distribution map and performing linear position encoding to generate an initial target query vector, the initial focus position of the decoder is consistent with the physical area where the droplet may appear. The initial target query vector is multiplied element-wise with the regional potential energy value, so that the high potential energy region is given higher weight, thereby enhancing the decoder's attention to the area below the filling needle. The weighted query vector performs cross-attention calculation with the encoded feature map layer by layer inside the decoder, so that the droplet position information gradually converges during the decoding process.
[0028] Through the above structural improvements, this invention improves the positioning accuracy of tiny droplets in complex filling scenarios and makes the model more consistent with the spatial distribution of droplets during the filling process, thereby improving the stability and repeatability of the droplet bounding box output.
[0029] In this embodiment, step five specifically includes: Based on the positional changes of the droplet bounding box in the fused image sequence, the coordinates of the center point of the droplet bounding box in consecutive frames are extracted, and the droplet motion trajectory is constructed in chronological order. Calculate the lateral and longitudinal displacements between points on the continuous droplet trajectory, and generate a velocity vector sequence based on the changing direction of the droplet bounding box center point between adjacent frames; The state of the droplet is determined according to a preset discrimination rule, which specifically includes: If the lateral displacement does not exceed the set displacement threshold in all consecutive frames, the vertical displacement continues to be downward, and the velocity vector change angle is less than or equal to the set amplitude threshold, the droplet is determined to be in a normal falling state. If the lateral displacement exceeds the set displacement threshold for three or more consecutive frames, and the velocity vector change angle is greater than the set amplitude threshold, the droplet state is determined to be a scattering and splashing state. If the droplet's trajectory is biased towards one side of the horizontal plane of the image, and the lateral displacement direction remains consistent across multiple consecutive frames while the cumulative displacement exceeds a set boundary distance, and the velocity vector direction deviates from the vertical direction by more than a set amplitude threshold across multiple consecutive frames, the droplet is determined to be in an offset and dissipation state.
[0030] In this embodiment, the multi-source environmental sensing data specifically includes ambient temperature data, ambient humidity data, and ambient air pressure data.
[0031] In this embodiment, step seven specifically includes: The droplet states in the multidimensional monitoring dataset are assigned different values according to their state types: 0 for normal falling state, 1 for scattering and splashing state, and 2 for offset and escaping state, to obtain the droplet state risk value. The lateral cumulative offset and the average change angle of the velocity vector of the droplet motion trajectory in the multidimensional monitoring dataset are normalized and then added together to obtain the trajectory risk value. The ambient temperature, ambient humidity, and ambient air pressure data from multi-source environmental sensors are compared with their respective set safety thresholds. For the portion exceeding the set safety threshold range, the deviation is multiplied by a preset proportional coefficient and then summed to obtain the environmental risk value. Specifically, the preset proportional coefficients are 0.4 for the temperature deviation, 0.3 for the humidity deviation, and 0.3 for the air pressure deviation. The environmental risk value is the weighted sum of the three product results, used to reflect the comprehensive impact of multi-source environmental factors on the residual risk in the filling process. The total risk score is calculated by adding the droplet state risk value, trajectory risk value, and environmental risk value. The residual risk level is obtained based on the total risk score. If the total risk score is less than a set first risk threshold, it is determined to be a low risk level; if the total risk score is greater than or equal to the set first risk threshold and less than or equal to a set second risk threshold, it is determined to be a medium risk level; if the total risk score is greater than the set second risk threshold, it is determined to be a high risk level.
[0032] In this embodiment, step eight specifically includes: When the residual risk level is low risk, the current filling process remains unchanged and the filling operation continues. When the residual risk level is medium risk level, the filling needle is automatically controlled to perform a cleaning action, including spraying cleaning fluid onto the outer wall of the filling needle and drying it with airflow, while briefly stopping the one-time filling operation, and resuming filling after cleaning is completed; When the residual risk level is high, the filling process should be immediately interrupted, and an audible and visual alarm should be triggered to alert the operator to conduct on-site investigation and handling.
[0033] Example 1
[0034] To verify the feasibility of this invention in practice, it was applied to an aseptic filling production line for a skin and mucous membrane disinfectant. This filling line is located in a Class 100 cleanroom core area and is equipped with a high-speed filling unit with a filling speed of 120 bottles / minute. It primarily targets applications in ophthalmology and gynecology, where microbial residue control is extremely sensitive, requiring high sensitivity and timely response to droplet states and environmental changes during the filling process. Traditional methods relying on manual sampling and simple pressure sensing cannot promptly detect risks such as droplet escaping and needle contamination, easily leading to trace residues that affect the product's sterility.
[0035] In actual deployment, three industrial cameras were installed at the filling station on the production line, located directly in front of the filling needle and diagonally above and to the left and right. These cameras used high-speed image acquisition equipment with a resolution of 4096×2160 and a frame rate of 150fps. The acquired raw image sequence was preprocessed by an image preprocessing module to construct a regional potential energy distribution map. Image enhancement focusing was performed on the center of the filling area. The regional potential energy distribution map was then fused with the preprocessed image sequence to generate a fused image sequence. This fused image sequence was fed into an improved Deformable DETR model incorporating a regional potential energy field focusing mechanism for droplet bounding box detection. This model demonstrated extremely high droplet boundary recognition accuracy and sensitivity to scattering in this scenario, achieving an accuracy rate of 98.7%, especially during the high-speed filling phase. In contrast, traditional image processing methods based on inter-frame difference and edge extraction, under the same test conditions, only achieved an average recognition accuracy of 91.2%, and frequently exhibited boundary blurring and missed scattering recognition during the high-speed droplet movement phase. The main reason is that traditional methods have limited use of spatial and contextual information of image features, making it difficult to accurately extract edge features of tiny droplets in complex backgrounds. However, the improved Deformable DETR model proposed in this invention integrates regional potential energy field focusing mechanism, which effectively improves the focusing ability of the core filling area. It still maintains high recognition accuracy in low contrast and high-speed motion backgrounds, verifying its advantages in high-precision filling scenarios.
[0036] During 72 hours of continuous operation, the system automatically collected droplet states, droplet trajectories, and multi-source environmental sensor data, completing a total of 8640 filling monitoring tasks. This data formed a multi-dimensional monitoring dataset, and the system was used to calculate the overall risk score and implement response control. The droplet states detected during system operation are shown in the table below.
[0037] Table 1. Droplet State Recognition Statistics Normal falling state 8422 97.48 1.2 3.5 Scattering and splashing state 152 1.76 6.9 24.3 Offset Dispersion State 66 0.76 12.5 31.7 As can be seen from the data in Table 1 above, the droplet state during the filling process was mainly in a normal falling state, accounting for as high as 97.48%, with a total of 8422 detections. This indicates that the filling system operated stably and reliably for the vast majority of the time. The average trajectory offset in this state was only 1.2 pixels, and the average velocity angle change was 3.5 degrees, indicating that the droplets maintained a basically stable falling state in the vertical direction, with small changes in trajectory and velocity, which meets the requirements of normal filling.
[0038] In contrast, the occurrence rates of scattering / splashing and drifting states were relatively low, at 1.76% and 0.76%, respectively. Although the frequency of occurrence was low, their characteristics were clearly abnormal. The average trajectory deviation of the scattering / splashing state reached 6.9 pixels, and the velocity angle change was as high as 24.3 degrees, indicating that the droplets were abnormally splashed due to external forces or airflow. The drifting state showed even greater trajectory deviation (12.5 pixels) and angular deviation (31.7 degrees), suggesting that the droplets may have deviated from the filling target area, posing a certain risk of residual contamination.
[0039] Through real-time acquisition of multi-source environmental sensor data, the system monitored temperature fluctuations ranging from 22.3 to 23.6℃, humidity fluctuations from 45.1% to 49.8%, and air pressure maintained between 101.2 and 102.3 kPa. Using set safety thresholds and preset proportional coefficients (temperature 0.4, humidity 0.3, air pressure 0.3), the system cumulatively assessed 18 high-risk levels, 47 medium-risk levels, and the remainder as low-risk states during 72 hours of operation. For different risk levels, the system automatically implemented different response control strategies. When a high-risk level occurred (mostly within 5 minutes after the cleaning cycle), the system immediately stopped filling, issued an audible and visual alarm, prompted the operator to check the needle status or environmental interference sources, and recorded the specific time point and environmental data.
[0040] This embodiment fully demonstrates the practical application value and technical advantages of the present invention in the production and filling process of skin and mucous membrane disinfectant preparations. By combining multi-angle image acquisition with regional potential energy distribution maps, the feature focusing capability of droplet images is effectively improved. An improved Deformable DETR model is introduced, achieving high-precision detection and intelligent state discrimination of droplet targets, maintaining strong robustness even in complex environments. Furthermore, through dynamic analysis of droplet trajectory and velocity vector changes, combined with multi-source environmental sensor data, a comprehensive multi-dimensional monitoring dataset is constructed. This further supports the quantitative scoring mechanism of the risk assessment module, the classification of risk levels, and the corresponding response control strategies. This enables the filling process to possess automated perception and response strategies, thereby effectively reducing filling deviation and residue risks, and improving the safety and stability of the production line.
[0041] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A deep learning-based skin mucous membrane disinfectant preparation production and filling monitoring system, characterized in that, Includes the following modules: The multi-angle image acquisition module is used to acquire image information of the filling needle during the filling cycle and obtain the original image sequence; The image preprocessing module is used to preprocess the original image sequence to obtain a preprocessed image sequence; The regional potential energy construction module is used to construct a regional potential energy distribution map, and then fuse the regional potential energy distribution map with the preprocessed image sequence to generate a fused image sequence. Specifically: Using the pixel coordinates of the lower end of the filling needle in the image as the center point, establish a two-dimensional matrix with the same size as the image, and calculate the Euclidean distance from each pixel in the matrix to the center point. Divide the Euclidean distance by the length of the image diagonal to obtain a normalized distance value, and subtract the normalized distance value from 1 to obtain the regional potential energy value of each pixel, thereby generating a regional potential energy distribution map. The gray value of each pixel in the corresponding frame image in the preprocessed image sequence is multiplied by the regional potential energy value at the same position in the regional potential energy distribution map to obtain the fused pixel value, which constitutes the fused image frame. All fused image frames are arranged in chronological order of acquisition time to form a fused image sequence; The droplet detection module is used to input the fused image sequence into the improved Deformable DETR model for droplet target detection. The improved Deformable DETR model introduces a regional potential energy field focusing mechanism and outputs droplet bounding boxes. The state discrimination module is used to extract the droplet motion trajectory based on the position change of the droplet bounding box in the fused image sequence, and to determine the droplet state according to a preset discrimination rule; The multi-source data synchronization module is used to collect multi-source environmental sensor data during each filling process and synchronize it with droplet trajectory and droplet state information in time to build a multi-dimensional monitoring dataset. The risk assessment module is used to calculate the total risk score based on the multidimensional monitoring dataset and obtain the residual risk level. The response control module is used to execute the corresponding response control strategy according to the residual risk level.
2. The deep learning-based skin-mucosa disinfectant preparation production and filling monitoring system according to claim 1, characterized by, The modules are connected in the following way: Step 1: Set up industrial cameras at multiple angles at the filling station to collect image information of the filling needle during the filling cycle and obtain the original image sequence; Step 2: Preprocess the original image sequence to obtain a preprocessed image sequence; Step 3: Construct a regional potential energy distribution map, and fuse the regional potential energy distribution map with the preprocessed image sequence to generate a fused image sequence; Step 4: Input the fused image sequence into the improved Deformable DETR model for droplet target detection. The improved Deformable DETR model introduces a regional potential energy field focusing mechanism and outputs droplet bounding boxes. Step 5: Extract the droplet trajectory based on the positional changes of the droplet bounding box in the fused image sequence, and determine the droplet state according to a preset discrimination rule; Step 6: Collect corresponding multi-source environmental sensor data during each filling process, synchronize the multi-source environmental sensor data with the droplet motion trajectory and droplet state in time, and construct a multi-dimensional monitoring dataset; Step 7: Calculate the total risk score based on the multidimensional monitoring dataset to obtain the residual risk level, including low risk level, medium risk level and high risk level; Step 8: Execute the corresponding response control strategy based on the residual risk level.
3. The deep learning-based skin-mucosa disinfectant preparation production and filling monitoring system according to claim 2, characterized by, Step one specifically involves: Industrial cameras are arranged at multiple angles relative to the filling needle at the filling station. The industrial cameras include a main camera located directly in front of the filling needle, and two auxiliary cameras respectively located on the left and right sides in front of the filling needle, forming an angle of 30° to 60° with the main camera in the horizontal direction. The installation height of the main camera and the auxiliary cameras is the same as the height of the filling needle nozzle. The industrial camera has an image resolution greater than 3840×2160 pixels, an image acquisition frame rate greater than 120 frames per second, and the acquired images are time-synchronized by an external trigger signal. The acquisition cycle corresponds to the filling cycle. An LED ring cold light source with an oblique incidence angle of 30° to 60° is set within the field of view of the industrial camera. The LED ring cold light source is installed above the filling needle to form a reflection enhancement area by oblique illumination. A background light shield and a linear backlight are set on the back of the filling needle. The surface reflectivity of the background light shield material is less than 15%, and the illuminance of the linear backlight ranges from 15,000 to 25,000 Lux. The acquired multi-angle image frames are arranged in chronological order to form an original image sequence.
4. The deep learning-based skin-mucosa disinfectant preparation production and filling monitoring system according to claim 2, characterized by, Step two specifically involves: Each frame of the original image sequence is subjected to two-dimensional convolution processing using a 3×3 Laplacian convolution operator to obtain an edge response image; The edge response images are reduced to 1 / 2 and 1 / 4 of their original size, respectively, to obtain two downsampled images, which together with the edge response images form a three-layer image pyramid; The two downsampled images are upsampled to the same size as the edge response image using bilinear interpolation. The gray values of each pixel in the three images are then weighted and averaged to output a fused scale image. The pixel grayscale values of the fused scale image are divided into 8 equal-width grayscale intervals within the range of 0 to 255. RGB three-channel values are set for each interval. Pixels in each grayscale interval are assigned preset combinations of red, green, and blue channel values to generate a pseudo-color image. All pseudo-color images are arranged in chronological order of image acquisition time to obtain a preprocessed image sequence.
5. The deep learning-based skin-mucosa disinfectant preparation production and filling monitoring system according to claim 2, characterized by, The improved Deformable DETR model specifically includes a backbone feature extraction network, an attention encoder, a region potential guided decoder, and a bounding box regression head. The backbone feature extraction network adopts the ResNet-50 structure, receives the fused image sequence and extracts image features through four stages of convolutional layers. The first stage outputs a shallow feature map with a size of 1 / 4 of the input image, the second stage outputs a mid-shallow feature map with a size of 1 / 8, the third stage outputs a mid-deep feature map with a size of 1 / 16, and the fourth stage outputs a deep feature map with a size of 1 / 32. The attention encoder is composed of a multi-layer Transformer encoder structure stacked together. Each layer includes a multi-head self-attention module and a feedforward neural network module, which encode the feature maps output from the four stages respectively to generate four sets of encoded feature maps. The regional potential energy guided decoder has several preset initial target query vectors. The initial target query vectors are generated by selecting several pixel positions with the highest regional potential energy values from the regional potential energy distribution map, extracting two-dimensional coordinates and performing linear position encoding, and serving as the initial input of the query vector. The regional potential energy guided decoder introduces a regional potential energy field focusing mechanism, which performs element-wise multiplication of the initial target query vector with the regional potential energy value at the corresponding position in the regional potential energy distribution map to obtain a weighted target query vector; the weighted target query vector interacts with the encoded feature map through cross-attention in the multi-layer decoder, and iteratively updates the output target representation vector layer by layer. The bounding box regression head consists of two cascaded fully connected networks. The first fully connected network maps the target representation vector to an intermediate feature vector of length 128, and the second fully connected network maps the intermediate feature vector to four-dimensional bounding box parameters, representing the center x-coordinate, center y-coordinate, width, and height, respectively, and outputs the droplet bounding box.
6. The deep learning-based skin-mucosa disinfectant preparation production and filling monitoring system according to claim 2, characterized by Step five specifically involves: Based on the positional changes of the droplet bounding box in the fused image sequence, the coordinates of the center point of the droplet bounding box in consecutive frames are extracted, and the droplet motion trajectory is constructed in chronological order. Calculate the lateral and longitudinal displacements between points on the continuous droplet trajectory, and generate a velocity vector sequence based on the changing direction of the droplet bounding box center point between adjacent frames; The state of the droplet is determined according to a preset discrimination rule, which specifically includes: If the lateral displacement does not exceed the set displacement threshold in all consecutive frames, the vertical displacement continues to be downward, and the velocity vector change angle is less than or equal to the set amplitude threshold, the droplet is determined to be in a normal falling state. If the lateral displacement exceeds the set displacement threshold for three or more consecutive frames, and the velocity vector change angle is greater than the set amplitude threshold, the droplet state is determined to be a scattering and splashing state. If the droplet's trajectory is biased towards one side of the horizontal plane of the image, and the lateral displacement direction remains consistent across multiple consecutive frames while the cumulative displacement exceeds a set boundary distance, and the velocity vector direction deviates from the vertical direction by more than a set amplitude threshold across multiple consecutive frames, the droplet is determined to be in an offset and dissipation state.
7. The deep learning-based monitoring system for the production and filling of skin and mucous membrane disinfectant preparations according to claim 2, characterized in that, The multi-source environmental sensing data specifically includes ambient temperature data, ambient humidity data, and ambient air pressure data.
8. The deep learning-based monitoring system for the production and filling of skin and mucous membrane disinfectant preparations according to claim 2, characterized in that, Step seven specifically involves: The droplet states in the multidimensional monitoring dataset are assigned different values according to their state types: 0 for normal falling state, 1 for scattering and splashing state, and 2 for offset and escaping state, to obtain the droplet state risk value. The lateral cumulative offset and the average change angle of the velocity vector of the droplet motion trajectory in the multidimensional monitoring dataset are normalized and then added together to obtain the trajectory risk value. The ambient temperature, ambient humidity, and ambient air pressure from the multi-source environmental sensor data are compared with their respective set safety thresholds. For the portion that exceeds the set safety threshold range, the deviation is multiplied by a preset proportional coefficient and then added to obtain the environmental risk value. The total risk score is calculated by adding the droplet state risk value, trajectory risk value, and environmental risk value. The residual risk level is obtained based on the total risk score. If the total risk score is less than a set first risk threshold, it is determined to be a low risk level. If the total risk score is greater than or equal to the set first risk threshold and less than or equal to the set second risk threshold, it is determined to be at a medium risk level. If the total risk score is greater than the set second risk threshold, it is determined to be a high-risk level.
9. The deep learning-based monitoring system for the production and filling of skin and mucous membrane disinfectant preparations according to claim 2, characterized in that, Step eight specifically involves: When the residual risk level is low risk, the current filling process remains unchanged and the filling operation continues. When the residual risk level is medium risk level, the filling needle is automatically controlled to perform a cleaning action, including spraying cleaning fluid onto the outer wall of the filling needle and drying it with airflow, while briefly stopping the one-time filling operation, and resuming filling after cleaning is completed; When the residual risk level is high, the filling process should be immediately interrupted, and an audible and visual alarm should be triggered to alert the operator to conduct on-site investigation and handling.
Citation Information
Patent Citations
Depth vision-based beverage bottle filling defect identification method
CN115330702A
Railway pedestrian warning system and method based on improved YOLOv7 algorithm
CN120299063A