A visual enhancement method and system for underground monorail based on dynamic modeling
By using dynamic modeling methods, the image instability problem caused by track impact and vibration during the operation of underground monorail cranes was solved, realizing accurate visual monitoring and enhanced stability of underground monorail cranes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI UNIV OF SCI & TECH
- Filing Date
- 2026-03-13
- Publication Date
- 2026-06-23
Smart Images

Figure CN122265088A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and more specifically, to a method and system for visual enhancement of underground monorail cranes based on dynamic modeling. Background Technology
[0002] Underground monorails are used for material transportation and inspection in tunnels. At the operation site of underground monorails, it is usually necessary to use camera equipment to obtain video frames of the operation of the underground monorail for track status observation and operation environment identification.
[0003] Existing technologies typically do not consider the dynamic relationship between track impact, vehicle vibration and load sway, making it difficult to effectively suppress the resulting motion blur and image instability, and thus failing to meet the need for accurate visual monitoring of the monorail's operating status in complex underground environments.
[0004] To address the aforementioned problems, a technical solution is provided. Summary of the Invention
[0005] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a method and system for visual enhancement of a downhole monorail based on dynamic modeling to solve the problems mentioned in the background art.
[0006] To achieve the above objectives, the present invention provides the following technical solution: A visual enhancement method for underground monorail cranes based on dynamic modeling includes the following steps: S1: Obtain video frames of the underground monorail crane operation and corresponding timestamp data and exposure time parameters, and construct exposure window time series data based on timestamp data and exposure time parameters; S2: Perform track joint impact feature detection on the exposure window time series data, identify the transient impact interval in the exposure window time series data, and generate transient impact interval identification data; S3: Based on transient impact interval identification data, establish a coupled dynamic model of the vibration of the underground monorail crane body and the swing of the load, calculate the time-varying pose response sequence within the exposure window, and generate time-varying pose response data; S4: Based on the time-varying pose response data, the exposure window time series data is piecewise integrated to construct a composite fuzzy kernel model corresponding to the transient impact interval and generate composite fuzzy kernel data. S5: Based on the composite fuzzy kernel data, constrained deconvolution is performed on the running video frames to generate image data with enhanced structural consistency; S6: Perform cross-frame consistency verification on the structural consistency enhancement image data and output the visual enhancement image of the underground monorail crane.
[0007] In a preferred embodiment, S1 specifically refers to: Acquire video frames of the underground monorail crane in operation, read the timestamp data corresponding to the video frames of the underground monorail crane in operation, and extract the exposure time parameters; The start and end times of exposure are determined based on timestamp data; The exposure window of the video frame of the underground monorail crane operation is extracted and aligned with the time according to the exposure time parameter to form the exposure window time series data.
[0008] In a preferred embodiment, S2 specifically refers to: Based on the exposure window time series data, the gray-level difference sequence and edge gradient change sequence between adjacent exposure windows are calculated, and the gray-level change amplitude curve and edge gradient fluctuation curve are constructed. Local extreme value search and duration determination are performed on the grayscale change amplitude curve and the edge gradient fluctuation curve. The interval that simultaneously satisfies the amplitude change and the duration being less than the exposure time parameter is selected to determine the transient impact interval. The exposure window time series data corresponding to the transient impact range is time-stamped to form transient impact range identification data.
[0009] In a preferred embodiment, S3 specifically refers to: Based on the transient impact interval identification data, the corresponding exposure window time series data is obtained, and the orbital impact position and impact time information in the exposure window time series data are extracted. A vibration excitation sequence for a downhole monorail crane body is constructed based on the information of track impact location and impact time. Using the vibration excitation sequence of the underground monorail crane body as input, a coupled dynamic model of the vibration of the underground monorail crane body and the swing of the load is established; The time-varying pose response sequence within the exposure window is obtained by solving the coupled dynamics model, thus generating time-varying pose response data.
[0010] In a preferred embodiment, S4 specifically refers to: The vibration velocity variation curve within the exposure window is determined based on time-varying pose response data; Based on the vibration velocity variation curve, the exposure window time series data is piecewise integrated to generate time-varying blurred trajectory data; By performing a two-dimensional convolution kernel transformation on time-varying fuzzy trajectory data, a composite fuzzy kernel model corresponding to the transient impact interval is constructed, generating composite fuzzy kernel data.
[0011] In a preferred embodiment, S5 specifically refers to: Using composite fuzzy kernel data as constraints, perform image domain deconvolution operation on running video frames, calculate image gray-level gradient distribution, and determine the deconvolution iteration convergence boundary; By constraining the edge stability of the image gray-level gradient distribution based on the convergence boundary of deconvolution iteration, image gray-level data that meets the structural consistency condition is obtained, and image data with enhanced structural consistency is generated.
[0012] In a preferred embodiment, S6 specifically refers to: Based on structural consistency enhancement image data, the spatial location distribution of edge features within consecutive multi-frame images is extracted, and the change amplitude of edge feature spatial location between adjacent frames is calculated. Using the stability threshold of the change amplitude as a cross-frame consistency constraint, structural consistency enhancement image data that meets the cross-frame consistency constraint is selected, and visual enhancement images of the underground monorail are output.
[0013] On the other hand, the present invention provides a vision enhancement system for a downhole monorail based on dynamic modeling, comprising: Data acquisition module: acquires video frames of the underground monorail crane operation and corresponding timestamp data and exposure time parameters, and constructs exposure window time series data based on the timestamp data and exposure time parameters; Impact detection module: performs track joint impact feature detection on exposure window time series data, identifies transient impact intervals in exposure window time series data, and generates transient impact interval identification data; Dynamic modeling module: Based on transient impact interval identification data, a coupled dynamic model of the vibration of the underground monorail crane body and the swing of the load is established, the time-varying pose response sequence within the exposure window is calculated, and time-varying pose response data is generated; Piecewise integration module: Based on the time-varying pose response data, the exposure window time series data is piecewise integrated to construct a composite fuzzy kernel model corresponding to the transient impact interval and generate composite fuzzy kernel data; Constrained Convolution Module: Based on composite fuzzy kernel data, constrained deconvolution is performed on running video frames to generate image data with enhanced structural consistency; Consistency verification module: Performs cross-frame consistency verification on the structural consistency enhancement image data and outputs visual enhancement images of the underground monorail crane.
[0014] The technical effects and advantages of the visual enhancement method and system for underground monorail cranes based on dynamic modeling, as described in this invention: By acquiring video frames and corresponding timestamps of the underground monorail crane's operation, along with exposure time parameters, and constructing exposure window time series data, the temporal boundaries of image degradation are uniformly expressed. By detecting track joint impact features on the exposure window time series data and generating transient impact interval identification data, abnormal imaging sections caused by track impacts are accurately located and distinguished from normal sections. Based on the transient impact interval identification data, a coupled dynamic model of the underground monorail crane's body vibration and load sway is established, and time-varying pose response data is generated, enabling a structured characterization of the composite motion during imaging and providing constraints for the blurring mechanism. Based on time-varying pose response data, the exposure window time series data is segmented and integrated to generate composite blur kernel data, so that the composite blur kernel model corresponds with the transient impact interval and improves the consistency between the blur kernel and degradation. Based on the composite blur kernel data, constrained deconvolution is performed on the running video frames to generate structure consistency enhancement image data, so that the deblurring process is physically constrained and false edges and ringing interference are reduced. By performing cross-frame consistency verification on the structure consistency enhancement image data and outputting the underground monorail visual enhancement image, the enhancement results are kept consistent in the time dimension and the stability and reliability of visual recognition of track defects and aerial obstacles are improved. Attached Figure Description
[0015] Figure 1 This is a schematic diagram of a visual enhancement method for a downhole monorail crane based on dynamic modeling according to the present invention. Figure 2 This is a schematic diagram of the structure of a visual enhancement system for a downhole monorail crane based on dynamic modeling, according to the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0017] Example 1 Figure 1 This invention presents a visual enhancement method for underground monorail cranes based on dynamic modeling, which includes the following steps: S1: Obtain video frames of the underground monorail crane operation and corresponding timestamp data and exposure time parameters, and construct exposure window time series data based on timestamp data and exposure time parameters; S2: Perform track joint impact feature detection on the exposure window time series data, identify the transient impact interval in the exposure window time series data, and generate transient impact interval identification data; S3: Based on transient impact interval identification data, establish a coupled dynamic model of the vibration of the underground monorail crane body and the swing of the load, calculate the time-varying pose response sequence within the exposure window, and generate time-varying pose response data; S4: Based on the time-varying pose response data, the exposure window time series data is piecewise integrated to construct a composite fuzzy kernel model corresponding to the transient impact interval and generate composite fuzzy kernel data. S5: Based on the composite fuzzy kernel data, constrained deconvolution is performed on the running video frames to generate image data with enhanced structural consistency; S6: Perform cross-frame consistency verification on the structural consistency enhancement image data and output the visual enhancement image of the underground monorail crane.
[0018] S1: Acquire video frames of the underground monorail crane operation and corresponding timestamp data and exposure time parameters. Construct exposure window time series data based on the timestamp data and exposure time parameters, including: Acquire video frames of the underground monorail crane in operation, read the timestamp data corresponding to the video frames of the underground monorail crane in operation, and extract the exposure time parameters; Specifically, industrial-grade cameras installed at fixed positions on the underground monorail crane body capture real-time video images of the monorail crane's operation at a set sampling frequency, obtaining continuous video frame images. The sampling frequency of the industrial-grade cameras is set as follows: The minimum image clarity requirement is determined based on the rated operating speed of the underground monorail crane body and the ambient lighting conditions. Then, the maximum exposure time that meets the motion blur tolerance range is determined, and the camera's sampling frequency is determined accordingly. For example, when the rated operating speed of the underground monorail crane is 3 meters per second, and the accuracy requirement for identifying track defects is within 10 millimeter, the exposure time parameter is set to 10 to 15 milliseconds, and the camera sampling frequency is determined to be 80 to 100 frames per second.
[0019] The timestamp data of each underground monorail crane operation video frame is read. Specifically, the frame synchronization timestamp data generated by the camera when each video image is acquired is extracted from the time synchronization module built into the camera. The extracted frame synchronization timestamp data is the absolute time information corresponding to each video frame image, with a time accuracy of milliseconds.
[0020] The method for extracting the exposure time parameter is as follows: read the exposure time parameter information corresponding to each video frame of the underground monorail crane operation from the exposure control module set in the camera; the exposure time parameter is the exposure duration of the camera image sensor when a single frame image is acquired, in milliseconds; when the industrial-grade camera adopts automatic exposure mode, the exposure time parameter is determined according to the real-time detection result of the underground ambient light intensity. The determination method is that the camera detects the overall light intensity of the scene before each frame is acquired in real time through the internal exposure control module and feeds back the light intensity, and then matches the corresponding exposure time parameter according to the pre-stored exposure lookup table based on the fed-back light intensity to ensure that the brightness of each video image frame is basically consistent.
[0021] The start and end times of exposure are determined based on timestamp data; Specifically, for each frame of video footage of the underground monorail crane in operation, the corresponding timestamp data is used as the exposure end time, and the exposure start time is calculated by subtracting the duration of the corresponding exposure time parameter from the exposure end time.
[0022] Based on the exposure time parameters, the video frames of the underground monorail crane operation are extracted into exposure windows and aligned with time to form exposure window time series data. Specifically, based on the exposure start and end times of each video frame, video image data corresponding to the exposure time window is extracted from the continuously acquired video frame sequence. That is, the image data located between the exposure start and end times is extracted as video image segments corresponding to the exposure window, and the exposure window time information corresponding to each video image segment is formed. Then, by using a unified exposure end time as a reference point, the video image segments of different exposure windows are time-axis aligned: using the exposure end time in the timestamp data as the standard reference time, the video image segments of different exposure windows are calibrated and aligned on the time axis, so that the exposure window time information corresponding to each video image segment maintains a temporal correspondence, thereby forming exposure window time series data.
[0023] Ultimately, the resulting exposure window time series data includes: video image segment data for each exposure window, the exposure start time and exposure end time corresponding to the video image segment, and the corresponding exposure time parameter information.
[0024] S2: Perform track joint impact feature detection on the exposure window time series data, identify transient impact intervals in the exposure window time series data, and generate transient impact interval identification data, including: Based on the exposure window time series data, the gray-level difference sequence and edge gradient change sequence between adjacent exposure windows are calculated, and the gray-level change amplitude curve and edge gradient fluctuation curve are constructed. Specifically, based on the exposure window time series data, a gray-level difference sequence is calculated between adjacent exposure windows. For two consecutive adjacent exposure window video image segments, the gray-level value corresponding to each pixel in the two exposure window video image segments is extracted. The gray-level difference is calculated for the corresponding pixels in the two video image segments, that is, the pixel gray-level value of the later exposure window video image segment is subtracted from the pixel gray-level value of the corresponding pixel at the same position in the earlier exposure window video image segment, resulting in a gray-level difference sequence. The gray-level difference value in each gray-level difference sequence represents the gray-level change of the corresponding pixel between two consecutive exposure windows, and the value range is from -255 to 255.
[0025] Simultaneously, the edge gradient change sequence between adjacent exposure windows is calculated based on the exposure window time series data. Spatial edge gradient operations are performed on video image segments from two adjacent exposure windows respectively. The operation method uses the Sobel operator to perform convolution operations on the video image segment data to obtain the edge gradient value distribution of the corresponding exposure window video image segment data. The Sobel operator's convolution kernel size is 3×3, and the numerical coefficients of the convolution kernel are: horizontal convolution kernel [[-1,0,1],[-2,0,2],[-1,0,1]], and vertical convolution kernel [[-1,-2,-1],[0,0,0],[1,2,1]]. Convolution is performed using the horizontal and vertical convolution kernels. The edge gradients in two directions are then obtained, and the square root of the sum of the squares of the gradients in the two directions is taken to obtain the edge gradient value, which represents the edge intensity at the corresponding pixel position in the video image segment data. The edge gradient values calculated for the video image segment data of two consecutive exposure windows are then subtracted pixel by pixel, that is, the edge gradient value of the video image segment data of the later exposure window is subtracted from the edge gradient value of the same position in the video image segment data of the earlier exposure window, to obtain the edge gradient change sequence. The range of data values in the edge gradient change sequence is determined according to the actual edge gradient difference.
[0026] Based on gray-level difference sequences and edge gradient change sequences, gray-level change amplitude curves and edge gradient fluctuation curves are constructed respectively. The gray-level difference sequences and edge gradient change sequences are processed by absolute value analysis of their spatial locations to eliminate sign differences in the gray-level difference and edge gradient change values, ensuring that the amplitude reflects the absolute degree of change. Then, the processed absolute value data are summed in the spatial domain for each exposure window to generate the corresponding gray-level change amplitude curve and edge gradient fluctuation curve. The gray-level change amplitude curve represents the cumulative degree of overall gray-level change between each exposure window and its adjacent exposure windows, while the edge gradient fluctuation curve represents the cumulative degree of edge feature change between each exposure window and its adjacent exposure windows.
[0027] Local extreme value search and duration determination are performed on the grayscale change amplitude curve and the edge gradient fluctuation curve. The interval that simultaneously satisfies the amplitude change and the duration being less than the exposure time parameter is selected to determine the transient impact interval. Specifically, for both the grayscale change amplitude curve and the edge gradient fluctuation curve, an extreme value search is performed point by point, starting from the curve's starting point. That is, for each curve data point, it is determined whether the values of the two preceding and following data points are both less than the current data point value. If so, the data point is determined to be a local maximum. For each identified local maximum location, the duration corresponding to multiple consecutive exposure windows before and after the local maximum location is analyzed in conjunction with timestamp data to determine the duration of the amplitude abrupt change in the grayscale change amplitude curve and the edge gradient fluctuation curve at the local maximum location. If the exposure window duration corresponding to the amplitude abrupt change at the local maximum location is less than the corresponding exposure time parameter, it is determined to be a valid transient impact feature location.
[0028] The exposure window interval that simultaneously satisfies the local extreme positions of both the grayscale change amplitude curve and the edge gradient fluctuation curve, and where the duration of the amplitude abrupt change is less than the corresponding exposure time parameter, is defined as the transient impact interval. For each determined valid transient impact feature position, the start and end positions of the exposure window corresponding to the grayscale change amplitude curve and the edge gradient fluctuation curve are determined respectively. Then, the intersection operation is performed on the exposure window positions corresponding to the grayscale change amplitude curve and the edge gradient fluctuation curve, that is, the exposure window interval that simultaneously satisfies the condition of the duration of the amplitude abrupt change in the grayscale change amplitude curve and the edge gradient fluctuation curve is defined as the transient impact interval.
[0029] The exposure window time series data corresponding to the transient impact range is time-stamped to form transient impact range identification data; Specifically, in the exposure window time series data, for each exposure window video image segment data that meets the transient impact interval condition, its corresponding start and end time information is marked as the time stamp information of the transient impact interval; the transient impact interval identification data includes the start and end times of the exposure window video image segment data, and the identification data format is "start time-end time", for example, the identification data is "2024-02-14 13:10:23.003-2024-02-1413:10:23.015", which indicates the time range of the start and end of the video image segment data corresponding to the transient impact interval.
[0030] S3: Based on transient impact interval identification data, establish a coupled dynamic model of the vibration of the underground monorail crane body and the swing of the load, calculate the time-varying pose response sequence within the exposure window, and generate time-varying pose response data, including: Based on the transient impact interval identification data, the corresponding exposure window time series data is obtained, and the orbital impact position and impact time information in the exposure window time series data are extracted. Specifically, based on the start and end times in the transient impact interval identifier data, video image segment data corresponding to the above time range is extracted from the exposure window time series data. This includes all pixel grayscale values, exposure start time, exposure end time, and exposure time parameter information in the video image segment data. The exposure window time series data corresponding to each transient impact interval is a two-dimensional array, where the horizontal dimension represents the spatial coordinates of the video image pixel position, and the vertical dimension represents the pixel grayscale value. The corresponding exposure start and end time information is associated as an independent data label.
[0031] The impact location and time information of the track were extracted from the time-series data of the exposure window. Spatial domain grayscale variation analysis was performed on the video image segment data in the two-dimensional array, and the spatial domain pixel grayscale variation variance analysis method was used to determine the track impact location. The standard deviation of the grayscale value of each column of pixels was calculated pixel by pixel along the running direction of the underground monorail, and the location with the largest change in standard deviation was taken as the track impact location. Based on the change trend of the pixel grayscale value corresponding to the track impact location with the exposure start to end time in the exposure window time-series data, the precise time when the maximum grayscale difference occurred was determined by frame-by-frame comparison, which was used as the impact time information. For example, the impact time information was determined by the time when the maximum pixel grayscale difference occurred between consecutive exposure windows.
[0032] A vibration excitation sequence for a downhole monorail crane body is constructed based on the information of track impact location and impact time. Specifically, the method for constructing the vibration excitation sequence of the underground monorail crane is as follows: The determined track impact position is used as the spatial excitation point, and the determined impact time information is used as the temporal excitation point. The sequence is constructed using a combination of a temporal impact function and a spatial position function. The temporal impact function is simulated using a Gaussian function, and the spatial position function is determined based on the track impact position. The Gaussian function is defined as: ;in, This is the time-domain impulse function expressed as a Gaussian function; The amplitude of the time-domain vibration shock excitation is determined based on the peak value of the grayscale difference in the time series data of the exposure window, for example, it is set as a proportional multiple of the maximum value of the grayscale difference; t is any time in the video image exposure window, in milliseconds; The time value determined for the impact time information, in milliseconds; The width parameter of the Gaussian function represents a control parameter indicating the duration of the impact excitation, for example, set to half of the exposure time parameter; Using the vibration excitation sequence of the underground monorail crane body as input, a coupled dynamic model of the vibration of the underground monorail crane body and the swing of the load is established; Specifically, the coupled dynamic model of the vibration of the underground monorail crane body and the oscillation of the load is a time-varying coupled equation set constructed by combining rigid body dynamics with the oscillation of the oscillator:
[0033] .in, This indicates the horizontal displacement response of the underground monorail crane body, expressed in meters. , These represent the first and second time derivatives of the horizontal displacement response of the underground monorail crane body, respectively. This indicates the mass of the underground monorail crane, expressed in kilograms. , These represent the damping coefficient and stiffness coefficient of the vehicle body, respectively, determined experimentally based on the mechanical structural characteristics of the underground monorail crane body. Indicates the load mass, expressed in kilograms; Indicates the suspension length of the swing load, in meters; This indicates the swing angle of the suspended load, in radians. , These represent the first and second time derivatives of the swing angle, respectively; The damping coefficient of the swing load is expressed in Newton-meter-second per radian; it is determined by fitting experimentally measured swing response data. This represents the acceleration due to gravity, with a value of 9.8 meters per second squared.
[0034] The time-varying pose response sequence within the exposure window is obtained by solving the coupled dynamics model, and time-varying pose response data is generated. Specifically, the solution method employs the fourth-order Runge-Kutta method to solve the dynamic coupling equations. The initial boundary conditions are set as follows: the underground monorail crane and the load are stationary before impact, meaning the initial displacement and velocity values are both zero. The calculation step size is set to one-tenth of the exposure time parameter. For example, when the exposure time parameter is 12 milliseconds, the calculation step size is set to 1.2 milliseconds.
[0035] Finally, by solving the equations of the coupled dynamics model using the fourth-order Runge-Kutta numerical integration method, the time-varying pose response sequence of the underground monorail crane body at each calculation step within the exposure window was obtained. The time-varying pose response data includes the horizontal displacement response. and swing angle response The data values and corresponding times.
[0036] S4: Based on the time-varying pose response data, the exposure window time series data is piecewise integrated to construct a composite fuzzy kernel model corresponding to the transient impact interval, generating composite fuzzy kernel data, including: The vibration velocity variation curve within the exposure window is determined based on time-varying pose response data; Specifically, using the horizontal displacement response data of the underground monorail crane body in the time-varying pose response data, the instantaneous vibration velocity is obtained by calculating the ratio of the displacement difference between adjacent sampling times to the corresponding time difference. Then, based on the timestamp data corresponding to the instantaneous vibration velocity, a continuous vibration velocity variation curve is generated with each sampling time within the exposure window as the horizontal axis and the instantaneous vibration velocity as the vertical axis. The calculation method for each instantaneous vibration velocity is as follows: subtract the displacement response data value of the underground monorail crane body at the previous time from the current time, and divide the difference by the time difference between the current time and the previous time to obtain the vibration velocity, in meters per second. For example, if the current time is 13:10:23.006 and the corresponding displacement response is 0.008 meters, and the previous time was 13:10:23.004 and the corresponding displacement response was 0.005 meters, then the instantaneous vibration velocity is calculated as: (0.008 meters - 0.005 meters) / (0.002 seconds) = 1.5 meters per second. The instantaneous vibration velocity is then assigned to the current time to form a vibration velocity change curve.
[0037] Based on the vibration velocity variation curve, the exposure window time series data is piecewise integrated to generate time-varying blurred trajectory data; Specifically, the sampling time corresponding to each instantaneous vibration velocity in the vibration velocity variation curve is taken as the integration starting point. The integration time length is determined according to the exposure time parameter corresponding to the exposure window. The trapezoidal integration method is used. For all instantaneous vibration velocities within each integration time length, the average value of two adjacent instantaneous vibration velocities is taken point by point, and then multiplied by the time difference between adjacent times to obtain the integration value between each cell. The integration values between each cell within the integration time length are then summed to form the total integration value within the integration time length. The integration result is the time-varying fuzzy trajectory data value corresponding to the exposure window, in meters. For example, if the exposure time parameter corresponding to the exposure window is 12 milliseconds, and the integration start time is 13:10:23.006, then the integration end time is 13:10:23.018. The trapezoidal integral of adjacent instantaneous vibration velocities is calculated within the above interval and summed to obtain the time-varying fuzzy trajectory data corresponding to the exposure window.
[0038] By performing two-dimensional convolution kernel transformation on time-varying fuzzy trajectory data, a composite fuzzy kernel model corresponding to the transient impact interval is constructed, and composite fuzzy kernel data is generated. Specifically, the method for two-dimensional convolution kernel transformation is as follows: a two-dimensional fuzzy kernel matrix is constructed based on the displacement fuzzy path represented by time-varying fuzzy trajectory data; the size of the fuzzy kernel matrix is determined by using the spatial displacement corresponding to the maximum time-varying fuzzy trajectory data value as the matrix side length, and converting the spatial displacement into the number of pixels based on the pixel spatial resolution of the camera. For example, if the maximum displacement fuzzy path data value is 0.02 meters and the pixel spatial resolution of the camera is 0.001 meters per pixel, then the side length of the fuzzy kernel matrix is determined to be 20 pixels; the initial values within the matrix are set to zero.
[0039] The two-dimensional convolution kernel matrix is assigned values as follows: Using the center point of the matrix as the initial point, the matrix is filled according to the displacement and direction of each integration point corresponding to the time-varying fuzzy trajectory data. The filling direction of the fuzzy trajectory is determined by the displacement direction of the time-varying fuzzy trajectory corresponding to each integration point, and the filling length of the fuzzy trajectory is determined by the magnitude of the integration value corresponding to each integration point. Matrix elements along the filling path are assigned weights corresponding to the integration values. The weights are determined using a normalization method, dividing the current integration value by the sum of all integration values to obtain weights between 0 and 1, which are then used as matrix elements along the filling path. Through this method, the two-dimensional fuzzy kernel matrix is filled.
[0040] Based on the generated two-dimensional fuzzy kernel matrix, and combined with the trajectory impact position and impact time information extracted from the video image segment data corresponding to the exposure window, the two-dimensional fuzzy kernel matrix is spatially translated to ensure that the composite fuzzy kernel model is aligned with the spatial coordinates of the actual impact position. The coordinate translation method is as follows: based on the spatial resolution of the video image segment data corresponding to the trajectory impact position and the exposure window, the pixel coordinate position of the trajectory impact position in the image space is determined; the determined pixel coordinate position is used as the center point of the two-dimensional fuzzy kernel matrix, and the original fuzzy kernel matrix centered at the matrix center point is translated as a whole to the pixel coordinate position corresponding to the trajectory impact position, ultimately forming the composite fuzzy kernel model corresponding to the transient impact interval.
[0041] Finally, based on the composite fuzzy kernel model, composite fuzzy kernel data is generated. The composite fuzzy kernel data is a two-dimensional array, where each element stores the weight of a normalized matrix element. Simultaneously, the composite fuzzy kernel data also stores the time stamp information of the corresponding transient impact interval and the exposure start and end times information of the exposure window.
[0042] S5: Based on the composite fuzzy kernel data, constrained deconvolution is performed on the running video frames to generate image data with enhanced structural consistency, including: Using composite fuzzy kernel data as constraints, perform image domain deconvolution operation on running video frames, calculate image gray-level gradient distribution, and determine the deconvolution iteration convergence boundary; Specifically, the video image segment data of the exposure window corresponding to the video frame of the underground monorail crane operation is extracted frame by frame in the form of a two-dimensional matrix to form a gray value matrix. Using the two-dimensional matrix corresponding to the composite blur kernel data as a constraint, the Richardson-Lucy deconvolution algorithm is used to perform iterative image restoration calculation on the gray value matrix. The Richardson-Lucy deconvolution algorithm is as follows: using the original gray value matrix as the input image, using the two-dimensional matrix corresponding to the composite blur kernel data as the point spread function, the convolution operation and deconvolution operation are repeatedly performed in an iterative manner until the preset convergence condition is met and the calculation is terminated.
[0043] The convergence boundary in the iteration process is determined as follows: In each iteration, the gray-level difference matrix between the current iteration image gray-level data and the previous iteration image gray-level data is calculated. The absolute value of all elements in the gray-level difference matrix is taken, and the mean of the absolute values of all elements is calculated. When the mean is less than a preset convergence threshold, it is determined that the deconvolution iteration convergence boundary has been reached, and the iteration process is terminated. The convergence threshold is set based on the gray-level noise level of the actual scene in which the underground monorail crane operation video frame is located. When the scene noise is low and the image gray-level uniformity is high, the convergence threshold is set to be small; when the scene noise is high or the image gray-level changes significantly, the convergence threshold is set to be large, for example, in the range of 0.01 to 0.1.
[0044] After the iterative calculation terminates, the image grayscale data is used to calculate the image grayscale gradient distribution. Specifically, the Sobel operator is used to perform pixel-by-pixel convolution operation on the image grayscale data obtained above through inverse convolution iteration to obtain the gradient value distribution of the image grayscale data in the horizontal and vertical directions respectively. The Sobel operator is a 3×3 size convolution kernel.
[0045] The horizontal convolution kernel is [[-1,0,1],[-2,0,2],[-1,0,1]], and the vertical convolution kernel is [[-1,-2,-1],[0,0,0],[1,2,1]]. After convolving the above two convolution kernels with the grayscale data of the iteratively converged image pixel by pixel, the horizontal and vertical gradient values corresponding to each pixel position are obtained. Then, the square root of the sum of the squares of the above horizontal and vertical gradient values is used to obtain the grayscale gradient value of each pixel position.
[0046] By constraining the edge stability of the image gray-level gradient distribution based on the convergence boundary of deconvolution iteration, image gray-level data that meets the structural consistency condition is obtained, and image data with enhanced structural consistency is generated. Specifically, the edge stability constraint method is as follows: Based on the gray-level gradient value corresponding to each pixel in the image gray-level gradient distribution, a dual-threshold edge gradient stability determination method is used to determine the edge region; two gradient thresholds, high and low, are set to determine the image edge stability, where the high threshold is determined as the statistical average of all gradient values in the image gray-level gradient distribution plus one standard deviation, and the low threshold is determined as the statistical average minus half a standard deviation; the gray-level gradient values of all pixels are judged: pixels with gray-level gradient values greater than the high threshold are determined as stable edge points, pixels with gray-level gradient values between the high and low thresholds are judged as whether they are adjacent to the determined stable edge points, if adjacent, they are determined as stable edge points, otherwise they are determined as non-edge points, and pixels with gray-level gradient values lower than the low threshold are determined as non-edge points; through the above method, stable edge region data that meets the edge stability constraint conditions are obtained.
[0047] Finally, based on the determined stable edge region data, structural consistency constraint processing is applied to the deconvolution iterative converged image grayscale data to generate image grayscale data that meets the structural consistency conditions, i.e., structural consistency enhanced image data. The structural consistency constraint processing method is as follows: for pixels determined to be non-edge points, grayscale smoothing is performed, that is, the average grayscale data of the surrounding 3×3 neighborhood pixels centered on the non-edge point is taken and replaced with the original non-edge point grayscale value to enhance structural consistency; for pixels determined to be stable edge points, the grayscale value is kept unchanged to maintain the integrity of edge structure information; the final processed image grayscale data is the structural consistency enhanced image data, and the data format is a two-dimensional grayscale matrix of the same size as the original exposure window video image segment data, where each matrix element is the grayscale value after structural consistency enhancement, and the grayscale value range is 0 to 255.
[0048] S6: Perform cross-frame consistency verification on the structural consistency enhancement image data, and output the visual enhancement image of the underground monorail crane, including: Based on structural consistency enhancement image data, the spatial location distribution of edge features within consecutive multi-frame images is extracted, and the change amplitude of edge feature spatial location between adjacent frames is calculated. Specifically, based on the two-dimensional grayscale matrix corresponding to the structure consistency enhancement image data, the Canny edge detection algorithm is used for edge feature extraction. Gaussian filtering is applied to the structure consistency enhancement image data, with a Gaussian kernel size of 5×5. The standard deviation of the Gaussian kernel is determined based on the lighting conditions and camera noise level of the underground monorail crane operation scene, for example, a standard deviation set to 1.4 to 2.0. The Sobel operator is used to calculate the gradient values of the image grayscale data in the horizontal and vertical directions respectively. The gradient magnitude matrix is then obtained by calculating the square root of the sum of the squares of the gradient values, and the gradient direction matrix for each pixel position is calculated. Non-maximum suppression is performed based on the gradient magnitude matrix and gradient direction matrix to refine the edge positions. Finally, a double thresholding method is used to determine the final spatial distribution of edge features. The method for setting the high and low thresholds in the dual threshold method is as follows: by statistically analyzing the global gradient magnitude distribution of the gradient magnitude matrix, the mean and standard deviation of all elements in the gradient magnitude matrix are calculated. The high threshold is set to the mean plus one standard deviation, and the low threshold is set to the mean minus half the standard deviation. The edge feature spatial location determined by the above method is the edge pixel coordinate position data. The spatial coordinates are expressed in pixel coordinates, and the range of coordinate values is determined according to the resolution of the structural consistency enhancement image data, for example, 640×480 pixels.
[0049] Edge features are extracted from consecutive frames of structure consistency enhancement image data as described above, forming a spatial distribution sequence of edge features within consecutive frames. Using the spatial distribution of edge features corresponding to each frame of structure consistency enhancement image data as the basic data unit, the change amplitude of edge feature spatial positions between adjacent frames is calculated. The calculation method for the change amplitude of edge feature spatial positions between adjacent frames is as follows: Determine the set of edge feature spatial positions for adjacent frames of structure consistency enhancement image data; perform nearest neighbor matching between the edge feature spatial position data of the previous frame and the corresponding edge feature spatial position data of the next frame, using the Euclidean distance minimum matching principle. That is, calculate the Euclidean distance between the edge feature positions of the previous frame and all edge feature positions of the next frame, and determine the position of the next frame with the smallest distance as the corresponding position; after completing the matching of all positions, calculate the Euclidean distance value between each pair of matched positions. The Euclidean distance value is the change amplitude of the edge feature spatial position. The overall change amplitude value is the average of all Euclidean distance values.
[0050] Using the stability threshold of the change amplitude as the cross-frame consistency constraint, the structural consistency enhancement image data that meets the cross-frame consistency constraint is selected, and the visual enhancement image of the underground monorail is output. Specifically, the stability threshold setting method is as follows: Statistical analysis is performed on the spatial position variation data of edge features in multiple consecutive frames of images, and the average and standard deviation of all variation data are calculated. The average minus one standard deviation is used as the stability threshold for the variation amplitude. If the spatial position variation of edge features between two adjacent frames is less than or equal to the stability threshold, the latter image data is determined to meet the cross-frame consistency constraint; if the variation amplitude exceeds the stability threshold, it is determined not to meet the cross-frame consistency constraint, and the image data of that frame is discarded. By filtering frame by frame using the above method, structurally consistent enhanced image data that meets the cross-frame consistency constraint is obtained.
[0051] Finally, temporal smoothing is performed on the structural consistency enhancement image data that has passed the cross-frame consistency screening to generate a visually enhanced image of the underground monorail. The temporal smoothing method is as follows: for structural consistency enhancement image data that meets the cross-frame consistency constraints for multiple consecutive frames, the temporal average value of the grayscale value of each pixel position within the consecutive frames is calculated and replaced with the original single-frame grayscale value to complete the temporal smoothing process; for example, for structural consistency enhancement image data that meets the cross-frame consistency constraints for 5 consecutive frames, the grayscale values of each pixel position for 5 consecutive frames are averaged and the average value is used as the final output grayscale value of the pixel position; through the above processing, a visually enhanced image of the underground monorail with a stable spatial position and smooth grayscale distribution is finally formed.
[0052] The visually enhanced image of the underground monorail crane is a two-dimensional grayscale matrix. Each matrix element represents the final grayscale value of the corresponding pixel position. The grayscale value ranges from 0 to 255. The matrix size and structure are consistent with the image data, for example, the resolution is 640×480 pixels. The visually enhanced image of the underground monorail crane is also associated with the timestamp data corresponding to each frame of the image. The timestamp data is absolute time information with millisecond precision. The value of the timestamp is obtained based on the frame synchronization timestamp data determined by the camera's built-in time synchronization module when the corresponding video frame is captured.
[0053] Example 2 The difference between Embodiment 2 and Embodiment 1 is that this embodiment introduces a visual enhancement system for a downhole monorail crane based on dynamic modeling.
[0054] Figure 2 A schematic diagram of a vision enhancement system for an underground monorail based on dynamic modeling is provided. The system includes: Data acquisition module: acquires video frames of the underground monorail crane operation and corresponding timestamp data and exposure time parameters, and constructs exposure window time series data based on the timestamp data and exposure time parameters; Impact detection module: performs track joint impact feature detection on exposure window time series data, identifies transient impact intervals in exposure window time series data, and generates transient impact interval identification data; Dynamic modeling module: Based on transient impact interval identification data, a coupled dynamic model of the vibration of the underground monorail crane body and the swing of the load is established, the time-varying pose response sequence within the exposure window is calculated, and time-varying pose response data is generated; Piecewise integration module: Based on the time-varying pose response data, the exposure window time series data is piecewise integrated to construct a composite fuzzy kernel model corresponding to the transient impact interval and generate composite fuzzy kernel data; Constrained Convolution Module: Based on composite fuzzy kernel data, constrained deconvolution is performed on running video frames to generate image data with enhanced structural consistency; Consistency verification module: Performs cross-frame consistency verification on the structural consistency enhancement image data and outputs visual enhancement images of the underground monorail crane.
[0055] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0056] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0057] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0058] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0059] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0060] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0061] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0062] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0063] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A visual enhancement method for underground monorail cranes based on dynamic modeling, characterized in that, Includes the following steps: S1: Obtain video frames of the underground monorail crane operation and corresponding timestamp data and exposure time parameters, and construct exposure window time series data based on timestamp data and exposure time parameters; S2: Perform track joint impact feature detection on the exposure window time series data, identify the transient impact interval in the exposure window time series data, and generate transient impact interval identification data; S3: Based on transient impact interval identification data, establish a coupled dynamic model of the vibration of the underground monorail crane body and the swing of the load, calculate the time-varying pose response sequence within the exposure window, and generate time-varying pose response data; S4: Based on the time-varying pose response data, the exposure window time series data is piecewise integrated to construct a composite fuzzy kernel model corresponding to the transient impact interval, and generate composite fuzzy kernel data. S5: Based on the composite fuzzy kernel data, constrained deconvolution is performed on the running video frames to generate image data with enhanced structural consistency; S6: Perform cross-frame consistency verification on the structural consistency enhancement image data and output the visual enhancement image of the underground monorail crane.
2. The visual enhancement method for a downhole monorail crane based on dynamic modeling according to claim 1, characterized in that, S1, specifically: Acquire video frames of the underground monorail crane in operation, read the timestamp data corresponding to the video frames of the underground monorail crane in operation, and extract the exposure time parameters; The start and end times of exposure are determined based on timestamp data; The exposure window of the video frame of the underground monorail crane operation is extracted and aligned with the time according to the exposure time parameter to form the exposure window time series data.
3. The method for visual enhancement of a downhole monorail crane based on dynamic modeling according to claim 2, characterized in that, S2, specifically: Based on the exposure window time series data, the gray-level difference sequence and edge gradient change sequence between adjacent exposure windows are calculated, and the gray-level change amplitude curve and edge gradient fluctuation curve are constructed. Local extreme value search and duration determination are performed on the grayscale change amplitude curve and the edge gradient fluctuation curve. The interval that simultaneously satisfies the amplitude change and the duration being less than the exposure time parameter is selected to determine the transient impact interval. The exposure window time series data corresponding to the transient impact range is time-stamped to form transient impact range identification data.
4. The method for visual enhancement of a downhole monorail crane based on dynamic modeling according to claim 3, characterized in that, S3, specifically: Based on the transient impact interval identification data, the corresponding exposure window time series data is obtained, and the orbital impact position and impact time information in the exposure window time series data are extracted. A vibration excitation sequence for a downhole monorail crane body is constructed based on the information of track impact location and impact time. Using the vibration excitation sequence of the underground monorail crane body as input, a coupled dynamic model of the vibration of the underground monorail crane body and the swing of the load is established; The time-varying pose response sequence within the exposure window is obtained by solving the coupled dynamics model, thus generating time-varying pose response data.
5. The method for visual enhancement of a downhole monorail crane based on dynamic modeling according to claim 4, characterized in that, S4, specifically: The vibration velocity variation curve within the exposure window is determined based on time-varying pose response data; Based on the vibration velocity variation curve, the exposure window time series data is piecewise integrated to generate time-varying blurred trajectory data; By performing a two-dimensional convolution kernel transformation on time-varying fuzzy trajectory data, a composite fuzzy kernel model corresponding to the transient impact interval is constructed, generating composite fuzzy kernel data.
6. The method for visual enhancement of a downhole monorail crane based on dynamic modeling according to claim 5, characterized in that, S5, specifically: Using composite fuzzy kernel data as constraints, perform image domain deconvolution operation on running video frames, calculate image gray-level gradient distribution, and determine the deconvolution iteration convergence boundary; By constraining the edge stability of the image gray-level gradient distribution based on the convergence boundary of deconvolution iteration, image gray-level data that meets the structural consistency condition is obtained, and image data with enhanced structural consistency is generated.
7. The method for visual enhancement of a downhole monorail crane based on dynamic modeling according to claim 6, characterized in that, S6, specifically: Based on structural consistency enhancement image data, the spatial location distribution of edge features within consecutive multi-frame images is extracted, and the change amplitude of edge feature spatial location between adjacent frames is calculated. Using the stability threshold of the change amplitude as a cross-frame consistency constraint, structural consistency enhancement image data that meets the cross-frame consistency constraint is selected, and visual enhancement images of the underground monorail are output.
8. A visual enhancement system for an underground monorail based on dynamic modeling, used to implement the visual enhancement method for an underground monorail based on dynamic modeling as described in any one of claims 1-7, characterized in that, include: Data acquisition module: acquires video frames of the underground monorail crane operation and corresponding timestamp data and exposure time parameters, and constructs exposure window time series data based on the timestamp data and exposure time parameters; Impact detection module: performs track joint impact feature detection on exposure window time series data, identifies transient impact intervals in exposure window time series data, and generates transient impact interval identification data; Dynamic modeling module: Based on transient impact interval identification data, a coupled dynamic model of the vibration of the underground monorail crane body and the swing of the load is established, the time-varying pose response sequence within the exposure window is calculated, and time-varying pose response data is generated; Piecewise integration module: Based on the time-varying pose response data, the exposure window time series data is piecewise integrated to construct a composite fuzzy kernel model corresponding to the transient impact interval and generate composite fuzzy kernel data; Constrained Convolution Module: Based on composite fuzzy kernel data, constrained deconvolution is performed on running video frames to generate image data with enhanced structural consistency; Consistency verification module: Performs cross-frame consistency verification on the structural consistency enhancement image data and outputs visual enhancement images of the underground monorail crane.