Coal mine coal mining face monitoring method and system based on image recognition
Patent Information
- Application Number
- CN202610654404.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-13
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]然而,在真实井下常态化工况中,采煤机的截割本质上是剧烈的物质剥离过程,煤壁在被截割时不断破碎、坠落,呈现出极度不规则的碎片化粗糙表面;与此同时,截割产生的局部高浓度粉尘与非均匀强光照射,会在该粗糙表面投射出大量随机分布的明暗阴影;若割裂了视觉图像与目标物体正在发生动态物理形变之间的因果联系,便会极易导致将煤层天然节理裂隙或粉尘光影误判为真实截割边界;这种误判会引发提取的虚拟轮廓产生剧烈的无规律跳变误差,致使自动化控制系统频繁失效
Smart Images

Figure CN122551238A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of intelligent coal mining and computer vision technology, specifically to a method and system for monitoring coal mining faces based on image recognition. Background Technology
[0002] In intelligent coal mining, automated control of the fully mechanized mining face is the core link. To achieve automatic height adjustment and cutting trajectory planning of the coal mining machine, explosion-proof cameras can be used to collect video for image recognition, thereby enabling non-contact monitoring of the working face status.
[0003] Traditional monitoring methods typically preprocess video images and then use edge detection operators or object detection networks to extract visual features. These methods directly fit the geometric boundary line of the coal-rock interface by finding areas of abrupt changes in pixel grayscale or color, and use this as a benchmark reference to guide the operation of the coal mining machine.
[0004] However, in real underground normal working conditions, the cutting of coal by the mining machine is essentially a violent material stripping process. The coal wall is constantly breaking and falling during cutting, presenting an extremely irregular, fragmented, and rough surface. At the same time, the local high concentration of dust and non-uniform strong light generated during cutting will project a large number of randomly distributed light and dark shadows on this rough surface. If the causal relationship between the visual image and the dynamic physical deformation of the target object is severed, it will be very easy to misjudge the natural joints and fissures of the coal seam or the light and shadow of dust as the real cutting boundary. This misjudgment will cause the extracted virtual contour to produce violent and irregular jump errors, causing the automated control system to fail frequently.
[0005] Therefore, how to accurately capture the physical phase transition boundary caused by real physical destruction processes under normalized and severe visual interference has become a technical bottleneck that urgently needs to be solved in this field. Summary of the Invention
[0006] To address the shortcomings of existing technologies, this invention provides a method and system for monitoring coal mining faces based on image recognition.
[0007] To achieve the above objectives, the technical solution of the present invention is as follows:
[0008] In a first aspect, the present invention discloses a method for monitoring coal mining faces based on image recognition, comprising the following steps:
[0009] Acquire a time-continuous sequence of working face video frames, and the physical space status data of the coal mining equipment corresponding to the timestamps of each video frame in the working face video frame sequence;
[0010] Based on the physical space state data and the pre-calibrated camera projection parameters, the target truncated area in the video frame sequence of the working surface is determined, and the target truncated area is divided into multiple image grids.
[0011] First image features and second image features are extracted from each image grid. The first image features characterize the differences in surface roughness distribution of pixels within the image grid, and the second image features characterize the local sinking expansion rate of pixels within the image grid in the direction of target gravity projection.
[0012] Combining the first image features, the second image features, and the baseline feature data determined based on the physical space state data, the first cost data for each image grid to independently transition to multiple preset physical evolution states is calculated respectively.
[0013] A motion direction constraint model for coal mining equipment is constructed based on physical space state data. In conjunction with the motion direction constraint model, the second cost data of state transition of each image grid in the spatiotemporal dimension is calculated.
[0014] A global integrated cost function is constructed based on the first cost data and the second cost data, and the target state distribution result that satisfies the condition of minimizing the global integrated cost function is obtained.
[0015] Based on the target state distribution results, the topological boundary lines between physical evolution states are extracted to generate the three-dimensional physical cutting contours of the corresponding coal mining equipment.
[0016] Secondly, this invention discloses a coal mine working face monitoring system based on image recognition, comprising:
[0017] The data acquisition module is used to acquire a time-continuous sequence of working face video frames, as well as the physical space status data of the coal mining equipment corresponding to the timestamps of each video frame in the working face video frame sequence.
[0018] The grid division module is used to determine the target truncating region in the video frame sequence of the working surface based on the physical space state data and the pre-calibrated camera projection parameters, and to divide the target truncating region into multiple image grids;
[0019] The feature extraction module is used to extract the first image feature and the second image feature of each image grid respectively; the first image feature represents the difference in surface roughness distribution of pixels in the image grid, and the second image feature represents the local sinking expansion rate of pixels in the image grid in the direction of target gravity projection.
[0020] The first cost calculation module is used to combine the first image features, the second image features, and the baseline feature data determined based on the physical space state data to calculate the first cost data for each image grid to independently transition to multiple preset physical evolution states.
[0021] The second cost calculation module is used to construct a motion direction constraint model of the coal mining equipment based on physical space state data, and to calculate the second cost data of state transition of each image grid in the spatiotemporal dimension in combination with the motion direction constraint model.
[0022] The global optimization solution module is used to construct a global comprehensive cost function based on the first cost data and the second cost data, and solve for the target state distribution that satisfies the condition of minimizing the global comprehensive cost function.
[0023] The cutting contour generation module is used to extract the topological boundary lines between physical evolution states based on the target state distribution results, and generate the three-dimensional physical cutting contour of the corresponding coal mining equipment.
[0024] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0025] 1. By extracting the first image feature representing the difference in surface roughness distribution of the coal wall and the second image feature representing the local subsidence expansion rate, a rigid correlation is established between the visual appearance and the real physical material stripping process. In addition, a motion direction constraint model is constructed by combining the physical space state data of the coal mining equipment. The state transition cost is calculated in the spatiotemporal dimension and a global comprehensive cost function is constructed. The false boundaries caused by dust obstruction and light spot shadows are stripped away. Under normal severe visual interference, the three-dimensional physical cutting contour generated by the real physical destruction process can be accurately and stably extracted, eliminating the risk of frequent failure of the automated control system.
[0026] 2. By calculating the consistency of the direction between the two-dimensional displacement vector of the image grid and the projection of the coal mining machine's traction speed, a motion wavefront reachability determination mask is constructed. This mask forces an infinitely large blocking constant cost to state transition attempts that violate physical laws (such as cutting in areas inaccessible to the coal mining machine). Combined with the graph cut optimization algorithm, the global comprehensive cost function containing data and smoothing terms is minimized, reducing the possibility of dense dust at the edge of the image masquerading as a cutting action and ensuring the rationality of the overall image state distribution. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 This is an overall block diagram of the method in Embodiment 1 of the present invention;
[0029] Figure 2 This is a flowchart illustrating the overall execution process of the method in Embodiment 1 of the present invention.
[0030] Figure 3 This is a schematic diagram of the meshing and motion constraint projection of the target truncating region in Embodiment 1 of the present invention;
[0031] Figure 4 This is an overall block diagram of the system in Embodiment 2 of the present invention. Detailed Implementation
[0032] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0033] In the field of intelligent coal mining, the automated control of fully mechanized mining faces relies heavily on explosion-proof cameras for non-contact status monitoring to guide the automatic height adjustment and cutting trajectory planning of the coal mining machine. Traditional solutions directly fit the geometric boundary line of the coal-rock interface by finding areas of pixel grayscale or color abrupt change.
[0034] However, existing technologies sever the causal link between visual images and the dynamic physical deformation of the target object; real cutting is essentially a violent material stripping process, with the coal wall breaking and presenting an extremely irregular, fragmented, and rough surface; localized high-concentration dust and non-uniform strong light will project a large number of randomly distributed light and dark shadows on this rough surface; due to the lack of rigid verification of the mechanical motion wavefront and physical evolution state, the system only relies on two-dimensional pixel gradients for judgment, resulting in the inability to establish a strict physical correspondence between visual features and physical phase transitions.
[0035] For example, in normal underground working conditions, the monitoring system can only capture grayscale changes in the image using conventional target detection networks, but cannot distinguish whether they represent actual coal block collapses. When dense dust is generated at the working face or strong light spots intersect, the system only records the appearance of pixel-level gradients, easily misjudging natural joints and fissures in the coal seam or dust shadows as actual cutting boundaries. Specifically, the system misjudges free dust or light spots as coal wall edges, failing to detect the sinking and expansion physical characteristics and kinematic spatial constraints that actual cutting must meet, leading to the continuous solidification of erroneous contour extraction patterns.
[0036] If the above problems are not solved, the monitoring system will continue to lose its ability to accurately capture and objectively distinguish the phase transition boundaries of entities under complex working conditions; the failure to remove false boundaries will cause the extracted virtual contours to produce drastic and irregular jump errors; at the same time, the lack of physical stripping features and mechanical kinematic constraints will cause the algorithm's anti-interference ability to drop sharply, resulting in frequent failures of the automated control system; ultimately, the inaccuracy of visual perception will systematically hinder the generation of precise cutting commands by the coal mining machine, affecting the achievement of the goal of intelligent coal mining.
[0037] Example 1:
[0038] like Figures 1-3 As shown, the image recognition-based coal mine face monitoring method includes the following steps:
[0039] Step S1: Obtain a time-continuous sequence of working face video frames, and the physical space status data of the coal mining equipment corresponding to the timestamps of each video frame in the working face video frame sequence;
[0040] In the actual deployment scenarios of intelligent coal mining, the longwall face is in an extremely harsh environment with high dust, low illumination, and strong mechanical vibration. This physical environment poses a significant engineering challenge to the acquisition and fusion of multi-source heterogeneous data. To ensure that subsequent image feature extraction and state transition cost inference are based on a solid physical foundation, the system must rely on a pre-configured cross-modal spatiotemporal synchronous data acquisition architecture before executing the core monitoring logic. Specifically, the underlying hardware of this architecture typically includes an explosion-proof camera array deployed on a hydraulic support, an industrial Ethernet connection to the main control module of the coal mining machine, and a Precision Time Protocol (PTP) clock synchronization server. This high real-time bus and clock synchronization architecture is adopted because subsequent processing requires extremely precise mechanical wavefront positions. Even a misalignment of tens of milliseconds between the visual image and the mechanical coordinates on the time axis will be magnified into a physical error of tens of centimeters during perspective projection.
[0041] Based on the aforementioned cross-modal acquisition architecture, the system first acquires a temporally continuous sequence of video frames from the working surface. In specific implementation, the video decoder extracts continuous image data from the real-time video stream transmitted from the explosion-proof camera according to a preset sampling frequency (e.g., 30 frames / second), defining the currently extracted video frame as... The video frame of the previous adjacent time moment is To eliminate high-frequency flicker noise from underground electrical equipment and optical scattering interference caused by aerosol dust, the system performs grayscale conversion on the extracted original RGB video frames and simultaneously performs spatial domain smoothing processing based on bilateral filtering. The bilateral filtering algorithm not only considers the spatial geometric distance between pixels but also introduces a similarity penalty for pixel grayscale values. Its weight allocation formula is as follows: ,in The center pixel coordinates, For the neighboring pixel coordinates, and These represent the grayscale values at the center pixel and the neighboring pixels in the current frame, respectively. Let the standard deviation of the Gaussian kernel in the spatial domain be . The standard deviation of the Gaussian kernel in the grayscale range is denoted as . By adjusting these two parameters, the system can effectively smooth out fine dust noise while perfectly preserving the sharp grayscale abrupt changes at the coal face stripping fault. After processing, the system will output a high signal-to-noise ratio baseline grayscale matrix sequence for the current moment. Compared with the reference grayscale matrix sequence of the previous time step This serves as the base data for subsequent image grid feature extraction.
[0042] While continuously capturing video frame sequences, the system acquires real-time physical space status data of the coal mining equipment by parsing the PLC (Programmable Logic Controller) control bus messages of the coal mining machine. In practical engineering applications, to accurately characterize the dynamic mechanical properties of the coal mining equipment, this physical space status data is strictly defined as including the current three-dimensional physical coordinates and three-dimensional traction speed vector of the coal mining equipment. Specifically, the system establishes a three-dimensional absolute world coordinate system with a specific location on the fully mechanized mining face (such as the end of the machine roadway) as the physical origin, and reads and calculates the three-dimensional physical coordinates of the center points of the left and right drums of the coal mining machine in this coordinate system in real time. ,in The system characterizes the current physical depth of the roller; simultaneously, it reads the speed and direction encoding information from the traction inverter and calculates the three-dimensional traction velocity vector characterizing the mechanical kinematics. .
[0043] To achieve absolute alignment between the visual and physical flows, the system further performs hardware-level interpolation binding based on timestamps. Assuming the data acquired at a specific moment... Video frames of resolution The corresponding timestamp is The system will retrieve two physical state data packets before and after the timestamp from the PLC bus's historical buffer queue, and use a linear interpolation algorithm to calculate the data packets that strictly correspond to the timestamp. Three-dimensional physical coordinates at time With three-dimensional traction velocity vector For example, if the coal mining machine is currently... The traction speed moves towards the tail, and the system is in The three-dimensional physical coordinates extracted at each moment may be: The meter, and the corresponding three-dimensional traction velocity vector is precisely quantized as Through this deep fusion spatiotemporal alignment mechanism, the system successfully rigidly binds the macroscopic mechanical motion state with the microscopic pixel flow sequence on the same time cross section, thus laying the data foundation for subsequent use of camera projection parameters to lock the target truncated area and construct motion direction masks.
[0044] Step S2: Based on the physical space state data and the pre-calibrated camera projection parameters, determine the target truncating region in the working surface video frame sequence, and divide the target truncating region into multiple image grids;
[0045] After completing the hardware-level synchronous acquisition of multimodal time-series data, the system faces a massive video data processing load. The full-frame high-definition video of the fully mechanized mining face contains a large amount of background information unrelated to the core cutting action, such as hydraulic support columns, distant roadways, and roof. Performing pixel-level calculations on the entire image would not only consume massive amounts of edge computing power but also introduce significant environmental noise interference. Therefore, the system must use a mathematical mapping mechanism to precisely constrain subsequent highly complex calculations within the actual local space where physical stripping occurs. To achieve this, the system needs to pre-configure camera projection parameters, which form a mathematical bridge connecting the three-dimensional physical world of the coal mine with the two-dimensional pixel plane of the explosion-proof camera. Specifically, the camera projection parameters are composed of an intrinsic parameter matrix and an extrinsic parameter matrix. The intrinsic parameter matrix characterizes the camera's inherent optical properties, including focal length and principal point offset, and is typically calculated offline using a standard Zhang Zhengyou calibration plate before the equipment is installed in the well. The extrinsic parameter matrix, on the other hand, is a rigid body transformation matrix containing rotation and translation vectors, used to describe the spatial transformation relationship from the absolute world coordinate system of the working surface to the camera's local coordinate system. This matrix is calculated after the camera support is fixed by measuring the physical and pixel coordinates of multiple marker points on site. By combining these two matrices, the system constructs a complete perspective projection model.
[0046] Based on the pre-constructed perspective projection model, the system begins to transform the input physical space state data into the coordinate domain. The system extracts the three-dimensional physical coordinates representing the current position of the coal mining equipment from the physical space state data, and calculates the perspective projection coordinates of these three-dimensional physical coordinates on the two-dimensional image pixel plane according to the pre-calibrated camera projection parameters. Assuming the three-dimensional physical coordinates of the coal mining equipment are... The camera's intrinsic parameter matrix is The extrinsic parameter matrix is The system uses matrix multiplication in homogeneous coordinates to solve this projection relationship, and its core calculation formula is as follows: In this formula, These are the perspective projection coordinates obtained by solving the problem. This is the depth scaling factor for that point in the camera coordinate system. After the above geometric algebraic operations, the system will calculate the perspective projection coordinates. The coordinates of the target center are determined. If at a certain moment, the three-dimensional physical coordinates of the front drum of the coal mining machine are located in the world coordinate system... At a distance of meters, the system substitutes this coordinate into the aforementioned perspective projection equation and solves for the coordinates at that location. The pixel coordinates corresponding to the resolution of the video frame are This coordinate point serves as the reference core for tracking mechanical equipment in the image domain at this moment.
[0047] After accurately locating the specific position of the coal mining equipment in the image, the system needs to further define the effective field of view for feature extraction. Specifically, using the previously calculated target center coordinates as a reference point, the system defines a rectangular bounding box of a preset size within the current single frame of the working face video frame sequence. This preset size is dynamically calculated based on the actual physical diameter of the coal mining machine drum and the camera imaging resolution at that physical depth, ensuring that the bounding box precisely covers the drum body and the coal wall areas in front of and above and below it that are being cut. Subsequently, the system formally determines the image area covered by this rectangular bounding box as the target cutting area. For example, if the preset size determined based on physical proportions is... Pixels, the system will use Centered on the image, the coordinates extend 200 pixels outwards in each direction, thus precisely cropping a coordinate range from the center of the large full-frame image. to Using a local portion of the image as the target cropping area, the invalid background noise, which accounts for more than 90% of the full-frame image, is completely blocked.
[0048] To further adapt to the computational architecture of subsequent Markov Random Field (MRF) and other probabilistic graphical models and further reduce computational dimensionality, the system does not stop at pixel-level observation of the target cut-off region. Instead, it spatially divides the determined target cut-off region into multiple non-overlapping image grids. The principle behind this gridding is that the fracturing, spalling, and dust dispersion of the coal face have a certain degree of coherence in physical space. Adjacent pixels often exhibit highly consistent physical state transitions. Packing them into a microscopic graph node for processing can achieve an exponential decrease in computational power without losing macroscopic physical characteristics. In actual implementation, the system typically divides the cut-off region into multiple non-overlapping image grids. Divide the size equally. Continuing with the size example above, if the target cut-off area is... Pixels, the system uses By dividing the target region into regular subdivisions with pixels as units, the original target region containing 160,000 discrete pixels was ultimately reduced in dimensionality and transformed into a smaller subdivision. A topological set consisting of 625 image grids. Through this series of data processing steps, from 3D projection locking and local region extraction to spatial grid partitioning, the system successfully established a strong coupling mapping between the image perception domain and the physical generation domain, constructing a microscopic data skeleton for the subsequent extraction of roughness differences and sinking expansion rate features of each image grid.
[0049] Step S3: Extract the first image feature and the second image feature of each image grid respectively; the first image feature represents the difference in surface roughness distribution of pixels in the image grid, and the second image feature represents the local sinking expansion rate of pixels in the image grid in the direction of target gravity projection;
[0050] After completing the precise coordinate mapping and micro-mesh division of the target truncated area, the system enters the physical feature extraction stage, which involves extracting the first and second image features of each image grid. Traditional machine vision monitoring solutions often fall into the technical pitfall of searching for pixel geometric gradients. However, in the harsh real-world working conditions of underground drilling, characterized by high dust and intense light spots, the signal-to-noise ratio of pure geometric gradients is extremely low. To fundamentally solve this problem, this embodiment completely abandons the feature extraction paradigm based on edge operators and instead extracts deep features with clear mechanical and optical properties from two orthogonal physical dimensions: "static surface material phase transition" and "dynamic gravity stripping motion." Specifically, the first image feature aims to characterize the surface roughness distribution differences of pixels within the image grid, while the second image feature aims to characterize the local sinking expansion rate of pixels within the image grid in the direction of target gravity projection.
[0051] For the extraction of the first image feature, the system introduces the entropy model from information theory to quantify the physical roughness of the coal wall surface. Specifically, during the extraction process, the system first counts the frequency of a preset number of gray levels within each target image grid in the current video frame. In practical applications, since video frame sequences typically use 8-bit grayscale image format, this preset number of gray levels is 256 grayscale levels from 0 to 255. The system then iterates through all pixels within the grid (e.g., in a single...). The system iterates through 256 pixels in a pixel grid, recording the total number of pixels in each grayscale level. Then, it calculates the local grayscale probability distribution of the corresponding image grid through normalization. Based on the normalization result, the system rigorously calculates the local texture information entropy of the corresponding image grid as the first image feature, according to information theory principles and the local grayscale probability distribution. The calculation formula is as follows:
[0052] ;in, Represents the local texture information entropy of the target image mesh. This represents the probability of the i-th gray level appearing within the image grid. To ensure computational continuity and prevent the program from throwing logarithmic singularity anomalies, according to the information theory limit law, when... At that time, clearly define the limitations .
[0053] The reason for using local texture information entropy as the characterization variable for roughness is that the surface of the original coal face, which has not been cut by the mining machine, is extremely rugged and accompanied by a large number of joints. Under the illumination of explosion-proof lights, it will produce complex diffuse reflection and small highlights and shadows, resulting in a gray-level probability distribution within the grid. Extremely uniform and diffuse, thus yielding extremely high information entropy values (e.g. Conversely, the newly formed coal face after being cut and stripped by the drum is relatively flat and homogeneous. Its local pixel grayscale distribution is highly concentrated at a few specific grayscale levels, causing a sharp drop in information entropy (e.g., This entropy-based measurement mechanism can filter out linear interference caused by local strong light or global brightness changes, and only strongly responds to changes in "roughness" itself.
[0054] While extracting the first image features characterizing the static material properties, the system simultaneously extracts the second image features to capture the instantaneous dynamic features generated when the coal face is mechanically damaged. This process relies on acquiring a high-precision pixel-level motion field. Specifically, during the extraction process, the system first calculates the pixel motion vector field based on the image grid grayscale data of two adjacent video frames in the working face video frame sequence. To accurately depict this dynamic process, this embodiment uses the Farnebäck dense optical flow algorithm model at the underlying level. This model receives the reference grayscale matrix from the previous time step output in step S1. Compared with the current time reference grayscale matrix As input layer features, a multi-scale image pyramid is constructed in the spatial domain, and a quadratic surface is fitted to the neighborhood of each pixel using polynomial expansion techniques. Finally, a dense two-dimensional displacement vector field covering the entire image is generated in the output layer, physically representing the displacement direction and velocity magnitude of each pixel between adjacent frames. It is worth noting that since explosion-proof cameras are typically rigidly fixed to hydraulic supports, the mechanical vibration of the supports causes high-frequency jitter in the camera itself. If the omnidirectional optical flow amplitude is used directly, it will be completely overwhelmed by this global translational noise. To achieve accurate noise reduction, the system combines the physical gravity projection direction calibrated in the camera's projection parameters to directionally extract the velocity component of the pixel motion vector field in the physical gravity projection direction from the optical flow output. Furthermore, the actual falling of coal is not a uniform motion, but an accelerated descent under the action of gravity. According to the principles of fluid mechanics and rigid body kinematics, an accelerated falling object will inevitably experience spatial stretching or expansion effects in its direction of motion. Based on this, the system uses discrete difference operators (such as forward difference) to calculate the spatial partial derivatives of the velocity components along the physical gravity projection direction (i.e., to calculate the divergence of the local flow field) for each image grid. To completely eliminate the artifact of uniform downward or upward translation caused by camera pitch vibration (the spatial partial derivatives of uniform translation approach zero or are negative due to compression), the system uses a one-sided truncation function to retain values greater than zero after calculating the partial derivatives, ultimately obtaining the downward motion dilation rate of the image grid as the second image feature. Its core calculation formula is configured as follows: ,in Indicates the center coordinates as The rate of expansion of the image grid's sinking motion. This indicates the extracted velocity component. The coordinate axes represent the direction of the physical gravity projection.
[0055] Through the aforementioned processing links, the system successfully decoupled the local texture information entropy matrix, which is resistant to illumination interference, and the subsidence motion expansion rate matrix, which is resistant to vibration interference, from the macroscopic video stream. These two microscopic feature data streams not only achieve complete coverage of the "static phase transition" and "dynamic damage" of the coal wall in terms of physical mechanisms, but also provide observational variable inputs for subsequent calculations of the first cost data for the transition of each image grid to different physical evolution states.
[0056] Step S4: Combining the first image features, the second image features, and the baseline feature data determined based on the physical space state data, calculate the first cost data for each image grid to independently transition to multiple preset physical evolution states;
[0057] After extracting the dual-modal physical stripping features (i.e., the first image feature representing surface roughness and the second image feature representing subsidence and expansion), the system needs to transform these continuous numerical features into probabilistic criteria for determining the macroscopic physical evolution state, that is, to calculate the first cost data for each image grid to independently transition to multiple preset physical evolution states. In implementing this scheme, to digitally model the complex cutting and destruction process, the system's underlying layer is configured with a univariate potential energy mapping model based on Markov random fields (MRF). This model defines three discrete preset physical evolution states: the "first static state," representing the original state of the coal wall before it is cut; the "second transition state," representing the coal wall being stripped and destroyed by the roller; and the "third new state," representing the completion of cutting and the exposure of a new coal wall. The core logic for calculating the first cost data is to construct a cost penalty mapping equation from a "two-dimensional feature space" to a "three-dimensional state space." Under this equation, the more the feature observations of a grid conform to the prior laws of a certain physical state, the lower the cost penalty it is assigned to that state.
[0058] Due to the varying coal seam materials and lighting conditions in different mines, there are significant differences in the absolute roughness between the original and newly formed coal faces. If a fixed, rigid threshold is used to evaluate features, the model will quickly become ineffective. Therefore, before calculating the matching cost, the system must perform dynamic benchmark anchoring, i.e., determine benchmark feature data based on physical space state data. Specifically, the system analyzes the three-dimensional physical coordinates and three-dimensional traction velocity vector of the mining equipment in real time. In the three-dimensional world coordinate system, it determines a first physical bounding box located in front of the mining equipment at a first preset distance (e.g., 5 meters in front), and a second physical bounding box located behind the mining equipment at a second preset distance (e.g., 5 meters behind). These two physical spaces represent the absolutely undisturbed original area and the absolutely truncated newly formed area, respectively. Subsequently, the system uses pre-calibrated camera projection parameters to rigorously project these two three-dimensional physical bounding boxes onto a two-dimensional mesh set, and extracts the average first image feature (average local texture information entropy) within the first projection area as the first benchmark feature data. And extract the average first image features within the second projection area as the second reference feature data. Assuming the original coal face ahead is extremely rugged under the current operating conditions, the system dynamically calculates... It could be as high as 7.2; while the newly cut surface behind it is relatively smooth, and the calculated value is... It may drop to 3.8. These two dynamic baseline values provide an adaptive scale for subsequent cost calculations.
[0059] After establishing the high and low benchmarks, the system begins to calculate the matching cost for each image grid as it transitions to the three states mentioned above. For the static state matching cost transitioning to the first static state, the physical judgment logic is: the first image feature (entropy value) of the current grid should be sufficiently large, at least close to the benchmark value of the original coal face. The system uses the degree to which this feature deviates unidirectionally from the first benchmark feature data as the cost metric, and the specific calculation formula is configured as follows: .in This represents the cost of matching in the static state. The first image feature of the current grid. The first preset weighting coefficient, This is a one-sided truncation function. If the entropy value of the current grid... Even higher than the benchmark value (That is, coarser), the truncation function outputs zero, indicating that the matching cost is 0, perfectly matching the static state; but once If the value is lower than the benchmark, the difference is multiplied by a weight and converted into a linearly increasing cost penalty. This means the matching cost is positively correlated with the difference between the first image feature and the first benchmark feature data. Similarly, for the matching cost of the newborn state transitioning to the third newborn state, the system uses the formula... Perform calculations, where The third preset weighting coefficient is used. In this way, the cost of classifying the current grid as a new state only approaches zero when the entropy value of the current grid drops to close to or even below the second baseline feature data; if the entropy value is still too high, the difference will be converted into a positively correlated matching cost, thereby rejecting incorrect classifications.
[0060] The most complex aspect is the cost assessment of the second transition state, as this state characterizes the phase transition moment of the coal face's instantaneous collapse and must satisfy a dual physical orthogonality constraint. The system first determines the arithmetic mean between the first and second baseline feature data, using this as an intermediate transition feature value (e.g., in the example above). Next, the system integrates the cost from two dimensions: on the one hand, the moment of cutting is inevitably accompanied by the falling of coal blocks, so the cost is extracted from the second image features (the rate of expansion of the sinking motion). The state decay term, which is negatively correlated, is expressed through a negative exponential function. The greater the divergence, the more exponentially the cost approaches zero; on the other hand, the surface roughness at the moment of phase transition should be in the transition period between high and low references, therefore the system extracts features from the first image. The absolute difference deviating from the intermediate transition eigenvalue is positively correlated with the state penalty term. Finally, the system fuses and sums the state decay term and the state penalty term to obtain the transition state matching cost for moving to the second transition state. The complete calculation formula is as follows:
[0061] (in, This is the second preset weighting coefficient. (This is a preset sinking expansion attenuation factor). For example, if a mesh is undergoing severe peeling and falling ( If the texture entropy drops to around 5.5, then both terms in the above formula will drop to extremely low levels, resulting in a very small transition state matching cost.
[0062] Ultimately, the system processes each image grid in parallel, encapsulating the static state matching cost, emerging state matching cost, and transition state matching cost corresponding to each grid into first cost data representing the primary data potential energy of that grid. Through this series of operations, the system completely transforms the chaotic optical flow and entropy values into a probability energy matrix that conforms to the laws of physical evolution, thus preparing the data for subsequent integration into the mechanical spatiotemporal constraint model.
[0063] Step S5: Construct a motion direction constraint model for the coal mining equipment based on physical space state data, and calculate the second cost data of state transition of each image grid in the spatiotemporal dimension in combination with the motion direction constraint model;
[0064] After completing the first-state cost calculation at the micro-grid level, the system can make preliminary inferences about the grid state based on local roughness and subsidence expansion rate characteristics, but it ignores the mechanical rigidity constraints of the coal wall stripping and failure process in the macroscopic physical space. In a real fully mechanized mining face, the coal wall will never break and collapse out of thin air. Any state transition from "uncut" to "cut" can only occur physically in front of the real-time movement trajectory of the mining equipment drum (i.e., the mechanical motion wavefront). If relying solely on visual features, dense dust blown up by ventilation at the far end of the working face can easily generate pseudo-features similar to the interweaving of high dispersion and high entropy values, which the system may misjudge as cutting in progress. To eliminate this false correlation, the system pre-configures and constructs a "motion direction constraint model" that integrates mechanical kinematics before further evaluating the state transition cost. This model is a physical prior gating mechanism based on three-dimensional spatial analytical geometry and equipment kinematic topology. Its core architectural logic lies in using the three-dimensional physical trajectory of the coal mining equipment as input parameters, and mapping it into a velocity vector-driven "influence cone" in two-dimensional pixel space through perspective transformation. This cone serves as a rigid spatial mask to screen whether a local image mesh is qualified to undergo a sudden change in physical state. The reason for adopting this deterministic geometric prior structure instead of a statistical learning model is that the rigid body motion law of mechanical equipment is a 100% deterministic physical inevitability.
[0065] Based on the above principles, the system begins to construct a motion direction constraint model for the coal mining equipment based on physical space state data. In the specific implementation process, the system first extracts the three-dimensional traction velocity vector (e.g., the velocity vector of the coal mining machine moving towards the tail end along the working face track at a specific speed) from the control bus in real time. Simultaneously, using pre-calibrated camera projection parameters (including the camera's intrinsic and extrinsic rotation and translation matrices) and the current three-dimensional physical coordinates of the coal mining equipment, the system accurately maps this three-dimensional traction velocity vector to a two-dimensional image plane. This projection process not only considers the camera's viewing angle deflection but also integrates the current physical depth information of the equipment, thereby calculating the corresponding two-dimensional projection velocity vector in the image pixel coordinate system. Assuming the three-dimensional traction velocity is projected onto a high-definition video frame, the resulting two-dimensional projection velocity vector... Possibly This indicates that the device is moving horizontally to the right in the image at a rate of 15 pixels per frame.
[0066] Subsequently, the system calculates the two-dimensional displacement vector of the center coordinates of any image grid within the target truncated area relative to the target center coordinates where the roller is located. Based on the extracted two-dimensional displacement vector and two-dimensional projected velocity vector, the system extracts the degree of directional consistency between the two (usually characterized by calculating the dot product or cosine of the included angle) and the spatial distance between the two-dimensional displacement vectors (i.e., the magnitude of the vector), and uses this information to determine the wavefront reachability of the image mesh. In practical applications, the system will preset a physical influence angle (e.g., front end angle). The system calculates the cone angle range and the maximum wavefront influence distance (e.g., a pixel distance of 1.5 meters in actual physical space). If the calculated directional consistency meets the preset angle threshold and the spatial distance meets the preset range condition, it means that the image grid is in the "coverage area" that the roller is about to sweep over, and the system sets the motion direction mask value of the corresponding image grid to a reachable flag (usually assigned a value of 1 in the computer's underlying logic); otherwise, for grids located behind the roller or too far away, the system sets the motion direction mask value of the corresponding image grid to an unreachable flag (usually assigned a value of 0). For example, for a grid located 100 pixels directly in front of the roller with an angle of only 1 / 2000 meters, the system sets the motion direction mask value of the corresponding image grid to an unreachable flag (usually assigned a value of 0). For a grid cell, the mask value is 1; while for a grid cell located 300 pixels behind or directly above the roller, the mask value is strictly locked to 0. This set of 0 or 1 mask values for all grid cells in the image constitutes the complete motion direction constraint model.
[0067] After successfully constructing a motion direction constraint model, similar to a physical filter, the system then combines this model to calculate the second cost data for the state transitions of each image grid in the spatiotemporal dimension. This second cost data aims to evaluate the continuity and rationality of the grid states in terms of spatial distribution and temporal passage. Specifically, the system first calculates the spatial smoothing penalty cost when spatially adjacent image grids within the same video frame are assigned different physical evolution states. For example, if a grid is determined to be in the first stationary state (original, uncut), while its eight surrounding adjacent grids are all determined to be in the third emerging state (cut), this isolated spatial transition severely violates the continuity law of physical coal wall peeling, and the system will assign an extremely high spatial smoothing penalty cost for this.
[0068] Furthermore, based on the aforementioned generated motion direction mask value, the system accurately calculates the temporal transition penalty cost for state transitions of the same image grid between two adjacent video frames. When the system detects that an image grid was in the first stationary state in the previous video frame, and the algorithm attempts to transition it to the second transitional state (i.e., the state of truncation and destruction) in the current video frame, the system will forcibly query the motion direction mask value corresponding to that grid. If the corresponding motion direction mask value is a reachability flag (value 1), it indicates that the area is indeed covered by the wavefront of the coal mining machine, and physical destruction is allowed. The system then allows the transition smoothly, assigning only a very small baseline constant cost (e.g., set to a constant). The normal time evolution decay is used; however, if the corresponding motion direction mask value is an unreachable flag (value 0), it indicates that the mesh is physically impossible to be cut by the roller. In this case, the system will intercept this physically illogical transfer attempt at the highest level, directly assigning an blocking constant cost. In the mathematical model, this blocking constant cost is usually set to positive infinity (…). Through this gating mechanism, the system completely eliminates the possibility of dense dust at the image edges masquerading as a truncating action. Ultimately, the system linearly superimposes and aggregates the spatial smoothing penalty cost and temporal transition penalty cost corresponding to each image grid, and the two together constitute the second cost data that constrains the macroscopic physical evolution law.
[0069] Step S6: Construct a global integrated cost function based on the first cost data and the second cost data, and solve for the target state distribution that satisfies the condition of minimizing the global integrated cost function;
[0070] Because the working video frames are divided into hundreds or even thousands of image grids, and each grid has three possible physical evolution states, the total number of state combinations in the entire image exhibits an exponential growth. To find a solution that best matches the current underlying image features within this vast solution space, the system needs to fuse isolated cost matrices into a unified mathematical evaluation model. Specifically, the system constructs a global comprehensive cost function based on the first and second cost data from the image grids.
[0071] In this embodiment, the calculation formula for the global synthesis cost function is configured as follows:
[0072] ;
[0073] This formula reveals the underlying logic of inferring the state of physical space. Among other things, This represents the global comprehensive cost function, which indicates the overall energy penalty of the current global graph state allocation scheme. The smaller the penalty, the more reasonable the scheme. The set of meshes representing the target cut-off region; and These represent the preset physical evolution states (i.e., the first static state, the second transitional state, or the third nascent state) assigned to image meshes p and q, respectively. In the data section of the formula, The first image feature (i.e., local texture information entropy) represents the image grid p. The second image feature (i.e., the rate of expansion of the sinking motion) represents the image grid p. Let p represent the first cost data of the image grid, which physically represents the degree of matching between the local visual features of the current grid and the assumed state. In the smoothing term of the formula, The second cost data represents the relationship between image grids p and q, and is used to constrain the coherence of the state space and time. It represents a set of spatiotemporal neighborhoods that includes spatial adjacency and the previous frame in time sequence. That is, it not only considers the state continuity of adjacent grids in the same frame, but also forces the legality of the state transition between the previous frame and the current frame of the same grid on the time axis. This represents the preset spatiotemporal constraint weighting coefficient, used to balance the proportion between local visual confidence and macroscopic physical laws. For example, in areas with extreme dust obscuration, the visual features of a certain grid are extremely blurry, leading to a lack of discriminative power in the first cost data. In this case, increasing the weighting coefficient... With a value (e.g., set to 1.5), the system will rely more on the state of the surrounding grid and the motion waves of the coal mining machine to dominate the state inference of that grid.
[0074] After constructing the global comprehensive cost function, the system's underlying configuration and invocation of the graph cuts optimization algorithm minimizes the global comprehensive cost function. The core of the graph cuts algorithm lies in transforming the discrete Markov energy minimization problem into a minimum-cut / maximum-flow problem in graph theory. In the specific execution flow, the system first dynamically maps a topology graph in memory. Ordinary nodes in the graph correspond to each image grid within the target cut region, while terminal nodes correspond to multiple preset physical evolution states. Subsequently, the system transforms the first cost data into the connection edge capacity (data link) between ordinary nodes and terminal nodes, and the second cost data into the connection edge capacity (neighborhood link) between adjacent ordinary nodes. After the network topology graph is constructed, the system executes the maximum-flow algorithm on the graph to find the cutting scheme (i.e., the minimum cut) that completely separates different terminal nodes and minimizes the total edge capacity. According to rigorous proof in graph theory, the cutting path of this minimum cut is mathematically equivalent to the global minimum (or strong local minimum) of the global comprehensive cost function.
[0075] Through the precise segmentation of the graph cut optimization algorithm described above, the system successfully outputs the optimal state allocation matrix that minimizes the global comprehensive cost function. In this matrix, each two-dimensional image grid, which was originally full of uncertainty, is assigned a unique and reasonable physical evolution state scalar. Even if a grid exhibits strong visual pseudo-features of truncation phase transition due to severe interference from local bright water mist, under the global view of graph cut optimization, since it is located in the unreachable region of the coal mining machine's mechanical kinematic mask, forcibly flipping its state would trigger a second cost penalty that tends to infinity on the neighborhood links. Therefore, the algorithm will decisively maintain it in its original static state. Finally, the system formally determines the optimal state allocation matrix obtained after global optimization game as the target state distribution result, thus completing the lossless cognitive leap from the low-level chaotic pixel features to the high-level structured physical state, providing a unique and high-confidence state decision basis for subsequent three-dimensional physical truncation contour reconstruction.
[0076] Step S7: Extract the topological boundary lines between physical evolution states based on the target state distribution results, and generate the three-dimensional physical cutting contour of the corresponding coal mining equipment;
[0077] After obtaining the target state distribution result reflecting global energy minimization, the system has transformed the originally chaotic pixel flow into a grid state matrix with clear macroscopic physical meaning. However, automated control systems (such as electro-hydraulic control systems for hydraulic supports) cannot directly read the discrete two-dimensional grid state for mechanical height adjustment control. To achieve the transition from visual perception to physical execution, the system must perform a series of inverse mapping reconstructions from discrete topology to continuous geometry and from two-dimensional pixels to three-dimensional physical space. Specifically, the core of this reconstruction process is to extract the topological boundary lines between physical evolution states based on the target state distribution result, and ultimately generate the corresponding three-dimensional physical cutting contour of the coal mining equipment.
[0078] In the specific implementation process, the system first performs connected component traversal on the output optimal state allocation matrix (i.e., the target state distribution result). Since the preceding spatiotemporal graph cut optimization has already assigned a definite state label to each image grid, the system focuses on finding the connected components of the grid in the first static state (representing the original coal face that has not yet been destroyed) and the connected components of the grid in the third emerging state (representing the new coal face exposed after cutting). The system performs edge tracking between these two connected components exhibiting significant physical phase transitions, accurately locating the topological boundary between them. It is worth noting that, since the underlying layer is based on... The calculation is performed using a sized image grid, and the extracted topological boundary lines appear visually as stepped, discrete, jagged lines. If this jagged trajectory is directly sent to mechanical equipment, it will cause extremely violent, high-frequency mechanical vibrations in the rocker arm lifting cylinder of the coal mining machine, severely shortening the equipment's lifespan.
[0079] Based on this, the system's underlying layer is configured with a preset spline curve fitting algorithm to smooth the topological boundary lines. This fitting algorithm is not a simple moving average, but rather employs a cubic B-spline (B-Spline) control model. To eliminate the disorder of the boundary point set output by the graph cut algorithm in memory and the multi-valued bifurcation caused by local grid classification, the system first extracts the disordered discrete boundary node coordinates along the current traction direction of the coal mining equipment (i.e., the horizontal axis pixel coordinates of the image). Perform a monotonically increasing sort; then, for the same horizontal axis coordinate... To address multiple vertical axis jumps (multi-valued overlapping noise), the system uses a mean filter for local fusion and dimensionality reduction, ensuring that each horizontal axis coordinate corresponds to a unique vertical axis coordinate. Its input layer receives the ordered discrete node coordinates after monotonically sorting and deduplication, using them as control vertices. The intermediate layer of the algorithm constructs basis function approximation equations and optimizes the curvature tensor of the curve using the least squares method, ensuring that the final generated curve not only strictly passes through the phase transition core region but also possesses continuous second derivatives. After processing by this spline curve fitting algorithm, the system successfully eliminates the quantization error caused by the discrete mesh, generating smooth, continuously truncated contour lines on a two-dimensional pixel plane that conform to the dynamics of mechanical motion.
[0080] Although the system has obtained high-quality two-dimensional continuous truncated contours, due to the inherent perspective dimensionality reduction characteristics of monocular camera imaging, the two-dimensional pixel coordinates... Reverse derivation of three-dimensional physical coordinates This is a mathematical problem lacking depth constraints. To overcome this physical blind spot in monocular vision, the system must introduce external spatial prior constraints. In practice, the system acquires in real time the local reference plane equation derived from the spatial coordinates of equipment surrounding the coal face. In real coal mining scenarios, multiple hydraulic supports are usually closely arranged near the explosion-proof camera. These supports typically have tilt and stroke sensors mounted on their side plates or bases. The system reads the spatial attitude coordinates of adjacent hydraulic supports through the industrial communication bus of the fully mechanized mining face and uses principal component analysis (PCA) or the spatial three-point surface method to fit the local reference plane equation of the current coal face. This equation can be expressed in three-dimensional space using standard analytical geometry formulas. ,in These are the normal vector components of the coal face plane. The intercept is given. The introduction of this reference plane perfectly fills the gap in the depth dimension constraint missing in two-dimensional inverse projection.
[0081] After establishing spatial plane constraints, the system constructs a spatial ray intersection model based on the inverse matrix of pre-calibrated camera projection parameters. The internal geometric logic of this intersection model follows the principle of pinhole imaging reverse ray tracing. Specifically, for any two-dimensional pixel coordinate on a continuously cut contour line... The system first converts it to homogeneous coordinates. Then, left-multiply by the joint inverse of the camera intrinsic and extrinsic parameter matrices. Calculate the ray direction vector in the three-dimensional world coordinate system. The physical optical center of the camera is used as the starting point of the ray. The system constructs a three-dimensional ray that strikes the real coal face, and its parameterized ray equation is expressed as follows: ,in This is an unknown parameter characterizing the depth. Furthermore, the system substitutes this ray equation into the previously obtained local reference plane equation. In the process, calculate the intersection of these two points. Solve analytically for the unique parameter. The system can then accurately calculate the set of three-dimensional physical intersection points between the spatial rays corresponding to the pixel coordinates on the continuous cut contour line and the equation of the local reference plane.
[0082] Assuming that after spline fitting, the pixel coordinates of a key point on the 2D contour line are... The system uses inverse matrix operations to determine the direction vector of the ray that travels from the camera's optical center through the pixel and onto the coal face. It then calculates the unique physical intersection of this ray with the coal face plane equation derived from the hydraulic support sensor (e.g., the plane is approximately 2.5 meters vertically from the camera). Finally, it calculates the absolute 3D coordinates of this point. The system traverses the entire two-dimensional curve, mapping and solving each pixel one by one, and finally determines the set of three-dimensional physical intersection points as the three-dimensional physical cutting contour of the corresponding coal mining equipment. Through this series of closed-loop reconstruction links from noise reduction fitting to depth completion and then to ray intersection, the system breaks down the physical barrier between "two-dimensional representation" and "three-dimensional entity". The output high-precision three-dimensional physical cutting contour will be directly sent to the underlying control bus via industrial Ethernet, becoming the command reference for driving the precise cutting of the coal mining machine drum.
[0083] In summary, this embodiment establishes a direct mapping between visual observation data and the actual material stripping and destruction process by simultaneously extracting dual-modal physical features characterizing the variation in surface roughness of the coal face and the expansion of local gravity subsidence. Based on this, the system further integrates the state transition probability of the microscopic image mesh with the mechanical kinematic wavefront of the macroscopic coal mining equipment, constructing a global comprehensive cost function strictly constrained by the laws of physical spatiotemporal continuity. A graph cut optimization algorithm is then used to solve for the global optimal solution in the vast state solution space. This physical deduction path enables the system to adaptively shield against random visual interference caused by high-concentration dust dispersion and local strong light spot intersections underground, accurately locating the actual physical phase transition boundary conforming to mechanical dynamics. Finally, by finding the intersection of reverse spatial rays, the two-dimensional continuous topological boundary line is transformed into a precise three-dimensional physical cutting contour. This provides the fully mechanized mining face automated control system with a decision-making benchmark possessing extremely high engineering robustness and confidence, effectively solving the technical problem of the easy failure of conventional visual monitoring under normal and harsh working conditions.
[0084] Example 2:
[0085] like Figure 4 As shown, the image recognition-based coal mine face monitoring system includes:
[0086] The data acquisition module is used to acquire a time-continuous sequence of working face video frames, as well as the physical space status data of the coal mining equipment corresponding to the timestamps of each video frame in the working face video frame sequence.
[0087] The grid division module is used to determine the target truncating region in the video frame sequence of the working surface based on the physical space state data and the pre-calibrated camera projection parameters, and to divide the target truncating region into multiple image grids;
[0088] The feature extraction module is used to extract the first image feature and the second image feature of each image grid respectively; the first image feature represents the difference in surface roughness distribution of pixels in the image grid, and the second image feature represents the local sinking expansion rate of pixels in the image grid in the direction of target gravity projection.
[0089] The first cost calculation module is used to combine the first image features, the second image features, and the baseline feature data determined based on the physical space state data to calculate the first cost data for each image grid to independently transition to multiple preset physical evolution states.
[0090] The second cost calculation module is used to construct a motion direction constraint model of the coal mining equipment based on physical space state data, and to calculate the second cost data of state transition of each image grid in the spatiotemporal dimension in combination with the motion direction constraint model.
[0091] The global optimization solution module is used to construct a global comprehensive cost function based on the first cost data and the second cost data, and solve for the target state distribution that satisfies the condition of minimizing the global comprehensive cost function.
[0092] The cutting contour generation module is used to extract the topological boundary lines between physical evolution states based on the target state distribution results, and generate the three-dimensional physical cutting contour of the corresponding coal mining equipment.
[0093] The above description is merely an example and illustration of the structure of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the structure of the invention or exceed the scope defined in the claims, all of which should fall within the protection scope of the present invention.
[0094] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0095] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A coal mine face monitoring method based on image recognition, characterized in that, Includes the following steps: Acquire a time-continuous sequence of working face video frames, and physical space status data of the coal mining equipment corresponding to the timestamps of each video frame in the working face video frame sequence; Based on the physical space state data and the pre-calibrated camera projection parameters, the target truncated region in the working surface video frame sequence is determined, and the target truncated region is divided into multiple image grids; First image features and second image features are extracted from each image grid. The first image feature represents the difference in surface roughness distribution of pixels within the image grid, and the second image feature represents the local sinking expansion rate of pixels within the image grid in the direction of target gravity projection. Combining the first image features, the second image features, and the baseline feature data determined based on the physical space state data, the first cost data for each image grid to independently transition to multiple preset physical evolution states is calculated respectively. Based on the physical space state data, a motion direction constraint model of the coal mining equipment is constructed, and combined with the motion direction constraint model, the second cost data of state transition of each image grid in the spatiotemporal dimension is calculated. Based on the first cost data and the second cost data, a global comprehensive cost function is constructed, and the target state distribution result that satisfies the minimization condition of the global comprehensive cost function is solved. Based on the target state distribution results, the topological boundary lines between physical evolution states are extracted to generate the three-dimensional physical cutting contour of the corresponding coal mining equipment.
2. The coal mine working face monitoring method based on image recognition according to claim 1, characterized in that: The physical space state data includes the three-dimensional physical coordinates and three-dimensional traction speed vector of the coal mining equipment; based on the physical space state data and pre-calibrated camera projection parameters, the target cutting area in the working face video frame sequence is determined, including: Calculate the perspective projection coordinates of the three-dimensional physical coordinates of the coal mining equipment on the two-dimensional image pixel plane based on the camera projection parameters, and determine the perspective projection coordinates as the target center coordinates. Using the target center coordinates as a reference point, a rectangular bounding box of a preset size is delineated in a single frame image of the working surface video frame sequence, and the image area covered by the rectangular bounding box is determined as the target truncated area.
3. The coal mine working face monitoring method based on image recognition according to claim 1, characterized in that: The extraction process of the first image feature includes: counting the frequency of occurrence of a preset number of gray levels within the corresponding image grid, and calculating the local gray level probability distribution of the corresponding image grid through normalization processing; and calculating the local texture information entropy of the corresponding image grid as the first image feature based on the principles of information theory and the local gray level probability distribution. The extraction process of the second image feature includes: calculating the pixel motion vector field based on the image grid grayscale data of two adjacent video frames in the working surface video frame sequence; extracting the velocity component of the pixel motion vector field in the physical gravity projection direction in combination with the physical gravity projection direction calibrated in the camera projection parameters; calculating the spatial partial derivative of the velocity component along the physical gravity projection direction for each image grid, retaining the value greater than zero using a one-sided truncation function, and obtaining the sinking motion expansion rate of the image grid as the second image feature.
4. The coal mine working face monitoring method based on image recognition according to claim 2, characterized in that: The reference feature data includes first reference feature data and second reference feature data, and its calculation process includes: Based on the three-dimensional physical coordinates and three-dimensional traction velocity vector in the physical space state data, a first physical enclosure box located in front of the coal mining equipment at a first preset distance and a second physical enclosure box located behind the coal mining equipment at a second preset distance are determined respectively. Using the camera projection parameters, the first physical bounding box and the second physical bounding box are projected into a two-dimensional grid set. The average first image features within the first projection area are extracted as the first reference feature data, and the average first image features within the second projection area are extracted as the second reference feature data.
5. The coal mine working face monitoring method based on image recognition according to claim 4, characterized in that: The multiple preset physical evolution states include a first static state, a second transitional state, and a third nascent state; The calculation process for the first cost data includes: The degree to which the first image feature deviates unidirectionally from the first reference feature data is determined as the static state matching cost for transferring to the first static state; the static state matching cost is positively correlated with the difference between the first image feature and the first reference feature data. The degree to which the first image feature deviates unidirectionally from the second reference feature data is determined as the new state matching cost for transitioning to the third new state; the new state matching cost is positively correlated with the difference between the first image feature and the second reference feature data. Determine the intermediate transition feature value between the first reference feature data and the second reference feature data; Extract a state decay term that is negatively correlated with the second image feature, and a state penalty term that is positively correlated with the absolute difference between the first image feature and the intermediate transition feature value. The state decay term and the state penalty term are fused together to obtain the transition state matching cost for transitioning to the second transition state. The static state matching cost, the emerging state matching cost, and the transition state matching cost corresponding to each image grid are collectively used to form the first cost data.
6. The coal mine working face monitoring method based on image recognition according to claim 5, characterized in that: The process of constructing the motion direction constraint model includes: Using the camera projection parameters and the current three-dimensional physical coordinates of the coal mining equipment, the three-dimensional traction velocity vector is mapped to the two-dimensional image plane to obtain the corresponding two-dimensional projected velocity vector; For any image grid, calculate the two-dimensional displacement vector of its center coordinates relative to the center coordinates of the target; Based on the degree of consistency between the directions of the two-dimensional displacement vector and the two-dimensional projected velocity vector, as well as the spatial distance between the two-dimensional displacement vectors, the motion wavefront reachability of the image grid is determined. In response to the directional consistency degree meeting a preset angle threshold and the spatial distance meeting a preset range condition, the motion direction mask value of the corresponding image grid is set as an reachable identifier; otherwise, it is set as an unreachable identifier. The motion direction constraint model is constituted by the motion direction mask values of each of the image grids.
7. The coal mine face monitoring method based on image recognition according to claim 6, characterized in that: The calculation process for the second cost data includes: Calculate the spatial smoothing penalty cost when spatially adjacent image grids within the same video frame are assigned different physical evolution states; Based on the motion direction mask value, calculate the temporal transition penalty cost for state transitions occurring between two adjacent video frames in the same image grid. The process of determining the temporal transition penalty cost includes: when the image grid is in the first static state in the previous video frame and transitions to the second transition state in the current video frame, if the corresponding motion direction mask value is the reachable identifier, then a base constant cost is assigned; if the corresponding motion direction mask value is the unreachable identifier, then a blocking constant cost is assigned. The spatial smoothing penalty cost and the temporal transition penalty cost corresponding to each image grid are combined to form the second cost data.
8. The coal mine working face monitoring method based on image recognition according to claim 7, characterized in that: Based on the first cost data and the second cost data, a global comprehensive cost function is constructed, and the target state distribution result satisfying the minimization condition of the global comprehensive cost function is solved, including: Based on the first cost data and the second cost data of the image mesh, a global comprehensive cost function is constructed, and its calculation formula is as follows: ; in, Represents the global synthesis cost function. The set of meshes representing the target cut-off region. and These represent the preset physical evolution states assigned to image grid p and image grid q, respectively. The first image feature represents the image grid p. The second image feature represents the image grid p. This represents the first cost data of the image grid p. This represents the second cost data between image grids p and q. This represents a set of spatiotemporal neighborhoods that includes spatially adjacent elements and the temporally preceding frame. This represents the preset spatiotemporal constraint weighting coefficients; The global integrated cost function is minimized using a graph cut optimization algorithm, and the optimal state allocation matrix that minimizes the global integrated cost function is output. The optimal state allocation matrix is then determined as the target state distribution result.
9. The coal mine working face monitoring method based on image recognition according to claim 8, characterized in that: Based on the target state distribution results, the topological boundary lines between physical evolution states are extracted to generate the corresponding three-dimensional physical cutting contour of the coal mining equipment, including: In the target state distribution results, the topological boundary line between the grid connected domain in the first static state and the grid connected domain in the third emerging state is located; The topological boundary line is smoothed using a preset spline curve fitting algorithm to generate a continuous cut contour line on a two-dimensional pixel plane. Obtain the local reference plane equation derived in advance from the spatial coordinates of the equipment around the coal mining face; A spatial ray intersection model is constructed based on the inverse matrix of the camera projection parameters. The set of three-dimensional physical intersection points between the spatial rays corresponding to the pixel coordinates on the continuous cut contour line and the equation of the local reference plane is calculated. The set of three-dimensional physical intersection points is determined as the three-dimensional physical cut contour.
10. A coal mine face monitoring system based on image recognition, characterized in that: The coal mine face monitoring method based on image recognition as described in any one of claims 1-9 includes: The data acquisition module is used to acquire a time-continuous sequence of working face video frames, and physical space status data of the coal mining equipment corresponding to the timestamps of each video frame in the working face video frame sequence. The grid division module is used to determine the target truncated region in the video frame sequence of the working surface based on the physical space state data and the pre-calibrated camera projection parameters, and to divide the target truncated region into multiple image grids; The feature extraction module is used to extract the first image feature and the second image feature of each image grid respectively; the first image feature characterizes the difference in surface roughness distribution of pixels in the image grid, and the second image feature characterizes the local sinking expansion rate of pixels in the image grid in the direction of target gravity projection; The first cost calculation module is used to combine the first image features, the second image features, and the baseline feature data determined based on the physical space state data to calculate the first cost data for each image grid to independently transition to multiple preset physical evolution states. The second cost calculation module is used to construct a motion direction constraint model of the coal mining equipment based on the physical space state data, and to calculate the second cost data of state transition of each image grid in the spatiotemporal dimension in combination with the motion direction constraint model. The global optimization solution module is used to construct a global comprehensive cost function based on the first cost data and the second cost data, and solve for the target state distribution result that satisfies the minimization condition of the global comprehensive cost function. The cutting contour generation module is used to extract the topological boundary lines between physical evolution states based on the target state distribution results, and generate the three-dimensional physical cutting contour of the corresponding coal mining equipment.