Method and apparatus for cross-modal coal gangue separation
By employing a cross-modal coal gangue sorting method and utilizing multimodal image fusion technology, the problems of false detection and difficult detection under a single detection method are solved, and more accurate coal gangue sorting is achieved.
Patent Information
- Application Number
- CN202111532203.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-14
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2041-12-14
AI Technical Summary
Existing coal gangue sorting equipment is prone to false detection or difficulty in detection when using a single detection method, resulting in high uncertainty in the detection results.
A cross-modal coal gangue sorting method is adopted, which acquires images of multiple modalities through different types of imaging devices, and uses a trained coal gangue sorting model to perform feature extraction, cross-modal feature fusion and multimodal image detection, and outputs the final coal gangue detection results.
It improves the certainty of coal gangue detection, reduces the defects of false detection and difficulty in detection, and enhances the accuracy of detection results.
Smart Images

Figure CN114519377B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of coal gangue sorting technology, and in particular to a cross-modal coal gangue sorting method and apparatus. Background Technology
[0002] In the traditional energy industry, coal gangue sorting is a crucial step in coal production. In recent years, the widespread application of automatic coal gangue sorting equipment has improved the efficiency of coal production. However, under current technology, most coal gangue sorting equipment uses a single detection method, such as X-ray detection or electromagnetic detection. These coal gangue sorting equipment using a single detection method are prone to false detection or failure to detect certain defects. Summary of the Invention
[0003] This invention provides a cross-modal coal gangue sorting method and apparatus to solve the defects of false detection or difficulty in detection when using a single detection method to detect coal gangue in the prior art, and to achieve more certain detection results.
[0004] This invention provides a cross-modal coal gangue sorting method, comprising: acquiring N modal images of the coal gangue to be sorted, wherein the N modal images are acquired by different types of imaging devices; inputting the N modal images into a trained coal gangue sorting model to obtain the sorting result of the coal gangue to be sorted; wherein the coal gangue sorting model is obtained by training based on the N modal coal gangue image samples and the actual detection results corresponding to the N modal coal gangue image samples, where N is a positive integer greater than or equal to 2.
[0005] According to the present invention, a cross-modal coal gangue sorting method is provided, wherein the coal gangue sorting model includes a feature map extraction layer, a cross-modal feature fusion layer, and a multimodal image detection layer; wherein, the feature map extraction layer is used to extract features from the images of the N modalities respectively, and outputs the convolutional feature maps corresponding to the images of each modality in the images of the N modalities; the cross-modal feature fusion layer is used to fuse the convolutional feature maps corresponding to the images of each modality in the images of the N modalities, and outputs the fused feature maps corresponding to the images of each modality in the images of the N modalities; the multimodal image detection layer is used to detect coal gangue in the images of each modality through the fused feature maps corresponding to the images of each modality in the images of the N modalities, and obtains the coal gangue detection results in the images of each modality, and outputs the coal gangue detection results corresponding to the images of the N modalities according to the mapping relationship between the images of the N modalities.
[0006] According to the present invention, a cross-modal coal gangue sorting method is provided, wherein the cross-modal fusion of convolutional feature maps corresponding to each modality in the images of N modalities, and outputting fused feature maps corresponding to each modality in the images of N modalities, includes: flattening the convolutional feature maps corresponding to each modality in the images of N modalities to obtain multi-scale feature vectors corresponding to each modality; fusing the multi-scale feature vectors corresponding to each modality based on a planar cross-attention mechanism to obtain fused feature vectors corresponding to each modality; inputting the fused feature vectors corresponding to each modality in combination with position encoding into a multi-head self-attention encoder, and fusing the feature vectors output by the multi-head self-attention encoder using a planar cross-attention mechanism and inputting them into a multi-head self-attention decoder; restoring the feature vectors output by the multi-head self-attention decoder to the size of the convolutional feature maps to obtain the fused feature maps corresponding to each modality.
[0007] According to the present invention, a cross-modal coal gangue sorting method is provided, the method further comprising: training the coal gangue sorting model; wherein, when N=2, training the coal gangue sorting model includes: acquiring a first image sample, a second image sample, and actual detection results corresponding to the first image sample and the second image sample, wherein the first image sample and the second image sample are coal gangue image samples obtained by taking pictures of the same coal gangue sample with different types of imaging devices; wherein the first image sample constitutes a first input sample set, the second image sample constitutes a second input sample set, and the actual detection results corresponding to the first image sample and the second image sample constitute an output sample set; randomly selecting a first image sample from the first input sample set as a first input training sample, and taking a second image sample corresponding to the first input training sample from the second input sample set as a second input training sample; inputting the first input training sample and the second input training sample into the coal gangue sorting model to obtain an output result; and taking the actual detection results corresponding to the first input training sample and the second input training sample from the output sample set. The detection results are used to calculate a loss value based on the output results and the actual detection results, and the parameters of the coal gangue sorting model are updated based on the loss value. It is then determined whether the training termination condition has been met. If so, the parameters of the coal gangue sorting model in the current iteration are saved to obtain a completed coal gangue sorting training model; otherwise, the next first input training sample and the next second input training sample are selected for training. The step of inputting the first input training sample and the second input training sample into the coal gangue sorting model to obtain the output result includes: inputting the first input training sample into... The first input convolutional feature map is obtained from the feature map extraction layer of the coal gangue sorting model. The second input training sample is input into the feature map extraction layer of the coal gangue sorting model to obtain the second input convolutional feature map. The first input convolutional feature map is input into the cross-modal feature fusion layer of the coal gangue sorting model to obtain the first feature map. The second convolutional feature map is input into the cross-modal feature fusion layer of the coal gangue sorting model to obtain the second feature map. The first feature map and the second feature map are input into the multimodal image detection layer of the coal gangue sorting model to obtain the output result.
[0008] According to a cross-modal coal gangue sorting method provided by the present invention, the method of inputting the first input convolutional feature map into the cross-modal feature fusion layer of the coal gangue sorting model to obtain a first feature map, and inputting the second convolutional feature map into the cross-modal feature fusion layer of the coal gangue sorting model to obtain a second feature map, specifically involves: in the cross-modal feature fusion layer, flattening the first input convolutional feature map and the second input convolutional feature map to obtain a first feature vector and a second feature vector respectively; fusing the first feature vector and the second feature vector through a planar cross-attention mechanism to obtain a third feature vector and a fourth feature vector respectively; and then fusing the third feature vector... The vector and the fourth feature vector are combined with positional encoding and then input into a multi-head self-attention encoder to obtain the fifth and sixth feature vectors, respectively. The fifth and sixth feature vectors are fused through a planar cross-attention mechanism to obtain the seventh and eighth feature vectors, respectively. The seventh and eighth feature vectors are then input into a multi-head self-attention decoder to obtain the ninth and tenth feature vectors, respectively. The ninth feature vector is restored to the size of the first input convolutional feature map to obtain the first feature map, and the tenth feature vector is restored to the size of the first input convolutional feature map to obtain the second feature map.
[0009] According to the cross-modal coal gangue sorting method provided by the present invention, the planar cross-attention mechanism specifically comprises: transforming the feature vector of an image in the current layer of one modality into a feature map through channel deformation and mapping it onto the feature map of the feature vector of the current layer of another modality after channel deformation; restoring each feature in the mapped feature map to the size of the convolutional feature map corresponding to the image of that modality, multiplying it with a preset weight, and then adding it element-wise to the feature map of the feature vector of the image of the other modality in the current layer after channel deformation, to obtain the fused feature map of the image of the other modality in the current layer; and flattening the fused feature map by channels to obtain the fused feature vector.
[0010] According to a cross-modal coal gangue sorting method provided by the present invention, the method involves inputting the first feature map and the second feature map into the multimodal image detection layer of the coal gangue sorting model to obtain a detection result, the detection result being the output result. Specifically, in the multimodal image detection layer, the first feature map is input into the detector to obtain a first center heat map, a first center offset map, and a first width-height map; the second feature map is input into the detector to obtain a second center heat map, a second center offset map, and a second width-height map; the category of the corresponding coal gangue target is determined based on the first center heat map as a first category; the category of the corresponding coal gangue target is determined based on the second center heat map as a second category; and the bounding box of the corresponding coal gangue target is determined based on the first center heat map, the first center offset map, and the first width-height map as a first... The bounding box is determined based on the second center heatmap, the second center offset map, and the second width and height map, and is used as the second bounding box. The first category, the second category, the first bounding box, and the second bounding box are fused through a detection result fusion mechanism to obtain the detection result, which is the output result. Specifically, the detection result fusion mechanism is as follows: the detected position and category in one modality image are mapped to another modality image, and it is determined whether the maximum overlap rate of the bounding boxes of the two image planes after mapping is greater than a preset threshold. If so, it is determined whether the categories of the mapped bounding boxes are the same; otherwise, the categories of coal gangue corresponding to the two modal images are different. If the categories of the mapped bounding boxes are the same, the detection result is the average value of the bounding boxes and the corresponding same category; otherwise, the detection result is an uncertain category.
[0011] According to a cross-modal coal gangue sorting method provided by the present invention, the step of calculating the loss value based on the output result and the actual detection result includes: calculating the loss value using a target loss function based on the output result and the actual detection result; wherein the target loss function is determined based on the basic detection loss function and the cyclic mapping loss function of the cross-modal detection result.
[0012] The present invention also provides a cross-modal coal gangue sorting device, comprising:
[0013] The image acquisition module is used to acquire images of N modalities of the coal gangue to be sorted, wherein the images of the N modalities are acquired by different types of imaging devices;
[0014] The sorting module is used to input the images of the N modalities into the trained coal gangue sorting model to obtain the sorting result of the coal gangue to be sorted; wherein, the coal gangue sorting model is obtained by training based on the coal gangue image samples of the N modalities and the actual detection results corresponding to the coal gangue image samples of the N modalities, where N is a positive integer greater than or equal to 2.
[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of any of the above-described cross-modal coal gangue sorting methods.
[0016] The cross-modal coal gangue sorting method and apparatus provided by the present invention adopts a new coal gangue sorting model and detects coal gangue based on multimodal images. To a certain extent, it eliminates the defects of false detection and difficulty in detection that are easy to occur in the detection process of single-modal images. At the same time, the images of multiple modalities can compensate each other for image information, making the final detection result more accurate than that of using single-modal images. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0018] Figure 1 This is a schematic flowchart of the cross-modal coal gangue sorting method provided by the present invention;
[0019] Figure 2 This is a schematic diagram of the structure of a coal gangue sorting system provided by the present invention;
[0020] Figure 3 This is a schematic diagram of the cross-modal fusion process provided by the present invention;
[0021] Figure 4 This is a schematic diagram of the structure of the coal gangue sorting model provided by the present invention;
[0022] Figure 5 This is a flowchart illustrating the training coal gangue sorting model provided by the present invention.
[0023] Figure 6 This is a schematic diagram of the structure of the cross-modal coal gangue sorting device provided in an embodiment of the present invention;
[0024] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0026] The terms "first," "second," etc., used in the specification and claims of this invention are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0027] It should be understood that the phrase "one embodiment" or "an embodiment" in the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the invention. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification do not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments.
[0028] The following is combined with Figures 1-5 This invention describes the cross-modal coal gangue sorting method provided by the present invention.
[0029] Figure 1 This is a schematic flowchart of the cross-modal coal gangue sorting method provided by the present invention. It should be noted that the executing entity in the embodiments of the present invention can be a coal gangue sorting device, which can be a functional module or functional entity in an electronic device capable of implementing the cross-modal coal gangue sorting method; the executing entity in the embodiments of the present invention can also be a data processing unit located in a processor or microprocessor or other device capable of implementing the cross-modal coal gangue sorting method; wherein, the coal gangue sorting device or data processing unit can be implemented through a combination of software and / or hardware. Figure 1 As shown, the method includes:
[0030] Step 100: Obtain images of N modes of the coal gangue to be sorted.
[0031] The images of the N modalities were acquired using different types of imaging devices;
[0032] It is important to understand that multimodal images refer to images acquired by devices using different imaging principles for the same target. For example, X-ray imaging devices, visible light cameras, and infrared imaging devices can be used to photograph the same coal gangue target.
[0033] For example, the N modal images can be visible light images and X-ray images, or X-ray images and infrared images, or visible light images and infrared images, or images of various different modalities such as visible light images, X-ray images and infrared images.
[0034] It should be noted that coal gangue is a type of rock mixed in with coal seams. It contains a small amount of combustible material and is not easily combustible. Sort coal gangue can improve the quality of coal.
[0035] It is understandable that since the images are acquired through different types of imaging devices, N is a positive integer greater than or equal to 2.
[0036] Step 101: Input the images of the N modalities into the trained coal gangue sorting model to obtain the sorting results of the coal gangue to be sorted.
[0037] The coal gangue sorting model is obtained by training N modal coal gangue image samples and the actual detection results corresponding to the N modal coal gangue image samples, where N is a positive integer greater than or equal to 2.
[0038] Specifically, the input of the coal gangue sorting model is multiple modal images of the coal gangue to be sorted that have been collected, and the output is the sorting result of the coal gangue to be sorted, that is, the category of the coal gangue.
[0039] The coal gangue sorting model has the function of predicting the location and category of coal gangue based on multimodal images of the coal gangue to be sorted.
[0040] It should be noted that before sorting coal gangue, it is necessary to determine in advance the mapping relationship between the planar coordinate system of the images of different modes and the conveyor belt coordinate system, the mapping relationship between the conveyor belt coordinate system and the planar coordinate system of the images of different modes, and then determine the mapping relationship between the planar coordinate systems of the images of different modes.
[0041] Understandably, a certain number of coal gangue images of different modalities can be collected as training samples according to the training requirements of the coal gangue model. The actual detection results corresponding to the coal gangue images can be used as the expected output to train the coal gangue sorting model. The trained coal gangue model can effectively sort coal gangue.
[0042] In one embodiment, the cross-modal coal gangue sorting method provided by the present invention can be applied to a coal gangue sorting system, which includes: a conveyor belt, a robotic arm, a multimodal imaging module, an X-ray generator, an X-ray camera, and an industrial control computer. The conveyor belt is responsible for transporting the coal gangue to the field of view of the X-ray camera or the multimodal imaging module for detection. The robotic arm is used to grab the coal gangue on the conveyor belt. The multimodal imaging module is responsible for acquiring images of the coal gangue on the conveyor belt within the imaging range. The X-ray generator is responsible for emitting X-rays. The X-ray camera is responsible for receiving the diffracted X-ray images. The industrial control computer is responsible for processing the coal gangue sorting based on multimodal images (e.g., a joint detection task based on visible light images and X-ray images) and the motion control of the robotic arm provided by the embodiments of the present invention.
[0043] Figure 2 This is a schematic diagram of the structure of a coal gangue sorting system provided by the present invention. Figure 2 As shown, the system includes a visible light camera, a robotic arm, an X-ray camera, an X-ray transmitter, and a conveyor belt. This system can sort coal / gangue on the conveyor belt. A spatial rectangular coordinate system is constructed for the visible light camera, shown in Figure O. R X represents the center point of the image plane of a visible light camera. R The X-axis and Y-axis represent the image plane of a visible light camera. R The Y-axis and Z-axis represent the image plane of a visible light camera. R The Z-axis represents the image plane of the visible light camera; a spatial rectangular coordinate system is constructed for the image plane of the X-ray camera. In the figure, OI represents the center point of the image plane of the X-ray camera, XI represents the X-axis of the image plane of the X-ray camera, YI represents the Y-axis of the image plane of the X-ray camera, and Z represents the center point of the image plane of the X-ray camera. R The Z-axis represents the image plane of the X-ray camera; a spatial rectangular coordinate system is constructed for the conveyor belt, O in the figure. T X represents the center point of the conveyor belt. T The X-axis and Y-axis of the conveyor belt are represented by... T The Y-axis and Z-axis of the conveyor belt are represented. T This indicates the Z-axis of the conveyor belt.
[0044] Assume that the pixel in the image plane coordinate system of the visible light camera is [u R v R ] T And assuming the corresponding three-dimensional spatial point in the conveyor belt coordinate system is [X T Y T Z T ] T At this point, the transformation relationship from the conveyor belt coordinate system to the image plane coordinate system of the visible light camera is as follows:
[0045]
[0046] Among them, T TR This represents the mapping relationship between the conveyor belt coordinate system and the image plane coordinate system of the visible light camera, s R K is the scaling factor. R For the intrinsic parameters of a visible light camera, R... TR and t TR These are the external parameters of the visible light camera (representing rotation and translation from the conveyor belt coordinate system to the camera coordinate system, respectively).
[0047] The intrinsic and extrinsic parameters of the visible light camera are calibrated using a checkerboard pattern. Because one calibration step defines the coordinate system and axis by using a fixed position on the conveyor belt plane as the origin of the conveyor belt coordinate system, the extrinsic parameters from this calibration can be used as a transformation from the conveyor belt coordinate system to the visible light camera coordinate system.
[0048] Assume that the pixel in the image plane coordinate system of the X-ray camera is [u I v I ] T At this point, the transformation relationship from the conveyor belt coordinate system to the image plane coordinate system of the X-ray camera is as follows:
[0049]
[0050] Among them, T TI This represents the mapping relationship between the conveyor belt coordinate system and the image plane coordinate system of the X-ray camera, s I K is the scaling factor. I R represents the intrinsic parameters of the X-ray camera. TI and t TI These are the external parameters of the X-ray camera.
[0051] Knowing the transformation relationships from the conveyor belt coordinate system to the image plane coordinate systems of the visible light camera and the X-ray camera respectively, we can then use [u] R v R ] T Calculate [X] T Y T Z T ] T , and according to [u I v I ] T Calculate [X] T Y T Z T ] T :
[0052]
[0053]
[0054] Among them, T RT This represents the mapping relationship between the image coordinate system of the visible light camera and the conveyor belt coordinate system; T IT This represents the mapping relationship between the X-ray camera's image coordinate system and the conveyor belt coordinate system.
[0055] At this point, the transformation relationship from the visible light camera coordinate system to the X-ray camera coordinate system is as follows:
[0056]
[0057] The transformation relationship from the X-ray camera coordinate system to the visible light camera coordinate system is as follows:
[0058]
[0059] In this embodiment of the invention, a new coal gangue sorting model is adopted to detect coal gangue based on multimodal images. This eliminates, to some extent, the defects of single-modal images that are prone to false detection and difficult detection during the detection process. At the same time, the images of multiple modalities can compensate each other for image information, making the final detection result more accurate than that of using single-modal images.
[0060] Optionally, the coal gangue sorting model includes a feature map extraction layer, a cross-modal feature fusion layer, and a multimodal image detection layer;
[0061] The feature map extraction layer is used to extract features from the images of the N modalities respectively, and output the convolutional feature map corresponding to each modality in the images of the N modalities;
[0062] The cross-modal feature fusion layer is used to fuse the convolutional feature maps corresponding to each modality in the N modal images, and output the fused feature map corresponding to each modality in the N modal images;
[0063] The multimodal image detection layer is used to detect coal gangue in each modal image by using the fused feature map corresponding to each modal image in the N modal images, to obtain the coal gangue detection result in each modal image, and to output the coal gangue detection result corresponding to the N modal images according to the mapping relationship between the N modal images.
[0064] It should be noted that the convolutional feature map in the embodiments of the present invention refers to the multi-scale convolutional feature map extracted by a high-resolution feature network. For example, HRNet-48 can be used as a high-resolution feature network to extract the convolutional feature map.
[0065] It's important to understand that the coal gangue sorting model consists of three layers: a feature map extraction layer, a cross-modal feature fusion layer, and a multimodal image detection layer. The image is first input to the feature map extraction layer. The output of the feature map extraction layer becomes the input to the cross-modal feature fusion layer, and the output of the cross-modal feature fusion layer becomes the input to the multimodal image detection layer. The multimodal image detection layer then outputs the final detection result, which is the sorting result.
[0066] It should be noted that the mappings mentioned in this invention are all based on pre-determined mapping relationships.
[0067] Specifically, images of N modalities are input into the feature map extraction layer. The high-resolution feature network extracts multi-scale convolutional feature maps for each of the N modalities, resulting in N multi-scale convolutional feature maps. Each modality image is convolved with a single feature map. Figure 1 One-to-one correspondence.
[0068] It should be noted that cross-modal fusion in this embodiment of the invention refers to the mutual fusion of convolutional feature maps of images from different modalities.
[0069] Figure 3 This is a schematic diagram of the cross-modal fusion process provided by the present invention, such as... Figure 3 As shown, cross-modal fusion is performed on the convolutional feature maps of visible light images and X-ray images. The convolutional feature map of the X-ray image is incorporated into the convolutional feature map of the visible light image to obtain the fused feature map of the visible light image, and the convolutional feature map of the visible light image is incorporated into the convolutional feature map of the X-ray image to obtain the convolutional feature map of the X-ray image.
[0070] The N multi-scale convolutional feature maps output from the feature map extraction layer are input into the cross-modal feature fusion layer. The N multi-scale convolutional feature maps are fused across modally to obtain a fused feature map corresponding to the image of each modality.
[0071] It should be noted that the multimodal image detection layer mainly calculates the location and category of the corresponding coal gangue target based on the fusion feature map corresponding to each modality image. Then, the calculated location and category of each coal gangue target are fused through the mapping relationship between different modal images to obtain the final coal gangue detection result.
[0072] The N fused feature maps output from the cross-modal feature fusion layer are input into the multimodal image detection layer. Each fused feature map is then detected, resulting in N detection results. These N results are then fused using the mapping relationships between the images of the N modalities to obtain the final detection result. This result includes the position and category of the coal gangue target within the corresponding modal image. Based on the pre-determined mapping relationship between the planar coordinate system of the corresponding modal image and the conveyor belt coordinate system, the position of the coal gangue target on the conveyor belt can be determined.
[0073] Optionally, the images of the N modalities are input into the trained coal gangue sorting model to obtain the sorting results of the coal gangue to be sorted, including:
[0074] Feature extraction is performed on the images of the N modalities respectively, and the convolutional feature map corresponding to each modality in the images of the N modalities is output;
[0075] Cross-modal fusion of convolutional feature maps corresponding to each modality in the images of the N modalities, and outputting the fused feature map corresponding to each modality in the images of the N modalities;
[0076] The coal gangue in each modality of the N modal images is detected by the fused feature map corresponding to each modality of the images, and the coal gangue detection result in each modality of the images is obtained. According to the mapping relationship between the N modal images, the coal gangue detection result corresponding to the N modal images is output.
[0077] In this embodiment of the invention, a new coal gangue sorting model is constructed. In this model, a feature map extraction layer and a cross-modal feature fusion layer are set up so that the image information of multiple modal images can compensate each other and enhance the image features. Then, the enhanced feature map is analyzed by the multimodal image detection layer, which can effectively sort coal gangue. This realizes the joint sorting of cross-modal coal gangue detection model under the condition that the relative poses of each imaging device remain unchanged.
[0078] Optionally, the cross-modal fusion of convolutional feature maps corresponding to each modality in the N modal images, and the output of the fused feature map corresponding to each modality in the N modal images, includes:
[0079] Flatten the convolutional feature maps corresponding to each modality in the N modal images to obtain the multi-scale feature vectors corresponding to each modality image;
[0080] The multi-scale feature vectors corresponding to the images of each modality are fused based on the planar cross-attention mechanism to obtain the fused feature vectors corresponding to the images of each modality;
[0081] The fused feature vectors corresponding to the images of each modality are combined with position encoding and then input into a multi-head self-attention encoder. The feature vectors output by the multi-head self-attention encoder are then fused using a planar cross-attention mechanism and input into a multi-head self-attention decoder. The feature vectors output by the multi-head self-attention decoder are restored to the size of the convolutional feature map to obtain the fused feature map corresponding to the images of each modality.
[0082] The planar cross-fusion mechanism specifically involves: transforming the feature vector of an image in the current layer of one modality into a feature map through channel deformation and mapping it onto the feature map of the feature vector of the current layer of another modality after channel deformation; restoring each feature in the mapped feature map to the size of the convolutional feature map corresponding to the image of that modality; multiplying it with a preset weight; and then adding it element-wise to the feature map of the feature vector of the image of the other modality in the current layer after channel deformation to obtain the fused feature map of the image of the other modality in the current layer; and flattening the fused feature map by channels to obtain the fused feature vector.
[0083] Figure 4 This is a schematic diagram of the structure of the coal gangue sorting model provided by the present invention, as shown below. Figure 4 The following explains the planar cross-fusion mechanism:
[0084] HRNet-48 was used to extract multi-scale convolutional features M from visible light images and X-ray images, respectively. R and M I The two convolutional feature maps are then flattened to obtain two feature vectors m. R and m I Then, the fused feature vector f of the visible light image and the X-ray image is obtained through planar cross-attention PA. R and f I (f is not marked in the figure) R and f I ). Using the fused feature vector f of X-ray images I Taking the calculation as an example: Given the feature map M of the X-ray image I The size of the feature vector m I Feature map G is obtained by channel deformation. vm (mI); Feature map M of a known visible light image R The size of the feature vector m R Feature map G is obtained by channel deformation. vm (m R ); through the planar mapping relationship T RI G vm (m R Projected from the visible light image plane onto the X-ray image plane, and bilinear interpolation B is used. RI reconstruct each feature in the feature map to the feature map M of the X-ray image. I The size; multiply the above results element by w. RI , and the feature map G of the X-ray image vm (m I The fusion feature F of X-ray images is obtained by adding elements together. I ; F I The fused feature vector G of the X-ray image is obtained by flattening it by channel. mv (F I ), that is, f I :
[0085] f I =PA(m) I m R ) = G mv (wR I B RI (T RI G vm (m R ))+G vm (m I ));
[0086] Similarly, the fused feature vector f of a visible light image can be obtained through planar cross-attention. R :
[0087] f R =PA(m) R m I ) = G mv (wI R B IR (T IR G vm (m I ))+G vm (m R )).
[0088] In one embodiment, bilinear interpolation can be used to restore each feature in the mapped feature map to the size of the convolutional feature map corresponding to the modality image. It is important to understand that bilinear interpolation can either shrink or enlarge the feature map.
[0089] In this embodiment of the invention, the feature vectors corresponding to the images of each modality are fused by using a planar cross-fusion mechanism multiple times, so that the images of multiple modalities compensate each other for image information, enhance image features, and thus enhance the detection effect of the subsequent multimodal image detection layer, making the detection results more reliable.
[0090] Optionally, the method further includes: training the coal gangue sorting model;
[0091] Figure 5 This is a flowchart illustrating the training coal gangue sorting model provided by the present invention, as shown below. Figure 5 As shown, when N=2, the coal gangue sorting model is trained and includes:
[0092] Step 501: Obtain a first image sample, a second image sample, and the actual detection results corresponding to the first image sample and the second image sample. The first image sample and the second image sample are coal gangue image samples obtained by taking pictures of the same coal gangue sample with different types of imaging devices. The first image sample forms a first input sample set, the second image sample forms a second input sample set, and the actual detection results corresponding to the first image sample and the second image sample form an output sample set.
[0093] Step 502: Randomly select a first image sample from the first input sample set as the first input training sample, and take a second image sample corresponding to the first input training sample from the second input sample set as the second input training sample;
[0094] Step 503: Input the first input training sample and the second input training sample into the coal gangue sorting model to obtain the output result;
[0095] Step 504: Extract the actual detection results corresponding to the first input training sample and the second input training sample from the output sample set, calculate the loss value based on the output results and the actual detection results, and update the parameters of the coal gangue sorting model based on the loss value;
[0096] Step 505: Determine whether the training termination condition has been met. If so, save the parameters of the coal gangue sorting model in the current iteration and obtain the trained coal gangue sorting training model. Otherwise, select the next first input training sample and the next second input training sample, and repeat step 503.
[0097] It should be noted that in this embodiment of the invention, the training of the coal gangue sorting model is only for two modal images. When there are N modal images, the training process when N is greater than 2 can refer to the training process when N=2.
[0098] Specifically, coal gangue image samples obtained by taking pictures of the same coal gangue sample with two different types of imaging devices are used as input to the coal gangue sorting model. The actual detection results corresponding to the coal gangue image samples are used as the expected output to train the coal gangue sorting model. The loss value is calculated based on the output result and the actual detection result, and the parameters of the coal gangue sorting model are updated based on the loss value. The coal gangue sorting model is iteratively trained until the training termination condition is met.
[0099] The termination condition refers to reaching a preset number of iterations, or the loss value continuously decreasing and stabilizing; the number of iterations can be customized as needed.
[0100] It should be noted that before training the coal gangue sorting model, it is necessary to initialize the coal gangue sorting model, such as initializing the number of iterations and parameters of the coal gangue sorting model.
[0101] Understandably, training a coal gangue sorting model involves inputting images of coal gangue in different modalities into the model for forward propagation. Then, the loss value is calculated based on the model's output and actual detection results. This loss value is then backpropagated to update the model's parameters. The trained coal gangue sorting model then possesses the ability to sort coal gangue.
[0102] Optionally, step 503 inputs the first input training sample and the second input training sample into the coal gangue sorting model to obtain the output result, including:
[0103] Step 5031: Input the first input training sample into the feature map extraction layer of the coal gangue sorting model to obtain the first input convolutional feature map, and input the second input training sample into the feature map extraction model of the coal gangue sorting model to obtain the second input convolutional feature map;
[0104] Step 5032: Input the first input convolutional feature map into the cross-modal feature fusion layer of the coal gangue sorting model to obtain the first feature map, and input the second convolutional feature map into the cross-modal feature fusion layer of the coal gangue sorting model to obtain the second feature map;
[0105] Step 5033: Input the first feature map and the second feature map into the multimodal image detection layer of the coal gangue sorting model to obtain the output result.
[0106] Optionally, step 5032 involves inputting the first input convolutional feature map into the cross-modal feature fusion layer of the coal gangue sorting model to obtain a first feature map, and inputting the second convolutional feature map into the cross-modal feature fusion layer of the coal gangue sorting model to obtain a second feature map. Specifically:
[0107] In the cross-modal feature fusion layer, the first input convolutional feature map and the second input convolutional feature map are flattened to obtain the first feature vector and the second feature vector respectively. The first feature vector and the second feature vector are fused through a planar cross attention mechanism to obtain the third feature vector and the fourth feature vector respectively.
[0108] The third and fourth feature vectors are combined with position encodings and then input into a multi-head self-attention encoder to obtain the fifth and sixth feature vectors respectively.
[0109] The fifth and sixth feature vectors are fused using a planar cross-attention mechanism to obtain the seventh and eighth feature vectors.
[0110] The seventh and eighth feature vectors are input into the multi-head self-attention decoder to obtain the ninth and tenth feature vectors respectively.
[0111] The ninth feature vector is restored to the size of the first input convolutional feature map to obtain the first feature map, and the tenth feature vector is restored to the size of the first input convolutional feature map to obtain the second feature map.
[0112] It is understandable that the first feature map and the first input convolutional feature map are of equal size, and the second feature map and the second input convolutional feature map are of equal size.
[0113] Optionally, in step 5033, the first feature map and the second feature map are input into the multimodal image detection layer of the coal gangue sorting model to obtain the output result, specifically as follows:
[0114] In the multimodal image detection layer, the first feature map is input into the detector to obtain a first center heat map, a first center offset map, and a first width and height map; the second feature map is input into the detector to obtain a second center heat map, a second center offset map, and a second width and height map.
[0115] The corresponding coal gangue target category is determined based on the first center heat map and designated as the first category; the corresponding coal gangue target category is determined based on the second center heat map and designated as the second category.
[0116] The bounding box of the corresponding coal gangue target is determined based on the first center heat map, the first center offset map, and the first width and height map, and is used as the first bounding box. The bounding box of the corresponding coal gangue target is determined based on the second center heat map, the second center offset map, and the second width and height map, and is used as the second bounding box.
[0117] The first category, the second category, the first bounding box, and the second bounding box are fused together using a detection result fusion mechanism to obtain the detection result, which is the output result.
[0118] The detection result fusion mechanism is specifically as follows:
[0119] The detected location and category in one modality image are mapped to another modality image. It is then determined whether the maximum overlap rate of the bounding boxes in the two image planes after mapping is greater than a preset threshold. If so, it is determined whether the categories of the bounding boxes after mapping are the same. Otherwise, the categories of coal gangue corresponding to the two modal images are different. If the categories of the bounding boxes after mapping are the same, the detection result is the average value of the bounding boxes and the corresponding same category. Otherwise, the detection result is an uncertain category.
[0120] It should be noted that the multimodal image detection layer can only process feature maps.
[0121] The multimodal image detection layer employs a target detection algorithm, utilizing bounding box classification and regression to obtain the location and category of coal gangue; the bounding boxes corresponding to the coal gangue need to be pre-defined.
[0122] It is important to understand that inputting a feature map into the detector will result in three different images: the center heatmap of the bounding box, the center offset map of the bounding box, and the width and height map of the bounding box.
[0123] The category of the bounding box is determined by obtaining the category of the coal gangue target within the bounding box based on the center heat map; the location of the bounding box is determined by obtaining the bounding box location based on the center heat map, center offset map and width and height map.
[0124] It should be noted that when training the coal gangue sorting model using images from two modalities, the multimodal image detection layer outputs two results. The location and category of the coal gangue target can be obtained from either output. For example, if visible light images and X-ray images are used as inputs to the coal gangue sorting model, one output from the multimodal image detection layer is an X-ray image detection result assisted by the visible light image detection result; the other output is a visible light image detection result assisted by the X-ray image detection result.
[0125] It is understandable that when the maximum overlap rate of the bounding boxes corresponding to the images of two modalities is greater than a preset threshold, it can be preliminarily determined that the two modal images are images of the same target.
[0126] In one embodiment, when the bounding box position and category of the coal gangue target are determined, the plane coordinate system of the image of different modalities can be transformed to the conveyor belt coordinate system to obtain the position of the coal gangue target on the conveyor belt, thereby facilitating the robotic arm to grasp the coal gangue target online.
[0127] In this embodiment of the invention, by employing a target detection algorithm in the multimodal image detection layer of the coal gangue sorting model, the location and category of coal gangue are obtained from the fusion features of each modality using bounding box regression and classification. At the same time, due to the adoption of a detection result fusion mechanism, the detection results of different modal images assist each other. When training the cross-modal coal gangue sorting model, the information of different modal images compensates for each other, further improving the accuracy of the coal gangue sorting model.
[0128] Optionally, calculating the loss value based on the output result and the actual detection result includes:
[0129] Based on the output results and the actual detection results, the loss value is calculated using the target loss function;
[0130] The target loss function is determined based on the basic detection loss function and the cyclic mapping loss function of the cross-modal detection results.
[0131] It should be noted that the cyclic mapping loss function for cross-modal detection results refers to the weighted value of the target parameter mapping loss when mapping images from multiple modalities to each other. Taking visible light image mapping and X-ray image mapping as an example, the target parameter cyclic mapping loss function refers to the weighted value of the target parameter mapping loss from the visible light image to the X-ray image and the target parameter mapping loss from the X-ray image to the visible light image. The target parameter mapping loss from the visible light image to the X-ray image refers to the difference between the mapped detection result of the coal gangue detection result in the visible light image and the actual detection result of the X-ray image when the coal gangue detection result in the X-ray image is mapped to the conveyor belt plane and then mapped to the X-ray image through the conveyor belt. Similarly, the target parameter mapping loss from the X-ray image to the visible light image refers to the difference between the mapped detection result of the visible light image and the actual detection result of the visible light image when the coal gangue detection result in the X-ray image is mapped to the conveyor belt plane and then mapped to the visible light image through the conveyor belt.
[0132] In one embodiment, the basic detection loss function L of the detector is defined. total,basic ,for:
[0133] L total,basic =∑ o∈{I,R} L ctr,fcl (M hm,o )+L ctr,l1 (M coff,o )+
[0134] L wh,l1 (M wh,o )+L cls,fcl (M hm,o );
[0135] Where I represents an image of one modality, R represents an image of another modality, and M... hm M represents the center heatmap of the bounding box. coff M represents the center offset map of the bounding box. wh A width and height plot representing the bounding box; L ctr,fcl This indicates information about the heatmap M. hm Focal loss at the target center, L ctr,l1 L1 loss represents the loss with respect to the center offset of the bounding box. wh,l1 L1 loss represents the loss with respect to the width and height of the bounding box. cls,fcl For the heat map M hm Focal loss for the target category;
[0136] Define the cyclic mapping loss function L for cross-modal detection results. ctr,cy c and L cls,cyc ,for:
[0137] L ctr,cyc =L ctr,fcl (T RI M hm,R )+L ctr,fcl (T IR M hm,I );
[0138] L cls,cyc =L cls,fcl (T RI M hm,R )+L cls,fcl (T IR M hm,I );
[0139] Among them, T RI This represents the mapping relationship from the modal image corresponding to R to the modal image corresponding to I, where T IR This represents the mapping relationship from the modal image corresponding to I to the modal image corresponding to R; L ctr,fcl (T RI M hm,R ) indicates information about the heatmap T RI M hm,R Focal loss at the target center, L ctr,fcl (T IR M hm,I ) indicates information about the heatmap T IR M hm,I Focal loss at the target center, L cls,fcl (T RI M hm,R ) indicates information about the heatmap T RI M hm,R Focal loss for the target category, Lcls,fcl (T IR M hm,I ) indicates information about the heatmap T IR M hm,I Focal loss for the target category.
[0140] Based on the above, the target loss function can be defined as Lt o tal:
[0141] L total =L total,basic +γ(L ctr,cyc +L cls,cyc );
[0142] Where γ represents the weight, which takes values in the range [0, 1], and is used to balance the basic detection loss function and the cyclic mapping loss function of the detection target parameters.
[0143] In this embodiment of the invention, by setting a cyclic mapping loss function for the detection target parameters, the parameters of the coal gangue detection model are updated as an aid during the training of the coal gangue detection model, thereby ensuring the consistency of cross-modal target localization.
[0144] The following example uses the input of visible light and X-ray images of the coal gangue to be detected into a trained coal gangue sorting model, combined with... Figure 4 The cross-modal coal gangue sorting method provided by the present invention will be further described as follows:
[0145] The mapping relationship T between the conveyor belt coordinate system and the image plane coordinate system of the visible light camera is determined in advance. TR Determine the mapping relationship T between the conveyor belt coordinate system and the image plane coordinate system of the X-ray camera. TI Using HRNet-48, multi-scale convolutional features M are extracted from visible light images and X-ray images respectively. R and M I The two convolutional feature maps are then flattened to obtain two feature vectors m. R and m I Then, the fused feature vector f of the visible light image and the X-ray image is obtained through a planar cross-attention mechanism. R and f I (f is not marked in the figure) R and f I ).
[0146] f I and f R Adding it to the position code p, we get f I +p and f R +p are processed by a multi-head self-attention encoder to obtain feature vector e. I and e R(e not marked in the figure) I and e R Then, through the planar cross-attention mechanism, e is obtained respectively. I and e R The corresponding fusion feature vector PA(e) I e R ) and PA(e R e I (PA(e) is not marked in the figure) I e R ) and PA(e R e I The two fused feature vectors are then processed by a multi-head self-attention decoder to obtain feature vector d. I and d R Finally, the feature vector d I and d R The feature maps D are obtained by restoring them to their original sizes according to the multi-scale convolutional feature maps. I and D R .
[0147] Feature map D I and D R The feature maps D can be obtained by inputting them into the detector respectively. R The corresponding center heatmap, center offset map, and width / height map, feature map D I The corresponding center heatmap, center offset map, and width / height map are based on feature map D. R The corresponding center heatmap is used to obtain the corresponding bounding box classification, thereby determining the corresponding coal gangue category as c. R According to feature map D R The corresponding center heatmap, center offset map, and width / height map are used to obtain the bounding box location, thereby determining the corresponding coal gangue bounding box as b. R Similarly, it can be seen from the feature map D I The corresponding center heatmap is used to obtain the corresponding bounding box classification, thereby determining the corresponding coal gangue category as c. I According to feature map D I The corresponding center heatmap, center offset map, and width / height map are used to obtain the bounding box location, thereby determining the corresponding coal gangue bounding box as b. I .
[0148] b R c R b I and c I The detection results are fused to obtain the output result.
[0149] When training the model, the target loss function is L, based on the output results and the actual detection results. total =L total,basic +γ(Lctr,cyc +L cls,cyc Calculate the loss value and use the loss value to update the parameters of the coal gangue sorting model.
[0150] Figure 6 This is a schematic diagram of the structure of the cross-modal coal gangue sorting device provided in an embodiment of the present invention, as shown below. Figure 6 As shown, this embodiment of the invention provides a cross-modal coal gangue sorting device, including an image acquisition module 601 and a sorting module 602. The image acquisition module 601 is used to acquire N modal images of the coal gangue to be sorted, wherein the N modal images are acquired by different types of imaging devices. The sorting module 602 is used to input the N modal images into a trained coal gangue sorting model to obtain the sorting result of the coal gangue to be sorted. The coal gangue sorting model is obtained by training based on the N modal coal gangue image samples and the actual detection results corresponding to the N modal coal gangue image samples, where N is a positive integer greater than or equal to 2.
[0151] It should be noted that the cross-modal coal gangue sorting device provided in this embodiment of the invention can realize all the method steps implemented in the cross-modal coal gangue sorting method embodiment and achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.
[0152] The cross-modal coal gangue sorting device provided by the present invention adopts a new coal gangue sorting model and detects coal gangue based on multimodal images. To a certain extent, it eliminates the defects of false detection and difficulty in detection that are easy to occur in the detection process of single-modal images. At the same time, the images of multiple modalities can compensate each other for image information, making the final detection result more accurate than that of using single-modal images.
[0153] Figure 7 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 7 As shown, the electronic device may include a processor 710, a communications interface 720, a memory 730, and a communication bus 740. The processor 710, communications interface 720, and memory 730 communicate with each other via the communication bus 740. The processor 710 can call logical instructions in the memory 730 to execute the aforementioned cross-modal coal gangue sorting method, which includes, for example, acquiring images of N modalities of the coal gangue to be sorted; inputting the images of the N modalities into a trained coal gangue sorting model to obtain the sorting result of the coal gangue to be sorted.
[0154] Furthermore, the logical instructions in the aforementioned memory 730 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0155] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the cross-modal coal gangue sorting method provided by the above methods, for example including: acquiring images of N modalities of the coal gangue to be sorted; inputting the images of the N modalities into a trained coal gangue sorting model to obtain the sorting result of the coal gangue to be sorted.
[0156] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0157] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0158] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for separating cross-modal coal gangue, characterized in that, include: Images of N modalities of coal gangue to be sorted are acquired, wherein the images of the N modalities are acquired by imaging devices of different types; The images of the N modalities are input into the trained coal gangue sorting model to obtain the sorting results of the coal gangue to be sorted; The coal gangue sorting model is obtained by training N modal coal gangue image samples and the actual detection results corresponding to the N modal coal gangue image samples, where N is a positive integer greater than or equal to 2. The coal gangue sorting model includes a feature map extraction layer, a cross-modal feature fusion layer, and a multimodal image detection layer. The feature map extraction layer is used to extract features from the images of the N modalities respectively, and output the convolutional feature map corresponding to each modality in the images of the N modalities; The cross-modal feature fusion layer is used to fuse the convolutional feature maps corresponding to each modality in the N modal images, and output the fused feature map corresponding to each modality in the N modal images; The multimodal image detection layer is used to detect coal gangue in each modal image by using the fused feature map corresponding to each modal image in the N modal images, to obtain the coal gangue detection result in each modal image, and to output the coal gangue detection result corresponding to the N modal images according to the mapping relationship between the N modal images; The cross-modal fusion of convolutional feature maps corresponding to each modality in the N modal images, and the output of the fused feature map corresponding to each modality in the N modal images, includes: Flatten the convolutional feature maps corresponding to each modality in the N modal images to obtain the multi-scale feature vectors corresponding to each modality image; The multi-scale feature vectors corresponding to the images of each modality are fused based on the planar cross-attention mechanism to obtain the fused feature vectors corresponding to the images of each modality; The fused feature vectors corresponding to the images of each modality are combined with position encoding and then input into a multi-head self-attention encoder. The feature vectors output by the multi-head self-attention encoder are then fused using a planar cross-attention mechanism and input into a multi-head self-attention decoder. The feature vectors output by the multi-head self-attention decoder are restored to the size of the convolutional feature map to obtain the fused feature map corresponding to the images of each modality.
2. The cross-modal coal gangue sorting method according to claim 1, characterized in that, The method further includes: training the coal gangue sorting model; Wherein, when N=2, the training to obtain the coal gangue sorting model includes: A first image sample, a second image sample, and actual detection results corresponding to the first image sample and the second image sample are obtained. The first image sample and the second image sample are coal gangue image samples obtained by taking pictures of the same coal gangue sample with different types of imaging devices. The first image sample forms a first input sample set, the second image sample forms a second input sample set, and the actual detection results corresponding to the first image sample and the second image sample form an output sample set. Randomly select a first image sample from the first input sample set as the first input training sample, and take a second image sample corresponding to the first input training sample from the second input sample set as the second input training sample; The first input training sample and the second input training sample are input into the coal gangue sorting model to obtain the output result; The actual detection results corresponding to the first input training sample and the second input training sample are extracted from the output sample set. The loss value is calculated based on the output result and the actual detection result, and the parameters of the coal gangue sorting model are updated based on the loss value. Determine whether the training termination condition has been met. If so, save the parameters of the coal gangue sorting model in the current iteration and obtain the trained coal gangue sorting training model. Otherwise, select the next first input training sample and the next second input training sample for training. The step of inputting the first input training sample and the second input training sample into the coal gangue sorting model to obtain the output result includes: The first input training sample is input into the feature map extraction layer of the coal gangue sorting model to obtain the first input convolutional feature map, and the second input training sample is input into the feature map extraction model of the coal gangue sorting model to obtain the second input convolutional feature map; The first input convolutional feature map is input into the cross-modal feature fusion layer of the coal gangue sorting model to obtain the first feature map, and the second input convolutional feature map is input into the cross-modal feature fusion layer of the coal gangue sorting model to obtain the second feature map; The first feature map and the second feature map are input into the multimodal image detection layer of the coal gangue sorting model to obtain the output result.
3. The cross-modal coal gangue sorting method according to claim 2, characterized in that, The step of inputting the first input convolutional feature map into the cross-modal feature fusion layer of the coal gangue sorting model to obtain the first feature map, and inputting the second input convolutional feature map into the cross-modal feature fusion layer of the coal gangue sorting model to obtain the second feature map, specifically involves: In the cross-modal feature fusion layer, the first input convolutional feature map and the second input convolutional feature map are flattened to obtain the first feature vector and the second feature vector respectively. The first feature vector and the second feature vector are fused through a planar cross attention mechanism to obtain the third feature vector and the fourth feature vector respectively. The third and fourth feature vectors are combined with position encodings and then input into a multi-head self-attention encoder to obtain the fifth and sixth feature vectors respectively. The fifth and sixth feature vectors are fused using a planar cross-attention mechanism to obtain the seventh and eighth feature vectors. The seventh and eighth feature vectors are input into the multi-head self-attention decoder to obtain the ninth and tenth feature vectors respectively. The ninth feature vector is restored to the size of the first input convolutional feature map to obtain the first feature map, and the tenth feature vector is restored to the size of the first input convolutional feature map to obtain the second feature map.
4. The cross-modal coal gangue sorting method according to claim 1 or 3, characterized in that, The planar cross-attention mechanism is specifically as follows: The feature vector of an image in the current layer of one modality is transformed into a feature map through channel deformation and mapped onto the feature map of the feature vector of the image in the current layer of another modality, which is then transformed through channel deformation. Each feature in the mapped feature map is restored to the size of the convolutional feature map corresponding to the image of that modality, and multiplied with a preset weight. Then, it is added element-wise to the feature map of the image in the current layer of the other modality, which is transformed through channel deformation, to obtain the fused feature map of the image of the other modality in the current layer. The fused feature map is then flattened through channels to obtain the fused feature vector.
5. The cross-modal coal gangue sorting method according to claim 2, characterized in that, The first feature map and the second feature map are input into the multimodal image detection layer of the coal gangue sorting model to obtain the detection result, which is the output result. Specifically: In the multimodal image detection layer, the first feature map is input into the detector to obtain a first center heat map, a first center offset map, and a first width and height map; the second feature map is input into the detector to obtain a second center heat map, a second center offset map, and a second width and height map. The corresponding coal gangue target category is determined based on the first center heat map and designated as the first category; the corresponding coal gangue target category is determined based on the second center heat map and designated as the second category. The bounding box of the corresponding coal gangue target is determined based on the first center heat map, the first center offset map, and the first width and height map, and is used as the first bounding box. The bounding box of the corresponding coal gangue target is determined based on the second center heat map, the second center offset map, and the second width and height map, and is used as the second bounding box. The first category, the second category, the first bounding box, and the second bounding box are fused together using a detection result fusion mechanism to obtain the detection result, which is the output result. The detection result fusion mechanism is specifically as follows: The detected location and category in one modality image are mapped to another modality image. It is then determined whether the maximum overlap rate of the bounding boxes in the two image planes after mapping is greater than a preset threshold. If so, it is determined whether the categories of the bounding boxes after mapping are the same. Otherwise, the categories of coal gangue corresponding to the two modal images are different. If the categories of the bounding boxes after mapping are the same, the detection result is the average value of the bounding boxes and the corresponding same category. Otherwise, the detection result is an uncertain category.
6. The cross-modal coal gangue sorting method according to claim 2, characterized in that, The calculation of the loss value based on the output result and the actual detection result includes: Based on the output results and the actual detection results, the loss value is calculated using the target loss function; The target loss function is determined based on the basic detection loss function and the cyclic mapping loss function of the cross-modal detection results.
7. A cross-modal coal gangue sorting device, characterized in that, include: The image acquisition module is used to acquire images of N modalities of the coal gangue to be sorted, wherein the images of the N modalities are acquired by different types of imaging devices; The sorting module is used to input the images of the N modalities into the trained coal gangue sorting model to obtain the sorting result of the coal gangue to be sorted; wherein, the coal gangue sorting model is obtained by training based on the coal gangue image samples of the N modalities and the actual detection results corresponding to the coal gangue image samples of the N modalities, where N is a positive integer greater than or equal to 2; The coal gangue sorting model includes a feature map extraction layer, a cross-modal feature fusion layer, and a multimodal image detection layer. The feature map extraction layer is used to extract features from the images of the N modalities respectively, and output the convolutional feature map corresponding to each modality in the images of the N modalities; The cross-modal feature fusion layer is used to fuse the convolutional feature maps corresponding to each modality in the N modal images, and output the fused feature map corresponding to each modality in the N modal images; The multimodal image detection layer is used to detect coal gangue in each modal image by using the fused feature map corresponding to each modal image in the N modal images, to obtain the coal gangue detection result in each modal image, and to output the coal gangue detection result corresponding to the N modal images according to the mapping relationship between the N modal images; The cross-modal fusion of convolutional feature maps corresponding to each modality in the N modal images, and the output of the fused feature map corresponding to each modality in the N modal images, includes: Flatten the convolutional feature maps corresponding to each modality in the N modal images to obtain the multi-scale feature vectors corresponding to each modality image; The multi-scale feature vectors corresponding to the images of each modality are fused based on the planar cross-attention mechanism to obtain the fused feature vectors corresponding to the images of each modality; The fused feature vectors corresponding to the images of each modality are combined with position encoding and then input into a multi-head self-attention encoder. The feature vectors output by the multi-head self-attention encoder are then fused using a planar cross-attention mechanism and input into a multi-head self-attention decoder. The feature vectors output by the multi-head self-attention decoder are restored to the size of the convolutional feature map to obtain the fused feature map corresponding to the images of each modality.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the cross-modal coal gangue sorting method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Coal recognition method and device
CN108564108A