Mine car track foreign matter detection method and system based on deep neural network
Through a multimodal perception method based on deep neural network, foreign objects on mine car tracks are detected in real time, and safety threats caused by foreign objects such as rolling stones on mine car tracks are solved, high-precision foreign object detection and risk warning are achieved to ensure the safe production of mines.
Patent Information
- Application Number
- CN202510582942.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-05-07
AI Technical Summary
The suddenness and irregularity of foreign objects such as rolling stones on mine truck tracks leads to accidents such as derailment and subversion of mine trucks. It is difficult for the existing technology to achieve timely and accurate detection of foreign objects on mine truck tracks, threatening mine production safety and operating efficiency.
The multimodal perception method based on deep neural network is adopted, and the real-time image and text prompt words of mine car tracks are obtained, and the hierarchical multimodal features are extracted. The multimodal deep neural network is used to fuse images and text information, and the composite loss function and gradient descent backpropagation training model is combined to achieve foreign object detection and position recognition.
Real-time high-precision target detection and risk warning of foreign objects on mine truck tracks, and improve the reliability and operating efficiency of mine safety production.
Smart Images

Figure CN120472426A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of target detection technology, and specifically relates to a mine car track foreign object detection method and system based on deep neural network. Background Art
[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.
[0003] Mine cars are narrow-gauge railway vehicles used to transport bulk materials such as coal, ore, and waste rock in mines. They are primarily used for rail transportation in underground mine tunnels, shafts, and on the surface. The intrusion of foreign objects such as boulders into mine car tracks is sudden and erratic. Foreign objects on the tracks can collide with moving mine cars, leading to serious accidents such as derailment and overturning. As the primary transportation channel for ore, the inability to detect and identify foreign objects on the tracks in a timely and accurate manner poses a significant threat to the safety and efficiency of the entire mine.
[0004] Therefore, in order to ensure the driving safety of mine cars and the production safety of mines, it is urgent to monitor foreign objects in the mine track area in real time. Summary of the Invention
[0005] To solve the above problems, the present invention proposes a mine car track foreign object detection method and system based on deep neural network. Based on multimodal perception and target detection, it monitors and identifies in real time whether there are foreign objects on the mine car track and determines the size of the foreign objects and the position of the foreign objects on the mine car track, thereby realizing real-time high-precision target detection and risk warning of foreign objects on the mine car track.
[0006] According to some embodiments, a first solution of the present invention provides a method for detecting foreign objects on mine car tracks based on a deep neural network, which adopts the following technical solutions: A method for detecting foreign objects on mine car tracks based on a deep neural network, comprising: Get real-time images and text prompts of the minecart track; extracting hierarchical multimodal features of the acquired real-time image; The hierarchical multimodal features and text prompt words obtained by fusing multimodal deep neural networks are used to obtain the image-text fusion features of the mine car track; Detect whether there are foreign objects on the mine car track based on the obtained image-text fusion features and the preset foreign object detection model; When there is foreign matter, the size of the foreign matter is identified and its location is determined, completing real-time monitoring of foreign matter on the mine car track.
[0007] As a further technical limitation, the preset foreign object detection model adopts a deep neural network detection model, and the deep neural network detection model adopts a composite loss function and gradient descent back propagation for model training.
[0008] Furthermore, the composite loss function for ,in, represents the weight of each loss function, and ; Represents different loss functions, where Represents the text loss function, using the cross entropy loss function, that is , Indicates the marked text information. Indicates that a multimodal large language model predicts text information; Represents the target detection loss function, using the intersection-over-union ratio of the predicted box and the real box, that is, , represents the target detection prediction box, Represents a label box; represents the first-layer joint loss function of image and text, represents the second-layer joint loss function of image and text, represents the third-layer joint loss function of image and text, 、 and Both use contrast loss function, , , ; in, is the number of samples in a batch, is the number of negative samples in a batch, where positive samples represent matching image and text pairs, and negative samples represent unmatched image and text pairs; and represents the first layer image feature vectors, Indicates the first layer of positive samples text feature vectors, Represents the first layer of negative samples text feature vectors; and represents the second layer image feature vectors, Indicates the first positive sample in the second layer text feature vectors, Represents the negative sample in the second layer text feature vectors; and represents the third layer image feature vectors, Indicates the first positive sample in the third layer text feature vectors, Represents the third layer of negative samples text feature vectors, represents the cosine similarity metric, Represents the temperature parameter used to adjust the similarity score distribution.
[0009] Furthermore, the contrast loss function is minimized based on gradient descent backpropagation, and the supervised learning of the deep neural network detection model is performed through the loss function to complete the training of the foreign object detection model.
[0010] As a further technical limitation, a foreign object detection model is used to perform target detection on the obtained mine car track image and text fusion features. The position and size of foreign objects are marked in the real-time image of the mine car track through the target detection frame to complete the real-time monitoring of foreign objects on the mine car track.
[0011] As a further technical limitation, the hierarchical multimodal features are extracted and fused based on the convolutional layer, and the text prompt words are mapped and matched to the same dimension in combination with the linear layer. The feature correlation between the hierarchical multimodal features and the text prompt words is extracted through the cross-self-attention mechanism and the self-attention mechanism, and the saliency features are fused. The hierarchical multimodal features and the features of the text prompt words are extracted by the bypass additional splicing operation to form a residual neural network. The image and text fusion features of the mine car track are obtained by adding the backbone network and the bypass residual network.
[0012] According to some embodiments, a second solution of the present invention provides a mine car track foreign object detection system based on a deep neural network, which adopts the following technical solutions: A mine car track foreign object detection system based on deep neural network, including: An acquisition module configured to acquire a real-time image and text prompt words of the mine car track; an extraction module configured to extract hierarchical multimodal features of the acquired real-time image; A fusion module is configured to use the hierarchical multimodal features obtained by fusing the multimodal deep neural network and the text prompt words to obtain the image-text fusion features of the mine car track; The detection module is configured to detect whether there are foreign objects on the mine car track based on the obtained image and text fusion features of the mine car track and a preset foreign object detection model; when foreign objects are present, the foreign object size is identified and the location of the foreign object is determined, thereby completing real-time monitoring of foreign objects on the mine car track.
[0013] According to some embodiments, a third solution of the present invention provides a computer-readable storage medium, which adopts the following technical solution: A computer-readable storage medium stores a program, which, when executed by a processor, implements the steps of the method for detecting foreign objects on a mine car track based on a deep neural network as described in the first embodiment of the present invention.
[0014] According to some embodiments, a fourth solution of the present invention provides an electronic device, which adopts the following technical solution: An electronic device includes a memory, a processor, and a program stored in the memory and running on the processor. When the processor executes the program, it implements the steps of the mine car track foreign body detection method based on deep neural network as described in the first embodiment of the present invention.
[0015] According to some embodiments, a fifth solution of the present invention provides a computer program product, which adopts the following technical solution: A computer program product includes software code, wherein the program in the software code executes the steps of the mine car track foreign object detection method based on deep neural network as described in the first embodiment of the present invention.
[0016] Compared with the prior art, the present invention has the following beneficial effects: Based on multimodal perception and target detection, the present invention monitors and identifies in real time whether there are foreign objects on the mine car track and determines the size of the foreign objects and the position of the foreign objects on the mine car track, thereby realizing real-time high-precision target detection and risk warning of foreign objects on the mine car track. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings constituting a part of the specification of this embodiment are used to provide a further understanding of this embodiment. The schematic embodiments and descriptions of this embodiment are used to explain this embodiment and do not constitute an improper limitation on this embodiment.
[0018] Figure 1 This is a flow chart of the method for detecting foreign objects on mine car tracks based on a deep neural network in Example 1 of the present invention; Figure 2 Schematic diagram of the hardware structure of the method for detecting foreign objects on mine car tracks based on a deep neural network in the first embodiment of the present invention; Figure 3 This is a schematic diagram of a method for detecting foreign matter on a mine car track based on a deep neural network in the first embodiment of the present invention; Figure 4 This is a flowchart of the image-text fusion in the first embodiment of the present invention; Figure 5 Detailed steps of the method for detecting foreign objects on mine car tracks based on a deep neural network in Example 1 of the present invention; Figure 6 This is a structural block diagram of a mine car track foreign object detection system based on a deep neural network in Example 2 of the present invention; Among them: 1. Mine car track; 2. Track rolling stones; 3. Visible light camera; 4. Infrared camera; 5. Computer; 6. Communication equipment. DETAILED DESCRIPTION
[0019] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0020] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present invention belongs.
[0021] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.
[0022] In the present invention, terms such as "upper", "lower", "left", "right", "front", "back", "vertical", "horizontal", "side", "bottom", etc. indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. They are relational words determined only for the convenience of describing the structural relationships of the various parts or elements of the present invention, and do not specifically refer to any part or element in the present invention, and should not be understood as limiting the present invention.
[0023] In the present invention, terms such as "fixed connection," "connected," and "connection" should be interpreted broadly to mean a fixed connection, an integral connection, or a detachable connection; a direct connection or an indirect connection through an intermediary. Relevant researchers or technicians in this field may determine the specific meanings of these terms in the present invention based on specific circumstances, and they should not be construed as limitations of the present invention.
[0024] In the absence of conflict, the embodiments of the present invention and the features thereof may be combined with each other.
[0025] Example 1 Embodiment 1 of the present invention introduces a method for detecting foreign objects on mine car tracks based on a deep neural network.
[0026] like Figure 1 The method for detecting foreign matter on a mine car track based on a deep neural network comprises: Get real-time images and text prompts of the minecart track; extracting hierarchical multimodal features of the acquired real-time image; The hierarchical multimodal features and text prompt words obtained by fusing multimodal deep neural networks are used to obtain the image-text fusion features of the mine car track; Detect whether there are foreign objects on the mine car track based on the obtained image-text fusion features and the preset foreign object detection model; When there is foreign matter, the size of the foreign matter is identified and its location is determined, completing real-time monitoring of foreign matter on the mine car track.
[0027] This embodiment takes track rolling stone 2 as an example to detect foreign objects on the mine car track 1. Figure 2 In this embodiment, the track rolling stone detection data is collected and detected by using a visible light camera 3, an infrared light camera 4, a computer 5 and a communication device 6; specifically: Visible light camera 3 deploys multiple wide-angle high-definition cameras (resolution 4K @30fps), covering key track areas (slopes, curves, and tunnel entrances) and capturing RGB video streams in real time. Infrared camera 4 deploys multiple wide-angle infrared cameras (resolution 1280 x 1024 @30fps), covering key track areas (slopes, curves, and tunnel entrances) and capturing infrared video streams in real time. Computer 5 is connected to visible light camera 3 and infrared camera 4 via a wired connection, acquires the captured video streams online, and runs the proposed multimodal deep neural network fusion model to implement track rockfall target detection. It is used to calculate and process the deep neural network computing framework and accelerate the deep neural network model calculation to meet the requirements of online track rockfall target detection. Communication device 6: When a rockfall is detected, the communication module sends the rockfall detection result's location and size information to the central server via the wireless communication module, completing the track rockfall warning.
[0028] This embodiment adopts Figure 2 The hardware architecture of the deep neural network-based mine car track foreign object detection method shown here collects 500 sets of different rockfall data samples, covering different lighting and location scenarios, and samples with different rockfall characteristics such as size, material (rock, ore), quantity, and location. Manual object annotation is performed using the LabelImg tool to form a rockfall detection training dataset for subsequent deep neural network model training. Simultaneously, manually annotated foreign objects similar to rockfall (such as dead wood and metal fragments) in the samples are used to improve the recognition rate of rockfall and effectively distinguish between rockfall and other similar foreign objects.
[0029] In the process of track rockfall detection, this embodiment adopts a multimodal hierarchical deep neural network fusion model based on target detection, which consists of text encoding, image encoding, image-text fusion and target detection. The specific detection proposed deep neural network model structure is as follows: Figure 3As shown in the figure, its input is a visible light image, an infrared light image and a text prompt word, and the output is a track rolling stone target detection frame based on the visible light image and the infrared light image. The target detection frame marks the position and size of the rolling stone appearing in the input image.
[0030] This embodiment adopts Figure 3 The deep neural network model shown includes: (1) Text encoding module The text encoding module in this embodiment is based on the Qianwen multimodal large language model, which takes as input a visible light image, an infrared image, and a text prompt. The text prompt is: "The current input is a visible light image and an infrared image. The target scene is a mine car track. Please output whether there are rolling stones or foreign objects in the scene." Leveraging the multimodal large language model's combined image and text reasoning capabilities, it generates richer and more comprehensive semantic information, improving the accuracy and generalization of object detection.
[0031] (2) Graph encoding module The image encoding module in this embodiment consists of a convolution layer and a pooling layer, which are used to extract image features from the input visible light image and infrared light image respectively, and extract convolution features layer by layer from shallow to deep layers.
[0032] (3) Graphics and text fusion module like Figure 4 As shown, this embodiment uses a cross-attention mechanism to fuse the visible light features and infrared fusion features extracted by the image encoding module with the text features extracted by the text encoder. The backbone first extracts visible light and infrared fusion image features using a convolutional layer, then uses a linear layer to map the text features to the same dimension. Cross-self-attention and self-attention mechanisms are then used to extract feature correlations between the two modalities and fuse significant features. Furthermore, an additional concatenation operation is used to extract both features together, which are then used to form a residual neural network. Finally, the backbone network and the bypass residual network are combined to produce the final fused features.
[0033] (4) Object Detection This embodiment performs target detection on the decoded fusion features, and the target detection head generates a final target detection frame; the constructed rolling stone detection training dataset is used to supervise learning through training data.
[0034] This embodiment adopts a composite loss function and uses the gradient descent back propagation method to complete the effective training of the proposed multi-modal hierarchical deep neural network fusion model. Composite loss function It consists of five loss functions, namely text loss function , target detection loss function , the first layer of image and text joint loss function , the second layer joint loss function of image and text And the third layer image and text joint loss function ; That is, the composite loss function for ,in, represents the weight of each loss function, and ; Text loss function The cross entropy loss function is used, that is, , Indicates the marked text information. Indicates that a multimodal large language model predicts text information; Object Detection Loss Function The intersection of the predicted box and the real box is used, that is, , represents the target detection prediction box, Represents a label box; The first layer joint loss function of image and text , the second layer joint loss function of image and text And the third layer image and text joint loss function Both use contrast loss function, , , ; in, is the number of samples in a batch, is the number of negative samples in a batch, where positive samples represent matching image and text pairs, and negative samples represent unmatched image and text pairs; and represents the first layer image feature vectors, Indicates the first layer of positive samples text feature vectors, Represents the first layer of negative samples text feature vectors; and represents the second layer image feature vectors, Indicates the first positive sample in the second layer text feature vectors, Represents the negative sample in the second layer text feature vectors; and represents the third layer image feature vectors, Indicates the first positive sample in the third layer text feature vectors, Represents the third layer of negative samples text feature vectors, represents the cosine similarity metric, Represents the temperature parameter used to adjust the similarity score distribution.
[0035] Based on gradient descent backpropagation to minimize the contrast loss function, supervised learning of the deep neural network detection model is performed through the loss function to complete the training of the foreign object detection model; the model can learn feature representations that make the similarity scores of matching image-text pairs higher and the similarity scores of mismatching image-text pairs lower.
[0036] After the deep neural network model training is completed in this embodiment, the model will be deployed in the track rolling stone detection data acquisition and detection system. The model will be called to perform forward propagation calculations using the online collected visible light and infrared light video streams to generate target detection results, determine the location and size of the rolling stone, and issue an abnormal rolling stone warning. The specific process is as follows: Figure 5 shown.
[0037] This embodiment establishes a visible light and infrared light data video stream acquisition system for track rolling stone detection, proposes a multimodal fusion model framework, and fully extracts potential effective features from multimodal input information through hierarchical feature fusion of text and image. It is used for effective visual detection of target objects in multimodal images, ultimately achieving real-time, online, high-precision track rolling stone target detection and risk warning. Based on multimodal perception and target detection, the presence of foreign objects on the mine car track is monitored and identified in real time, and the size and position of foreign objects on the mine car track are determined, thereby achieving real-time high-precision target detection and risk warning of foreign objects on the mine car track.
[0038] Example 2 The second embodiment of the present invention introduces a mine car track foreign object detection system based on deep neural network.
[0039] like Figure 6 The system for detecting foreign objects on a mine car track based on a deep neural network is shown, comprising: An acquisition module configured to acquire a real-time image and text prompt words of the mine car track; an extraction module configured to extract hierarchical multimodal features of the acquired real-time image; A fusion module is configured to use the hierarchical multimodal features obtained by fusing the multimodal deep neural network and the text prompt words to obtain the image-text fusion features of the mine car track; The detection module is configured to detect whether there are foreign objects on the mine car track based on the obtained image and text fusion features of the mine car track and a preset foreign object detection model; when foreign objects are present, the foreign object size is identified and the location of the foreign object is determined, thereby completing real-time monitoring of foreign objects on the mine car track.
[0040] The detailed steps are the same as those of the mine car track foreign body detection method based on deep neural network provided in Example 1, and will not be repeated here.
[0041] Example 3 A third embodiment of the present invention provides a computer-readable storage medium.
[0042] A computer-readable storage medium stores a program, which, when executed by a processor, implements the steps of the method for detecting foreign objects on a mine car track based on a deep neural network as described in the first embodiment of the present invention.
[0043] The detailed steps are the same as those of the mine car track foreign body detection method based on deep neural network provided in Example 1, and will not be repeated here.
[0044] Example 4 A fourth embodiment of the present invention provides an electronic device.
[0045] An electronic device includes a memory, a processor, and a program stored in the memory and running on the processor. When the processor executes the program, it implements the steps of the method for detecting foreign objects on mine car tracks based on a deep neural network as described in Example 1 of the present invention.
[0046] The detailed steps are the same as those of the mine car track foreign body detection method based on deep neural network provided in Example 1, and will not be repeated here.
[0047] Example 5 A fifth embodiment of the present invention provides a computer program product.
[0048] A computer program product includes software code, wherein the program in the software code executes the steps of the method for detecting foreign objects on mine car tracks based on a deep neural network as described in Example 1 of the present invention.
[0049] The detailed steps are the same as those of the mine car track foreign body detection method based on deep neural network provided in Example 1, and will not be repeated here.
[0050] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk drives, CD-ROMs, optical storage devices, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention may be implemented using various computer languages, such as the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0051] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0052] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0053] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0054] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0055] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
[0056] The above description is merely a preferred embodiment of this embodiment and is not intended to limit this embodiment. Those skilled in the art will readily appreciate that this embodiment may be modified and varied in various ways. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of this embodiment shall be within the scope of protection of this embodiment.
Claims
1. A mine car track foreign body detection method based on deep neural network, characterized in that: include: Get real-time images and text prompts of the minecart track; extracting hierarchical multimodal features of the acquired real-time image; The hierarchical multimodal features and text prompt words obtained by fusing multimodal deep neural networks are used to obtain the image-text fusion features of the mine car track; Detect whether there are foreign objects on the mine car track based on the obtained image-text fusion features and the preset foreign object detection model; When there is foreign matter, the size of the foreign matter is identified and its location is determined, completing real-time monitoring of foreign matter on the mine car track.
2. A method for detecting foreign matter on a mine car track based on a deep neural network as claimed in claim 1, characterized in that: The preset foreign body detection model adopts a deep neural network detection model, and the deep neural network detection model adopts a composite loss function and gradient descent back propagation for model training.
3. A method for detecting foreign matter on mine car tracks based on a deep neural network as described in claim 2, characterized in that: The composite loss function for ,in, represents the weight of each loss function, and ; Represents different loss functions, where Represents the text loss function, using the cross entropy loss function, that is , Indicates the marked text information. Indicates that a multimodal large language model predicts text information; Represents the target detection loss function, using the intersection-over-union ratio of the predicted box and the real box, that is, , represents the target detection prediction box, Represents a label box; represents the first-layer joint loss function of image and text, represents the second-layer joint loss function of image and text, represents the third-layer joint loss function of image and text, 、 and Both use contrast loss function, , , ; in, is the number of samples in a batch, is the number of negative samples in a batch, where positive samples represent matching image and text pairs, and negative samples represent unmatched image and text pairs; and represents the first layer image feature vectors, Indicates the first layer of positive samples text feature vectors, Represents the first layer of negative samples text feature vectors; and represents the second layer image feature vectors, Indicates the first positive sample in the second layer text feature vectors, Represents the negative sample in the second layer text feature vectors; and represents the third layer image feature vectors, Indicates the first positive sample in the third layer text feature vectors, Represents the third layer of negative samples text feature vectors, represents the cosine similarity metric, Represents the temperature parameter used to adjust the similarity score distribution.
4. A method for detecting foreign matter on mine car tracks based on a deep neural network as claimed in claim 3, characterized in that: Based on gradient descent back propagation to minimize the contrast loss function, the loss function is used to perform supervised learning of the deep neural network detection model to complete the training of the foreign object detection model.
5. A method for detecting foreign matter on mine car tracks based on a deep neural network as claimed in claim 1, characterized in that: The foreign object detection model is used to perform target detection on the obtained mine car track image and text fusion features. The position and size of foreign objects are marked in the real-time image of the mine car track through the target detection frame, completing the real-time monitoring of foreign objects on the mine car track.
6. A method for detecting foreign matter on a mine car track based on a deep neural network as claimed in claim 1, characterized in that: The convolutional layer extracts and fuses hierarchical multimodal features, and combines the linear layer to map the text prompt words to the same dimension. The cross-self-attention mechanism and the self-attention mechanism are used to extract the feature correlation between the hierarchical multimodal features and the text prompt words, and fuse the saliency features. The hierarchical multimodal features and text prompt word features are extracted by bypassing the additional splicing operation contract, which are used to construct a residual neural network. The image-text fusion features of the mine car track are obtained by adding the backbone network and the bypass residual network.
7. A mine car track foreign body detection system based on deep neural network, characterized in that: include: An acquisition module configured to acquire a real-time image and text prompt words of the mine car track; an extraction module configured to extract hierarchical multimodal features of the acquired real-time image; A fusion module is configured to use the hierarchical multimodal features obtained by fusing the multimodal deep neural network and the text prompt words to obtain the image-text fusion features of the mine car track; The detection module is configured to detect whether there are foreign objects on the mine car track based on the obtained image and text fusion features of the mine car track and a preset foreign object detection model; when foreign objects are present, the foreign object size is identified and the location of the foreign object is determined, thereby completing real-time monitoring of foreign objects on the mine car track.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the mine car track foreign body detection method based on deep neural network as described in any one of claims 1 to 6 are implemented.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the steps of the mine car track foreign object detection method based on deep neural network as described in any one of claims 1 to 6 are implemented.
10. A computer program product comprising software code, characterized in that The program in the software code executes the steps of the mine car track foreign body detection method based on deep neural network as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Method and system for realizing anomaly identification of mine data based on artificial intelligence
CN117953313A
Belt tearing detection method and system based on discrete state selectable space model
CN118333979A
Target detection method and device based on multi-mode prompt, electronic equipment and medium
CN118628846A
Tippler track foreign matter detection and removal system and method based on time-space separation
CN118823734A
Open domain image recognition method based on combinable text prompt framework
CN119046722A