Floating object identification method and device, computer program product and electronic equipment
By combining multimodal input of railway images and meteorological data, the target model is trained, and the problem of poor accuracy of floating objects recognition is solved, and efficient identification is achieved when the data volume is insufficient to ensure the safety of railway facilities.
Patent Information
- Application Number
- CN202510807696.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-06-17
AI Technical Summary
In the prior art, the accuracy of floating objects recognition is poor, especially when the amount of training data is insufficient, which affects the safety of train operation.
By combining railway images and meteorological data, the target model is trained using multi-modal input, including convolutional neural networks and encoding codecs, fusing meteorological word vectors and feature maps, optimizing fusion parameters, and improving recognition accuracy.
Even when the amount of training data is insufficient, the accuracy of floating objects can be ensured, false alarms and missed reports can be reduced, and the safe operation of railway facilities can be ensured.
Smart Images

Figure CN120318504A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of artificial intelligence and big data, and in particular, to a floating object recognition method, device, computer program product, and electronic device. Background Technique
[0002] With the development of the times, the area of railway laying has also increased rapidly, and the safe operation of trains is the key concern. The catenary provides stable power supply for trains. The shear, tensile, friction and other forces of the catenary cause damage to the catenary, which will threaten the safe operation of trains. Floating objects are the main cause of catenary damage. Under the action of wind, if not removed in time, they will have a great impact on the operating state of the catenary. At the lightest, it will cause catenary faults and affect catenary power supply. At the heaviest, it will cause catenary disconnection, affect the safe operation of trains, and even cause train derailment and overturning.
[0003] In the related art, a single-modal railway image is used as the input of the artificial intelligence algorithm. When performing target type recognition, it is easily affected by the training process. If the amount of training data is insufficient, it will affect the accuracy of the algorithm, resulting in poor accuracy of floating object recognition.
[0004] Aiming at the problem of poor accuracy of floating object recognition in the related art, no effective solution has been proposed yet. Summary of the Invention
[0005] The main purpose of the present application is to provide a floating object recognition method, device, computer program product, and electronic device to solve the problem of poor accuracy of floating object recognition in the related art.
[0006] To achieve the above object, according to one aspect of the present application, a floating object recognition method is provided. The method includes: obtaining target railway images by acquiring railway images collected by a monitoring device at each acquisition moment within a target time period; determining meteorological data at each acquisition moment to obtain target meteorological data corresponding to each frame of the target railway image; inputting the target railway images and the target meteorological data into a target model to obtain a floating object recognition result at each acquisition moment, where the target model is trained by multiple groups of training samples, and each group of training samples includes historical railway images, historical meteorological data corresponding to the historical railway images, and historical floating object recognition results.
[0007] Optionally, the target model is obtained in the following manner: determine an initial convolutional processing layer, an initial encoder, and an initial decoder; obtain a first training sample set, and train the initial convolutional processing layer based on the first training sample set to obtain a trained convolutional processing layer, where the first training sample set includes historical railway images and historical floating object recognition results; obtain a second training sample set, and train the initial encoder and the initial decoder based on the second training sample set to obtain a trained encoder and a trained decoder, where the second training sample set includes historical superimposed feature maps, historical encoded features, and historical floating object recognition results; obtain multiple groups of training samples, and jointly train the trained convolutional processing layer, the trained encoder, and the trained decoder based on the multiple groups of training samples to obtain the target model.
[0008] Optionally, inputting the target railway image and the target meteorological data into the target model to obtain the floating object recognition result at each acquisition moment includes: extracting the three primary color pixel values of the target railway image to obtain a three-primary color pixel matrix; inputting the three-primary color pixel matrix into the convolutional processing layer of the target model for processing to obtain multiple feature maps; extracting keywords from the target meteorological data to obtain meteorological keywords, and converting the meteorological keywords into meteorological word vectors through a word vector conversion tool, where the dimension of the meteorological word vector is the same as that of the feature map; for each feature map, superimpose the meteorological word vector and the feature map through the fusion parameter of the target model to obtain a superimposed feature map; encode the superimposed feature map through the encoder of the target model to obtain an encoded feature, and decode the encoded feature through the decoder of the target model to obtain the floating object recognition result.
[0009] Optionally, the fusion parameter includes a meteorological word vector coefficient and a feature map coefficient. Superimposing the meteorological word vector and the feature map through the fusion parameter of the target model to obtain a superimposed feature map includes: determining each first element in the meteorological word vector, multiplying the meteorological word vector coefficient by each first element to obtain an intermediate word vector; determining each second element in the feature map, multiplying the feature map coefficient by each second element to obtain an intermediate feature map; summing the intermediate word vector and the intermediate feature map to obtain the superimposed feature map.
[0010] Optionally, obtaining the target railway image by acquiring the railway images collected by the monitoring device at each acquisition moment within the target time period includes: receiving the railway images collected by the monitoring device at the acquisition moment; comparing the railway images with a preset background image to obtain a differential pixel region, where the background image is an image without floating objects collected by the monitoring device at the same shooting angle as the railway images; determining the image corresponding to the differential pixel region as the target railway image.
[0011] Optionally, after obtaining the floating object recognition result at each acquisition moment, the method further includes: when the floating object recognition result indicates that there is a floating object in the railway image, determining the railway image as the initial image where the floating object appears; determining multiple frames of railway images within a preset time period collected after the initial image, and determining the movement trajectory of the floating object based on the multiple frames of railway images, and determining the predicted movement trajectory of the floating object within the target time period based on the movement trajectory of the floating object; determining the position information of the railway power grid coverage area, and judging whether the floating object falls into the railway power grid coverage area based on the position information and the predicted movement trajectory; when it is determined that the floating object will fall into the railway power grid coverage area, controlling the air blowing device to adjust the movement trajectory of the floating object.
[0012] Optionally, controlling the air blowing device to adjust the movement trajectory of the floating object includes: determining the position information of the air blowing device, and determining the starting time and air blowing parameters of the air blowing device based on the position information of the air blowing device and the predicted movement trajectory, where the air blowing parameters include at least one of the following: wind force level and wind direction; controlling the air blowing device to operate according to the air blowing parameters at the starting time to adjust the movement trajectory of the floating object.
[0013] To achieve the above object, according to another aspect of the present application, there is provided a floating object recognition device. The device includes: an acquisition unit, configured to acquire railway images collected by a monitoring device at each acquisition moment within a target time period to obtain target railway images; a first determination unit, configured to determine meteorological data at each acquisition moment to obtain target meteorological data corresponding to each frame of the target railway image; an input unit, configured to input the target railway images and the target meteorological data into a target model to obtain a floating object recognition result at each acquisition moment, where the target model is obtained by training with multiple groups of training samples, and each group of training samples includes historical railway images, historical meteorological data corresponding to the historical railway images, and historical floating object recognition results.
[0014] To achieve the above object, according to another aspect of the present application, there is provided a computer program product, including a computer program, where when the computer program is executed by a processor, the steps of the floating object recognition method described in various embodiments of the present application are implemented.
[0015] Through this application, the following steps are adopted: obtaining railway images collected by a monitoring device at each acquisition moment within a target time period to obtain target railway images; determining meteorological data at each acquisition moment to obtain target meteorological data corresponding to each frame of the target railway image; inputting the target railway image and the target meteorological data into a target model to obtain a floating object recognition result at each acquisition moment, where the target model is trained by multiple groups of training samples, and each group of training samples includes historical railway images, historical meteorological data corresponding to the historical railway images, and historical floating object recognition results, which solves the problem of poor accuracy in floating object recognition in the related art. Since floating objects are greatly affected by weather, by combining railway images and meteorological data, multi-modal data is input into the target model for floating object recognition. Even when the amount of training data is insufficient, the recognition accuracy can still be guaranteed, thereby achieving the effect of improving the accuracy of floating object recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0017] Figure 1 is a flowchart of a floating object recognition method provided according to an embodiment of this application;
[0018] Figure 2 is a flowchart of a training method of a target model provided according to an embodiment of this application;
[0019] Figure 3 is a schematic diagram of an optional floating object recognition method provided according to an embodiment of this application;
[0020] Figure 4 is a schematic structural diagram of a recognition algorithm adopted by a floating object recognition device provided according to an embodiment of this application;
[0021] Figure 5 is a schematic diagram of a floating object recognition device provided according to an embodiment of this application;
[0022] Figure 6 is a schematic diagram of an electronic device provided according to an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.
[0024] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of this application.
[0025] It should be noted that the terms "first", "second", etc. in the description and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to describe the embodiments of this application here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0026] The present invention will be described below in conjunction with preferred implementation steps. Figure 1 is a flowchart of a floating object recognition method provided according to an embodiment of this application. As Figure 1 shown, the method includes the following steps:
[0027] Step S101: Obtain railway images collected by a monitoring device at each acquisition moment within a target time period to obtain target railway images.
[0028] In step S101, a monitoring device is installed near the railway. The monitoring device can be a camera, which is communicatively connected to the floating object recognition device. The camera captures railway images and sends each frame of railway image and the acquisition moment of each frame of railway image to the floating object recognition device. The floating object recognition device saves each frame of railway image and the acquisition moment of each frame of railway image.
[0029] Step S102: Determine the meteorological data at each acquisition moment to obtain the target meteorological data corresponding to each frame of the target railway image.
[0030] In step S102, the floating object recognition device is also communicatively connected to the meteorological system. The floating object recognition device receives the meteorological data sent by the meteorological system, counts the acquisition times of multiple frames of railway images, generates a target time period, encodes the target time period using a preset communication protocol to generate a data request, and sends the data request to the meteorological system. After receiving the data request, the meteorological system decodes the data request to obtain the target time period, retrieves the meteorological data of the target time period from local, and sends the meteorological data to the floating object recognition device, so as to obtain the target meteorological data corresponding to each frame of the target railway image.
[0031] Step S103: Input the target railway image and the target meteorological data into the target model to obtain the floating object recognition result at each acquisition time. The target model is obtained by training with multiple groups of training samples. Each group of training samples includes a historical railway image, the historical meteorological data corresponding to the historical railway image, and the historical floating object recognition result.
[0032] In step S103, the floating object recognition result may include the probability that the target railway image contains a floating object, and the probability of which type of floating object the floating object belongs to. The floating object types may include: leaves, plastic films, branches, straws, weeds, etc. The target model may be a convolutional neural network model. The convolutional neural network model is trained by collecting multiple groups of training samples. Each group of training samples may include: a historical railway image, which may or may not contain a floating object. Historical meteorological data, the weather information corresponding to the acquisition time of the historical railway image, including but not limited to wind speed, wind direction, temperature, etc. Historical floating object recognition result, a label marked by an expert or a previous model on whether there is a floating object in the image, which can be binary classification (yes or no), or multi-class classification (for example, identifying the specific type of the floating object). The convolutional neural network model is trained with multiple groups of training samples to obtain the trained target model. The target railway image and the target meteorological data are input into the target model, and the target model outputs the floating object recognition result of the target railway image.
[0033] The floating object recognition method provided by the embodiments of the present application obtains target railway images by acquiring railway images collected by a monitoring device at each acquisition moment within a target time period; determines meteorological data at each acquisition moment to obtain target meteorological data corresponding to each frame of the target railway image; and inputs the target railway image and the target meteorological data into a target model to obtain a floating object recognition result at each acquisition moment. The target model is obtained by training with multiple groups of training samples, and each group of training samples includes historical railway images, historical meteorological data corresponding to the historical railway images, and historical floating object recognition results, thereby solving the problem of poor accuracy in floating object recognition in the related art. By combining railway images and meteorological data, multi-modal data is input into the target model for floating object recognition. Even when the amount of training data is insufficient, the recognition accuracy can still be guaranteed, thus achieving the effect of improving the accuracy of floating object recognition.
[0034] To improve the accuracy of floating object recognition, a target model needs to be trained. Optionally, in the floating object recognition method provided by the embodiments of the present application, the target model is obtained in the following manner: determine an initial convolutional processing layer, an initial encoder, and an initial decoder; obtain a first training sample set and train the initial convolutional processing layer based on the first training sample set to obtain a trained convolutional processing layer, where the first training sample set includes historical railway images and historical floating object recognition results; obtain a second training sample set and train the initial encoder and the initial decoder based on the second training sample set to obtain a trained encoder and a trained decoder, where the second training sample set includes historical superimposed feature maps, historical encoded features, and historical floating object recognition results; obtain multiple groups of training samples and jointly train the trained convolutional processing layer, the trained encoder, and the trained decoder based on the multiple groups of training samples to obtain the target model.
[0035] In some embodiments, the target model may include multiple layers of convolutional processing layers, an encoder, and a decoder. Figure 2 It is a flowchart of the training method of the target model provided by the embodiments of the present application. As Figure 2 shown, the training method includes:
[0036] First, obtain a first training sample set and use the first training sample set to train the convolutional processing layer to optimize the parameters in the convolutional processing layer and obtain a trained convolutional processing layer.
[0037] Specifically, a convolutional neural network is constructed based on the convolutional processing layer by adding fully connected layers. Obtain the first training sample set, which includes historical railway images and historical floating object recognition results. Use the first training sample set to train the convolutional neural network, optimize the parameters of the convolutional processing layer and the fully connected layers, and obtain the trained convolutional neural network. Obtain the number of layers of the convolutional processing layer and the parameters of each layer of the convolutional processing layer from the trained convolutional neural network, so as to obtain the trained multi-layer convolutional processing layer.
[0038] By adding fully connected layers to the convolutional processing layer, a complete convolutional neural network is constructed and trained, which enables the convolutional processing layer to learn an effective representation of image features in an independent environment. This independent training method allows the convolutional processing layer to focus on image feature extraction without being affected by other components (such as encoders and decoders), thereby improving the accuracy and efficiency of feature extraction.
[0039] Then, obtain the second training sample set, and use the second training sample set to train the encoder and the decoder, optimize the parameters in the encoder and the decoder, and obtain the trained encoder and the trained decoder.
[0040] Specifically, use the trained multi-layer convolutional processing layer as the input layer of the encoder. Use the image input layer to process the historical railway image to obtain the first training input feature, input the first training input feature into the trained multi-layer convolutional processing layer, and output multiple training feature maps. Concatenate the multiple training feature maps to generate the second training input feature, that is, the historical superimposed feature map. Use the multi-level encoder to encode the historical superimposed feature map to obtain the historical encoded feature, use the multi-level decoder to decode the historical encoded feature, and output the probability that the historical railway image contains floating objects. Calculate the loss function according to the probability that the historical railway image contains floating objects and the historical floating object recognition result, and use the loss function to optimize the parameters in the encoder and the decoder. By continuously optimizing the parameters in the encoder and the decoder, obtain the trained encoder and the trained decoder.
[0041] After obtaining the trained convolutional processing layer, directly use it as the input layer of the encoder for subsequent training, without the need for additional training of the input layer of the encoder. This effectively avoids repeated training, reduces the training time, and improves the overall training efficiency. Directly using the trained convolutional processing layer as the input layer of the encoder ensures the continuity of feature representation during training. This continuity helps the encoder better understand and utilize the features extracted by the convolutional processing layer, thereby improving the recognition performance of the entire model. In addition, since the interface between the convolutional processing layer and the encoder remains consistent, it also reduces the complexity of the model during training and deployment.
[0042] Finally, the trained convolutional layer, the trained encoder, and the trained decoder are jointly trained with multiple sets of training samples, and the fusion parameters are optimized according to the loss function, so as to obtain the trained target model.
[0043] Specifically, the image input layer is used to process the historical railway images in each set of training samples to obtain the first training input features, and the first training input features are input into the trained multi-layer convolutional processing layer to output multiple training feature maps. Keyword extraction is performed on the historical meteorological data to obtain training meteorological keywords, and the word vector conversion tool is used to convert the training meteorological keywords into training meteorological word vectors, so that the dimension of the training meteorological word vectors is the same as that of the training feature maps. For each training feature map, the training meteorological word vectors and the training feature maps are superimposed using the fusion parameters to obtain multiple superimposed training feature maps, and the multiple superimposed training feature maps are concatenated to obtain the second training input features. The trained multi-level encoder is used to encode the second training input features to obtain the encoded features, and the trained multi-level decoder is used to decode the encoded features to output the probability that the historical railway image contains floating objects. The loss function is calculated according to the probability that the historical railway image contains floating objects and the historical floating object recognition results, and the fusion parameters are optimized using the loss function. The trained target model is obtained through continuous optimization.
[0044] In the joint training stage, since the convolutional processing layer has been fully trained, the encoder and the decoder can adapt faster and learn how to work with the convolutional processing layer. This collaborative work not only improves the recognition accuracy of the model, but also accelerates the convergence speed of the training process, further improving the training efficiency.
[0045] In this embodiment, by first training the convolutional processing layer, the encoder, and the decoder separately, and then jointly training the trained convolutional layer, the trained encoder, and the trained decoder to optimize the fusion parameters, compared with training the parameters of the convolutional processing layer, the parameters of the encoder, the parameters of the decoder, and the fusion parameters simultaneously, the step-by-step training method can reduce the number of parameters to be trained simultaneously, reduce the model training time, simplify the training process, and reduce the complexity of model training. Since the step-by-step training strategy decomposes the training process into multiple stages, and each stage only focuses on the optimization of part of the parameters, the training objective of each stage is more clear and the training process is more efficient. In the preliminary training stage, by separately training the convolutional processing layer, the encoder, and the decoder, these components can have better initial performance before joint training, thus accelerating the convergence speed of the subsequent joint training stage and improving the overall training efficiency.
[0046] The step-by-step training strategy first trains each component separately, enabling each component to learn certain feature representation capabilities during individual training. During the joint training phase, these components can work better together to jointly improve the generalization ability of the model. By optimizing the fusion parameters, the meteorological word vectors and feature maps can be better fused during the superposition process, thereby enhancing the accuracy of floating object recognition. When facing different meteorological conditions and railway images, the model can more accurately identify floating objects, reducing false alarms and missed detections. The step-by-step training strategy makes the model training process more modular and flexible. When new components need to be added or existing components need to be improved, adjustments and optimizations can be made more conveniently. At the same time, since each component has been individually trained and tested, the maintainability of the model is enhanced, and potential problems can be more easily located and solved.
[0047] After training the target model, the target railway image and target meteorological data are input into the target model to obtain the floating object recognition result. Optionally, in the floating object recognition method provided in the embodiments of the present application, inputting the target railway image and target meteorological data into the target model to obtain the floating object recognition result at each acquisition moment includes: extracting the three primary color pixel values of the target railway image to obtain a three primary color pixel matrix; inputting the three primary color pixel matrix into the convolutional processing layer of the target model for processing to obtain multiple feature maps; extracting keywords from the target meteorological data to obtain meteorological keywords, and converting the meteorological keywords into meteorological word vectors through a word vector conversion tool, where the dimension of the meteorological word vector is the same as that of the feature map; for each feature map, superimposing the meteorological word vector and the feature map through the fusion parameters of the target model to obtain a superimposed feature map; encoding the superimposed feature map through the encoder of the target model to obtain an encoded feature, and decoding the encoded feature through the decoder of the target model to obtain the floating object recognition result.
[0048] In some embodiments, by extracting the three primary color pixel values of the target railway image, a three primary color pixel matrix is obtained. The three primary color pixel matrix is used as the first input feature. The three primary color pixel matrix is input into the convolutional processing layer of the target model for processing. Each convolutional processing layer includes a convolutional layer and a pooling layer connected to the output end of the convolutional layer. The convolutional layer performs convolutional processing on the input three primary color pixel matrix, and the pooling layer performs downsampling on the data after convolutional processing and then outputs a feature map. Each convolutional processing layer processes the feature map output by the previous convolutional processing layer and outputs a feature map. Each convolutional processing layer has multiple convolutional kernels, and each convolutional kernel processes the feature map output by the previous convolutional processing layer, so that multiple features can be output. If the dimensions of the feature maps are different, the multiple feature maps can be unified to the same dimension through linear transformation.
[0049] In the image processing stage, the embodiments of the present application adopt a multi-layer convolutional processing layer to process railway images and output multiple feature maps. The convolutional layer extracts local features in the image through convolutional operations, such as edges, corners, textures, etc. The introduction of the multi-layer convolutional processing layer enables the floating object recognition device to gradually extract richer and more abstract feature representations, further enhancing the feature extraction ability. It makes the device more sensitive and specific to the features of floating objects, thus improving the recognition accuracy.
[0050] The floating object recognition device uses a preset keyword extraction method to extract keywords from the target meteorological data and obtains meteorological keywords. Then, it uses a word vector conversion tool to perform conversion processing on the meteorological keywords and outputs meteorological word vectors. When setting the parameters of the word vector conversion tool, it can be set based on the size of the feature map to ensure that the dimension of the output meteorological word vector is the same as that of the feature map. The meteorological word vector and the feature map are matrices with the same dimension. After processing the meteorological word vector and each feature map using the fusion parameter, the processed feature map and the processed meteorological word vector are added matrix by matrix to obtain the superimposed feature map. Subsequently, the superimposed feature maps are concatenated to obtain the second input feature.
[0051] The embodiments of the present application obtain meteorological word vectors by performing keyword extraction and word vector conversion on meteorological data and fuse them with feature maps. The introduction of meteorological word vectors enables the floating object recognition device to consider the influence of meteorological factors on floating object recognition, thus improving the recognition robustness. Especially under complex and changeable meteorological conditions, such as strong winds, rain, snow and other weather, the fusion of meteorological word vectors and feature maps can significantly improve the recognition ability of the floating object recognition device for floating objects and reduce the situations of false alarms and missed detections.
[0052] The second input feature is input into the multi-level encoder of the target model, and the encoder encodes the input data and then outputs it. Among them, the first-level encoder encodes the second input feature, the second-level encoder encodes the encoded data output by the first-level encoder, the third-level encoder encodes the encoded data output by the second-level encoder, and so on. Each encoder includes a first multi-head attention layer, a first residual normalization layer, a first fully connected layer, and a second residual normalization layer. The input end and the output end of the first multi-head attention layer are both connected to the first residual normalization layer, the output of the first residual normalization layer is connected to the input end of the first fully connected layer, and the output end and the input end of the first fully connected layer are both connected to the second residual normalization layer.
[0053] For the first - layer encoder, the first multi - head attention layer is used to process the second input features to obtain the first multi - head attention matrix. The first residual normalization layer includes a residual layer and a normalization layer. The first residual normalization layer is used to process the second input features and the first multi - head attention matrix and output the first residual normalization matrix. The input end of the first fully - connected layer receives the first residual normalization matrix, and the first fully - connected layer processes the first residual normalization matrix and outputs the first fully - connected layer matrix. The input end of the second residual normalization layer receives the first residual normalization matrix and the first fully - connected matrix, and processes the first residual normalization matrix and the first fully - connected matrix to output the first encoded matrix.
[0054] For other encoders, the first multi - head attention layer is used to process the encoded matrix output by the previous encoder to obtain the second multi - head attention matrix. The first residual normalization layer is used to process the encoded matrix output by the previous encoder and the second multi - head attention matrix and output the second residual normalization matrix. The first fully - connected layer receives the second residual normalization matrix, processes the second residual normalization matrix and outputs the second fully - connected matrix. The second residual normalization layer receives the second residual normalization matrix and the second fully - connected matrix, and processes the second residual normalization matrix and the second fully - connected matrix to output an encoded matrix.
[0055] Since the recognition of railway floating objects highly depends on meteorological conditions (such as wind speed, wind direction, etc.), through the multi - head attention mechanism, the encoder can more effectively integrate meteorological data and image features, thereby improving the recognition accuracy. This feature - fusion ability enables the model to make more accurate judgments when facing complex and changeable railway environments. The encoder includes a first multi - head attention layer, a first residual normalization layer, a first fully - connected layer, and a second residual normalization layer. Since meteorological data is added to the second input features, the first multi - head attention layer can focus on the impact of meteorological data on each feature map. Subsequently, the first fully - connected layer is used for processing to achieve the fusion of each feature map, taking into account both local and global features, and information integrity can be retained. During the encoding process, the introduction of the first residual normalization layer and the second residual normalization layer effectively avoids information loss. The residual normalization layer, by adding residual connections, enables the output of each layer to contain the information of the previous layer, which helps the model to still retain the important features of the original input during deep encoding. Retaining this information integrity can improve the recognition accuracy of floating objects. The use of the residual normalization layer not only helps to retain information but also can improve the training efficiency of the model. By introducing residual connections, the model is more likely to find the optimal solution during training, thus accelerating convergence, significantly reducing the training time, and lowering the computational cost.
[0056] Each decoder decodes the encoded data output by the last encoder and the data output by the previous decoder, and finally outputs the recognition result of the floating object. Among them, the first-level decoder decodes the encoded data output by the last encoder, and the second-level decoder decodes the decoded data output by the first-level decoder according to the encoded data output by the last encoder, and so on. Until the last decoder decodes the decoded data output by the penultimate decoder according to the encoded data output by the last encoder. In addition, the decoder also needs to connect the Softmax (activation) function, so as to output the probability of whether the target railway image contains floating objects.
[0057] For example, each decoder includes a second multi-head attention layer, a third residual normalization layer, a second fully-connected layer, and a fourth residual normalization layer. The second multi-head attention layer receives the last encoded matrix, and the second multi-head attention layer receives the previous decoded matrix. The output end of the second multi-head attention layer is connected to the input end of the third residual normalization layer, and the input end of the third residual normalization layer is also connected to the output end of the previous decoder. The output end of the third residual normalization layer is connected to the input end of the second fully-connected layer, and the output end and the input end of the second fully-connected layer are connected to the input end of the fourth residual normalization layer.
[0058] Use the second multi-head attention layer to receive the last encoded matrix and the decoded matrix output by the previous decoder, process the last encoded matrix and the decoded matrix output by the previous decoder, and output the second multi-head attention matrix. The third residual normalization layer receives the decoded matrix output by the previous decoder and the second multi-head attention matrix, and processes the third residual normalization matrix and the second multi-head attention matrix to output the third residual normalization matrix. If the decoder is the first-level decoder, the second multi-head attention layer receives the last encoded matrix, processes the last encoded matrix, and outputs the second multi-head attention matrix. The second fully-connected layer receives the third residual normalization matrix, processes the third residual normalization matrix, and outputs the third fully-connected matrix. The fourth residual normalization layer receives the third fully-connected matrix and the third residual normalization matrix, and processes the third fully-connected matrix and the third residual normalization matrix to output a decoded matrix.
[0059] The decoder includes a second multi-head attention layer, a third residual normalization layer, a second fully-connected layer, and a fourth residual normalization layer. The second multi-head attention layer processes the previous decoding matrix and the last encoding matrix, which can effectively preserve information during decoding and improve decoding accuracy. The second multi-head attention layer in the decoder can simultaneously consider the output decryption matrix of the previous decoder and the last encoding matrix. This design enables the decoding process to preserve information more effectively. The multi-head attention mechanism can capture the dependencies between different positions, thus enabling more accurate restoration of the original information during decoding. For railway floating object recognition, this improvement in information preservation ability helps ensure the accuracy of the decoding results and reduce the cases of misjudgment and missed judgment. The introduction of the third residual normalization layer and the fourth residual normalization layer effectively avoids information loss during decoding by adding residual connections. The residual connections enable each layer output of the decoder to contain the information of the previous layer, which helps maintain the integrity of information during decoding.
[0060] The use of the residual normalization layer not only helps with information retention but also improves the computational efficiency of the decoder. By introducing residual connections, the decoder can find the optimal solution faster during the decoding process, thus accelerating the computational process. This is particularly important for tasks of processing large-scale railway image data and meteorological data, which can significantly reduce the decoding time and improve the overall system response speed. The second fully-connected layer in the decoder is responsible for mapping the third residual matrix to the output space to achieve non-linear transformation of features. This design of the fully-connected layer helps the model capture complex non-linear relationships and improve the robustness of the model. For railway floating object recognition, the shape and position of the floating object may vary due to different meteorological conditions and environmental factors. Through the processing of the fully-connected layer, the decoder can better adapt to this variation and achieve accurate decoding.
[0061] In the feature encoding and decoding stages, the embodiments of this application adopt a multi-level encoder-decoder to process the fused features. The encoder encodes the input features through a multi-level attention mechanism to extract more advanced and abstract feature representations; the decoder decodes the encoded features through a multi-level attention mechanism and outputs the floating object recognition results. The introduction of the multi-level encoder and the multi-level decoder enables the floating object recognition device to better process complex feature representations, improving the accuracy and efficiency of recognition. At the same time, by training the encoder-decoder and optimizing its parameters and fusion parameters, the floating object recognition device can maintain stable recognition performance in different scenarios.
[0062] In the embodiments of the present application, by introducing meteorological data and performing multimodal fusion with railway images, the deficiencies of single-modal data are effectively made up. Meteorological data, such as wind speed, wind direction, temperature, etc., have a significant impact on the motion state and form of floating objects. By fusing meteorological data, the floating object recognition device can more comprehensively understand the characteristics and states of floating objects, so that even when the model training data is insufficient, a relatively high recognition accuracy can still be maintained.
[0063] In order to more accurately identify whether a railway image contains floating objects, the meteorological word vector and the feature map are superimposed through fusion parameters. Optionally, in the floating object recognition method provided by the embodiments of the present application, the fusion parameters include a meteorological word vector coefficient and a feature map coefficient. By using the fusion parameters of the target model to superimpose the meteorological word vector and the feature map, the superimposed feature map obtained includes: determining each first element in the meteorological word vector, multiplying the meteorological word vector coefficient by each first element to obtain an intermediate word vector; determining each second element in the feature map, multiplying the feature map coefficient by each second element to obtain an intermediate feature map; and summing the intermediate word vector and the intermediate feature map to obtain the superimposed feature map.
[0064] In some embodiments, the meteorological word vector and the feature map are matrices with the same dimension. The fusion parameters may include a meteorological word vector coefficient Kq and a feature map coefficient Kt. Multiply the meteorological word vector coefficient by each first element in the meteorological word vector to obtain an intermediate word vector. The meteorological word vector is a matrix Q of size m×n, and qij represents any element in the meteorological word vector. The intermediate word vector Q1 is Kq×qij, where 1≤i≤m and 1≤j≤n. i, j, m, and n are all positive integers. Multiply the feature map coefficient by each element in the feature map to obtain an intermediate feature map. The feature map is a matrix T of size m×n, and tij represents any element in the feature map. The intermediate feature map T1 is Kt×tij. Calculate the sum of the intermediate word vector and the intermediate feature map according to the following formula to obtain the superimposed feature map T2:
[0065] t2ij = Kt×tij + Kq×qij;
[0066] where t2ij is the element in the i-th row and j-th column of the superimposed feature map T2.
[0067] For example: there is one 5×4 meteorological word vector and ten 5×4 feature maps. For each 5×4 feature map, use the fusion parameters to perform superimposition processing on the 5×4 meteorological word vector and the 5×4 feature map, and ten 5×4 superimposed feature maps can be obtained. Connect the ten 5×4 superimposed feature maps to obtain the second input feature of 50×4. The second input feature can be characterized as:
[0068] ;
[0069] Among them, T20 is the first superimposed feature map matrix of 5×4, T21 is the second superimposed feature map matrix of 5×4, T28 is the ninth superimposed feature map matrix of 5×4, and T29 is the tenth superimposed feature map matrix of 5×4.
[0070] In this embodiment, the meteorological word vector coefficient is multiplied by each element in the meteorological word vector to make the meteorological word vector share a coefficient, and the information in the meteorological word vector is completely retained. Correspondingly, the feature map coefficient is multiplied by each element in the feature map to make the feature map share a coefficient, and the information in the feature map is completely retained. Then, the intermediate word vector and the intermediate feature map are summed to obtain the superimposed feature map, realizing the addition of meteorological data to each feature map. In addition, when fusing the meteorological word vector and the feature map, only two parameters are required, which can reduce the model training process, thereby reducing the complexity of model training and improving the training efficiency. By introducing the meteorological word vector coefficient and the feature map coefficient, the embodiments of the present application achieve the effective fusion of the meteorological word vector and the feature map. This fusion method not only retains the respective information of the meteorological word vector and the feature map, but also through the adjustment of the coefficient, enables the two to complement each other during the superimposition process and jointly improve the recognition accuracy. Since meteorological data has a significant impact on the shape and motion state of floating objects, integrating meteorological data into the recognition process can significantly improve the robustness and adaptability of recognition. Especially under complex and changeable meteorological conditions, such as strong winds, rain, snow and other weather, the fusion of the meteorological word vector and the feature map can enable the floating object recognition device to better cope with environmental changes, reduce false alarms and missed detections, and improve the recognition accuracy.
[0071] To improve the image processing efficiency, the differential pixel region can be extracted before the image is input into the target model. Optionally, in the floating object recognition method provided by the embodiments of the present application, obtaining the railway images collected by the monitoring device at each acquisition moment within the target time period to obtain the target railway images includes: receiving the railway images collected by the monitoring device at the acquisition moment; comparing the railway images with the preset background image to obtain the differential pixel region, where the background image is an image without floating objects collected by the monitoring device at the same shooting angle as the railway images; and determining the image corresponding to the differential pixel region as the target railway image.
[0072] In some embodiments, in the railway images captured by cameras installed near the railway, most of the objects are backgrounds. By comparing the reference image and the railway image, the objects newly entered into the image can be extracted, and the objects in the newly entered image can be recognized, reducing the amount of data processing. The background image can be a pre-collected railway image, and the background image and the railway image are images captured at the same angle, so as to facilitate the comparison between the railway image and the background image. Each pixel in the railway image is compared with the corresponding pixel in the background image. If they are the same, the next pixel is selected for comparison. If they are different, the pixel position is saved. By traversing each pixel in the railway image, the differential pixel region is obtained.
[0073] According to the position information of the differential pixel region, the pixel values of the differential pixel region are extracted from the railway image, and the pixel regions outside the differential pixel region are filled with supplementary pixel values to obtain the target railway image. For example: The railway image is 5×5 in size, and the rectangular region with the pixels in the first row and first column, the first pixel in the third row, the first pixel in the first row of the third column, and the pixel in the third row and third column as the four vertices is the differential pixel region. The pixel values within this rectangular region are extracted, and the pixels in the fourth and fifth rows are filled with supplementary pixel values, and the pixels in the fourth and fifth columns are filled with supplementary pixel values to obtain the target railway image.
[0074] In this embodiment, by comparing the railway image with the background image and only extracting the differential pixel region for processing, the amount of data to be processed can be significantly reduced. This makes the subsequent image feature extraction and recognition process more efficient and improves the overall processing efficiency. The differential pixel region contains the objects newly entered into the railway image, and these objects are the key information that the floating object recognition device needs to focus on. By comparing the background image, these key information can be accurately located and extracted, avoiding the interference of the background region. This enables the floating object recognition device to more accurately identify target objects such as floating objects in the railway image, improving the accuracy and reliability of the recognition. The environment along the railway is complex and changeable, and the background image may change due to factors such as weather and lighting. By comparing the differential pixel region, the floating object recognition device can automatically adapt to these changes and only focus on the objects newly entered into the image. This preprocessing strategy enhances the robustness of the system, enabling it to maintain stable recognition performance in a complex and changeable environment.
[0075] After obtaining the floating object recognition result, the floating object can be adjusted based on the floating object recognition result. Optionally, in the floating object recognition method provided in the embodiments of the present application, after obtaining the floating object recognition result at each acquisition moment, the method further includes: when the floating object recognition result indicates that there is a floating object in the railway image, determining the railway image as the initial image where the floating object appears; determining multiple frames of railway images within a preset time period collected after the initial image, determining the movement trajectory of the floating object based on the multiple frames of railway images, and determining the predicted movement trajectory of the floating object within the target time period based on the movement trajectory of the floating object; determining the position information of the railway grid coverage area, and judging whether the floating object falls into the railway grid coverage area based on the position information and the predicted movement trajectory; when it is determined that the floating object will fall into the railway grid coverage area, controlling the air blowing device to adjust the movement trajectory of the floating object.
[0076] In some embodiments, after determining that there is a floating object in the railway image, the floating object recognition device obtains the recognition results of multiple frames of railway images within a preset time period collected after the initial image, and obtains the actual position of the floating object at each acquisition moment according to the recognition results of each frame of railway image, so as to determine the movement trajectory of the floating object.
[0077] For example, for a certain frame of railway image, after the decoding result output by the decoder indicates that there is a floating object in the railway image, the floating object recognition device obtains the recognition results of multiple frames of railway images collected after this frame of railway image, and extracts the position of the floating object in the image from the recognition results of each frame of railway image. After converting the position of the floating object in the image using the internal and external parameters of the camera, the actual position of the floating object at each acquisition moment is obtained. According to the actual position of the floating object at each acquisition moment, the speed of the floating object at each acquisition moment is calculated, the speed at a future moment is predicted according to the speed at each acquisition moment, and the position of the floating object at the future moment is obtained according to the speed at the future moment, so as to determine the predicted movement trajectory of the floating object within the target time period.
[0078] An air blowing device is installed near the railway grid. The position information of the area covered by the railway grid is determined according to the installation data of the railway grid. According to the position information of the area covered by the railway grid and the predicted movement trajectory of the floating object, it is judged whether the trajectory of the floating object coincides with the area covered by the railway grid. If it coincides, it is determined that the floating object will fall on the railway grid. When the floating object passes through the air blowing device, control the air blowing device to operate to change the trajectory of the floating object to prevent the floating object from falling on the railway grid.
[0079] In this embodiment, after identifying a floating object from a railway image, the camera is controlled to continuously capture multiple frames of railway images, the positions of the floating object in each frame of railway image are extracted, and the trajectory of the floating object is predicted based on the positions of the floating object and the acquisition times of each frame of railway image, so as to determine whether the floating object will fall onto the railway power grid. When it is determined that the floating object may fall onto the railway power grid, the air blowing device is controlled to change the trajectory of the floating object, so as to prevent the floating object from falling onto the power grid and damaging the railway facilities. Based on the recognition results of multiple frames of railway images and the acquisition time information, the trajectory of the floating object is predicted in a data-driven manner. The movement trend and possible influence range of the floating object can be judged more accurately. This embodiment can eliminate potential hazards under the condition of normal operation of railway facilities and ensure the robustness of the operation of railway facilities. By predicting the trajectory of the floating object in real time, potential hazards can be detected in advance before the floating object approaches the railway power grid, providing a valuable time window for subsequent active intervention. This embodiment reduces the need for manual intervention through an automated processing flow and improves the processing efficiency. At the same time, intervening at the beginning of the potential hazard also avoids the expansion of equipment damage or safety hazards caused by the floating object hanging on the power grid for a long time.
[0080] After determining that the floating object will fall within the coverage area of the railway power grid, the movement trajectory of the floating object is changed by the air blowing device. Optionally, in the floating object recognition method provided in the embodiment of the present application, controlling the air blowing device to adjust the movement trajectory of the floating object includes: determining the position information of the air blowing device, and determining the opening time and air blowing parameters of the air blowing device based on the position information of the air blowing device and the predicted movement trajectory, where the air blowing parameters include at least one of the following: wind force level and wind direction; controlling the air blowing device to operate according to the air blowing parameters at the opening time to adjust the movement trajectory of the floating object.
[0081] In some embodiments, if it is determined that the floating object will fall onto the railway power grid, the opening time and air blowing parameters of the air blowing device are determined according to the trajectory of the floating object and the installation position of the air blowing device; the air blowing parameters include the wind force magnitude and the wind direction. Calculate the distances between each trajectory point of the floating object and the installation position of the air blowing device, and select the time corresponding to the trajectory point with the shortest distance as the opening time of the air blowing device. Calculate the wind force magnitude and the wind direction according to the speed of the floating object at the trajectory point with the shortest distance and the distance between this trajectory point and the railway power grid, so that the floating object will not fall onto the railway power grid. For example: the air blowing device provides an upward buoyancy force to make the floating object cross over the railway power grid. The magnitude of the upward buoyancy force is related to the wind speed of the air blowing device. The wind speed of the air blowing device can be determined according to the speed of the floating object in the vertical direction.
[0082] When it is determined in this embodiment that the floating object will fall onto the railway power grid, the optimal opening time and optimal blowing parameters are estimated based on the trajectory of the floating object, meteorological data, and the installation position of the blowing device, so as to change the trajectory of the floating object with less energy consumption, reduce the operating cost of the blowing device, and save energy. After determining that the floating object has a risk of falling into the railway power grid, this embodiment can determine the optimal opening time of the blowing device. The blowing device can be started at a critical moment before the floating object approaches the power grid to change its trajectory in the most efficient way, thereby ensuring the safe operation of railway facilities. The potential damage of the floating object to the railway power grid is avoided, and the normal operation of the train is guaranteed.
[0083] Meteorological data plays an important role in the calculation of blowing parameters. By considering real-time meteorological conditions, this embodiment can make the control of the blowing device more flexible and adapt to different environmental conditions. For example, in strong wind weather, the blowing parameters can be automatically adjusted to offset the influence of external wind force to ensure the accuracy of changing the trajectory of the floating object. This adaptability enables the system to maintain efficient operation in various complex environments.
[0084] According to another embodiment of the present application, an optional floating object recognition method is also provided. Figure 3 It is a schematic diagram of the optional floating object recognition method provided according to the embodiment of the present application. As Figure 3 shown, the method includes:
[0085] S301. The floating object recognition device obtains the railway image captured by the camera and the acquisition time of the railway image.
[0086] Among them, the camera is installed near the railway. The railway image is captured by the camera, and each frame of the railway image and the acquisition time of each frame of the railway image are sent to the floating object recognition device. The floating object recognition device saves each frame of the railway image and the acquisition time of each frame of the railway image.
[0087] S302. The floating object recognition device generates a data request according to the acquisition time of the railway image, and sends the data request to the meteorological system, so that the meteorological system issues the meteorological data at the acquisition time to the floating object recognition device.
[0088] Among them, the floating object recognition device counts the acquisition times of multiple frames of railway images to generate an observation time period. After encoding the observation time period using a preset communication protocol, a data request is generated and sent to the meteorological system. After receiving the data request, the meteorological system decodes the data request to obtain the observation time period, obtains the meteorological data of the observation time period from the local area, and issues the meteorological data to the floating object recognition device.
[0089] S303. The floating object recognition device processes the railway image using the image input layer to obtain the first input feature, and inputs the first input feature into the multi-layer convolutional processing layer to output multiple feature maps.
[0090] Among them, by extracting the tri-color pixel values of the railway image, a tri-color pixel matrix is obtained. The tri-color pixel matrix is used as the first input feature. Each convolutional processing layer includes a convolutional layer and a pooling layer connected to the output end of the convolutional layer. The convolutional layer performs convolutional processing on the input matrix, and the pooling layer downsamples the data after convolutional processing and then outputs the feature map. Each convolutional processing layer processes the feature map output by the previous convolutional processing layer and outputs the feature map. Each convolutional processing layer has multiple convolutional kernels, and each convolutional kernel processes the feature map output by the previous convolutional processing layer, so that multiple features can be output. If the dimensions of the feature maps are different, the multiple feature maps can be unified to the same dimension through linear transformation.
[0091] For example, Figure 4 is a schematic structural diagram of the recognition algorithm adopted by the floating object recognition device provided in the embodiment of the present application. As Figure 4 shown, the recognition algorithm includes: a multi-layer convolutional processing layer, a word vector conversion tool, an encoder, and a decoder. The multi-layer convolutional processing layer processes the railway image to obtain the feature map, and the word vector conversion tool processes the meteorological data to obtain the meteorological word vector. The feature map and the meteorological word vector are superimposed through the fusion parameter to obtain the superimposed feature. The encoder encodes the superimposed feature to obtain the encoded feature, the decoder decodes the encoded feature to obtain the decoded feature, and finally the Softmax function outputs the probability that the decoded feature is of the floating object type.
[0092] S304. The floating object recognition device extracts meteorological keywords from the meteorological data, and uses the word vector conversion tool to convert the meteorological keywords into meteorological word vectors so that the dimension of the meteorological word vector is the same as that of the feature map.
[0093] Among them, the floating object recognition device uses the existing keyword extraction method to extract keywords from the meteorological data to obtain meteorological keywords. Then, the existing word vector conversion tool is used to perform conversion processing on the meteorological keywords to output the meteorological word vector. When setting the parameters of the word vector conversion tool, it can be set based on the size of the feature map, so as to ensure that the dimension of the output meteorological word vector is the same as that of the feature map.
[0094] S305. For each feature map, the meteorological word vector and the feature map are superimposed using the fusion parameter to obtain multiple superimposed feature maps, and the multiple superimposed feature maps are connected to obtain the second input feature.
[0095] Among them, the meteorological word vector and the feature map are matrices with the same dimension. After processing the meteorological word vector and each feature map using the fusion parameter, the processed feature map and the processed meteorological word vector are added matrix-wise to obtain the superimposed feature map. Subsequently, the superimposed feature maps are concatenated to obtain the second input feature. For example: there is a 5×4 meteorological word vector and 10 5×4 feature maps. For each 5×4 feature map, the 5×4 meteorological word vector and the 5×4 feature map are superimposed using the fusion parameter, and 10 5×4 superimposed feature maps can be obtained. By concatenating the 10 5×4 superimposed feature maps, a 50×4 second input feature can be obtained.
[0096] The second input feature is characterized as:
[0097] ;
[0098] Among them, T20 is the matrix of the 1st 5×4 superimposed feature map, T21 is the matrix of the 2nd 5×4 superimposed feature map, T28 is the matrix of the 9th 5×4 superimposed feature map, and T29 is the matrix of the 10th 5×4 superimposed feature map.
[0099] S306. Use a multi-level encoder to encode the second input feature to obtain the encoded feature, and use a multi-level decoder to decode the encoded feature to output whether the target is a light floating object type.
[0100] Among them, after obtaining the second input feature in the above manner, the second input feature is input into the multi-level encoder, and the encoder encodes the input data and then outputs it. More specifically, the first-level encoder encodes the second input feature, the second-level encoder encodes the encoded data output by the first-level encoder, the third-level encoder encodes the encoded data output by the second-level encoder, and so on. Each decoder decodes the encoded data output by the last encoder and the data output by the previous decoder and then outputs, finally outputting the target object type. More specifically, the first-level decoder decodes the encoded data output by the last encoder, and the second-level decoder decodes the decoded data output by the first-level decoder according to the encoded data output by the last encoder, and so on. Until the last decoder decodes the decoded data output by the penultimate decoder according to the encoded data output by the last encoder. In addition, the decoder also needs to connect the Softmax function, so as to output the probability that the target is a light floating object.
[0101] In this embodiment, the floating object recognition device receives the railway image sent by the camera. After receiving the railway image, it sends a data request to the meteorological system to obtain the meteorological data issued by the meteorological system. The floating object recognition device processes the railway image through a multi-layer convolutional layer to obtain multiple feature maps, processes the meteorological data through keyword extraction and word vector conversion to obtain a meteorological word vector, and processes the meteorological word vector and the feature map using a fusion parameter, so as to combine the meteorological word vector with each feature map, and then uses a multi-attention mechanism for encoding and decoding, enabling the attention mechanism to consider the influence of meteorological data on the recognition of feature maps, realizing the recognition of light floating objects based on the multi-modal input of meteorological data and railway images, and improving the accuracy of light floating object recognition.
[0102] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0103] The embodiment of the present application also provides a floating object recognition device. It should be noted that the floating object recognition device of the embodiment of the present application can be used to execute the floating object recognition method provided by the embodiment of the present application. The following introduces the floating object recognition device provided by the embodiment of the present application.
[0104] Figure 5 is a schematic diagram of the floating object recognition device provided according to the embodiment of the present application. As Figure 5 shown, the device includes:
[0105] An acquisition unit 501, configured to acquire the railway images collected by the monitoring device at each acquisition moment within the target time period to obtain the target railway image;
[0106] A first determination unit 502, configured to determine the meteorological data at each acquisition moment to obtain the target meteorological data corresponding to each frame of the target railway image;
[0107] An input unit 503, configured to input the target railway image and the target meteorological data into the target model to obtain the floating object recognition result at each acquisition moment, where the target model is trained by multiple sets of training samples, and each set of training samples includes a historical railway image, the historical meteorological data corresponding to the historical railway image, and the historical floating object recognition result.
[0108] The floating object recognition device provided by the embodiment of the present application obtains the railway images collected by the monitoring device at each acquisition moment within the target time period through the acquisition unit 501 to obtain the target railway images; the first determination unit 502 determines the meteorological data at each acquisition moment to obtain the target meteorological data corresponding to each frame of the target railway image; the input unit 503 inputs the target railway images and the target meteorological data into the target model to obtain the floating object recognition results at each acquisition moment. The target model is obtained by training with multiple groups of training samples, and each group of training samples includes historical railway images, historical meteorological data corresponding to the historical railway images, and historical floating object recognition results. This solves the problem of poor accuracy in floating object recognition in the related art. By combining railway images and meteorological data, multi-modal data is input into the target model for floating object recognition. Even when the amount of training data is insufficient, the recognition accuracy can still be guaranteed, thereby achieving the effect of improving the accuracy of floating object recognition.
[0109] Optionally, in the floating object recognition device provided by the embodiment of the present application, the device further includes: a second determination unit for determining the initial convolutional processing layer, the initial encoder, and the initial decoder; a first training unit for obtaining the first training sample set and training the initial convolutional processing layer based on the first training sample set to obtain the trained convolutional processing layer, where the first training sample set includes historical railway images and historical floating object recognition results; a second training unit for obtaining the second training sample set and training the initial encoder and the initial decoder based on the second training sample set to obtain the trained encoder and the trained decoder, where the second training sample set includes historical superimposed feature maps, historical encoded features, and historical floating object recognition results; a third training unit for obtaining multiple groups of training samples and jointly training the trained convolutional processing layer, the trained encoder, and the trained decoder based on the multiple groups of training samples to obtain the target model.
[0110] Optionally, in the floating object recognition device provided by the embodiment of the present application, the input unit 503 includes: a first extraction module for extracting the three-primary-color pixel values of the target railway image to obtain a three-primary-color pixel matrix; an input module for inputting the three-primary-color pixel matrix into the convolutional processing layer of the target model for processing to obtain multiple feature maps; a second extraction module for extracting keywords from the target meteorological data to obtain meteorological keywords, and converting the meteorological keywords into meteorological word vectors through a word vector conversion tool, where the dimension of the meteorological word vectors is the same as that of the feature maps; a superimposing module for superimposing the meteorological word vectors and the feature maps for each feature map through the fusion parameters of the target model to obtain the superimposed feature maps; an encoding and decoding module for encoding the superimposed feature maps through the encoder of the target model to obtain encoded features, and decoding the encoded features through the decoder of the target model to obtain the floating object recognition results.
[0111] Optionally, in the floating object recognition device provided in the embodiments of the present application, the superimposing module includes: a first determination sub-module, configured to determine each first element in the meteorological word vector, multiply the meteorological word vector coefficient by each first element to obtain an intermediate word vector; a second determination sub-module, configured to determine each second element in the feature map, multiply the feature map coefficient by each second element to obtain an intermediate feature map; and a summation sub-module, configured to sum the intermediate word vector and the intermediate feature map to obtain a superimposed feature map.
[0112] Optionally, in the floating object recognition device provided in the embodiments of the present application, the acquisition unit 501 includes: a receiving module, configured to receive a railway image collected by a monitoring device at a collection moment; a comparison module, configured to compare the railway image with a preset background image to obtain a differential pixel region, where the background image is an image without floating objects collected by the monitoring device at the same shooting angle as the railway image; and a first determination module, configured to determine the image corresponding to the differential pixel region as a target railway image.
[0113] Optionally, in the floating object recognition device provided in the embodiments of the present application, the device further includes: a third determination unit, configured to determine the railway image as an initial image where a floating object appears when the floating object recognition result indicates that there is a floating object in the railway image; a fourth determination unit, configured to determine multiple frames of railway images within a preset time period collected after the initial image, determine the movement trajectory of the floating object based on the multiple frames of railway images, and determine the predicted movement trajectory of the floating object within a target time period based on the movement trajectory of the floating object; a fifth determination unit, configured to determine the position information of the railway power grid coverage area, and determine whether the floating object falls into the railway power grid coverage area based on the position information and the predicted movement trajectory; and a control unit, configured to control the air blower device to adjust the movement trajectory of the floating object when it is determined that the floating object will fall into the railway power grid coverage area.
[0114] Optionally, in the floating object recognition device provided in the embodiments of the present application, the control unit includes: a second determination module, configured to determine the position information of the air blower device, and determine the start time and air blowing parameters of the air blower device based on the position information of the air blower device and the predicted movement trajectory, where the air blowing parameters include at least one of the following: wind force gear and wind direction; and a control module, configured to control the air blower device to operate according to the air blowing parameters at the start time to adjust the movement trajectory of the floating object.
[0115] The floating object recognition device includes a processor and a memory. The above-mentioned acquisition unit 501, first determination unit 502, input unit 503, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to implement corresponding functions.
[0116] The processor contains a kernel, which retrieves the corresponding program unit from the memory. One or more kernels can be set, and the accuracy of floating object recognition can be improved by adjusting the kernel parameters.
[0117] The memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0118] An embodiment of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, a floating object recognition method is implemented.
[0119] An embodiment of the present invention provides a processor for running a program, and when the program runs, a floating object recognition method is executed.
[0120] Figure 6 It is a schematic diagram of an electronic device provided according to an embodiment of the present application. As Figure 6 shown, the electronic device 601 includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, the following steps are implemented: obtaining railway images collected by a monitoring device at each acquisition moment within a target time period to obtain target railway images; determining meteorological data at each acquisition moment to obtain target meteorological data corresponding to each frame of the target railway image; inputting the target railway image and the target meteorological data into a target model to obtain a floating object recognition result at each acquisition moment, where the target model is trained by multiple groups of training samples, and each group of training samples includes historical railway images, historical meteorological data corresponding to the historical railway images, and historical floating object recognition results. The device in this article can be a server, a PC, a PAD, a mobile phone, etc.
[0121] The present application also provides a computer program product, which is suitable for executing a program initialized with the following method steps when executed on a data processing device: obtaining railway images collected by a monitoring device at each acquisition moment within a target time period to obtain target railway images; determining meteorological data at each acquisition moment to obtain target meteorological data corresponding to each frame of the target railway image; inputting the target railway image and the target meteorological data into a target model to obtain a floating object recognition result at each acquisition moment, where the target model is trained by multiple groups of training samples, and each group of training samples includes historical railway images, historical meteorological data corresponding to the historical railway images, and historical floating object recognition results.
[0122] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0123] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0124] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implement the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0125] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0126] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0127] The memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.
[0128] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0129] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0130] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0131] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A floating object recognition method, characterized in that, Including: Obtaining railway images collected by a monitoring device at each acquisition moment within a target time period to obtain target railway images; Determining meteorological data at each of the acquisition moments to obtain target meteorological data corresponding to each frame of the target railway images; Inputting the target railway images and the target meteorological data into a target model to obtain a floating object recognition result at each of the acquisition moments, where the target model is obtained by training with multiple groups of training samples, and each group of training samples includes historical railway images, historical meteorological data corresponding to the historical railway images, and historical floating object recognition results.
2. The method according to claim 1, wherein The target model is obtained by the following method: Determining an initial convolutional processing layer, an initial encoder, and an initial decoder; Obtaining a first training sample set and training the initial convolutional processing layer based on the first training sample set to obtain a trained convolutional processing layer, where the first training sample set includes historical railway images and historical floating object recognition results; Obtaining a second training sample set and training the initial encoder and the initial decoder based on the second training sample set to obtain a trained encoder and a trained decoder, where the second training sample set includes historical superimposed feature maps, historical encoded features, and historical floating object recognition results; Obtaining the multiple groups of training samples and jointly training the trained convolutional processing layer, the trained encoder, and the trained decoder based on the multiple groups of training samples to obtain the target model.
3. The method according to claim 1, characterized in that, Inputting the target railway images and the target meteorological data into the target model to obtain a floating object recognition result at each of the acquisition moments includes: Extracting the three primary color pixel values of the target railway images to obtain a three primary color pixel matrix; Inputting the three primary color pixel matrix into the convolutional processing layer of the target model for processing to obtain multiple feature maps; Extracting keywords from the target meteorological data to obtain meteorological keywords, and converting the meteorological keywords into meteorological word vectors through a word vector conversion tool, where the dimension of the meteorological word vectors is the same as that of the feature maps; For each of the feature maps, superimposing the meteorological word vectors and the feature maps through the fusion parameters of the target model to obtain a superimposed feature map; Encoding the superimposed feature map through the encoder of the target model to obtain an encoded feature, and decoding the encoded feature through the decoder of the target model to obtain the floating object recognition result.
4. The method according to claim 3, wherein The fusion parameters include a meteorological word vector coefficient and a feature map coefficient. Superimposing the meteorological word vectors and the feature maps through the fusion parameters of the target model includes: Determining each first element in the meteorological word vectors, multiplying the meteorological word vector coefficient by each first element to obtain an intermediate word vector; Determining each second element in the feature maps, multiplying the feature map coefficient by each second element to obtain an intermediate feature map; Summing the intermediate word vector and the intermediate feature map to obtain the superimposed feature map.
5. The method according to claim 1, wherein Obtaining railway images collected by a monitoring device at each acquisition moment within a target time period to obtain target railway images includes: Receive the railway image collected by the monitoring device at the collection moment; Compare the railway image with a preset background image to obtain a differential pixel region, where the background image is an image without floating objects collected by the monitoring device at the same shooting angle as the railway image; Determine the image corresponding to the differential pixel region as the target railway image.
6. The method according to claim 1, wherein After obtaining the floating object recognition result at each collection moment, the method further includes: When the floating object recognition result indicates that there is a floating object in the railway image, determine the railway image as the initial image where the floating object appears; Determine multiple frames of railway images within a preset time period collected after the initial image, determine the movement trajectory of the floating object based on the multiple frames of railway images, and determine the predicted movement trajectory of the floating object within the target time period based on the movement trajectory of the floating object; Determine the position information of the railway power grid coverage area, and judge whether the floating object falls into the railway power grid coverage area based on the position information and the predicted movement trajectory; When it is determined that the floating object will fall into the railway power grid coverage area, control the air blowing device to adjust the movement trajectory of the floating object.
7. The method according to claim 6, wherein Controlling the air blowing device to adjust the movement trajectory of the floating object includes: Determine the position information of the air blowing device, and determine the starting time and air blowing parameters of the air blowing device based on the position information of the air blowing device and the predicted movement trajectory, where the air blowing parameters include at least one of the following: wind force level and wind direction; Control the air blowing device to operate according to the air blowing parameters at the starting time to adjust the movement trajectory of the floating object.
8. A floating object recognition device, characterized in that, Includes: An acquisition unit for acquiring railway images collected by the monitoring device at each collection moment within the target time period to obtain target railway images; A first determination unit for determining the meteorological data at each collection moment to obtain the target meteorological data corresponding to each frame of the target railway image; An input unit for inputting the target railway image and the target meteorological data into a target model to obtain the floating object recognition result at each collection moment, where the target model is trained by multiple groups of training samples, and each group of training samples includes a historical railway image, the historical meteorological data corresponding to the historical railway image, and a historical floating object recognition result.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the floating object recognition method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, Includes one or more processors and a memory, where the memory is used to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the floating object recognition method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Tide image denoising processing method, terminal and computer readable storage medium
CN113744152A
Risk early warning method and device for floating objects along railway
CN113850482A
Encoder training method and device and storage medium
CN114418069A
Speech recognition model training method and device, equipment and storage medium
CN118016053A
Road floating object sensing and processing method and device, electronic equipment, medium and vehicle
CN118887635A