Communication facility fault detection method and device under low light intensity, equipment and medium
By acquiring multimodal images in low-light environments and performing modality preference evaluation and feature fusion, the problem of insufficient detection accuracy in low-light environments is solved, and higher accuracy in detecting communication facility faults is achieved.
Patent Information
- Application Number
- CN202511096922.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-06
- Publication Date
- 2025-11-21
AI Technical Summary
In low light intensity environments, existing fault detection methods for communication facilities suffer from insufficient detection accuracy and high false detection and false negative rates due to the static and singular modal fusion strategy.
Multimodal images (including multi-channel images and infrared images) are acquired, and first and second image features are obtained through feature extraction. Preference weights are obtained based on modality preference evaluation, and feature fusion is performed to dynamically adjust the modality fusion strategy.
It improves the accuracy of fault detection in communication facilities under low light intensity and effectively enhances detection performance.
Smart Images

Figure CN120997157A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer application, and in particular to a communication facility fault detection method and device under low light intensity, equipment and medium. BACKGROUND
[0002] In the construction and maintenance process of communication engineering, if the environmental light is insufficient (for example, at night, underground or in a closed space, etc.), there are significant technical challenges in communication facility fault detection. Such scenes are generally characterized by limited visual perception, complex environmental structure (for example, cable interlacing or pipeline shielding, etc.) and harsh sensing conditions, which seriously restrict the accuracy of communication equipment fault detection.
[0003] In related technologies, in order to improve the performance of communication facility fault detection under low light intensity, multi-modal perception technology is usually introduced. A simple and equal feature fusion strategy (such as direct weighted average, feature splicing or early fusion, etc.) is generally used for multiple modal images, ignoring the fact that different light intensity environments have significant differences in the discriminative nature of different modal information. This single and static modal fusion strategy can easily lead to key discriminative information being submerged or noise being amplified, ultimately resulting in insufficient detection accuracy in complex low-light scenes and high false detection and missed detection rates. SUMMARY
[0004] The present application provides a communication facility fault detection method and device under low light intensity, equipment and medium to solve the technical problem of insufficient detection accuracy caused by static and single modal fusion strategy in the prior art under insufficient light environment.
[0005] According to an aspect of the present application, a communication facility fault detection method under low light intensity is provided, which comprises:
[0006] In the case where the actual light intensity is lower than the light intensity threshold, multi-modal images of the target detection area are collected; wherein the multi-modal images include multi-channel images and infrared images;
[0007] The multi-channel images are subjected to feature extraction to obtain first image features, and the infrared images are subjected to feature extraction to obtain second image features;
[0008] The multi-modal images are subjected to modal preference evaluation based on the first image features and the second image features to obtain a first preference weight corresponding to the multi-channel images and a second preference weight corresponding to the infrared images;
[0009] The first image features and the second image features are subjected to feature fusion according to the first preference weight and the second preference weight to obtain a fused image feature;
[0010] determine a communication facility fault detection result corresponding to the target detection region according to the fused image features.
[0011] According to another aspect of the present application, there is provided a communication facility fault detection method under low light intensity, comprising:
[0012] In a case where the actual light intensity is lower than the light intensity threshold, a multi-modal image of a target detection region is collected; wherein the multi-modal image comprises a multi-channel image and an infrared image;
[0013] feature extraction is performed on the multi-channel image to obtain first image features, and feature extraction is performed on the infrared image to obtain second image features;
[0014] modal preference evaluation is performed on the multi-modal image based on the first image features and the second image features to obtain a first preference weight corresponding to the multi-channel image and a second preference weight corresponding to the infrared image;
[0015] feature fusion is performed on the first image features and the second image features according to the first preference weight and the second preference weight to obtain fused image features;
[0016] determine a communication facility fault detection result corresponding to the target detection region according to the fused image features.
[0017] According to another aspect of the present application, there is provided an electronic device, comprising:
[0018] at least one processor; and
[0019] a memory in communication connection with the at least one processor; wherein,
[0020] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the communication facility fault detection method under low light intensity according to any one of the embodiments of the present application.
[0021] According to another aspect of the present application, there is provided a computer readable storage medium storing computer instructions for enabling a processor to implement the communication facility fault detection method under low light intensity according to any one of the embodiments of the present application when executed by the processor.
[0022] The technical scheme of the embodiment of the present application comprises the following steps: collecting a multi-modal image of a target detection area when the actual light intensity is lower than the light intensity threshold; the multi-modal image comprises a multi-channel image and an infrared image; performing feature extraction on the multi-channel image to obtain a first image feature; performing feature extraction on the infrared image to obtain a second image feature; performing modal preference evaluation on the multi-modal image based on the first image feature and the second image feature to obtain a first preference weight corresponding to the multi-channel image and a second preference weight corresponding to the infrared image; performing feature fusion on the first image feature and the second image feature according to the first preference weight and the second preference weight to obtain a fused image feature; and determining a communication facility fault detection result corresponding to the target detection area according to the fused image feature. Based on the present application, the multi-modal image features are fused based on the modal preference corresponding to the actual light environment in the insufficient light environment, which is more conducive to communication facility fault detection, and the accuracy of communication facility fault detection under low light intensity is effectively improved.
[0023] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0025] Figure 1 is a flow chart of a communication facility fault detection method under low light intensity according to the first embodiment of the present application;
[0026] Figure 2 is a flow chart of a communication facility fault detection method under low light intensity according to the second embodiment of the present application;
[0027] Figure 3 is an architecture diagram of a communication facility fault detection system under low light intensity according to the present application;
[0028] Figure 4 is a structural schematic diagram of a communication facility fault detection device under low light intensity according to the third embodiment of the present application;
[0029] Figure 5is a structural schematic diagram of an electronic device for implementing a communication facility fault detection method under low light intensity according to an embodiment of the present application. DETAILED DESCRIPTION
[0030] In order for those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should fall within the scope of protection of the present application.
[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0032] Embodiment one
[0033] Figure 1 A flowchart of a communication facility fault detection method under low light intensity is provided for the first embodiment of the present application. The present embodiment can be applicable to the case of communication facility fault detection under low light intensity based on multi-modal images. The method can be executed by a communication facility fault detection device under low light intensity, which can be realized in the form of hardware and / or software, and can be configured in a computer. As shown in the figure, the method comprises: Figure 1
[0034] S110, in the case that the actual light intensity is lower than the light intensity threshold value, multi-modal images of the target detection area are collected.
[0035] The actual light intensity can be understood as the light intensity in the actual environment. It can be understood that the light intensity in different time periods of the same day, different weather, or different geographical locations is usually different, but there are also the same cases. For example, the light intensity in the noon period of the same day is high, and the light intensity in the night period is low. In the embodiment of the application, the specific value of the actual light intensity is related to the actual application scene, which is not limited here.
[0036] The light intensity threshold value can be understood as a preset threshold value related to the light intensity. In the embodiment of the application, the light intensity threshold value can be related to the actual demand, which is not limited here.
[0037] It should be noted that in the case where the actual light intensity is higher than the light intensity threshold value, it is considered that the light intensity is strong, and the image of a single mode can be collected to realize relatively accurate communication facility fault detection, without the need to collect multi-modal images. In the case where the actual light intensity is lower than the light intensity threshold value, it is considered that the light intensity is weak, and it is difficult to realize relatively accurate communication facility fault detection by collecting an image of a single mode, and multi-modal images need to be collected to realize accurate communication facility fault detection.
[0038] The target detection area can be understood as an area to be detected for communication facility fault. The target detection area includes a communication device to be detected for fault. In the embodiment of the application, the specific communication device can be preset according to the scene demand, which is not limited here. The communication device to be detected for fault is a communication device that may have a fault.
[0039] The multi-modal image can include images of at least two modes. In the embodiment of the application, the multi-modal image includes a multi-channel image and an infrared image. The multi-channel image can be a three-color channel (Red Green Blue, RGB) image. On the basis of the above embodiment, the multi-modal image can be collected based on the following manner. Optionally, in the case where the actual light intensity is lower than the light intensity threshold value, the multi-modal image of the target detection area is collected, comprising:
[0040] In the case where the actual light intensity is lower than the light intensity threshold value, the target working parameter of the target sensor is determined according to the actual light intensity;
[0041] So that the target sensor collects the multi-modal image of the target detection area based on the target working parameter.
[0042] The target sensor can have an image acquisition function. In the embodiment of the present application, the target sensor can be set according to the scene requirement, which is not specifically limited here. For example, the target sensor includes a complementary metal-oxide-semiconductor (CMOS) sensor and an infrared thermal imager. Specifically, the CMOS sensor is used to acquire a multi-channel image; and
[0043] Based on the above embodiment, the working parameters of the sensor are dynamically adjusted according to the actual light intensity, so as to acquire a multi-modal image for high-accuracy communication facility fault detection.
[0044] The infrared thermal imager is used to acquire an infrared image.
[0045] The target working parameter can be understood as a working parameter of the target sensor. For example, the target working parameter can be an exposure degree. In the embodiment of the present application, the working parameter of the target sensor can be different under different actual light intensities. For example, the weaker the actual light intensity is, the greater the exposure degree is.
[0046] S120, feature extraction is performed on the multi-channel image to obtain a first image feature; and feature extraction is performed on the infrared image to obtain a second image feature.
[0047] The first image feature can be understood as an image feature corresponding to the multi-channel image. Alternatively, the first image feature can include a texture feature and a shape feature.
[0048] The second image feature can be understood as an image feature corresponding to the infrared image. Alternatively, the second image feature can include a temperature distribution feature and a thermal radiation feature.
[0049] On the basis of the above embodiment, the multi-modal image can be feature-extracted based on the following manner.
[0050] Alternatively, the feature extraction on the multi-channel image to obtain the first image feature includes:
[0051] The input multi-channel image is feature-extracted by a first residual network to obtain a texture feature and a shape feature, and the first image feature is determined according to the texture feature and the shape feature;
[0052] The feature extraction on the infrared image to obtain the second image feature includes:
[0053] The infrared image is input into a second residual network for feature extraction to obtain a temperature distribution feature and a thermal radiation feature, and the second image feature is determined according to the temperature distribution feature and the thermal radiation feature.
[0054] The first residual network and the second residual network can be understood as two deep learning models with image feature extraction functions. In the embodiment of the present application, the first residual network can be different from the second residual network. The first residual network is used to extract the image features of the multi-channel image. The second residual network is used to extract the image features of the infrared image. The first residual network can be obtained by training a ResNet-50 model based on first training samples. The second residual network can be obtained by training a ResNet-50 model based on second training samples.
[0055] Based on the above embodiment, the accuracy of image feature extraction can be improved by using the residual network trained based on the ResNet-50 model to extract features from the multi-modal image.
[0056] Based on the above embodiment, the first image feature can be determined according to the texture feature and the shape feature, including:
[0057] The texture feature and the shape feature are normalized respectively, and the normalized texture feature and the normalized shape feature are taken as the first image feature.
[0058] The second image feature can be determined according to the temperature distribution feature and the thermal radiation feature, including:
[0059] The temperature distribution feature and the thermal radiation feature are normalized respectively, and the normalized temperature distribution feature and the normalized thermal radiation feature are taken as the first image feature.
[0060] Based on the above embodiment, the texture, shape, temperature distribution and thermal radiation image features are normalized to ensure that the image features of different modalities have comparability in the numerical range.
[0061] S130, modal preference evaluation is performed on the multi-modal image based on the first image feature and the second image feature to obtain a first preference weight corresponding to the multi-channel image and a second preference weight corresponding to the infrared image.
[0062] The first preference weight can represent the discriminability of the multi-channel image for communication facility fault detection. The second preference weight can represent the discriminability of the infrared image for communication facility fault detection.
[0063] It needs to be understood that the image features extracted from the multi-modal images are different under different communication facility fault detection environments, and accordingly, the discriminativeness of different modal images for communication facility fault detection is also different. In the embodiments of the present application, the greater the preference weight corresponding to the modal image with higher discriminativeness for communication facility fault detection. For example, in the A detection environment, the first preference weight is greater than the second preference weight; in the B detection environment, the first preference weight is less than the second preference weight; in the C detection environment, the first preference weight is equal to the second preference weight.
[0064] S140, according to the first preference weight and the second preference weight, the first image feature and the second image feature are fused to obtain a fused image feature.
[0065] Among them, the fused image feature can be understood as the total feature obtained after the fusion of the multi-modal image features.
[0066] S150, according to the fused image feature, a communication facility fault detection result corresponding to the target detection area is determined.
[0067] Among them, the communication facility fault detection result can represent whether the communication equipment in the target detection area exists fault. Optionally, the communication facility fault detection result at least includes a detection conclusion and a detection conclusion corresponding accuracy probability. Among them, the detection conclusion can be whether the communication equipment exists fault. For example, the communication equipment exists A type fault, the communication equipment exists B type fault or the communication equipment does not exist fault, etc. The accuracy probability can represent the probability of the detection conclusion accurate. The detection accuracy probability is related to the actual application scene, which is not limited here. Exemplarily, the detection accuracy probability is 70%, 80% or 95%, etc. For example, the probability of the communication equipment existing A type fault is 70%; or the probability of the communication equipment not existing fault is 95%, etc.
[0068] The technical scheme of the embodiment of the present application comprises the following steps: collecting a multi-modal image of a target detection area when actual light intensity is lower than a light intensity threshold; wherein the multi-modal image comprises a multi-channel image and an infrared image; performing feature extraction on the multi-channel image to obtain a first image feature; and performing feature extraction on the infrared image to obtain a second image feature; performing modal preference evaluation on the multi-modal image based on the first image feature and the second image feature to obtain a first preference weight corresponding to the multi-channel image and a second preference weight corresponding to the infrared image; performing feature fusion on the first image feature and the second image feature according to the first preference weight and the second preference weight to obtain a fused image feature; and determining a communication facility fault detection result corresponding to the target detection area according to the fused image feature. Based on the present application, the multi-modal image features are fused based on the modal preference corresponding to the actual light environment in the insufficient light environment, which is more conducive to communication facility fault detection, and the accuracy of communication facility fault detection under low light intensity is effectively improved.
[0069] Embodiment two
[0070] Figure 2 A flowchart of a communication facility fault detection method under low light intensity provided by the second embodiment of the present application is provided. The present embodiment refines the modal preference evaluation on the multi-modal image based on the first image feature and the second image feature to obtain the first preference weight corresponding to the multi-channel image and the second preference weight corresponding to the infrared image in the above-mentioned embodiment. As shown in the figure, the method comprises the following steps: Figure 2
[0071] S210, collecting a multi-modal image of a target detection area when actual light intensity is lower than a light intensity threshold.
[0072] S220, performing feature extraction on the multi-channel image to obtain a first image feature; and performing feature extraction on the infrared image to obtain a second image feature.
[0073] S230, evaluating the input first image feature by a first expert network to obtain a first preliminary weight corresponding to the multi-channel image; and evaluating the input second image feature by a second expert network to obtain a second preliminary weight corresponding to the infrared image.
[0074] The first expert network can have the function of discriminative evaluation on the features of the multi-channel image. The second expert network can have the function of discriminative evaluation on the features of the infrared image. The above-mentioned discriminative evaluation can represent whether the image of the target modal is more helpful for communication facility fault detection.
[0075] In the embodiments of the present application, the higher the discriminativeness of the first image feature evaluated by the first expert network, the greater the first preliminary weight output. Similarly, the higher the discriminativeness of the second image feature evaluated by the second expert network, the greater the second preliminary weight output.
[0076] S240, cross-validation of the first preliminary weight and the second preliminary weight is performed by a cross-validation network to obtain a cross-validation result, and the first preference weight and the second preference weight are determined according to the first preliminary weight, the second preliminary weight and the cross-validation result.
[0077] The cross-validation network can have the function of cross-validating the preliminary weight of the multi-modal image. The cross-validation result can represent whether the preliminary weight determined based on the expert network is accurate. Alternatively, the first preference weight and the second preference weight are determined according to the first preliminary weight, the second preliminary weight and the cross-validation result, including:
[0078] The first preliminary weight and the second preliminary weight are adjusted based on the cross-validation result to obtain the first preference weight and the second preference weight.
[0079] S250, the first image feature and the second image feature are fused according to the first preference weight and the second preference weight to obtain a fused image feature.
[0080] On the basis of the above embodiments, the fusion of multi-modal features can be performed based on the following manner. Alternatively, the first image feature and the second image feature are fused according to the first preference weight and the second preference weight to obtain a fused image feature, including:
[0081] The first preference weight, the second preference weight, the first image feature and the second image feature are weighted and summed to obtain a first fused feature;
[0082] The first fused feature is interacted through a cross-attention mechanism to obtain a second fused feature;
[0083] The second fused feature is spatially aligned through a gating mechanism to obtain a fused image feature.
[0084] The first fusion feature can be understood as a fusion feature obtained by weighted summation of multi-modal image features. The cross-attention mechanism can be used to process the relationship between two different modal image features. The second fusion feature can be understood as a fusion feature obtained by further interaction of the fusion feature based on the cross-attention mechanism on the basis of the first fusion feature. The gating mechanism can be understood as a technology for controlling information flow in a neural network. In the embodiment of the present application, the gating mechanism can control the information flow between multi-modal features, adjust the fusion ratio between different modal features, eliminate the problem of spatial misalignment between different modal features, so as to ensure that the finally obtained fusion feature has accurate positional correspondence. The fusion image feature can be understood as a target fusion feature obtained by three times of feature fusion processing of multi-modal features based on weighted summation, cross-attention mechanism and gating mechanism.
[0085] Based on the above embodiment scheme, three times of feature fusion processing of multi-modal features based on weighted summation, cross-attention mechanism and gating mechanism, realizes the effect of dynamically fusing multi-modal features based on individual preference of image features, avoids the technical problem of poor communication facility fault detection performance caused by equal treatment of multi-modal features and simple fusion of multi-modal features. Based on the embodiment scheme of the present application, more adaptive feature fusion of multi-modal features for communication facility fault detection can be realized to effectively improve the performance of communication facility fault detection.
[0086] S260, determining a communication facility fault detection result corresponding to the target detection area according to the fusion image feature.
[0087] The technical scheme of the embodiment of the present application evaluates the input first image feature through the first expert network to obtain the first preliminary weight corresponding to the multi-channel image, and evaluates the input second image feature through the second expert network to obtain the second preliminary weight corresponding to the infrared image. The first preliminary weight and the second preliminary weight are cross-validated through the cross-validation network to obtain a cross-validation result, and the first preference weight and the second preference weight are determined according to the first preliminary weight, the second preliminary weight and the cross-validation result. The present application introduces individual preference perception to perform discriminative evaluation of multi-modal image features to determine the preference weight corresponding to different modal features. The present application first determines the preliminary weight through preliminary discriminative evaluation of the expert network, and then detects the consistency of the preliminary weight determined by the expert network based on the cross-validation network to adjust the preliminary weight, thereby ensuring the accuracy of the determined preference weight.
[0088] The application solves the problem of insufficient detection accuracy caused by single modal fusion strategy in insufficient light environment for the existing communication facility fault detection method. The application dynamically evaluates the discriminative difference of different targets in RGB and infrared modalities, and adaptively adjusts the modal fusion strategy, thereby significantly improving the detection performance under low light conditions.
[0089] Figure 3 is a low light intensity communication facility fault detection system architecture provided by an embodiment of the application. As follows Figure 3 The low light intensity communication facility fault detection system is described. It should be noted that the low light intensity communication facility fault detection system can be used to implement the low light intensity communication facility fault detection method. Specifically, the communication facility fault detection system can include an image acquisition module, a feature extraction module, a preference learning module, a cross-modal fusion module, and a detection output module.
[0090] The image acquisition module is the data input core of the system, responsible for acquiring multi-modal images under complex lighting conditions. This module uses a combination of high-sensitivity CMOS sensors and infrared thermal imagers, which ensures image quality under normal lighting conditions and ensures imaging capability in complete darkness. The system supports automatic exposure adjustment and dynamic range optimization, which can automatically switch working modes according to environmental light intensity to adapt to various low-light scenes from dusk to complete darkness.
[0091] The feature extraction module uses a dual-branch deep neural network architecture that can process RGB images and infrared images in parallel. Specifically, based on the improved ResNet-50 model, attention mechanisms and feature calibration layers are added to the two branches. The RGB image branch focuses on extracting texture and shape features of the target, while the infrared branch focuses on temperature distribution and thermal radiation features. The two branches interact through a cross-modal feature sharing layer to lay the foundation for subsequent preference learning. All feature maps are normalized to ensure that features from different modalities are comparable in numerical range.
[0092] The preference learning module uses a cascaded expert network structure to realize dynamic evaluation of modal preference. In the first stage, two independent expert networks score the importance of RGB and infrared features, respectively, to generate preliminary modal preference weights. In the second stage, the cross-validation network checks the consistency of the scoring results of the two modalities, and corrects unreasonable preference allocation through an attention mechanism. The final preference weight is output through a differentiable softmax function, ensuring that the entire system can be trained end-to-end. The module also introduces a memory mechanism that can record the preference distribution of historical samples to guide the preference prediction of new samples.
[0093] The cross-modal fusion module adopts a multi-level feature fusion strategy. First, feature fusion is achieved through weighted summation based on preference weights; further, feature interaction is achieved using a preference-guided cross-attention mechanism; further, a gated fusion unit is used to dynamically adjust the fusion ratio. Based on this way, the spatial misalignment problem between different modalities can be eliminated, and the fused features have accurate position correspondence. The module also includes a lightweight feature enhancement subnetwork to improve the discriminability of the fused features.
[0094] The detection output module is responsible for converting the fused features into the final target detection results. Specifically, an anchor-free based detection architecture is used, which includes three parallel branches: a classification branch to predict target classes, a regression branch to predict target positions, and a preference verification branch to evaluate the detection results. The outputs of the three branches are integrated through non-maximum suppression to obtain the final detection boxes and class labels. The system also supports outputting the modal preference score of each target to provide a reference for subsequent analysis.
[0095] Based on the technical scheme of the present application, the problem of insufficient detection accuracy of existing communication facility fault detection methods in low light environments due to single modal fusion strategy is solved. The present application dynamically evaluates the discriminative difference of different targets in RGB and infrared modalities, and adaptively adjusts the modal fusion strategy, thereby significantly improving the detection performance in low light conditions.
[0096] Embodiment three
[0097] Figure 4 A structure diagram of a communication facility fault detection device under low light intensity provided by the third embodiment of the present application is shown in FIG. 3. As shown in FIG. 3, the device includes an image acquisition module 310, a feature extraction module 320, a preference evaluation module 330, a feature fusion module 340, and a result determination module 350. Figure 4
[0098] The image acquisition module 310 is configured to acquire a multi-modal image of a target detection area when the actual light intensity is lower than the light intensity threshold value, wherein the multi-modal image comprises a multi-channel image and an infrared image.
[0099] The technical scheme of the embodiment of the application comprises the following steps: acquiring a multi-modal image of a target detection area when the actual light intensity is lower than the light intensity threshold value, wherein the multi-modal image comprises a multi-channel image and an infrared image; performing feature extraction on the multi-channel image to obtain first image features; performing feature extraction on the infrared image to obtain second image features; performing modal preference evaluation on the multi-modal image based on the first image features and the second image features to obtain a first preference weight corresponding to the multi-channel image and a second preference weight corresponding to the infrared image; performing feature fusion on the first image features and the second image features according to the first preference weight and the second preference weight to obtain fused image features; and determining a communication facility fault detection result corresponding to the target detection area according to the fused image features.
[0100] Optionally, the preference evaluation module 330 is specifically configured to:
[0101] perform evaluation on the input first image features through a first expert network to obtain a first preliminary weight corresponding to the multi-channel image; and perform evaluation on the input second image features through a second expert network to obtain a second preliminary weight corresponding to the infrared image.
[0102] cross-validation network to cross-validate the first preliminary weight and the second preliminary weight, to obtain a cross-validation result, and determine the first preference weight and the second preference weight according to the first preliminary weight, the second preliminary weight, and the cross-validation result.
[0103] Optionally, the feature fusion module 340 is specifically configured to:
[0104] perform weighted summation on the first preference weight, the second preference weight, the first image feature, and the second image feature, to obtain a first fusion feature;
[0105] perform feature interaction on the first fusion feature through a cross-attention mechanism, to obtain a second fusion feature;
[0106] perform spatial alignment on the second fusion feature through a gating mechanism, to obtain a fusion image feature.
[0107] Optionally, the feature extraction module 320 comprises a multi-channel feature extraction unit configured to perform feature extraction on the input multi-channel image through a first residual network, to obtain texture features and shape features, and determine the first image feature according to the texture features and the shape features.
[0108] The feature extraction module 320 comprises an infrared feature extraction unit configured to perform feature extraction on the input infrared image through a second residual network, to obtain temperature distribution features and thermal radiation features, and determine the second image feature according to the temperature distribution features and the thermal radiation features.
[0109] Optionally, the multi-channel feature extraction unit comprises a first normalization subunit configured to perform normalization processing on the texture features and the shape features respectively, and take the normalized texture features and the normalized shape features as the first image feature.
[0110] The infrared feature extraction unit comprises a second normalization subunit configured to perform normalization processing on the temperature distribution features and the thermal radiation features respectively, and take the normalized temperature distribution features and the normalized thermal radiation features as the first image feature.
[0111] Optionally, the image acquisition module 310 comprises a working parameter determination unit and an image acquisition unit.
[0112] The working parameter determination unit is configured to determine a target working parameter of a target sensor according to an actual light intensity when the actual light intensity is lower than a light intensity threshold.
[0113] The image acquisition unit is configured to cause the target sensor to acquire a multi-modal image of a target detection region based on the target working parameter.
[0114] Optionally, the communication facility fault detection result at least includes a detection conclusion and an accuracy probability corresponding to the detection conclusion.
[0115] The communication facility fault detection device at low light intensity provided by the embodiment of the present application can execute the communication facility fault detection method at low light intensity provided by any embodiment of the present application, and has the corresponding function modules and beneficial effects of the execution method.
[0116] Embodiment four
[0117] Figure 5 A structural schematic diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed in this document.
[0118] As shown in Figure 5 The electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is in communication with the at least one processor 11, wherein the memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or loaded into the random access memory (RAM) 13 from the storage unit 18. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0119] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.
[0120] The processor 11 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, and the like. The processor 11 performs various methods and processes described above, such as the communication facility failure detection method under low light intensity.
[0121] In some embodiments, the communication facility failure detection method under low light intensity can be implemented as a computer program tangibly embodied in a computer readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more steps of the communication facility failure detection method under low light intensity described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the communication facility failure detection method under low light intensity by any other appropriate means, such as by means of firmware.
[0122] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0123] Computer programs for implementing the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer program, when executed, enables the functions / acts specified in the flowcharts and / or block diagrams to be implemented. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as a standalone software package and partially on a remote machine or entirely on a remote machine or server.
[0124] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0125] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0126] The systems and techniques described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described herein, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0127] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and VPS service.
[0128] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be executed in parallel, executed in sequence, or executed in a different order, as long as the desired results of the technical solutions of the present disclosure are achieved, and the present disclosure is not limited herein.
[0129] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A method for detecting a failure of a communication facility at low light levels, characterized by, Comprise: In the case that the actual light intensity is lower than the light intensity threshold, a multi-modal image of the target detection area is collected; wherein the multi-modal image comprises a multi-channel image and an infrared image; Feature extraction is performed on the multi-channel image to obtain a first image feature, and feature extraction is performed on the infrared image to obtain a second image feature; Modal preference evaluation is performed on the multi-modal image based on the first image feature and the second image feature to obtain a first preference weight corresponding to the multi-channel image and a second preference weight corresponding to the infrared image; Feature fusion is performed on the first image feature and the second image feature according to the first preference weight and the second preference weight to obtain a fused image feature; A communication facility fault detection result corresponding to the target detection area is determined according to the fused image feature.
2. The method of claim 1, wherein, The modal preference evaluation on the multi-modal image based on the first image feature and the second image feature to obtain the first preference weight corresponding to the multi-channel image and the second preference weight corresponding to the infrared image comprises: The first image feature input is evaluated by a first expert network to obtain a first preliminary weight corresponding to the multi-channel image, and the second image feature input is evaluated by a second expert network to obtain a second preliminary weight corresponding to the infrared image; The first preliminary weight and the second preliminary weight are cross-validated by a cross-validation network to obtain a cross-validation result, and the first preference weight and the second preference weight are determined according to the first preliminary weight, the second preliminary weight, and the cross-validation result.
3. The method of claim 1, wherein, The feature fusion on the first image feature and the second image feature according to the first preference weight and the second preference weight to obtain a fused image feature comprises: The first preference weight, the second preference weight, the first image feature, and the second image feature are weighted and summed to obtain a first fused feature; The first fused feature is interacted by a cross-attention mechanism to obtain a second fused feature; The second fused feature is spatially aligned by a gating mechanism to obtain a fused image feature. The feature extraction on the multi-channel image to obtain a first image feature comprises:
4. The method of claim 1, wherein, The multi-channel image input is extracted by a first residual network to obtain a texture feature and a shape feature, and the first image feature is determined according to the texture feature and the shape feature; The feature extraction on the infrared image to obtain a second image feature comprises: The infrared image input is extracted by a second residual network to obtain a temperature distribution feature and a thermal radiation feature, and the second image feature is determined according to the temperature distribution feature and the thermal radiation feature. The determination of the first image feature according to the texture feature and the shape feature comprises:
5. The method of claim 4, wherein, The texture feature and the shape feature are normalized respectively, and the normalized texture feature and the normalized shape feature are taken as the first image feature; The determining the second image feature according to the temperature distribution feature and the thermal radiation feature comprises: The temperature distribution feature and the thermal radiation feature are normalized respectively, and the normalized temperature distribution feature and the normalized thermal radiation feature are taken as the first image feature.
6. The method of claim 1, wherein, The acquiring the multi-modal image of the target detection area under the condition that the actual light intensity is lower than the light intensity threshold value comprises: The target working parameter of the target sensor is determined according to the actual light intensity under the condition that the actual light intensity is lower than the light intensity threshold value; So that the target sensor acquires the multi-modal image of the target detection area based on the target working parameter.
7. The method of claim 1, wherein, The communication facility fault detection result at least comprises a detection conclusion and an accurate probability corresponding to the detection conclusion.
8. A low light level communication facility fault detection apparatus, characterised by, Comprise: The image acquisition module is used for acquiring the multi-modal image of the target detection area under the condition that the actual light intensity is lower than the light intensity threshold value; wherein the multi-modal image comprises a multi-channel image and an infrared image; The feature extraction module is used for extracting features from the multi-channel image to obtain a first image feature, and extracting features from the infrared image to obtain a second image feature; The preference evaluation module is used for evaluating the modal preference of the multi-modal image based on the first image feature and the second image feature to obtain a first preference weight corresponding to the multi-channel image and a second preference weight corresponding to the infrared image; The feature fusion module is used for fusing the first image feature and the second image feature according to the first preference weight and the second preference weight to obtain a fused image feature; The result determination module is used for determining the communication facility fault detection result corresponding to the target detection area according to the fused image feature.
9. An electronic device, comprising: The electronic device comprises: At least one processor; and The memory is in communication connection with the at least one processor; wherein The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the low-light-intensity communication facility fault detection method in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to execute the low-light-intensity communication facility fault detection method in any one of claims 1-7 when executed.