Cross-modal data equipment infrared image recognition method and system, medium and equipment
The cross-modal data method improves infrared image recognition in substations by aligning visible and infrared images using affine transformation and semantic segmentation, addressing the limitations of single infrared camera methods and enhancing recognition accuracy and reliability.
Patent Information
- Application Number
- CN202510391483.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-15
AI Technical Summary
The existing infrared recognition methods of substation equipment rely on single-light infrared camera images, resulting in insufficient recognition accuracy and reliability, difficulty in capturing subtle features of the equipment, and the direct use of visible-light image network to process infrared images is poor.
By acquiring visible light and infrared images, the device in the visible light image is identified using the object detection model, and combined with the affine transformation matrix and the semantic segmentation model, the object recognition box of the visible light image is mapped into the infrared image, and the device component mask information is extracted.
It improves the accuracy and reliability of infrared image recognition, realizes the precise recognition of equipment and its components, overcomes the limitations of single-light infrared images, and improves the recognition accuracy and stability.
Smart Images

Figure CN120318757A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of power equipment status monitoring and fault diagnosis, and relates to a method, system, medium and device for identifying infrared images of equipment with cross-modal data. Background Art
[0002] Substations are key components in the power system, responsible for voltage conversion, distribution, and control. With the expansion of the power grid scale and the increase in the complexity of power equipment, higher requirements are put forward for the status monitoring and fault diagnosis of equipment. Infrared image recognition technology has been widely used in the power system due to its characteristics such as fast speed, safety, and effectiveness. However, traditional infrared image recognition methods rely on images captured by single-light infrared cameras, which limit their recognition accuracy and reliability in some cases.
[0003] Existing infrared recognition methods for substation equipment mainly rely on images captured by single-light infrared cameras. These images have low resolution and are difficult to capture the fine features of the equipment, thus affecting the performance of deep convolutional neural networks in detail recognition. In addition, due to significant differences in texture, contrast, and noise between infrared images and visible light images, directly using deep convolutional neural networks designed for visible light images to process infrared images cannot achieve the best results. These problems lead to limitations in the existing technology for high-precision equipment recognition. Summary of the Invention
[0004] In view of the deficiencies of the prior art, the present application provides a method, system, medium and device for identifying infrared images of equipment with cross-modal data, which solves the problem of poor recognition ability of single-light infrared images.
[0005] To achieve the above object, in a first aspect, the present invention provides a method for identifying infrared images of equipment with cross-modal data, including:
[0006] Obtain visible light images and infrared images of each substation equipment;
[0007] Identify each substation equipment in the visible light image according to a preset target detection model to obtain corresponding target recognition frames;
[0008] Map each of the target recognition frames to the infrared image according to a preset affine transformation matrix to obtain effective recognition frames; wherein, the affine transformation matrix is obtained by fitting and solving the feature point matching relationship; the feature point matching relationship is constructed according to the feature points and feature descriptors of the visible light image, and the feature points and feature descriptors of the infrared image;
[0009] Process the region map corresponding to the effective recognition frame in the visible light image according to a preset semantic segmentation model to obtain component mask information of the substation equipment corresponding to each of the effective recognition frames;
[0010] Map each of the component mask information to the infrared image according to the affine transformation matrix, and complete the recognition of the device component mask in the infrared image.
[0011] Compared with the prior art, the embodiments of the present application have the following beneficial effects: obtaining visible light images and infrared images of each power transformation device, making full use of data in different modalities to improve the recognition accuracy; extracting the target recognition frames corresponding to each power transformation device in the visible light image through the model, taking advantage of the high resolution of the visible light image to improve the accuracy of target detection; mapping these target recognition frames to the infrared image through the affine transformation matrix to obtain effective recognition frames, ensuring the accurate conversion from the visible light image to the infrared image and effectively realizing the spatial alignment of data in two modalities; in particular, the affine transformation matrix is obtained by constructing and fitting the feature point matching relationship between the visible light image and the infrared image, making full use of the information of feature points and feature descriptors, avoiding errors caused by image differences or noise interference, and at the same time providing a reliable mathematical basis for the subsequent mapping of component mask information; processing the regional map corresponding to the effective recognition frame through the semantic segmentation model, extracting the component mask information of the power transformation device, accurately recognizing the component details of the device in the visible light image, and mapping them to the infrared image to achieve accurate mask recognition of the device components in the infrared image; this method effectively combines the advantages of visible light and infrared images, significantly improving the recognition accuracy and reliability of power transformation devices and their components.
[0012] In some embodiments of the first aspect of the present application, the feature point matching relationship is constructed according to the feature points and feature descriptors of the visible light image, and the feature points and feature descriptors of the infrared image, including:
[0013] Input the visible light image and the infrared image into a preset feature point extraction model respectively, and extract each first feature point and first feature descriptor of the visible light image, and each second feature point and second feature descriptor of the infrared image;
[0014] Input each of the first feature points, first feature descriptors, second feature points and second feature descriptors into a preset feature point matching model for matching to obtain the original feature point matching relationship;
[0015] According to a preset robust estimation method, screen the feature matching point pairs in the original feature point matching relationship that meet the preset reprojection error range to obtain the feature point matching relationship.
[0016] Compared with the prior art, the above embodiments have the following beneficial effects: By using the feature point extraction model to extract the feature points and feature descriptors of visible light and infrared images respectively, it provides high-quality data support for subsequent matching; With the help of the feature point matching model, the original feature point matching relationship is generated, and the robust estimation method is combined to screen the feature matching point pairs that meet the reprojection error range, effectively removing the wrong matches and improving the reliability of the matching relationship; This series of steps ensures the high-precision construction of the affine transformation matrix, lays a solid foundation for the subsequent mapping of the target recognition frame, and enhances the stability and accuracy of cross-modal data processing.
[0017] In some embodiments of the first aspect of the present application, the affine transformation matrix is obtained by fitting and solving the feature point matching relationship, including:
[0018] According to the preset original affine transformation matrix, a linear equation system is constructed for each feature matching point pair in the feature point matching relationship;
[0019] Solve each of the linear equation systems to obtain the affine transformation matrix.
[0020] Compared with the prior art, the above embodiments have the following beneficial effects: According to the preset original affine transformation matrix, a linear equation system is constructed for each feature matching point pair in the feature point matching relationship, clarifying the specific mathematical form of the affine transformation and ensuring the scientificity and accuracy of matrix solving; By solving these linear equation systems to obtain the final affine transformation matrix, this process not only ensures the high precision of the matrix, but also improves the accuracy of the target recognition frame mapping, providing strong support for cross-modal data processing.
[0021] In some embodiments of the first aspect of the present application, the structure of the original affine transformation matrix is:
[0022] where m 00 and m 11 respectively represent the scaling in the x direction and the y direction, m 01 represents the vertical shear parameter, m 02 and m 12 respectively represent the translation in the x direction and the y direction, m 10 represents the horizontal shear parameter;
[0023] The linear equation system for each feature matching point pair in the feature point matching relationship is:
[0024]
[0025] where and respectively represent the x and y coordinates of the infrared image feature points in the feature matching point pair, and respectively represent the x and y coordinates of the visible light image feature points in the feature matching point pairs.
[0026] Compared with the prior art, the above embodiments have the following beneficial effects: When defining the structure of the original affine transformation matrix, the scaling parameter, the horizontal and vertical shear parameters, and the translation parameter are clarified, so that the matrix can flexibly adapt to the geometric changes between different images, and at the same time provide a clear framework for subsequent calculations; Then, a system of linear equations is constructed for each feature matching point pair in the feature point matching relationship, and the coordinate relationship between the visible light image feature points and the infrared image feature points is transformed into a specific equation form through mathematical modeling, ensuring that the affine transformation matrix can accurately capture the geometric transformation law between the two modalities, improving the interpretability of the affine transformation matrix, and thus providing a reliable theoretical basis for the subsequent mapping of the target recognition frame.
[0027] In some embodiments of the first aspect of the present application, the solving of each of the systems of linear equations to obtain the affine transformation matrix includes:
[0028] Transform each of the systems of linear equations into a matrix equation, expressed as follows: A*X = B; where the design matrix A, the parameter vector X, and the target vector B are respectively represented as:
[0029]
[0030] where n is twice the number of feature matching point pairs;
[0031] According to the least squares method, solve the parameter vector in the matrix equation to obtain the affine transformation matrix; the solution method is:
[0032] X * = arg min X ||A·X - B|| 2 ; where X * is the affine transformation matrix.
[0033] Compared with the prior art, the above embodiments have the following beneficial effects: By transforming the system of linear equations into a matrix form and solving it using the least squares method, this transformation simplifies the calculation process and avoids the computational burden brought by complex optimization algorithms; at the same time, the solution method based on the least squares method has high numerical stability and can still maintain high accuracy under noise interference, thus providing a reliable mathematical guarantee for the subsequent mapping of the target recognition frame and further improving the practicality and efficiency of the entire method.
[0034] In some embodiments of the first aspect of the present application, the mapping of each of the target recognition frames to the infrared image according to the preset affine transformation matrix to obtain the effective recognition frames includes:
[0035] Map each of the target recognition frames into the infrared image according to the affine transformation matrix;
[0036] Calculate the actual coverage area of each of the target recognition frames on the infrared image respectively, and use the target recognition frames whose actual coverage areas meet a preset threshold as the effective recognition frames.
[0037] Compared with the prior art, the above embodiments have the following beneficial effects: Mapping the target recognition frames into the infrared image according to the affine transformation matrix ensures the accurate positioning of the recognition frames and reduces errors; Calculating the actual coverage area of each target recognition frame on the infrared image respectively and setting a threshold to screen the effective recognition frames avoid the influence of invalid or low-quality recognition frames on the final result and improve the effectiveness of recognition.
[0038] In some embodiments of the first aspect of the present application, the processing of the region map corresponding to the effective recognition frame in the visible light image according to a preset semantic segmentation model to obtain the component mask information of the substation equipment corresponding to each effective recognition frame includes:
[0039] Crop the corresponding region map in the visible light image according to each of the effective recognition frames;
[0040] Perform semantic segmentation on each of the region maps according to the semantic segmentation model to obtain the component mask information of the substation equipment corresponding to each effective recognition frame.
[0041] Compared with the prior art, the above embodiments have the following beneficial effects: When processing the region map corresponding to the effective recognition frame, first crop the corresponding region map in the visible light image according to each effective recognition frame, reducing redundant calculations and improving the operation efficiency; Then use the semantic segmentation model to perform semantic segmentation on these region maps to generate the component mask information of the substation equipment, which can accurately extract the component details of the equipment and provide high-quality component information for subsequent analysis.
[0042] In a second aspect, the present invention further provides a device infrared image recognition system for cross-modal data, including: an image acquisition module, a recognition frame extraction module, a first mapping module, a mask extraction module, and a second mapping module;
[0043] Among them, the image acquisition module is used to acquire the visible light image and the infrared image of each substation equipment;
[0044] The recognition frame extraction module is used to recognize each substation equipment in the visible light image according to a preset target detection model to obtain the corresponding target recognition frame;
[0045] The first mapping module is configured to map each of the target recognition frames to the infrared image according to a preset affine transformation matrix to obtain effective recognition frames; wherein, the affine transformation matrix is obtained by fitting and solving the feature point matching relationship; the feature point matching relationship is constructed according to the feature points and feature descriptors of the visible light image and the feature points and feature descriptors of the infrared image;
[0046] The mask extraction module is configured to process the region map corresponding to the effective recognition frame in the visible light image according to a preset semantic segmentation model to obtain the component mask information of the substation equipment corresponding to each of the effective recognition frames;
[0047] The second mapping module is configured to map each of the component mask information to the infrared image according to the affine transformation matrix to complete the recognition of the equipment component mask in the infrared image.
[0048] Compared with the prior art, the embodiments of the present application have the following beneficial effects: obtaining the visible light image and the infrared image of each substation equipment, making full use of data of different modalities to improve the recognition accuracy; extracting the target recognition frames corresponding to each substation equipment in the visible light image through the model, taking advantage of the high resolution of the visible light image to improve the accuracy of target detection; mapping these target recognition frames to the infrared image through the affine transformation matrix to obtain effective recognition frames, ensuring the accurate conversion from the visible light image to the infrared image and effectively realizing the spatial alignment of data of the two modalities; in particular, the affine transformation matrix is obtained by constructing and fitting the feature point matching relationship between the visible light image and the infrared image, making full use of the information of the feature points and feature descriptors, avoiding errors caused by image differences or noise interference, and at the same time providing a reliable mathematical basis for the subsequent mapping of the component mask information; processing the region map corresponding to the effective recognition frame through the semantic segmentation model to extract the component mask information of the substation equipment, accurately recognizing the component details of the equipment in the visible light image and mapping them to the infrared image to achieve the accurate mask recognition of the equipment components in the infrared image; this method effectively combines the advantages of the visible light and infrared images, significantly improving the recognition accuracy and reliability of the substation equipment and its components.
[0049] In a third aspect, the present invention further provides a device for recognizing the infrared image of equipment with cross-modal data, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and the steps of the method for recognizing the infrared image of equipment with cross-modal data are implemented when the computer program is loaded into the processor.
[0050] Fourthly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the infrared image recognition method for cross-modal data equipment are implemented. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 : A schematic flowchart of an infrared image recognition method for cross-modal data equipment provided in some embodiments of the present invention.
[0052] Figure 2 : A schematic structural diagram of an infrared image recognition system for cross-modal data equipment provided in some embodiments of the present invention.
[0053] Figure 3 : A structural diagram of an infrared image recognition device for cross-modal data equipment provided in some embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0054] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] Embodiment 1:
[0056] Please refer to Figure 1 , an infrared image recognition method for cross-modal data equipment provided in an embodiment of the present invention, including steps S1 to S5:
[0057] Step S1: Obtain visible light images and infrared images of each substation equipment.
[0058] In this embodiment, in step S1, visible light images and infrared images of each substation equipment are obtained, making full use of data in different modalities to provide a data basis for subsequent recognition and analysis.
[0059] Step S2: Identify each substation equipment in the visible light image according to a preset target detection model to obtain corresponding target recognition frames.
[0060] Compared with infrared images, visible light images often have higher resolution and more texture information. When the model is used to identify and extract target detection frames, it is easier to accurately identify and locate substation equipment and generate high-precision target recognition frames. In specific implementation, the target detection model can be a YOLOv8 target detection model pre-trained with historical manually annotated data, or other similar target detection models, which are not limited herein.
[0061] In this embodiment, in step S2, the target recognition frames corresponding to each substation device in the visible light image are extracted by the model, taking advantage of the high resolution of the visible light image to improve the accuracy of target detection.
[0062] Step S3: Map each of the target recognition frames to the infrared image according to a preset affine transformation matrix to obtain effective recognition frames.
[0063] Preferably, the affine transformation matrix is obtained by fitting and solving the feature point matching relationship; the feature point matching relationship is constructed based on the feature points and feature descriptors of the visible light image, and the feature points and feature descriptors of the infrared image.
[0064] Specifically, the feature point matching relationship can be obtained through the following preferred implementation manner, including steps S31 - S33, as follows:
[0065] S31: Input the visible light image and the infrared image into a preset feature point extraction model respectively to extract each first feature point and first feature descriptor of the visible light image, and each second feature point and second feature descriptor of the infrared image.
[0066] In specific implementation, the feature point extraction model can be the SuperPoint model, and the extraction results can be expressed as: the first feature point P vis 、the first feature descriptor D vis 、the second feature point P inf and the second feature descriptor D vis ; where the feature point P = (x, y, c) represents the feature point with pixel coordinates (x, y), c is the confidence level, and the descriptor D is a 256 - dimensional feature vector; in addition, other similar feature point extraction models can also be used, which are not limited here.
[0067] S32: Input each of the first feature points, first feature descriptors, second feature points and second feature descriptors into a preset feature point matching model for matching to obtain the original feature point matching relationship.
[0068] In specific implementation, step S32 can be implemented by using the SuperGlue model. After inputting the above four kinds of data into the SuperGlue model, the output original feature point matching relationship can be expressed as:
[0069] where and In the matching relationship, it is the two-dimensional feature point set of the visible light image and the two-dimensional feature point set of the infrared image, and the two are in one-to-one correspondence and matching. In addition, the model used in this step can also be implemented using other similar matching models, and there is no restriction either.
[0070] S33: According to the preset robust estimation method, screen the feature matching point pairs that meet the preset reprojection error range in the original feature point matching relationship to obtain the feature point matching relationship.
[0071] In specific implementation, the RANSAC algorithm can be used as a robust estimation method, and the operation is as follows: First, randomly select the minimum number of sample points (such as 4 pairs of non-collinear points) to calculate the initial homography transformation matrix, then use this matrix to project all points and calculate the reprojection error, mark the points with error less than the set threshold as inliers, repeat the above process until the maximum number of iterations is reached, and record the situation with the largest number of inliers in each iteration; finally, output the optimal homography transformation matrix and the corresponding inlier set, that is, the feature point matching relationship. Among them, the threshold can be set to 3, the maximum number of iterations can be set to 2000, and the reprojection error can be calculated using the Euclidean distance.
[0072] In this preferred embodiment, steps S31 - S33 respectively extract the feature points and feature descriptors of the visible light and infrared images through the feature point extraction model, providing high-quality data support for subsequent matching; generate the original feature point matching relationship with the help of the feature point matching model, and combine the robust estimation method to screen the feature matching point pairs that meet the reprojection error range, effectively removing the wrong matches and improving the reliability of the matching relationship; this series of steps ensures the high-precision construction of the affine transformation matrix, laying a solid foundation for the subsequent mapping of the target recognition frame, and enhancing the stability and accuracy of cross-modal data processing.
[0073] Further, the affine transformation matrix can be obtained through the following preferred implementation, including steps S34 - S35, specifically as follows:
[0074] S34: According to the preset original affine transformation matrix, construct a linear equation system for each feature matching point pair in the feature point matching relationship.
[0075] Further, the structure of the original affine transformation matrix is:
[0076] where, m 00 and m 11 respectively represent the scaling in the x direction and the y direction. If it is greater than 1, it means a magnification operation is performed. If it is less than 1, it means a reduction operation is performed. If it is equal to 1, it means no scaling operation is performed. m 01 represents the vertical shear parameter, m 02 and m12 represent translations in the x and y directions respectively, and m 10 represents the horizontal shearing parameter.
[0077] The system of linear equations for each feature matching point pair in the feature point matching relationship is:
[0078]
[0079] where and represent the x and y coordinates of the infrared image feature points in the feature matching point pair respectively, and represent the x and y coordinates of the visible light image feature points in the feature matching point pair respectively.
[0080] In this preferred embodiment, when defining the structure of the original affine transformation matrix in step S34, the scaling parameter, the horizontal and vertical shearing parameters, and the translation parameters are specified, enabling the matrix to flexibly adapt to geometric changes between different images and providing a clear framework for subsequent calculations. Then, a system of linear equations is constructed for each feature matching point pair in the feature point matching relationship, and the coordinate relationship between the visible light image feature points and the infrared image feature points is transformed into a specific equation form through mathematical modeling, ensuring that the affine transformation matrix can accurately capture the geometric transformation law between the two modalities, improving the interpretability of the affine transformation matrix, and thus providing a reliable theoretical basis for subsequent target recognition box mapping.
[0081] S35: Solve each of the systems of linear equations to obtain the affine transformation matrix.
[0082] Further, step S35 can be implemented through the following preferred implementation manner, including steps S351 - S352: Specifically as follows:
[0083] S351: Transform each of the systems of linear equations into a matrix equation, expressed as: A*X = B; where the design matrix A, the parameter vector X, and the target vector B are respectively represented as:
[0084]
[0085] where n is twice the number of feature matching point pairs;
[0086] S352: Solve the parameter vector in the matrix equation according to the least squares method to obtain the affine transformation matrix; where the solution method is: X * = arg min X ||A·X - B|| 2 ; where X * is the affine transformation matrix.
[0087] In this preferred embodiment, steps S351 - S352 are carried out by transforming the system of linear equations into matrix form and using the least squares method for solution. This transformation simplifies the calculation process and avoids the computational burden brought by complex optimization algorithms. At the same time, the solution method based on the least squares method has high numerical stability and can still maintain high accuracy under noise interference, thus providing a reliable mathematical guarantee for the subsequent mapping of the target recognition frame and further enhancing the practicability and efficiency of the entire method.
[0088] In summary, in this preferred embodiment, steps S34 - S35 construct a system of linear equations for each feature matching point pair in the feature point matching relationship according to the preset original affine transformation matrix, clarify the specific mathematical form of the affine transformation, and ensure the scientificity and accuracy of matrix solution. By solving these systems of linear equations, the final affine transformation matrix is obtained. This process not only ensures the high accuracy of the matrix but also improves the accuracy of target recognition frame mapping, providing strong support for cross-modal data processing.
[0089] Furthermore, step S3 can be implemented through the following preferred implementation manner, including steps S36 - S37, specifically as follows:
[0090] S36: Map each of the target recognition frames to the infrared image according to the affine transformation matrix.
[0091] S37: Calculate the actual coverage area of each of the target recognition frames on the infrared image respectively, and regard the target recognition frames whose actual coverage areas meet the preset threshold as the effective recognition frames.
[0092] In specific implementation, after mapping, calculate the actual coverage area of each target recognition frame on the infrared image. It is possible to only retain the target recognition frames whose actual coverage area exceeds half of their own complete area as the effective recognition frames, so as to avoid invalid or low-quality recognition frames from affecting the final result.
[0093] In this preferred embodiment, steps S36 - S37 map the target recognition frames to the infrared image according to the affine transformation matrix. This process ensures the precise positioning of the recognition frames and reduces errors. Calculate the actual coverage area of each target recognition frame on the infrared image respectively, and set a threshold to screen the effective recognition frames, avoiding invalid or low-quality recognition frames from affecting the final result and improving the effectiveness of recognition.
[0094] Step S4: Process the regional map corresponding to the effective recognition frame in the visible light image according to the preset semantic segmentation model to obtain the component mask information of the substation equipment corresponding to each of the effective recognition frames.
[0095] Preferably, step S4 can be implemented by the following preferred embodiments, including steps S41 - S42, specifically as follows:
[0096] S41: Crop the corresponding regional map in the visible light image according to each of the effective recognition frames;
[0097] S42: Perform semantic segmentation on each of the regional maps according to the semantic segmentation model to obtain the component mask information of the substation equipment corresponding to each of the effective recognition frames.
[0098] In specific implementation, the semantic segmentation model can be the YOLOv8 semantic segmentation model pre-trained with historical manually annotated data, or other similar semantic segmentation models, which are not limited herein.
[0099] In this preferred embodiment, when steps S41 - S42 process the regional maps corresponding to the effective recognition frames, first crop the corresponding regional maps in the visible light image according to each effective recognition frame, reducing redundant calculations and improving the operation efficiency; then use the semantic segmentation model to perform semantic segmentation on these regional maps to generate the component mask information of the substation equipment, which can accurately extract the component details of the equipment and provide high-quality component information for subsequent analysis.
[0100] Step S5: Map each of the component mask information to the infrared image according to the affine transformation matrix to complete the recognition of the equipment component mask in the infrared image.
[0101] In this embodiment, step S5 realizes the accurate mask recognition of the equipment components in the infrared image by mapping the component mask information into the infrared image.
[0102] In summary, compared with the prior art, the above embodiments of the present application have the following beneficial effects: obtaining visible light images and infrared images of each power transformation device, making full use of data in different modalities to improve the recognition accuracy; extracting the target recognition frames corresponding to each power transformation device in the visible light image through the model, taking advantage of the high resolution of the visible light image to improve the accuracy of target detection; mapping these target recognition frames into the infrared image through the affine transformation matrix to obtain effective recognition frames, ensuring accurate conversion from the visible light image to the infrared image and effectively realizing spatial alignment of data in two modalities; in particular, the affine transformation matrix is obtained by constructing and fitting the feature point matching relationship between the visible light image and the infrared image, making full use of the information of feature points and feature descriptors, avoiding errors caused by image differences or noise interference, and at the same time providing a reliable mathematical basis for the mapping of subsequent component mask information; processing the regional map corresponding to the effective recognition frame through the semantic segmentation model to extract the component mask information of the power transformation device, accurately identifying the component details of the device in the visible light image and mapping them into the infrared image, realizing accurate mask recognition of the device components in the infrared image; this method effectively combines the advantages of visible light and infrared images, significantly improving the recognition accuracy and reliability of power transformation devices and their components.
[0103] Embodiment 2:
[0104] Please refer to Figure 2 , based on the same inventive concept, an infrared image recognition system for devices with cross-modal data disclosed in an embodiment of the present invention includes: an image acquisition module M1, a recognition frame extraction module M2, a first mapping module M3, a mask extraction module M4, and a second mapping module M5;
[0105] Among them, the image acquisition module M1 is used to obtain visible light images and infrared images of each power transformation device.
[0106] In this embodiment, the image acquisition module M1 obtains visible light images and infrared images of each power transformation device, making full use of data in different modalities to provide a data basis for subsequent recognition and analysis.
[0107] The recognition frame extraction module M2 is used to identify each power transformation device in the visible light image according to a preset target detection model to obtain corresponding target recognition frames.
[0108] In this embodiment, the recognition frame extraction module M2 extracts the target recognition frames corresponding to each power transformation device in the visible light image through the model, taking advantage of the high resolution of the visible light image to improve the accuracy of target detection.
[0109] The first mapping module M3 is configured to map each of the target recognition frames into the infrared image according to a preset affine transformation matrix to obtain effective recognition frames. The affine transformation matrix is obtained by fitting and solving the feature point matching relationship, and the feature point matching relationship is constructed based on the feature points and feature descriptors of the visible light image and the feature points and feature descriptors of the infrared image.
[0110] Further, the first mapping module M3 includes: a feature extraction unit, a matching unit, and a screening unit.
[0111] The feature extraction unit is configured to input the visible light image and the infrared image into a preset feature point extraction model respectively, and extract each first feature point and first feature descriptor of the visible light image, and each second feature point and second feature descriptor of the infrared image.
[0112] The matching unit is configured to input each of the first feature points, first feature descriptors, second feature points, and second feature descriptors into a preset feature point matching model for matching to obtain an original feature point matching relationship.
[0113] The screening unit is configured to screen, according to a preset robust estimation method, the feature matching point pairs that meet the preset reprojection error range in the original feature point matching relationship to obtain a feature point matching relationship.
[0114] In this preferred embodiment, the first mapping module M3 extracts the feature points and feature descriptors of the visible light and infrared images respectively through the feature point extraction model, providing high-quality data support for subsequent matching. With the help of the feature point matching model, the original feature point matching relationship is generated, and combined with the robust estimation method, the feature matching point pairs that meet the reprojection error range are screened, effectively removing the wrong matches and improving the reliability of the matching relationship. This series of steps ensures the high-precision construction of the affine transformation matrix, laying a solid foundation for the subsequent mapping of the target recognition frame and enhancing the stability and accuracy of cross-modal data processing.
[0115] Further, the first mapping module M3 further includes: an equation construction unit and a solving unit.
[0116] The equation construction unit is configured to construct a linear equation system for each feature matching point pair in the feature point matching relationship according to a preset original affine transformation matrix.
[0117] Further, the structure of the original affine transformation matrix is:
[0118] where m 00 and m 11 respectively represent the scaling in the x direction and the y direction, m01 Represents the vertical shear parameter, m 02 and m 12 represent the translations in the x - direction and y - direction respectively, m 10 Represents the horizontal shear parameter;
[0119] The system of linear equations for each feature matching point pair in the feature point matching relationship is:
[0120]
[0121] where and represent the x - and y - coordinates of the infrared image feature points in the feature matching point pair respectively, and represent the x - and y - coordinates of the visible light image feature points in the feature matching point pair respectively.
[0122] In this preferred embodiment, when defining the structure of the original affine transformation matrix, the equation construction unit clarifies the scaling parameter, horizontal and vertical shear parameters, and translation parameters, enabling the matrix to flexibly adapt to geometric changes between different images and providing a clear framework for subsequent calculations; then constructs a system of linear equations for each feature matching point pair in the feature point matching relationship, transforming the coordinate relationship between the visible light image feature points and the infrared image feature points into a specific equation form through mathematical modeling, ensuring that the affine transformation matrix can accurately capture the geometric transformation rules between the two modalities, improving the interpretability of the affine transformation matrix, and thus providing a reliable theoretical basis for subsequent target recognition box mapping.
[0123] The solving unit is used to solve each system of linear equations to obtain the affine transformation matrix.
[0124] Furthermore, the solving unit includes: a transformation sub - unit and a solving sub - unit;
[0125] wherein, the transformation sub - unit is used to transform each system of linear equations into a matrix equation, expressed as: A*X = B; where the design matrix A, parameter vector X, and target vector B are respectively represented as:
[0126]
[0127] where n is twice the number of feature matching point pairs;
[0128] The solving sub - unit is used to solve the parameter vector in the matrix equation according to the least - squares method to obtain the affine transformation matrix; the solving method is:
[0129] X * = arg minX ||A·X - B|| 2 ; where X * is an affine transformation matrix.
[0130] In this preferred embodiment, the solving unit simplifies the calculation process by transforming the linear equations into matrix form and using the least squares method for solving, avoiding the computational burden brought by complex optimization algorithms. At the same time, the solving method based on the least squares method has high numerical stability and can maintain high accuracy under noise interference, thus providing a reliable mathematical guarantee for the subsequent mapping of the target recognition frame and further improving the practicality and efficiency of the entire method.
[0131] Furthermore, the first mapping module M3 further includes: a first mapping unit and a filtering unit;
[0132] Among them, the first mapping unit is used to map each of the target recognition frames to the infrared image according to the affine transformation matrix;
[0133] The filtering unit is used to calculate the actual coverage area of each of the target recognition frames on the infrared image respectively, and regard the target recognition frames whose actual coverage areas meet the preset threshold as the effective recognition frames.
[0134] In this preferred embodiment, the first mapping module M3 maps the target recognition frames to the infrared image according to the affine transformation matrix, which ensures the accurate positioning of the recognition frames and reduces errors. Calculate the actual coverage area of each target recognition frame on the infrared image respectively, and set a threshold to screen the effective recognition frames, avoiding the influence of invalid or low-quality recognition frames on the final result and improving the effectiveness of recognition.
[0135] The mask extraction module M4 is used to process the regional map corresponding to the effective recognition frame in the visible light image according to a preset semantic segmentation model, and obtain the component mask information of the substation equipment corresponding to each of the effective recognition frames.
[0136] Furthermore, the mask extraction module M4 includes: a cropping unit and a semantic segmentation unit;
[0137] Among them, the cropping unit is used to crop the regional map corresponding to the visible light image according to each of the effective recognition frames.
[0138] The semantic segmentation unit is used to perform semantic segmentation on each of the regional maps according to the semantic segmentation model, and obtain the component mask information of the substation equipment corresponding to each of the effective recognition frames.
[0139] In this preferred embodiment, when the mask extraction module M4 processes the regional map corresponding to the effective recognition box, it first crops the corresponding regional map in the visible light image according to each effective recognition box, reducing redundant calculations and improving the operation efficiency. Then, it uses a semantic segmentation model to perform semantic segmentation on these regional maps to generate component mask information of the substation equipment, which can accurately extract the component details of the equipment and provide high-quality component information for subsequent analysis.
[0140] The second mapping module M5 is configured to map each of the component mask information to the infrared image according to the affine transformation matrix, so as to complete the recognition of the equipment component mask in the infrared image.
[0141] In this embodiment, the second mapping module M5 realizes the accurate mask recognition of the equipment components in the infrared image by mapping the component mask information into the infrared image.
[0142] In summary, compared with the prior art, the embodiments of the present application have the following beneficial effects: obtaining the visible light image and the infrared image of each substation equipment, making full use of data in different modalities to improve the recognition accuracy; extracting the target recognition box corresponding to each substation equipment in the visible light image through the model, taking advantage of the high resolution of the visible light image to improve the accuracy of target detection; mapping these target recognition boxes to the infrared image through the affine transformation matrix to obtain effective recognition boxes, ensuring the accurate conversion from the visible light image to the infrared image and effectively realizing the spatial alignment of the two-modal data; in particular, the affine transformation matrix is obtained by constructing and fitting the feature point matching relationship between the visible light image and the infrared image, making full use of the information of the feature points and feature descriptors, avoiding errors caused by image differences or noise interference, and at the same time providing a reliable mathematical basis for the subsequent mapping of the component mask information; processing the regional map corresponding to the effective recognition box through the semantic segmentation model, extracting the component mask information of the substation equipment, accurately identifying the component details of the equipment in the visible light image, and mapping them to the infrared image, realizing the accurate mask recognition of the equipment components in the infrared image; this method effectively combines the advantages of the visible light and infrared images, significantly improving the recognition accuracy and reliability of the substation equipment and its components.
[0143] The division of the above-described modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system.
[0144] Embodiment 3:
[0145] Figure 3 The structural diagram of an equipment infrared image recognition device for cross-modal data of the present application is shown. As Figure 3As shown in the figure, the device infrared image recognition device for cross-modal data may include: a processor N1, a memory N2, a data interface N3, and a communication bus N4.
[0146] Among them: the processor N1, the memory N2, and the data interface N3 communicate with each other through the communication bus N4; the data interface N3 is used for data communication with other devices such as input devices or output devices; the processor N1 is used to execute the program N5, and specifically can execute the relevant steps in the above-mentioned embodiments of the method for device infrared image recognition of cross-modal data.
[0147] Specifically, the program N5 may include program code, and the program code includes computer-executable instructions.
[0148] The processor N1 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the device infrared image recognition device for cross-modal data may be of the same type of processor, such as one or more CPUs, or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0149] The memory N2 is used to store the program N5. The memory N2 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.
[0150] The algorithms or displays provided herein are not inherently related to any specific computer, virtual system, or other device. In addition, the embodiments of the present application are not directed to any specific programming language.
[0151] Embodiment 4:
[0152] The embodiments of the present invention also provide a computer-readable storage medium. The storage medium stores at least one executable instruction. When the executable instruction runs on the device infrared image recognition device / system for cross-modal data, the device infrared image recognition device / system for cross-modal data is enabled to execute the method for device infrared image recognition of cross-modal data in any of the above method embodiments.
[0153] In the specification provided herein, a large number of specific details are set forth. However, it will be understood that embodiments of the present application may be practiced without these specific details. Similarly, in order to streamline the present application and assist in understanding one or more of the various inventive aspects, in the above description of exemplary embodiments of the present application, the various features of the embodiments of the present application are sometimes grouped together in a single embodiment, figure, or description thereof. Among them, the claims following the specific implementation manner are hereby expressly incorporated into the specific implementation manner, and each claim itself serves as a separate embodiment of the present application.
[0154] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive.
Claims
1. An infrared image recognition method for cross-modal data of a device, characterized in that, Including: Obtain visible light images and infrared images of each substation equipment; According to a preset target detection model, identify each substation equipment in the visible light image to obtain corresponding target recognition frames; According to a preset affine transformation matrix, map each of the target recognition frames to the infrared image to obtain effective recognition frames; wherein, the affine transformation matrix is obtained by fitting and solving the feature point matching relationship; the feature point matching relationship is constructed based on the feature points and feature descriptors of the visible light image, and the feature points and feature descriptors of the infrared image; According to a preset semantic segmentation model, process the region map corresponding to the effective recognition frame in the visible light image to obtain component mask information of the substation equipment corresponding to each effective recognition frame; According to the affine transformation matrix, map each of the component mask information to the infrared image to complete the recognition of the equipment component mask in the infrared image.
2. The infrared image recognition method for cross-modal data devices according to claim 1, characterized in that, The feature point matching relationship is constructed based on the feature points and feature descriptors of the visible light image, and the feature points and feature descriptors of the infrared image, including: Input the visible light image and the infrared image into a preset feature point extraction model respectively, and extract each first feature point and first feature descriptor of the visible light image, and each second feature point and second feature descriptor of the infrared image; Input each of the first feature points, first feature descriptors, second feature points and second feature descriptors into a preset feature point matching model for matching to obtain an original feature point matching relationship; According to a preset robust estimation method, screen the feature matching point pairs in the original feature point matching relationship that meet the preset reprojection error range to obtain the feature point matching relationship.
3. The infrared image recognition method for cross-modal data devices according to claim 2, characterized in that, The affine transformation matrix is obtained by fitting and solving the feature point matching relationship, including: According to a preset original affine transformation matrix, construct a linear equation system for each feature matching point pair in the feature point matching relationship; Solve each of the linear equation systems to obtain the affine transformation matrix.
4. The infrared image recognition method of cross-modal data device according to claim 3, characterized in that, The structure of the original affine transformation matrix is: Among them, m 00 and m 11 respectively represent the scaling in the x - direction and the y - direction, m 01 represents the vertical shear parameter, m 02 and m 12 respectively represent the translation in the x - direction and the y - direction, m 10 represents the horizontal shear parameter; The linear equation system of each feature matching point pair in the feature point matching relationship is: wherein and respectively represent the x and y coordinates of the infrared image feature points in the feature matching point pairs, and respectively represent the x and y coordinates of the visible light image feature points in the feature matching point pairs.
5. The infrared image recognition method of a cross-modal data device according to claim 3, wherein, The step of solving each of the linear equation systems to obtain the affine transformation matrix includes: Convert each of the linear equation systems into a matrix equation, expressed as: A*X = B; wherein, the design matrix A, the parameter vector X and the target vector B are respectively represented as: where n is twice the number of feature matching point pairs; According to the least squares method, solve the parameter vector in the matrix equation to obtain the affine transformation matrix; wherein the solution method is: X * = arg min X ||A·X - B|| 2 ; where X * is an affine transformation matrix.
6. A method for infrared image recognition of a cross-modal data device according to any one of claims 1 to 5, characterized in that, The step of mapping each of the target recognition frames to the infrared image according to a preset affine transformation matrix to obtain effective recognition frames includes: Map each of the target recognition frames to the infrared image according to the affine transformation matrix; Calculate the actual coverage area of each target recognition frame on the infrared image respectively, and use the target recognition frames with each actual coverage area meeting the preset threshold as the effective recognition frames.
7. The infrared image recognition method for cross-modal data device according to claim 6, wherein The step of processing the region map corresponding to the effective recognition frame in the visible light image according to a preset semantic segmentation model to obtain component mask information of the substation equipment corresponding to each effective recognition frame includes: Crop the corresponding regional images in the visible light image according to each of the effective recognition frames; Perform semantic segmentation on each of the regional images according to the semantic segmentation model to obtain the component mask information of the substation equipment corresponding to each of the effective recognition frames.
8. An infrared image recognition system for cross-modal data of a device, characterized in that It includes: An image acquisition module, a recognition frame extraction module, a first mapping module, a mask extraction module, and a second mapping module; Among them, the image acquisition module is used to acquire the visible light images and infrared images of each substation equipment; The recognition frame extraction module is used to recognize each substation equipment in the visible light image according to a preset target detection model to obtain corresponding target recognition frames; The first mapping module is used to map each of the target recognition frames to the infrared image according to a preset affine transformation matrix to obtain effective recognition frames; wherein, the affine transformation matrix is obtained by fitting and solving the feature point matching relationship; the feature point matching relationship is constructed according to the feature points and feature descriptors of the visible light image and the feature points and feature descriptors of the infrared image; The mask extraction module is used to process the regional images corresponding to the effective recognition frames in the visible light image according to a preset semantic segmentation model to obtain the component mask information of the substation equipment corresponding to each of the effective recognition frames; The second mapping module is used to map each of the component mask information to the infrared image according to the affine transformation matrix to complete the recognition of the equipment component masks in the infrared image.
9. An infrared image recognition device for cross-modal data, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is loaded into the processor, the steps of a method for identifying the infrared image of equipment with cross-modal data according to any one of claims 1-7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of a method for identifying the infrared image of equipment with cross-modal data according to any one of claims 1-7 are implemented.