A multi-modal spatio-temporal frequency visual fusion reconstruction method based on acoustic image-infrared-visible light
Patent Information
- Application Number
- CN202410861627.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-28
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2044-06-28
AI Technical Summary
[0002]在电力领域,现代电力系统电力设备种类多,应用环境复杂,设备状态信息量庞大,在设备实际运行过程中,由于工作环境复杂、背景噪音干扰大等问题,缺陷信号易被淹没难以发现,大大增加了运维难度,传统的故障识别方法难以满足电力设备监测要求,因此必须结合基于智能信息融合处理的方法进行研究
[0027] 1. This invention, through the fusion of multimodal data, can more comprehensively and accurately identify equipment status, thereby promoting the improvement of fault target detection and processing capabilities from the source, and achieving a more comprehensive understanding and prevention of fault formation and development.
Smart Images

Figure CN118918422B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal fusion of power equipment, and in particular to a multimodal spatiotemporal-frequency visual fusion reconstruction technology based on acoustic-infrared-visible light. Background Technology
[0002] In the power sector, modern power systems have a wide variety of power equipment, complex application environments, and a massive amount of equipment status information. During actual operation, due to the complex working environment and strong background noise interference, defect signals are easily submerged and difficult to detect, greatly increasing the difficulty of operation and maintenance. Traditional fault identification methods are insufficient to meet the monitoring requirements of power equipment. Therefore, it is necessary to combine research with methods based on intelligent information fusion processing.
[0003] With advancements in various sensing and imaging technologies, digital image processing technologies, and other related technologies, intelligent analysis technologies based on multimodal fusion data have been increasingly widely applied in power equipment fault diagnosis in recent years. Common applications include the combined use of visible and infrared imaging to detect temperature anomalies, the combined use of visible and ultraviolet spectral bands to detect the external corona discharge status of high-voltage electrical equipment, and increasingly mature acoustic imaging technologies for detecting discharge, abnormal noises, and vibrations, which are playing a vital role in various scenarios within the power grid.
[0004] In the field of intelligent power data analysis, the fusion of multimodal data has become a key technology, with the fusion method of visible light, infrared, and acoustic data receiving widespread attention. A track-mounted, mobile, real-time monitoring scheme using visible light and infrared sensors can dynamically perceive the operating status of power equipment components. Combined with acoustic imaging monitoring technology, the spatiotemporal and frequency-based fusion analysis effectively filters out background interference, forming a comprehensive perception of the overall situation. Comparing the synchronous detection of different signals reduces the problems of inaccurate detection or decreased sensitivity caused by interference with single-spectral-band signals. Furthermore, considering the dynamic changes in the development of electrical equipment faults, the spatiotemporal and frequency trends of power equipment status data can accurately reflect the fault development status, which is of great significance for comprehensive fault diagnosis and early warning. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the existing technology and provide a multimodal spatiotemporal-frequency visual fusion reconstruction method based on sound-image-infrared-visible light.
[0006] The objective of this invention can be achieved through the following technical solutions:
[0007] This invention provides a multimodal spatiotemporal-frequency visual fusion reconstruction method based on acoustic-infrared-visible light, comprising the following steps:
[0008] S1. Select a visible light camera as the reference imaging camera, calibrate the parameters of the visible light camera, infrared camera, and audio-visual device, as well as the pose parameters between the cameras, and perform the first mapping of the infrared camera and audio-visual device to the visible light camera.
[0009] S2. Project each integer pixel coordinate in the infrared image, visible light image and acoustic image forward to the corresponding 3D coordinate, and perform a second mapping of the 3D coordinates of the infrared camera and the audio-visual device to the 3D coordinates of the visible light camera.
[0010] S3. Using the 3D coordinates of the visible light camera as world coordinates, the world coordinates are projected backwards onto the infrared camera and the audio-visual device to obtain the pixel coordinates of the world coordinates projected backwards onto the infrared camera and the audio-visual device.
[0011] S4. Using the forward projection in step S2 and the reverse projection in step S3 of the infrared camera and the audio-visual instrument, the third mapping from the gray values of the pixels of the original infrared image to the gray values of the pixels of the visible light image and the third mapping from the RGB values of the pixels of the original acoustic image to the RGB values of the pixels of the visible light image are obtained, and the registered infrared image and the registered acoustic image are generated respectively.
[0012] S5. The original visible light image is fused with the registered infrared image and the registered acoustic image from step S4 to obtain a multimodal registration and reconstruction image.
[0013] The parameters used to calibrate the visible light camera, infrared camera, and audio-visual instrument include the internal and external parameters for calibrating the visible light camera, infrared camera, and audio-visual instrument.
[0014] The intrinsic parameters of a visible light camera are determined by the camera's horizontal focal length, vertical focal length, horizontal offset, and vertical offset.
[0015] The intrinsic parameters of the optical camera are expressed by the following formula:
[0016]
[0017] Among them, f x f is the horizontal focal length of a visible light camera. y c is the vertical focal length of the visible light camera. x c is the horizontal offset of the visible light camera's optical axis in the image coordinate system, expressed in pixels. y This represents the vertical offset of the visible light camera's optical axis in the image coordinate system, expressed in pixels.
[0018] The optical camera extrinsic parameters include the pose parameters between the cameras.
[0019] The pose parameters between the cameras are determined by the rotation matrix vector and the horizontal matrix vector.
[0020] The pose parameters between the cameras are expressed by the following formula:
[0021]
[0022] in, c M0 represents the pose parameters between cameras, r ij (i, j = 1, 2, 3) are vectors of the rotation matrix, t x t x t x These are vectors representing the translation matrix.
[0023] Convert the pixels in the infrared camera and audio-visual equipment to the coordinates corresponding to the pixels in the visible light camera image.
[0024] Preferably, the pixels in the infrared camera and the audio-visual unit are converted to the coordinates corresponding to the pixels in the visible light camera image through matrix multiplication.
[0025] Secondly, the present invention provides a multimodal spatiotemporal-frequency visual fusion reconstruction system based on audio-visual-infrared-visible light, comprising: a visible light camera, an infrared camera, an audio-visual unit, a mapping module, a projection module, a registration module, and an image fusion module, wherein the system is used to implement any of the methods described above.
[0026] Compared with the prior art, the present invention has the following beneficial effects:
[0027] 1. This invention, through the fusion of multimodal data, can more comprehensively and accurately identify equipment status, thereby promoting the improvement of fault target detection and processing capabilities from the source, and achieving a more comprehensive understanding and prevention of fault formation and development.
[0028] 2. This invention has significant application potential in the field of power equipment fault detection, and solves the problems of low data dimensionality, weak cross-domain analysis and dynamic evolution capabilities in equipment status perception and acquisition, thereby effectively improving the accuracy and reliability of power equipment fault detection.
[0029] 3. This invention adopts key technologies for visual cross-modal fusion analysis, and provides a prototype deployment strategy for future intelligent sensing data fusion analysis combining visible light, infrared and acoustic images, providing dataset collection support for further unmanned intelligent power fault location.
[0030] 4. The present invention designs a multimodal spatiotemporal-frequency visual fusion reconstruction system based on sound-image-infrared-visible light. By sensing the relative position of the terminal in space, the system projects the position onto the world coordinate system of the three-dimensional model of the equipment, performs coordinate transformation, target scaling and offset processing, and then performs multimodal data fusion. This solves the problems of low data dimensionality, weak cross-domain analysis capability, and weak dynamic evolution capability in the state perception of power grid equipment. Attached Figure Description
[0031] Figure 1 Registration process for back projection transformation algorithms in existing technologies;
[0032] Figure 2 This invention provides a registration process based on a multimodal spatiotemporal-frequency visual fusion reconstruction technology using acoustic-infrared-visible light.
[0033] Figure 3 The intrinsic parameters, distortion coefficients, and pose parameters between cameras correspond to the three modes of sound-image, infrared, and visible light.
[0034] Figure 4 The registration result is for the fusion and reconstruction of three types of images: audio-visual, infrared, and visible light. Detailed Implementation
[0035] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0036] This embodiment provides a multimodal spatiotemporal-frequency visual fusion reconstruction technology based on acoustic-image, infrared, and visible light. For the fusion reconstruction of these three modalities, a visible light camera is used as the reference imaging camera, and its corresponding visible light RGB image is used as the reference image. An infrared camera and an acoustic-image unit are used as the cameras to be registered. The specific implementation is as follows: Figure 2 As shown, it includes the following steps:
[0037] Step 1: Pre-calibration.
[0038] A visible light camera is selected as the reference imaging camera, and the intrinsic parameters, distortion coefficients, and pose parameters between the visible light camera, infrared camera, and audio-visual unit are obtained (mapping 1).
[0039] The output is a rotation vector and a translation vector.
[0040] The relationships between the intrinsic parameters of the camera modules, distortion coefficients, and pose parameters between the cameras for the three modes of sound-image, infrared, and visible light are expressed by the following formula:
[0041]
[0042] Where s is the distortion factor, u is the u-axis coordinate of the visible light camera pixel, v is the v-axis coordinate of the visible light camera pixel, and f is the distortion factor. x f is the horizontal focal length of a visible light camera. y c is the vertical focal length of a visible light camera. xc represents the horizontal offset of the visible light camera's optical axis in the image coordinate system, expressed in pixels. y r is the vertical offset of the visible light camera's optical axis in the image coordinate system, in pixels. ij (i, j = 1, 2, 3) are vectors of the rotation matrix, t x t x t x These are vectors representing translation matrices, X0, Y0, and Z0 are the coordinates in the camera coordinate system of the visible light camera, and K is the intrinsic parameter of the visible light camera module. c M0 represents the pose parameters between the cameras.
[0043] Step 2: Project the camera pixel coordinates into the corresponding camera 3D coordinates.
[0044] Using the parameters of the infrared camera, visible light camera, and acoustic imager from step 1, each integer pixel coordinate in the infrared image, visible light image, and acoustic image is projected onto its corresponding 3D coordinates. Since the visible light camera is selected as the reference camera, its 3D coordinates are used as the world coordinates.
[0045] Step 3: Convert world coordinates to individual camera pixel coordinates.
[0046] Using the transformation parameters obtained in step 1 between the visible light camera, infrared camera, and audio-visual unit, the 3D coordinates of the infrared camera and audio-visual unit coordinate systems are transformed to (Mapping 2) the 3D coordinates (world coordinates) of the visible light camera coordinate system. Combining this with the projection of the infrared camera and audio-visual unit pixel coordinates onto the camera's 3D coordinates in step 2, the pixel coordinates of the infrared camera and audio-visual unit are then projected onto the world coordinates. Finally, the back-projection is solved to obtain the pixel coordinates of the infrared camera and audio-visual unit projected onto the world coordinates.
[0047] c M0 can be represented in a homogeneous form, and the points in the corresponding 3D coordinates can be transformed to the corresponding camera image. The coordinate relationship between the points in the 3D coordinates of the three modal images (sound-image, infrared, and visible light) and the corresponding camera image is shown in the following formula:
[0048]
[0049] Step 4: Generate the target image.
[0050] By using matrix multiplication, the pixels in the infrared camera and audio-visual unit are transformed into the coordinates corresponding to the pixels in the visible light camera image, such as... Figure 3 As shown.
[0051] As shown in the following formula, the 3D points projected in the image to be registered are converted into points corresponding to the registered image based on the coordinate relationship between the reference imaging camera image and the image to be registered:
[0052]
[0053] C1 M0 is the camera to be registered; C2 M0 is the reference imaging camera.
[0054] Next, the visible light image remains unchanged, while the original image coordinates of the infrared and acoustic images need to be mapped to the new pixel coordinates in step 3, so that the source pixels are transformed into the coordinates corresponding to the pixels in the visible light camera image.
[0055] Using the orthographic projection in step 2 and the back projection in step 3 of the infrared camera and audio-visual instrument, the mapping 3 (usually requiring interpolation) of the gray values of the pixels in the infrared image source image to the gray values of the pixels in the target image (which can be the light image) and the mapping 3 (usually requiring interpolation) of the RGB values of the pixels in the acoustic image source image to the RGB values of the pixels in the target image are obtained, and the registered infrared image and the registered acoustic image are generated respectively.
[0056] Step 5: Image Fusion and Reconstruction
[0057] The original visible light image is fused with the registered infrared image and the registered acoustic image obtained through mapping in step 4, as follows: Figure 4 As shown, this yields audio-visual-infrared-visible light registration fusion and reconstruction image materials for further unmanned intelligent power fault location.
[0058] Step 6: Deploy the application
[0059] The above algorithms and logical steps included in this fusion and reconstruction technology are deployed on the corresponding hardware devices. Based on the images that can be transmitted back from the light camera, infrared camera and audio-visual device, the technology will automatically perform registration, fusion and reconstruction work and generate the required audio-visual-infrared-visible light registration, fusion and reconstruction image materials.
[0060] Example 2
[0061] This embodiment provides a multimodal spatiotemporal-frequency visual fusion reconstruction system based on audio-visual-infrared-visible light, including: a visible light camera, an infrared camera, an audio-visual unit, a mapping module, a projection module, a registration module, and an image fusion module. The system is used to implement the methods described in this embodiment, and the similarities will not be repeated.
[0062] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A multimodal spatiotemporal-frequency visual fusion reconstruction method based on acoustic-infrared-visible light, characterized in that, Includes the following steps: S1. Select a visible light camera as the reference imaging camera, calibrate the parameters of the visible light camera, infrared camera, and audio-visual device, as well as the pose parameters between the cameras, and perform the first mapping of the infrared camera and audio-visual device to the visible light camera. S2. Project each integer pixel coordinate in the infrared image, visible light image and acoustic image forward to the corresponding 3D coordinate, and perform a second mapping of the 3D coordinates of the infrared camera and the audio-visual device to the 3D coordinates of the visible light camera. S3. Using the 3D coordinates of the visible light camera as world coordinates, the world coordinates are projected backwards onto the infrared camera and the audio-visual device to obtain the pixel coordinates of the world coordinates projected backwards onto the infrared camera and the audio-visual device. S4. Using the forward projection in step S2 and the reverse projection in step S3 of the infrared camera and the audio-visual instrument, the third mapping from the gray values of the pixels of the original infrared image to the gray values of the pixels of the visible light image and the third mapping from the RGB values of the pixels of the original acoustic image to the RGB values of the pixels of the visible light image are obtained, and the registered infrared image and the registered acoustic image are generated respectively. S5. The original visible light image is fused with the registered infrared image and the registered acoustic image from step S4 to obtain a multimodal registration and reconstruction image.
2. The method for multimodal spatiotemporal-frequency visual fusion reconstruction based on acoustic-infrared-visible light according to claim 1, characterized in that, The parameters used to calibrate the visible light camera, infrared camera, and audio-visual instrument include the internal and external parameters for calibrating the visible light camera, infrared camera, and audio-visual instrument.
3. The multimodal spatiotemporal-frequency visual fusion reconstruction method based on acoustic-infrared-visible light according to claim 2, characterized in that, The intrinsic parameters of a visible light camera are determined by the camera's horizontal focal length, vertical focal length, horizontal offset, and vertical offset.
4. The multimodal spatiotemporal-frequency visual fusion reconstruction method based on acoustic-infrared-visible light according to claim 3, characterized in that, The intrinsic parameters of the optical camera are expressed by the following formula: Among them, f x f is the horizontal focal length of a visible light camera. y c is the vertical focal length of a visible light camera. x c is the horizontal offset of the visible light camera's optical axis in the image coordinate system, expressed in pixels. y This represents the vertical offset of the visible light camera's optical axis in the image coordinate system, expressed in pixels.
5. The method for multimodal spatiotemporal-frequency visual fusion reconstruction based on acoustic-infrared-visible light according to claim 2, characterized in that, The optical camera extrinsic parameters include the pose parameters between the cameras.
6. The method for multimodal spatiotemporal-frequency visual fusion reconstruction based on acoustic-infrared-visible light according to claim 5, characterized in that, The pose parameters between the cameras are determined by the rotation matrix vector and the horizontal matrix vector.
7. The multimodal spatiotemporal-frequency visual fusion reconstruction method based on acoustic-infrared-visible light according to claim 6, characterized in that, The pose parameters between the cameras are expressed by the following formula: in, c M0 represents the pose parameters between cameras, r ij (i, j = 1, 2, 3) are vectors of the rotation matrix, t x t x t x These are vectors representing the translation matrix.
8. The multimodal spatiotemporal-frequency visual fusion reconstruction method based on acoustic-infrared-visible light according to claim 1, characterized in that, Convert the pixels in the infrared camera and audio-visual equipment to the coordinates corresponding to the pixels in the visible light camera image.
9. The multimodal spatiotemporal-frequency visual fusion reconstruction method based on acoustic-infrared-visible light according to claim 8, characterized in that, Matrix multiplication is used to convert the pixels in the infrared camera and audio-visual equipment to the coordinates corresponding to the pixels in the visible light camera image.
10. A multimodal spatiotemporal-frequency visual fusion reconstruction system based on acoustic-infrared-visible light, comprising: A visible light camera, an infrared camera, an audio-visual device, a mapping module, a projection module, a registration module, and an image fusion module, characterized in that the system is used to implement the method described in any one of claims 1-9.