A multi-modal feature extraction method, system, terminal device and storage medium
By extracting and fusing global and local features from infrared images of substations, the problem of missing detailed information in Fourier descriptors when the overall contours are similar is solved, thus achieving robustness and efficiency in infrared image feature extraction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- STATE GRID ZHEJIANG ELECTRIC POWER CO LTD HANGZHOU POWER SUPPLY CO
- Filing Date
- 2023-09-12
- Publication Date
- 2026-08-04
AI Technical Summary
In image feature extraction, Fourier descriptors lead to a significant loss of detailed information when the overall contours are similar, thus reducing the effectiveness of feature extraction.
Shape reconstruction is performed by acquiring the boundaries of infrared images of substations, global and local features are extracted, and global and local features are fused using a multimodal shape descriptor by combining Fourier descriptors, principal component analysis and Gaussian filters.
It improves the robustness of feature extraction from infrared images, reduces redundant computation, enhances the fusion effect of global and local features, and adapts to image translation, rotation, and scaling transformations.
Smart Images

Figure CN117274624B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image feature and data reconstruction technology, and in particular to a method, system, terminal device and storage medium for multimodal feature extraction of infrared images of substation equipment. Background Technology
[0002] Whether it's a substation inspection robot or a handheld infrared thermal imaging detector, the target equipment area still exists as a high-dimensional image. Due to the low data density and large amount of redundant information in the pixels, the differences between different types of equipment are distributed over a wider area, making it difficult to identify the type of target equipment. Feature extraction technology can represent the device pixel set in a more suitable way for processing. By compressing the image pixel set with a small number of vector sequences, a lower-dimensional feature space can be constructed, simplifying repetitive calculations and improving the equipment recognition rate.
[0003] The key to feature extraction of critical equipment areas in substation main grids lies in capturing the differences between different types of defects using low-dimensional features to improve classification accuracy and ensure the real-time performance of the algorithm. Common features include color features, texture features, and shape features. Color features are the basic features for human visual object discrimination, possessing invariance to rotation, translation, and scaling. Texture features reflect the interrelationships and structural information between pixels, possessing both local and global characteristics, and have strong noise resistance. However, both color and texture features are sensitive to illumination intensity and shooting conditions; changes in the image background and surface differences in the defect target area can all make feature extraction difficult. Since the illumination conditions for acquiring infrared images of different surfaces vary significantly during the equipment's infrared image acquisition process, the color and texture features of target areas on different surfaces exhibit visually discernible differences. Therefore, color- or texture-based extraction methods are insufficient to meet the requirements of defect area feature extraction, i.e., representing a large range of target objects with a small amount of data. The shape of a target area can be considered as a closed region enclosed by its contour lines. Utilizing shape descriptors and contour detail differences can effectively distinguish the differences between different target areas at a very low cost, making it suitable for application scenarios where infrared image shape distortion is controllable in substation equipment identification.
[0004] In an infrared image, the target region containing a device is known to be a set of pixels enclosed by a closed contour line. Using the Fourier descriptor, the boundary information of the target object can be transformed from the time domain to the frequency domain, starting from this closed contour line. Frequency domain information is extracted using Fourier transform, and the device shape can be reconstructed through vector reconstruction. The Fourier descriptor can capture the general features of the contour line based on a small number of vectors, and it is independent of rotation, displacement, and the choice of starting point, making it a robust image shape feature, often used in scenarios where image processing has high timeliness requirements. However, as a global shape feature descriptor, the Fourier descriptor focuses on the overall characterization of the target image, leading to the loss of a large amount of detailed information. Due to the lack of capture of local image information, the effectiveness of feature extraction is reduced when the overall contour is similar. Summary of the Invention
[0005] The technical problem this invention aims to solve is how to overcome the shortcomings of Fourier descriptors, which focus on the overall characterization of an image and thus suffer from a significant loss of detailed information, resulting in reduced effectiveness of image feature extraction when overall contours are similar. To address this technical problem, this invention provides a multimodal feature extraction method, system, terminal device, and storage medium.
[0006] In a first aspect, embodiments of the present invention provide a multimodal feature extraction method, including:
[0007] Acquire infrared image datasets of target equipment in the substation; the infrared image datasets include infrared images of target equipment in the main network and distribution network within the substation.
[0008] The shape of the infrared image is reconstructed to obtain the reconstructed digital boundary, and the global features of the digital boundary are extracted.
[0009] Obtain the curvature of all boundary points on the digital boundary;
[0010] Local features of the digital boundary are extracted based on the curvature;
[0011] The global features and local features are fused to obtain the multimodal shape descriptor of the digital boundary;
[0012] The infrared image is subjected to feature extraction based on the multimodal shape descriptor to obtain the multimodal features of the infrared image.
[0013] Preferably, the acquisition of the infrared image dataset of the target equipment in the substation includes:
[0014] Infrared images of target equipment in the main and distribution networks of the substation are collected using inspection robots and handheld infrared detectors to form an infrared image dataset of the target equipment in the substation.
[0015] Preferably, the step of reconstructing the shape of the infrared image boundary to obtain the reconstructed digital boundary, and extracting the global features of the digital boundary, includes:
[0016] The shape of the infrared image boundary is reconstructed based on the Fourier descriptor to obtain the reconstructed digital boundary.
[0017] The number of Fourier coefficients used in the Fourier descriptor is determined based on principal component analysis, and the eigenvectors corresponding to the Fourier coefficients are represented as global features of the digital boundary.
[0018] Preferably, obtaining the curvature of all boundary points on the digital boundary includes:
[0019] The digital boundary is smoothed using a Gaussian filter;
[0020] On the digital boundary, a point is taken at a preset distance in a clockwise direction as a boundary point, and the curvature of each boundary point is calculated.
[0021] Preferably, the step of extracting local features of the digital boundary based on the curvature includes:
[0022] Construct a curvature histogram based on the curvature;
[0023] The left and right boundary values of each segment of the curvature histogram, along with the number of boundary points contained therein, are combined to form a three-dimensional row vector.
[0024] All three-dimensional row vectors are combined to obtain the feature vector group of the curvature histogram, and the feature vector group is represented as the local feature of the digital boundary.
[0025] Preferably, fusing the global features and local features to obtain the multimodal shape descriptor of the digital boundary includes:
[0026] Based on canonical correlation analysis, the global features and local features are fused to obtain multiple pairs of canonical correlation vectors arranged in order, and the multiple pairs of canonical correlation vectors are represented as multimodal shape descriptors of the digital boundary.
[0027] Secondly, embodiments of the present invention also provide a multimodal feature extraction system, comprising:
[0028] The data acquisition module is used to acquire infrared image datasets of target equipment in the substation; the infrared image datasets include infrared images of target equipment in the main network and distribution network within the substation.
[0029] A global feature extraction module is used to reconstruct the shape of the boundary of the infrared image to obtain the reconstructed digital boundary and extract the global features of the digital boundary.
[0030] The curvature acquisition module is used to acquire the curvature of all boundary points on the digital boundary;
[0031] A local feature extraction module is used to extract local features of the digital boundary based on the curvature;
[0032] The feature fusion module is used to fuse the global features and local features to obtain the multimodal shape descriptor of the digital boundary;
[0033] The feature extraction module is used to extract features from the infrared image based on the multimodal shape descriptor to obtain the multimodal features of the infrared image.
[0034] Preferably, the feature fusion module includes:
[0035] A multimodal shape descriptor construction unit is used to fuse the global features and local features based on canonical correlation analysis to obtain multiple pairs of canonical correlation vectors arranged in order, and to represent the multiple pairs of canonical correlation vectors as the multimodal shape descriptor of the digital boundary.
[0036] Thirdly, embodiments of the present invention also provide a terminal device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the multimodal feature extraction method as described above.
[0037] Fourthly, embodiments of the present invention also provide a computer-readable storage medium, the computer-readable storage medium including a stored computer program, wherein, when the computer program is executed, it controls the device where the computer-readable storage medium is located to perform the multimodal feature extraction method as described above.
[0038] Compared with existing technologies, the multimodal feature extraction method, system, terminal device, and storage medium of this invention have the following advantages: After inputting the infrared images of the target equipment in the main network and distribution network inside the substation, the shape reconstruction map of the infrared image can be output through the relevant algorithm of embedding multimodal shape descriptors, providing a good application foundation for the further application of infrared images; combining Fourier descriptors, principal component analysis, and Gaussian filters ensures that the infrared image has good robustness to image transformations such as translation, rotation, and scale space transformation, and the local contour information of the infrared image is represented based on the curvature histogram, reducing redundant calculations and helping to enhance the fusion effect of global and local features. Attached Figure Description
[0039] Figure 1 This is a flowchart illustrating a multimodal feature extraction method according to an embodiment of the present invention;
[0040] Figure 2 This is a schematic diagram of infrared image boundary digitization based on Fourier descriptors according to an embodiment of the present invention;
[0041] Figure 3 This is a schematic diagram of the fitting of the boundary point on the digital boundary of the infrared image according to an embodiment of the present invention;
[0042] Figure 4 This is a schematic diagram illustrating the representation of local features based on curvature histograms in an embodiment of the present invention;
[0043] Figure 5 This is a schematic diagram illustrating the effect of infrared image boundaries under different dimensional Fourier descriptors according to an embodiment of the present invention;
[0044] Figure 6 This is a schematic diagram of the structure of a multimodal feature extraction system according to an embodiment of the present invention;
[0045] Figure 7 This is a schematic diagram of the structure of the multimodal shape descriptor building unit according to an embodiment of the present invention;
[0046] Figure 8 This is a schematic diagram of the structure of a terminal device according to an embodiment of the present invention. Detailed Implementation
[0047] The specific embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.
[0048] like Figure 1 As shown, this embodiment of the invention provides a multimodal feature extraction method, including the following steps:
[0049] S1. Obtain the infrared image dataset of the target equipment in the substation;
[0050] Specifically, an inspection robot and a handheld infrared detector are used to collect infrared images of target equipment in the main and distribution networks within the substation, forming an infrared image dataset of the substation's target equipment. In this embodiment, the inspection robot's inspection route, sampling points for each time period, and sampling frequency are rationally set according to the distribution of target equipment in the main and distribution networks within the substation. An initial infrared image dataset of the substation's target equipment is formed based on the infrared images collected by the inspection robot. Furthermore, due to the possibility of sampling failures such as unreasonable sampling, missampling, and missed sampling by the inspection robot, this embodiment also utilizes a handheld infrared detector to supplement the collection of point-based infrared images, expanding the initial infrared image dataset to obtain the final infrared image dataset. By using an inspection robot in conjunction with a handheld infrared detector for image acquisition, the comprehensiveness of the collected data can be ensured, thus laying a solid data foundation for subsequent feature extraction.
[0051] S2. Reconstruct the shape of the infrared image boundary to obtain the reconstructed digital boundary, and extract the global features of the digital boundary.
[0052] Specifically, the shape of the infrared image boundary is reconstructed based on the Fourier descriptor to obtain the reconstructed digital boundary; the number of Fourier coefficients used in the Fourier descriptor is determined based on principal component analysis, and the feature vectors corresponding to the Fourier coefficients are characterized as global features of the digital boundary.
[0053] like Figure 2 As shown, in this embodiment, K contour points of the target area are selected as the digital boundary of the infrared image in the two-dimensional space of the infrared image. Taking any point (x0, y0) as the starting point, moving counterclockwise along the digital boundary will sequentially pass through all contour points. The horizontal coordinate of the infrared image is taken as the real axis of the complex coordinate system, and the vertical coordinate is taken as the imaginary axis of the complex coordinate system. The coordinates of the contour points of the digital boundary are represented in complex form as follows:
[0054] s(k) = x(k) + jy(k)
[0055] A major advantage of the above complex number representation method is that it transforms a two-dimensional problem into a one-dimensional problem. Furthermore, performing a discrete Fourier transform on s(k) yields the following expression:
[0056]
[0057] Here, a(u) represents the complex coefficients, which are the Fourier descriptors of the digital boundary. By performing an inverse Fourier transform on the Fourier descriptors, the original appearance of the digital boundary can be recovered, as shown in the following equation:
[0058]
[0059] Furthermore, if only the first P terms of Fourier coefficients are used to recover the original appearance of the digital boundary, then the following approximation of s(k) is obtained:
[0060]
[0061] in, This represents the reconstructed digital boundary. By using the first P terms of Fourier coefficients to describe shape features, it is not necessary to reconstruct every contour point on the digital boundary, which can greatly reduce the computational load of feature extraction and subsequent image applications.
[0062] It is worth noting that changes in the starting point's position will affect the descriptor in different but known ways, and these differences can be eliminated by adjusting the starting point's position. In this embodiment, an initialization coefficient k0 is added to s(k) to obtain the corresponding digital boundary S. p and Fourier descriptor a p (u) is represented by the following formula:
[0063] S p =x(k-k0)+jy(k-k0)=s(k-k0)
[0064]
[0065] To further extract global features of the numerical boundaries, it is necessary to determine the value of P. This embodiment transforms the problem of selecting the number of terms P into an optimization problem of the optimal value of an orthogonal basis. The specific process is as follows:
[0066] Suppose we abstract a sample space, where each Fourier coefficient in the numerical boundary s(k) is a sample point in the sample space. We then establish an orthogonal coordinate system {w1, w2, ..., w...}. d}, where w i Let w represent an orthonormal vector that satisfies ||w||. i ||2=1, Reduce the dimension of the sample space to d′ dimension, and the sample point x i The projection in the new coordinate system is:
[0067] z i =(z i1 ;z i2 ;…z id′ )
[0068] in, z ij Represents sample point x i The coordinates of the j-th dimension in the low-dimensional coordinate system. Based on the reconstructed sample points, the distance between the original sample points and the reconstructed sample points in the sample space corresponding to the digital boundary is obtained, as expressed by the following formula:
[0069]
[0070] Where α is a constant, W = (w1, w2, ..., w d ), It is the covariance matrix. To ensure the minimum value of the above equation, we should ensure that:
[0071] min-tr(W T XX T W)
[0072] Furthermore, principal component analysis is used to solve the optimization problem of the orthogonal basis. Regarding XX in the above equation... T Eigenvalue analysis is performed to find the correlation between Fourier coefficients. An orthogonal transformation is then used to convert the Fourier descriptors into a set of linearly independent eigenvectors. Key eigenvectors are selected as principal components to achieve the optimal representation of the shape features of the infrared image. The specific process is as follows:
[0073] Let I be the Fourier coefficient in the shape feature. i The numerical boundary can then be represented as a matrix consisting of K random vectors, as shown in the following equation:
[0074] A = [I1-L] avg I2-L avg , ..., I K -L avg ]
[0075] Therefore, the covariance matrix corresponding to the digital boundary s(k) can be represented by the following formula:
[0076]
[0077] Calculate the eigenvector μ of the covariance matrix using the singular value decomposition theorem. i Specifically, it is represented by the following formula:
[0078]
[0079] Where, λ i Let F(m) represent the eigenvalues of the eigenvectors. The eigenvectors are arranged in descending order of their eigenvalues. The contribution rate F(m) of the first m vectors is calculated as follows:
[0080]
[0081] The m value corresponding to the contribution rate not lower than the preset threshold is taken as the P value, and the first m feature vectors are represented as global features of the digital boundary s(k). In this embodiment, the preset threshold is preferably set to 0.85. This preset threshold can be adaptively adjusted according to the requirements for global feature extraction. If the requirements are high, a higher preset threshold can be set, and vice versa.
[0082] S3. Obtain the curvature of all boundary points on the digital boundary;
[0083] Specifically, the digital boundary is smoothed using a Gaussian filter. Points are selected at preset intervals along the clockwise direction on the digital boundary as boundary points, and the curvature of each boundary point is calculated. This embodiment uses a Gaussian filter to smooth the digital boundary, and the specific process is as follows:
[0084] Using the normal distribution function to select weights, and assuming the arc length parameter is u, the digital boundary curve is represented by the following formula:
[0085] r(u) = (x(u), y(u))
[0086] The corresponding Gaussian function is expressed as follows:
[0087]
[0088] The smoothing of digital boundaries is achieved by convolving the digital boundaries with their corresponding Gaussian functions. The convolution operation is expressed as follows:
[0089]
[0090] Based on the distance between the current pixel and the center point, the Gaussian function assigns different weights to each contour point. Noise is eliminated through a weighted average of the contour points, achieving a smooth transition of the digital boundary. The digital boundary curve smoothed by the Gaussian filter is represented by the following formula:
[0091] s(u, δ)=(X(u, δ), Y(u, δ))
[0092] like Figure 3 As shown, in this embodiment, an arbitrary point n is selected on the digital boundary. i As boundary points, n is selected at positions d on either side of this point. i-d and n i+d Two points are used to fit an arc, where o is the center point of the fitted arc, r is the support radius, and α is the support angle.
[0093] The boundary point n is obtained based on the curvature calculation formula. i The curvature K(i) at point i is expressed by the following formula:
[0094]
[0095] Take a point at every clockwise interval d along the digital boundary as a boundary point, and calculate the curvature of each boundary point.
[0096] S4. Extract local features of digital boundaries based on curvature;
[0097] Specifically, a curvature histogram is constructed based on the curvature. The left boundary value, right boundary value, and number of boundary points contained in each segment of the curvature histogram are combined into a three-dimensional row vector. All three-dimensional row vectors are combined to obtain the feature vector group of the curvature histogram, and the feature vector group is represented as the local features of the digital boundary.
[0098] For infrared images with similar global features, local features are needed for differentiation. The curvature changes and distribution of boundary points on the digital boundary are important manifestations of shape features and contain rich local feature information.
[0099] like Figure 4 As shown, in this embodiment, the maximum curvature K is selected from the curvatures of all boundary points on the digital boundary s(k). max (i) and the minimum curvature K min (i) Divide the curvature histogram into H equal parts and determine the range interval Ra of the curvature histogram, as shown in the following formula:
[0100]
[0101] In this embodiment, the preferred value of H is 36. Of course, H can be set according to the curvature calculation. If the difference between the maximum and minimum curvature values is large, a larger H value can be set; conversely, the H value can be adjusted accordingly. Based on the range interval Ra, the curvature histogram is divided into a series of intervals, resulting in the following range of the curvature histogram:
[0102] [K min (i), K min (i)+Ra,...K max (i)-Ra,K max (i)]
[0103] Each segment of the curvature histogram has a left boundary value Bo. left、 Right boundary value Bo right and the number of boundary points Co node Divide each range into segments of Bo left、 Bo right and Co nodeA 1*3 three-dimensional row vector is formed. All three-dimensional row vectors are combined to obtain the feature vector group of the curvature histogram. The feature vector group is then represented as the local features of the digital boundary, realizing the abstract description of the local feature curvature corresponding to all boundary points.
[0104] S5. Fuse global features with local features to obtain a multimodal shape descriptor for the digital boundary;
[0105] Specifically, based on canonical correlation analysis, global features and local features are fused to obtain multiple pairs of canonical correlation vectors arranged in order, and these multiple pairs of canonical correlation vectors are represented as multimodal shape descriptors of digital boundaries.
[0106] In this embodiment, the global variable corresponding to the Fourier descriptor is set to Ve. x The local variable based on the curvature histogram is Ve. y For two correlated random vectors Ve x and Ve y These can be combined to form several representative canonical correlation vectors U. i and V i Specifically, it is represented by the following formula:
[0107] Ve x = (X1, X2, ..., X p )′
[0108] Ve y = (Y1, Y2, ..., Y) q )′
[0109] U i =a i1 X1+a i2 X2 + ... + a ip X p ≡a′Ve x
[0110] V i =b i1 Y1+b i2 Y2+…+b iq Y q ≡b′Ve y
[0111] Solving for a and b using the Lagrange multiplier method, such that the current U i and V i The correlation coefficient reached its maximum value, U i and V i The correlation coefficient is expressed by the following formula:
[0112]
[0113] Ensure the correlation coefficient ρ is within U i and V i The variance of all elements reaches its maximum when all variances are 1, thus yielding the current random vector Ve. x and Ve y The corresponding maximum canonical correlation vector is used to obtain canonical correlation vectors sorted from high to low. The multiple pairs of canonical correlation vectors obtained are then represented as multimodal shape descriptors of digital boundaries, realizing the abstract description of key feature information under different modalities.
[0114] S6. Extract features from the infrared image based on the multimodal shape descriptor to obtain the multimodal features of the infrared image.
[0115] like Figure 5 As shown, this embodiment demonstrates the effectiveness of infrared image boundary descriptors for substation target equipment under different dimensional Fourier descriptors, resolving the issue of feature differences in defect images under different shooting angles. Compared to existing color-based feature descriptor processing methods, this embodiment, combining Fourier descriptors, exhibits superior performance.
[0116] This invention provides a multimodal feature extraction method. By inputting infrared images of target equipment in the main and distribution networks of a substation, the method can output a shape reconstruction map of the infrared image using a related algorithm that embeds multimodal shape descriptors. This provides a solid foundation for further applications of the infrared image. Combining Fourier descriptors, principal component analysis, and Gaussian filters ensures the infrared image's robustness to image transformations such as translation, rotation, and scale space transformations. Furthermore, by characterizing the local contour information of the infrared image based on curvature histograms, redundant computation is reduced, which is beneficial for enhancing the fusion effect of global and local features.
[0117] like Figure 6 As shown, based on the above-described multimodal feature extraction method, this embodiment of the invention also provides a multimodal extraction system, including:
[0118] Data acquisition module 1 is used to acquire infrared image datasets of target equipment in the substation; the infrared image dataset includes infrared images of target equipment in the main network and distribution network within the substation;
[0119] Global feature extraction module 2 is used to reconstruct the shape of the boundary of the infrared image, obtain the reconstructed digital boundary, and extract the global features of the digital boundary;
[0120] Curvature acquisition module 3 is used to acquire the curvature of all boundary points on the digital boundary;
[0121] Local feature extraction module 4 is used to extract local features of digital boundaries based on curvature;
[0122] Feature fusion module 5 is used to fuse global features with local features to obtain a multimodal shape descriptor for the digital boundary;
[0123] Feature extraction module 6 is used to extract features from infrared images based on multimodal shape descriptors to obtain multimodal features of infrared images.
[0124] In one specific embodiment, such as Figure 7 As shown, feature fusion module 5 includes:
[0125] The multimodal shape descriptor construction unit 51 is used to fuse global features and local features based on canonical correlation analysis to obtain multiple pairs of canonical correlation vectors arranged in order, and to represent multiple pairs of canonical correlation vectors as multimodal shape descriptors of digital boundaries.
[0126] It should be noted that each module in the aforementioned multimodal feature extraction system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module. For specific limitations regarding a multimodal feature extraction system, please refer to the limitations regarding a multimodal feature extraction method above; both have the same function and role, and will not be repeated here.
[0127] This invention also provides a terminal device, which includes:
[0128] Processor, memory, and bus;
[0129] The bus is used to connect the processor and the memory;
[0130] The memory is used to store operation instructions;
[0131] The processor is configured to execute operations corresponding to the multimodal feature extraction method described above in this application by invoking the operation instructions.
[0132] In one alternative embodiment, a terminal device is provided, such as Figure 8 As shown, Figure 8 The terminal device 5000 shown includes a processor 5001 and a memory 5003. The processor 5001 and the memory 5003 are connected, for example, via a bus 5002. Optionally, the terminal device 5000 may also include a transceiver 5004. It should be noted that in practical applications, the transceiver 5004 is not limited to one type, and the structure of this terminal device 5000 does not constitute a limitation on the embodiments of this application.
[0133] Processor 5001 may be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 5001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0134] Bus 5002 may include a path for transmitting information between the aforementioned components. Bus 5002 may be a PCI bus or an EISA bus, etc. Bus 5002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 8 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0135] The memory 5003 may be a ROM or other type of static storage device capable of storing static information and instructions, RAM or other type of dynamic storage device capable of storing information and instructions, or it may be an EEPROM, CD-ROM or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but is not limited thereto.
[0136] The memory 5003 is used to store application code that executes the scheme of this application, and its execution is controlled by the processor 5001. The processor 5001 is used to execute the application code stored in the memory 5003 to implement the content shown in any of the foregoing method embodiments.
[0137] Terminal devices include, but are not limited to: mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), and in-vehicle terminals (such as in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers.
[0138] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the multimodal feature extraction method described above.
[0139] Another embodiment of this application provides a computer-readable storage medium storing a computer program that, when run on a computer, enables the computer to execute the corresponding content in the aforementioned method embodiments.
[0140] Furthermore, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.
[0141] In summary, the multimodal feature extraction method, system, terminal device, and storage medium of this invention, after inputting infrared images of target equipment in the main and distribution networks within a substation, can output a shape reconstruction map of the infrared image through a related algorithm embedding multimodal shape descriptors, providing a good application foundation for further applications of infrared images. By combining Fourier descriptors, principal component analysis, and Gaussian filters, the infrared image is ensured to have good robustness to image transformations such as translation, rotation, and scale space transformations. Furthermore, by characterizing the local contour information of the infrared image based on curvature histograms, redundant computation is reduced, which is beneficial for enhancing the fusion effect of global and local features.
[0142] The various embodiments in this specification are described in a progressive manner. For directly identical or similar parts of the embodiments, refer to each other. Each embodiment focuses on its differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0143] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and substitutions can be made without departing from the technical principles of the present invention, and these improvements and substitutions should also be considered within the scope of protection of the present invention.
Claims
1. A multimodal feature extraction method, characterized in that, include: Acquire infrared image datasets of target equipment in the substation; the infrared image datasets include infrared images of target equipment in the main network and distribution network within the substation. The shape of the infrared image is reconstructed to obtain the reconstructed digital boundary, and the global features of the digital boundary are extracted. Obtain the curvature of all boundary points on the digital boundary; Local features of the digital boundary are extracted based on the curvature; The global features and local features are fused to obtain the multimodal shape descriptor of the digital boundary; Based on the multimodal shape descriptor, feature extraction is performed on the infrared image to obtain the multimodal features of the infrared image; The process of reconstructing the shape of the infrared image boundary to obtain the reconstructed digital boundary, and extracting the global features of the digital boundary, includes: The shape of the infrared image boundary is reconstructed based on the Fourier descriptor to obtain the reconstructed digital boundary. The number of Fourier coefficients used in the Fourier descriptor is determined based on principal component analysis, and the feature vectors corresponding to the Fourier coefficients are characterized as global features of the digital boundary. The step of extracting local features of the digital boundary based on the curvature includes: Construct a curvature histogram based on the curvature; The left and right boundary values of each segment of the curvature histogram, along with the number of boundary points contained therein, are combined to form a three-dimensional row vector. All three-dimensional row vectors are combined to obtain the feature vector group of the curvature histogram, and the feature vector group is represented as the local feature of the digital boundary; The step of fusing the global features and local features to obtain the multimodal shape descriptor of the digital boundary includes: Based on canonical correlation analysis, the global features and local features are fused to obtain multiple pairs of canonical correlation vectors arranged in order, and the multiple pairs of canonical correlation vectors are represented as multimodal shape descriptors of the digital boundary.
2. The multimodal feature extraction method according to claim 1, characterized in that, The acquisition of the infrared image dataset of the target equipment in the substation includes: Infrared images of target equipment in the main and distribution networks of the substation are collected using inspection robots and handheld infrared detectors to form an infrared image dataset of the target equipment in the substation.
3. The multimodal feature extraction method according to claim 1, characterized in that, The step of obtaining the curvature of all boundary points on the digital boundary includes: The digital boundary is smoothed using a Gaussian filter; On the digital boundary, a point is taken at a preset distance in a clockwise direction as a boundary point, and the curvature of each boundary point is calculated.
4. A multimodal feature extraction system, characterized in that, include: The data acquisition module is used to acquire infrared image datasets of target equipment in the substation; the infrared image datasets include infrared images of target equipment in the main network and distribution network within the substation. A global feature extraction module is used to reconstruct the shape of the boundary of the infrared image to obtain the reconstructed digital boundary and extract the global features of the digital boundary. The curvature acquisition module is used to acquire the curvature of all boundary points on the digital boundary; A local feature extraction module is used to extract local features of the digital boundary based on the curvature; The feature fusion module is used to fuse the global features and local features to obtain the multimodal shape descriptor of the digital boundary; The feature extraction module is used to extract features from the infrared image based on the multimodal shape descriptor to obtain the multimodal features of the infrared image; The process of reconstructing the shape of the infrared image boundary to obtain the reconstructed digital boundary, and extracting the global features of the digital boundary, includes: The shape of the infrared image boundary is reconstructed based on the Fourier descriptor to obtain the reconstructed digital boundary. The number of Fourier coefficients used in the Fourier descriptor is determined based on principal component analysis, and the feature vectors corresponding to the Fourier coefficients are characterized as global features of the digital boundary. The step of extracting local features of the digital boundary based on the curvature includes: Construct a curvature histogram based on the curvature; The left and right boundary values of each segment of the curvature histogram, along with the number of boundary points contained therein, are combined to form a three-dimensional row vector. All three-dimensional row vectors are combined to obtain the feature vector group of the curvature histogram, and the feature vector group is represented as the local feature of the digital boundary; The feature fusion module includes: A multimodal shape descriptor construction unit is used to fuse the global features and local features based on canonical correlation analysis to obtain multiple pairs of canonical correlation vectors arranged in order, and to represent the multiple pairs of canonical correlation vectors as the multimodal shape descriptor of the digital boundary.
5. A terminal device, characterized in that, The method includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the multimodal feature extraction method as described in any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform the multimodal feature extraction method as described in any one of claims 1 to 3.