Image content rapid screening method and system based on multilayer semantic modeling
A rapid image content filtering method based on multi-layer semantic modeling, utilizing frequency domain filtering and wave function interferometry techniques, combined with geometric topological skeleton classification, solves the problem of pantograph detection in high-speed motion and complex backgrounds, achieving efficient and accurate crack detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN TIANBAOLAI INFORMATION TECH CO LTD
- Filing Date
- 2026-01-08
- Publication Date
- 2026-04-24
AI Technical Summary
In high-speed motion and complex backgrounds, existing pantograph defect detection methods suffer from low real-time performance and accuracy. Traditional image processing methods are prone to misjudgment, while deep learning methods are computationally intensive and difficult to meet real-time requirements.
A fast image content screening method based on multi-layer semantic modeling is adopted. It suppresses motion blur noise by using frequency domain filtering and dynamic bandpass filter, quickly locates suspected defect areas by using wave function interferometry, and performs high-level semantic classification by combining geometric topological skeleton and multi-constraint operators to achieve multi-layer semantic fusion decision.
It significantly improves image clarity and crack feature recognition, reduces false detection rate, achieves high detection accuracy and real-time performance, and meets the intelligent detection requirements of pantographs on high-speed railways.
Smart Images

Figure CN121921699A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition, and in particular to a method and system for rapid image content filtering based on multi-layer semantic modeling. Background Technology
[0002] With the rapid development of high-speed railways and the continuous increase in train speeds, the pantograph, as a key power supply component between the train and the overhead contact line, directly affects the safe operation of the train. During high-speed operation, the pantograph's sliding plate surface is prone to wear, cracks, and other defects due to prolonged frictional contact with the contact line. If these micro-cracks are not detected and addressed in time, they may gradually expand, eventually leading to pantograph failure, causing power outages, and seriously threatening train operation safety. Therefore, real-time and accurate detection of pantograph surface defects is of great significance.
[0003] Currently, pantograph defect detection mainly employs vision-based image recognition technology. In practical applications, detection systems are typically equipped with high-speed cameras to capture images of the pantograph during high-speed operation. However, when trains are traveling at high speeds, the images captured by the cameras face significant challenges: on the one hand, high-speed motion causes severe motion blur, making the detailed features of the pantograph unclear; on the other hand, the shooting background is extremely complex and variable, including overhead contact lines, clouds, and light reflections within tunnels, all of which generate a large amount of interference.
[0004] In existing technologies, traditional image recognition methods are mainly divided into two categories: The first category is based on traditional image processing methods, which extract crack features through edge detection, morphological operations, and other techniques. However, these methods are prone to misidentifying background textures as cracks or submerging real, minute cracks in noise when faced with motion blur and complex backgrounds, resulting in low detection accuracy. The second category is based on deep learning methods, which identify cracks by training deep neural network models. Although the recognition accuracy is high, it requires complete network inference for each frame of the image, resulting in a huge computational burden. In high frame rate video streaming scenarios, such as when high-speed cameras capture images at a rate of hundreds of frames per second, performing deep network inference for each frame would impose an unbearable computational burden, causing a severe lag in processing speed and failing to meet real-time requirements.
[0005] In summary, under the dual interference of strong motion blur caused by high-speed motion and complex and variable background, the real-time performance and accuracy of existing technologies are both low. Summary of the Invention
[0006] This application provides a method and system for rapid image content filtering based on multi-layer semantic modeling, which effectively ensures real-time performance and improves detection accuracy.
[0007] To achieve the above objectives, the embodiments of this application adopt the following technical solutions: Firstly, a method for rapid image content filtering based on multi-layer semantic modeling is provided for an image content rapid filtering system, which includes a high-speed camera, a speed sensor, and a backend server. The method includes: The system acquires video streams from high-speed cameras and real-time velocity data from velocity sensors, maps each frame of the video stream to the frequency domain, and generates an initial spectral distribution map. Based on the real-time speed, the spectral offset factor of the background texture is calculated, and a dynamic bandpass filter with the center frequency adjusted for blue shift correction by the spectral offset factor is constructed. The initial spectral distribution map is filtered by the dynamic bandpass filter to remove low-frequency aliasing noise caused by motion blur, and then inversely transformed back to the spatial domain to obtain a sharpened feature map representing the semantic features of the underlying texture. Convolutional layers are used to transform the sharpened feature map into a high-dimensional feature space tensor, and the numerical value of the high-dimensional feature space tensor is transformed into the amplitude, and the gradient direction of the high-dimensional feature space tensor is transformed into the phase, in order to construct a real-time object wave function and load a pre-constructed standard reference wave function. Perform the conjugate multiplication operation of the real-time object wave function and the standard reference wave function in the feature space to calculate the interference intensity field and extract the region above the preset energy threshold as the candidate defect substrate region with mid-level shape semantic features; For each candidate defect substrate region, a geometric topological skeleton is extracted, high-level semantic category recognition is performed based on the geometric topological skeleton, and the activation response value is determined based on the high-level semantic category recognition result. If the activation response value of any candidate defect substrate region exceeds the critical response threshold, then the current frame is determined to be a key frame containing cracks based on the recognition results of the bottom layer texture semantic features, the middle layer shape semantic features and the high layer semantic category. If the current frame is a keyframe containing cracks, then output the current frame to the backend server.
[0008] In one possible implementation of the first aspect, the spectral offset factor of the background texture is calculated based on the real-time speed, including: Obtain the image acquisition frame rate; Calculate the spectral compression coefficient caused by motion blur based on the real-time speed and image acquisition frame rate; Based on the spectral compression coefficient and the preset background texture reference frequency, the virtual Doppler frequency shift of the background texture in the frequency domain is calculated as the spectral offset factor.
[0009] In another possible implementation of the first aspect, a dynamic bandpass filter with a center frequency that undergoes spectral blue shift correction with a spectral offset factor is constructed, comprising: Determine the initial center frequency and bandwidth parameters of the reference bandpass filter; Calculate the blue shift correction amount for the center frequency based on the spectral shift factor; The initial center frequency is added to the blue shift correction to obtain the dynamically adjusted center frequency; A dynamic bandpass filter is constructed based on the dynamically adjusted center frequency and bandwidth parameters.
[0010] In another possible implementation of the first aspect, the numerical magnitude of the high-dimensional feature space tensor is converted into amplitude, and the gradient direction of the high-dimensional feature space tensor is converted into phase, to construct a real-time object wave function, including: Calculate the numerical value of each feature point in the high-dimensional feature space tensor, and use it as the amplitude component of the wave function; Calculate the gradient vector of each feature point in the high-dimensional feature space tensor, and extract the direction angle of the gradient vector as the phase component of the wave function; By combining the amplitude and phase components, a real-time object wave function in complex form is constructed.
[0011] In another possible implementation of the first aspect, the conjugate multiplication of the real-time object wavefunction and the standard reference wavefunction is performed in the characteristic space to calculate the interference intensity field, including: Perform a conjugate operation on the standard reference wavefunction to obtain the conjugate reference wavefunction; The real-time object wavefunction is added to the standard reference wavefunction by a complex number to obtain the superimposed wavefunction; The interference intensity field is obtained by calculating the square of the modulus of the superimposed wave functions; The phase matching degree distribution map is obtained by complex multiplying the real-time object wave function with the conjugate reference wave function. Based on the phase matching degree distribution map and the numerical distribution of the interference intensity field, constructive interference regions and destructive interference regions are identified in the interference intensity field. Regions with interference intensity values higher than the average value are constructive interference regions, while regions with interference intensity values lower than the average value are destructive interference regions.
[0012] In another possible implementation of the first aspect, regions with energy levels above a preset energy threshold are extracted as candidate defect substrate regions with mid-level shape semantic features, including: Iterate through all pixels in the interference intensity field and obtain the interference intensity value of each pixel; The interference intensity value is compared with a preset energy threshold, and all pixels with interference intensity values higher than the preset energy threshold are marked. Connectivity analysis is performed on the marked pixels to extract continuous high-energy regions; The extracted continuous high-energy regions are used as candidate defect substrate regions with mid-level shape semantic features.
[0013] In another possible implementation of the first aspect, for each candidate defect substrate region, a geometric topological skeleton is extracted, including: Binarization is performed on the candidate defect substrate region; A morphological thinning algorithm is applied to gradually erode the boundary of the candidate defect substrate region after binarization until a skeleton structure with a single pixel width is obtained. Extract the endpoints, branching points, and connection paths of the skeleton structure; Based on endpoints, bifurcation points, and connection paths, a geometric topological skeleton of the candidate defect substrate region is constructed as a structural feature for high-level semantic classification.
[0014] In another possible implementation of the first aspect, high-level semantic category recognition is performed based on a geometric topological skeleton, and activation response values are determined based on the high-level semantic category recognition results, including: A high-level semantic classifier is constructed, which includes a crack bifurcation rate constraint operator, a curvature continuity constraint operator, and an aspect ratio constraint operator, to classify candidate defect substrates into crack categories or non-crack categories. Based on the geometric topological skeleton, the ratio of the number of bifurcation points in the candidate defect substrate region to the total length is calculated as the actual bifurcation rate. Calculate the matching degree between the actual bifurcation rate and the standard bifurcation rate range defined by the crack bifurcation rate constraint operator to obtain the bifurcation rate matching component; Calculate the rate of curvature change of the candidate defect substrate region based on the geometric topological skeleton; Calculate the degree of matching between the rate of curvature change and the standard curvature continuity defined by the curvature continuity constraint operator to obtain the curvature fitting component; Calculate the aspect ratio of the candidate defect substrate region based on the geometric topological skeleton; Calculate the matching degree between the aspect ratio and the standard aspect ratio range defined by the aspect ratio constraint operator to obtain the aspect ratio matching component; High-level semantic category recognition is performed on the bifurcation rate matching component, curvature matching component and aspect ratio matching component to obtain induced matching degree, which is used to characterize the high-level semantic category recognition result. The reduction in activation energy is calculated based on the induced fit, where the higher the induced fit, the greater the reduction in activation energy. The actual activation energy is obtained by subtracting the reduction in activation energy from the preset baseline activation energy threshold. Substitute the actual activation energy into the exponential decay function to calculate the reaction rate constant; The reaction rate constant is output as the activation response value.
[0015] Secondly, this application provides an image processing server, comprising: The memory is configured to store instructions; and The processor is configured to retrieve the instructions from the memory and, when executing the instructions, to implement the aforementioned method for rapid image content filtering based on multi-layer semantic modeling.
[0016] Thirdly, this application provides an image content rapid filtering system, comprising: High-speed camera; Speed sensor; The backend server is connected to both the high-speed camera and the speed sensor. An image processing server, connected to a backend server, is used to execute the operation steps described above.
[0017] The above technical solutions effectively address the motion blur and complex background issues encountered in pantograph crack detection during high-speed motion scenarios by constructing a rapid image content screening method based on multi-layer semantic modeling. Utilizing frequency domain filtering and dynamic bandpass filter techniques, the filtering parameters are adaptively adjusted according to real-time speed, successfully suppressing low-frequency aliasing noise caused by motion blur and significantly improving image clarity and crack feature identifiability. The pattern matching problem in feature space is transformed into a wave function interference problem, and suspected defect areas are quickly located by calculating the interference intensity field. Compared to traditional point-by-point comparison methods, this significantly improves computational efficiency and matching accuracy. By extracting the geometric topological skeleton and combining it with multi-constraint operators for high-level semantic classification, the topological and geometric features of the crack are fully utilized, effectively distinguishing real cracks from background interference and reducing the false detection rate. A multi-layer semantic fusion decision mechanism is adopted, integrating feature information from three levels: bottom-layer texture, middle-layer shape, and high-layer category, achieving more comprehensive and accurate crack judgment. The overall solution, while ensuring high detection accuracy, significantly reduces the amount of data requiring deep processing through keyframe screening, thus meeting real-time requirements and providing effective technical support for intelligent detection of pantographs on high-speed railways.
[0018] Other features and advantages of the embodiments of this application will be described in detail in the following detailed description section. Attached Figure Description
[0019] Figure 1 A flowchart illustrating a method for rapid image content filtering based on multi-layer semantic modeling, provided in an embodiment of this application; Figure 2 A flowchart illustrating a standard reference wavefunction construction method provided in this application embodiment; Figure 3 A flowchart illustrating multi-layer semantic fusion provided in an embodiment of this application; Figure 4This is a schematic diagram of the structure of an image content rapid filtering system provided in an embodiment of this application. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are only for illustration and explanation of the embodiments of this application and are not intended to limit the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0021] It should be noted that if the embodiments of this application involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.
[0022] Furthermore, if the embodiments of this application involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, features defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed in this application.
[0023] Figure 1 The illustration shows a flowchart of a method for rapid image content filtering based on multi-layer semantic modeling according to an embodiment of this application. Figure 1 As shown in the figure, this application provides a method for fast image content filtering based on multi-layer semantic modeling, which is applied to an image content fast filtering system. The image content fast filtering system includes a high-speed camera, a speed sensor, and a back-end server. The method may include the following steps.
[0024] S110. Acquire the video stream captured by the high-speed camera and the real-time speed data captured by the speed sensor, map each frame of the video stream to the frequency domain space, and generate an initial spectrum distribution map. S120. Based on the real-time speed, calculate the spectral offset factor of the background texture, construct a dynamic bandpass filter with the center frequency adjusted for blue shift correction by the spectral offset factor, filter the initial spectral distribution map through the dynamic bandpass filter to remove low-frequency aliasing noise caused by motion blur, and inversely transform back to the spatial domain to obtain the sharpened feature map of the semantic feature representation of the underlying texture. S130. The sharpened feature map is transformed into a high-dimensional feature space tensor using a convolutional layer. The magnitude of the high-dimensional feature space tensor is transformed into the amplitude, and the gradient direction of the high-dimensional feature space tensor is transformed into the phase, so as to construct a real-time object wave function and load a pre-constructed standard reference wave function. S140. Perform the conjugate multiplication operation of the real-time object wave function and the standard reference wave function in the feature space, calculate the interference intensity field, and extract the region with a preset energy threshold as a candidate defect substrate region with mid-level shape semantic features. S150. For each candidate defect substrate region, extract the geometric topological skeleton, perform high-level semantic category recognition based on the geometric topological skeleton, and determine the activation response value based on the high-level semantic category recognition result. S160. If the activation response value of any candidate defect substrate region exceeds the critical response threshold, then determine whether the current frame is a key frame containing cracks based on the bottom layer texture semantic features, the middle layer shape semantic features and the high layer semantic category recognition results. S170. If the current frame is a key frame containing cracks, output the current frame to the backend server.
[0025] In the pantograph inspection scenario on high-speed railways, high-speed cameras continuously capture images of the pantograph's operation at a rate of hundreds of frames per second, forming a continuous video stream. Simultaneously, speed sensors installed on the train monitor the train's speed in real time, providing crucial parameters for subsequent motion compensation.
[0026] Upon receiving the video stream, each frame is preprocessed, including removing lens distortion and adjusting brightness and contrast, to ensure image quality meets the requirements of subsequent analysis. Next, a two-dimensional Fast Fourier Transform (FFT) technique is used to convert the spatial domain image to the frequency domain.
[0027] Specifically, for a picture The pixel-by-pixel image is processed using an FFT algorithm to generate a complex matrix of the same size. Each element of this matrix represents the amplitude and phase information of a specific frequency component. The frequency domain representation decomposes the image into a superposition of sine waves of different frequencies. Low-frequency components correspond to the overall outline and slowly changing regions of the image, while high-frequency components correspond to details such as edges and textures. The generated initial spectral distribution map is stored as a two-dimensional matrix, with most of the low-frequency energy concentrated in the central region and high-frequency information distributed in the edge regions.
[0028] Spectral analysis provides a direct observation of the spectral diffusion phenomenon caused by motion blur: due to high-speed motion, the originally concentrated frequency components will tail and broaden along the direction of motion. This spectral characteristic provides a theoretical basis for subsequent motion blur compensation. This step converts the time-domain image into a frequency-domain representation, laying the foundation for eliminating motion blur using frequency-domain filtering techniques, while preserving complete frequency information and avoiding information loss that may occur during spatial domain processing.
[0029] Based on real-time speed data provided by the speed sensor and combined with the fixed frame rate parameters of the high-speed camera, the actual displacement of the pantograph during each frame acquisition can be accurately calculated. When the train is traveling at a speed of 300 kilometers per hour, if the camera frame rate is 500 frames per second, the exposure time of each frame is approximately 2 milliseconds, during which the pantograph moves approximately 16.7 centimeters. This high-speed motion causes a spectral compression phenomenon in the frequency domain, similar to the Doppler effect, in the background textures of the image (such as overhead contact lines and tunnel walls).
[0030] In the specific calculation, the blur kernel length is first obtained by multiplying the motion speed by the exposure time. Then, the spectral compression coefficient is calculated using the frequency domain transfer function characteristics of motion blur. This coefficient reflects the proportional relationship between the original frequency of the background texture and the observed frequency. Furthermore, by using a pre-calibrated background texture reference frequency and combining it with the spectral compression coefficient, the virtual frequency shift of the background texture in the current motion state, i.e., the spectral shift factor, is calculated. The background texture reference frequency can be obtained through static scene sampling. Based on this shift factor, when constructing a dynamic bandpass filter, the center frequency of the standard bandpass filter is shifted to a higher frequency direction by a corresponding amount, achieving spectral blue shift correction.
[0031] This dynamic adjustment mechanism ensures that the filter is always aligned with the true crack characteristic frequency band, rather than the frequency band distorted by motion blur. The filtering process is performed directly in the frequency domain, multiplying the initial spectrum distribution map with the filter transfer function through a dot multiplication operation, effectively suppressing low-frequency aliasing noise and high-frequency random noise, while preserving the crack characteristic signal in the mid-frequency band.
[0032] Finally, the filtered spectrum is converted back to the spatial domain using inverse Fourier transform to obtain a sharpened feature map. This feature map significantly suppresses motion blur, greatly reduces interference from background textures, and makes the edges of defects such as cracks clearer and sharper, providing a high-quality underlying texture semantic representation for subsequent feature extraction.
[0033] After being processed by a multi-layer convolutional neural network, the sharpened feature map is mapped to a high-dimensional feature space, forming a multi-channel feature tensor. Each spatial location of this tensor contains hundreds of feature dimensions, which encode rich semantic information from the underlying texture to the intermediate shape.
[0034] To simulate interference phenomena in wave optics in the feature space, the feature tensor needs to be converted into a complex wave function representation. Specifically, the Euclidean norm of each eigenvector in the tensor is first calculated. The magnitude of this norm reflects the intensity of the feature response, and after normalization, it is used as the amplitude component of the wave function. A larger amplitude indicates a higher probability of a significant feature existing at that location.
[0035] The gradient field of the feature tensor in the spatial dimension is calculated. For each feature point, its gradient vector indicates the direction of the most dramatic feature change. The direction angle of the gradient vector is extracted using the arctangent function and mapped to... The interval serves as the phase component of the wave function. Phase information encodes the spatial distribution pattern and directionality of the feature. Combining amplitude and phase, a complex real-time object wave function is constructed, mathematically represented by amplitude multiplied by a phase exponent. This wave function representation transforms abstract features in a high-dimensional feature space into physical quantities capable of wave operations. Simultaneously, a pre-constructed standard reference wave function is loaded, representing the wave pattern of the ideal crack feature in the feature space.
[0036] The reference wavefunction also takes the form of a complex number, with its amplitude distribution reflecting the intensity characteristics of a typical crack and its phase distribution reflecting the morphological characteristics of the crack. By performing interferometry between the real-time object wavefunction and the standard reference wavefunction, pattern matching can be achieved in the feature space, effectively identifying regions similar to the characteristics of a standard crack.
[0037] Specifically, in this embodiment, the standard reference wavefunction is pre-constructed in the following manner: S1. Collect multiple sets of standard crack sample images and perform semantic annotation on the standard crack sample images; S2. Extract low-level semantic features and model mid-level semantic features from the labeled standard crack sample images to obtain the mid-level semantic feature tensor of the standard crack. S3. Encode the amplitude and phase of the mid-level semantic feature tensor of the standard crack to generate the mid-level semantic wave function representation of the standard crack. S4. Statistically average the mid-layer semantic wavefunction representations of multiple sets of standard cracks to obtain the standard reference wavefunction corresponding to the mid-layer semantic features.
[0038] When performing wavefunction interferometry in the characteristic space, the standard reference wavefunction is first conjugated, i.e., the amplitude is kept constant while the phase is inverted. Then, the real-time object wavefunction and the reference wavefunction are complexly added to simulate the superposition of two coherent light waves. According to the superposition principle, when the phase difference between the two wavefunctions approaches zero or... When the phase difference is an integer multiple of the phase difference, constructive interference occurs, and the amplitude of the superimposed amplitude is enhanced; when the phase difference is close to the phase difference, constructive interference occurs. When the amplitude is an odd multiple of the modulus, destructive interference occurs, and the amplitude decreases. The spatial distribution of the interference intensity field is obtained by calculating the square of the superimposed wavefunction modulus. In this intensity field, high-intensity regions correspond to locations where the real-time characteristics and standard crack characteristics highly match, while low-intensity regions indicate significant differences in characteristics.
[0039] To more accurately assess the matching degree, the complex multiplication of the real-time object wavefunction and the conjugate reference wavefunction is calculated to obtain a phase matching degree distribution map. The real part of this distribution map reflects phase consistency, while the imaginary part reflects phase orthogonality. By combining the interference intensity field and the phase matching degree distribution, constructive and destructive interference regions can be accurately identified.
[0040] Specifically, constructive interference regions represent structures in the feature space that are highly similar to crack modes, with interference intensity values significantly higher than the average. By setting a preset energy threshold, all continuous regions with interference intensities exceeding this threshold are extracted. These high-energy regions often exhibit elongated and curved morphological features in space, matching the geometry of cracks, and are therefore marked as candidate defect substrate regions. This method utilizes the interference principle of wave optics to achieve efficient pattern matching in the feature space. Compared to traditional point-by-point comparison methods, the interferometric method can simultaneously consider amplitude and phase information, resulting in a more comprehensive and accurate judgment of feature similarity and effectively reducing the false detection rate.
[0041] For each candidate defect substrate region, binarization is first performed, dividing the pixels within the region into foreground and background categories. Then, a morphological thinning algorithm is applied. This algorithm iteratively erodes the outer pixels of the region, ensuring that the region's connectivity and topological structure are not destroyed. The thinning process continues until the region shrinks to a skeleton structure with a single pixel width. This skeleton preserves the main shape features and topological relationships of the original region while significantly simplifying the geometric representation.
[0042] In the skeleton structure, three types of key topological nodes are identified and extracted: endpoints, bifurcation points, and connection paths. Endpoints are the termination points of the skeleton, typically corresponding to the start or end of a crack; bifurcation points are the intersections of multiple skeleton branches, reflecting the bifurcation or intersection characteristics of the crack; connection paths are skeleton segments between endpoints and bifurcation points, or between bifurcation points themselves, representing the direction and length of crack extension. Based on these topological elements, a complete geometric topological skeleton description is constructed. This skeleton contains not only geometric information (such as length and curvature) but also topological information (such as connectivity and the number of branches).
[0043] Next, a pre-trained high-level semantic classifier is used to analyze the skeleton. This classifier contains three core constraint operators: a crack bifurcation rate constraint operator to evaluate whether the density of bifurcation points conforms to the statistical characteristics of a real crack. Real cracks usually have a moderate bifurcation rate; too high or too low a rate may be background interference. A curvature continuity constraint operator to check whether the curvature change of the skeleton is smooth and continuous. The extension path of a crack usually follows the stress distribution law, and the curvature change is relatively gentle, while background textures often show irregular curvature jumps. An aspect ratio constraint operator to determine whether the geometry of the region exhibits slender characteristics. The aspect ratio of a crack is usually much greater than 1, while the aspect ratio of circular or square interference regions is close to 1.
[0044] By calculating the matching degree between the actual skeleton features and the standard ranges defined by each constraint operator, the bifurcation rate matching component, curvature matching component, and aspect ratio matching component are obtained. These three matching components are combined into an induced fit, the higher of which indicates that the candidate region better matches the high-level semantic features of a crack. Based on the activation energy theory of chemical reaction kinetics, regions with high induced fit are equivalent to reactants in a high-energy state, making it easier to overcome the activation energy barrier and react. Therefore, the activation energy reduction is calculated based on the induced fit, and this reduction is subtracted from the baseline activation energy to obtain the actual activation energy. The actual activation energy is substituted into an exponential decay function in the form of the Arrhenius equation to calculate the reaction rate constant, which is output as the activation response value. The activation response value quantifies the confidence that the candidate region is determined to be a real crack.
[0045] After obtaining the activation response values of all candidate defect substrate regions, a traversal check is performed to see if the activation response value of any candidate region exceeds a preset critical response threshold. The critical response threshold is determined through statistical analysis of a large amount of historical data and represents the boundary between the actual crack and background interference.
[0046] If candidate regions exceeding the threshold exist, it indicates that the current frame image is likely to contain real crack defects, requiring further comprehensive judgment. In this case, multi-layer semantic fusion decision-making is performed by integrating low-level texture semantic features, mid-level shape semantic features, and high-level semantic category recognition results, such as... Figure 3 As shown. The bottom-layer texture semantic features are derived from the sharpened feature map obtained in step S120, reflecting the local texture pattern and edge intensity of the image, and are used to verify whether the candidate region has texture features unique to cracks. The middle-layer shape semantic features are derived from the interference intensity field analysis in step S140, reflecting the overall shape and spatial distribution of the candidate region, and are used to confirm whether the region morphology conforms to the geometric features of a crack. The high-layer semantic category recognition results are derived from the topological skeleton analysis and activation response value calculation in step S150, providing semantic category judgment and confidence evaluation for the candidate region.
[0047] By designing multi-layer semantic fusion rules, the strength of evidence at three levels is comprehensively considered. For example, if a candidate region has clear edge texture at the bottom layer, exhibits a slender shape in the middle layer, and is identified as a crack with a high activation response value at the top layer, then the region is highly likely to be identified as a real crack, and the current frame is marked as a keyframe containing a crack. Conversely, if a candidate region has a high activation response value but blurry texture at the bottom layer or an irregular shape in the middle layer, it may be a false detection, and its confidence level needs to be reduced or it should be directly excluded. This multi-layer semantic modeling and fusion decision mechanism fully utilizes multi-scale feature information from the bottom to the top layers, significantly improving the accuracy and robustness of crack detection and effectively avoiding the false detection and missed detection problems that may be caused by single-level feature judgment.
[0048] Once the current frame is determined to be a keyframe containing a crack, the image of that frame and its related detection results are immediately packaged and output to the backend server. The output data includes multi-dimensional information such as the original image, sharpened feature map, position coordinates of the candidate defect substrate area, geometric topological skeleton description, and activation response values.
[0049] After receiving keyframe data, the backend server can perform more in-depth analysis and processing. For example, it can call a more accurate but computationally intensive deep learning model for secondary confirmation, or store the detection results in a database for professional review. Ordinary frames not identified as keyframes are discarded without further processing, significantly reducing the amount of data requiring in-depth analysis. This hierarchical filtering mechanism optimizes the allocation of computing resources: the front-end fast filtering module filters out a large number of irrelevant frames at extremely low computational cost, only passing a small number of suspected defective keyframes to the backend for detailed analysis. In high frame rate video stream scenarios, this strategy increases the overall processing speed by tens of times while maintaining high detection accuracy. Furthermore, the output keyframe information can trigger a real-time alarm mechanism. When a serious crack is detected, maintenance personnel are immediately notified for emergency response, effectively ensuring train operation safety.
[0050] This embodiment effectively solves the problems of motion blur and complex background in pantograph crack detection in high-speed moving scenes by constructing a rapid image content screening method based on multi-layer semantic modeling. Utilizing frequency domain filtering and dynamic bandpass filter techniques, the filtering parameters are adaptively adjusted according to real-time speed, successfully suppressing low-frequency aliasing noise caused by motion blur and significantly improving image clarity and crack feature discernibility. The pattern matching problem in feature space is transformed into a wave function interference problem, and suspected defect areas are quickly located by calculating the interference intensity field. Compared with traditional point-by-point comparison methods, the computational efficiency is significantly improved and the matching accuracy is higher. By extracting the geometric topological skeleton and combining it with multi-constraint operators for high-level semantic classification, the topological and geometric features of the crack are fully utilized, effectively distinguishing real cracks from background interference and reducing the false detection rate. A multi-layer semantic fusion decision mechanism is adopted, integrating feature information from three levels: bottom-layer texture, middle-layer shape, and high-layer category, achieving more comprehensive and accurate crack judgment. The overall scheme, while ensuring high detection accuracy, significantly reduces the amount of data requiring deep processing through keyframe screening, thus meeting real-time requirements and providing effective technical support for intelligent detection of pantographs on high-speed railways.
[0051] In one embodiment of this invention, the spectral offset factor of the background texture is calculated based on the real-time speed, including the following steps: S210, Obtain the image acquisition frame rate; S220. Calculate the spectral compression coefficient caused by motion blur based on the real-time speed and image acquisition frame rate. S230. Based on the spectral compression coefficient and the preset background texture reference frequency, calculate the virtual Doppler frequency shift of the background texture in the frequency domain, which is used as the spectral offset factor.
[0052] High-speed cameras are configured with fixed image acquisition frame rate parameters at the factory, which determine the number of image frames the camera can capture per second. In pantograph detection applications, the commonly used frame rate range for high-speed cameras is 200 frames per second to 1000 frames per second. A higher frame rate means finer temporal resolution, but it also generates a larger amount of data.
[0053] The process of obtaining the image acquisition frame rate is typically achieved by reading the camera's configuration registers or calling the camera driver's API. In actual deployment, the appropriate frame rate needs to be selected based on the train's maximum operating speed and the required detection accuracy. For example, when the train is traveling at 350 km / h, if the detection system is required to capture millimeter-level crack details, an acquisition frame rate of at least 500 frames per second is needed to ensure that the pantograph's displacement in the image sequence is sufficiently small, avoiding the loss of critical features between frames.
[0054] The acquired frame rate parameter will serve as the base time unit for subsequent calculations, used to estimate the exposure time of each frame and the pantograph's movement distance within a single frame. The accuracy of this parameter directly affects the effect of motion blur compensation; therefore, rigorous parameter verification is required during system initialization to ensure that the read frame rate value matches the camera's actual operating state.
[0055] Based on the acquired real-time speed data and image acquisition frame rate, the actual displacement of the pantograph during a single frame exposure can be accurately calculated. Specifically, the train speed is first converted from kilometers per hour to meters per second, then divided by the frame rate to obtain the inter-frame time interval, which is the exposure time of a single frame. Multiplying the speed by the exposure time yields the displacement distance of the pantograph along the direction of motion during the exposure; this distance is the length of the motion blur kernel.
[0056] For example, when the train speed is 300 kilometers per hour (approximately 83.3 meters per second) and the frame rate is 500 frames per second, the single frame exposure time is 0.002 seconds, during which the pantograph moves approximately 0.167 meters, or 16.7 centimeters. This high-speed movement causes each object point in the image to form a 16.7-centimeter-long trail on the sensor, corresponding to the number of pixels.
[0057] From a frequency domain perspective, motion blur is equivalent to performing a directional low-pass filter on the original sharp image. Its frequency domain transfer function exhibits a sinc function shape, remaining constant on the frequency axis perpendicular to the motion direction, while showing oscillating decay with increasing frequency on the frequency axis parallel to the motion direction. By analyzing the location of the first zero point of this transfer function, the spectral compression coefficient can be calculated. This coefficient reflects the degree to which spectral energy is concentrated in the low-frequency band. The faster the motion speed or the longer the exposure time, the more severe the spectral compression, and the greater the loss of high-frequency detail information. The calculation of the spectral compression coefficient provides a quantitative basis for subsequent spectral shift compensation.
[0058] The background texture reference frequency is pre-calibrated through spectral analysis of the background area in static or low-speed scenes. It represents the dominant frequency component of background elements such as overhead contact lines and tunnel walls under normal conditions. This reference frequency is typically located in the mid-frequency band, corresponding to the periodic structural features of the background texture.
[0059] When a train travels at high speed, the relative motion causes a shift in the frequency of the background texture observed by the camera, a phenomenon similar to the Doppler effect in sound or electromagnetic waves. Although there is no true Doppler frequency shift in optical imaging, the spectral compression caused by motion blur can be mathematically equivalent to a virtual frequency shift.
[0060] In the specific calculation, the background texture reference frequency is multiplied by the spectral compression factor to obtain the difference between the observed frequency and the reference frequency. This difference is the virtual Doppler frequency shift. Because motion causes the spectrum to compress to lower frequencies, the observed frequency is lower than the reference frequency, so the frequency shift is negative, corresponding to the redshift phenomenon in the spectrum. To compensate for this frequency shift during filtering, the center frequency of the filter needs to be adjusted to a higher frequency direction by a corresponding amount, i.e., spectral blueshift correction is performed.
[0061] The calculated virtual Doppler frequency shift is output as a spectral offset factor, which dynamically changes with train speed; the higher the speed, the greater the offset, thus achieving adaptive adjustment of filter parameters. This speed-based spectral offset compensation mechanism solves the performance degradation problem of traditional fixed-parameter filters in variable-speed scenarios, ensuring that the filter can accurately locate and retain the crack feature frequency band regardless of the train's speed. Simultaneously, it suppresses low-frequency noise introduced by motion blur, improving the restoration quality of motion-blurred images and the accuracy of subsequent crack detection.
[0062] This embodiment achieves dynamic adaptive adjustment of motion blur compensation parameters by establishing a spectral offset factor calculation mechanism based on real-time speed. The acquired image frame rate provides a precise time reference for subsequent calculations, ensuring the accuracy of motion parameter estimation. Based on the real-time speed and frame rate, the spectral compression coefficient caused by motion blur is calculated, establishing a quantitative relationship between motion state and frequency domain characteristics, revealing the impact of high-speed motion on the image's spectral distribution. Based on the spectral compression coefficient and a preset background texture reference frequency, the virtual Doppler frequency shift is calculated, equating the frequency domain effect of motion blur to a frequency shift phenomenon, providing a theoretical basis and calculation method for frequency domain compensation. This spectral offset factor dynamically changes with train speed, realizing real-time correlation between filter parameters and motion state, enabling the filter to automatically adapt to changes in spectral characteristics under different speed conditions. The overall scheme, through a complete calculation chain from speed measurement to spectral analysis, accurately quantifies the impact of motion blur on the image's frequency domain characteristics, providing key parameters for the subsequent construction of dynamic bandpass filters, ensuring the accuracy and adaptability of motion blur compensation, and improving the image restoration quality in high-speed motion scenes.
[0063] In one embodiment of this invention, constructing a dynamic bandpass filter with a center frequency that undergoes spectral blue shift correction based on a spectral shift factor includes the following steps: S310. Determine the initial center frequency and bandwidth parameters of the reference bandpass filter; S320. Calculate the blue shift correction amount of the center frequency based on the spectral shift factor; S330. Add the initial center frequency to the blue shift correction amount to obtain the dynamically adjusted center frequency; S340. Construct a dynamic bandpass filter based on the dynamically adjusted center frequency and bandwidth parameters.
[0064] The design of the reference bandpass filter requires spectral characteristic analysis based on a large number of pantograph crack samples. During the system development phase, hundreds of pantograph images containing real cracks were acquired, and Fourier transforms were performed on these images to statistically analyze the energy distribution patterns of crack features in the frequency domain.
[0065] Research has found that pantograph surface cracks, due to their elongated and continuous geometry, primarily exhibit mid-to-high frequency components in the frequency domain, with energy peaks typically concentrated within a specific frequency range. Statistical analysis of the spectra of all crack samples was conducted to calculate the mean and variance of the frequency energy distribution, thus determining the typical frequency range of crack characteristics. The center of this range is the initial center frequency of the reference bandpass filter, usually located in the mid-frequency band of the spectrum; the specific value depends on the image resolution and the typical crack width.
[0066] For example, for an image with a resolution of 2048×2048 pixels, if the typical crack width is 2-5 pixels, the corresponding center frequency is approximately 0.3-0.5 times the image's Nyquist frequency. The determination of the bandwidth parameter is also based on statistical analysis, requiring a balance between preserving sufficient crack information and suppressing background noise. Too narrow a bandwidth will result in the filtering out of some crack features, while too wide a bandwidth will fail to effectively suppress noise. By analyzing the standard deviation of the crack spectral energy distribution, the bandwidth is set to cover a frequency range that covers 95% of the crack energy, typically 0.4-0.6 times the center frequency. The determined initial center frequency and bandwidth parameters are stored as core parameters of the baseline filter, providing a benchmark reference for subsequent dynamic adjustments.
[0067] The spectral shift factor reflects the degree to which background texture frequencies are compressed to lower frequencies due to motion blur. This shift needs to be compensated for by adjusting the center frequency of the filter in reverse. When calculating the blue shift correction, the physical meaning of the spectral shift factor is first analyzed. Because high-speed motion causes the observed frequencies of the background texture to be lower than their true frequencies, the overall spectrum shifts to lower frequencies, exhibiting a redshift characteristic. To enable the filter to accurately locate the true crack feature frequency band, the center frequency of the filter needs to be adjusted to higher frequencies, i.e., blue shift correction.
[0068] The blue shift correction is calculated using a linear compensation strategy, multiplying the spectral shift factor by a correction coefficient. This correction coefficient, obtained through experimental calibration, reflects the effectiveness of frequency domain filtering in compensating for motion blur. In practical applications, the correction coefficient is typically set to 1.2-1.5; a setting slightly greater than 1 is used to overcompensate for the effects of motion blur, ensuring that crack features fall entirely within the filter's passband.
[0069] For example, if the spectral offset factor is -50 Hz (a negative value indicates a redshift) and the correction coefficient is 1.3, then the blueshift correction is 65 Hz. This overcompensation strategy can address the biases caused by velocity sensor measurement errors and the simplified assumptions of the motion fuzzy model in practice, improving the robustness of the filter. The calculated blueshift correction is a positive value, indicating the magnitude by which the center frequency needs to be adjusted towards higher frequencies.
[0070] Dynamically adjusting the center frequency is a key step in achieving adaptive filtering. The initial center frequency is algebraically added to the blue shift correction value to obtain the optimized center frequency for the current motion state. This dynamic adjustment mechanism allows the filter to track changes in train speed in real time and automatically adapt to different levels of motion ambiguity.
[0071] For example, if the initial center frequency is 200 Hz and the blue shift correction is 65 Hz, the dynamically adjusted center frequency will be 265 Hz. This frequency value is shifted 32.5% towards higher frequencies compared to the static filter. This significant adjustment ensures that the filter can accurately cover the true frequency range of crack characteristics even when high-speed motion causes severe spectral compression.
[0072] The dynamically adjusted center frequency changes in real time as the train accelerates, decelerates, or moves at a constant speed, forming a time-varying sequence of filter parameters. During acceleration, as the speed gradually increases, the absolute value of the spectral offset factor increases, and the center frequency adjusts accordingly to a higher frequency; during deceleration, the adjustment is reversed. This dynamic tracking capability ensures that the filter maintains optimal filtering characteristics throughout the entire operation, preventing degradation of filtering performance due to speed changes.
[0073] Based on the dynamically adjusted center frequency and preset bandwidth parameters, a complete dynamic bandpass filter transfer function is constructed. A bandpass filter behaves as a window function of a specific shape in the frequency domain, typically employing a Gaussian or Butterworth design. A Gaussian bandpass filter has a smooth frequency response curve and a wide transition band, avoiding the ringing effect caused by frequency domain truncation; a Butterworth filter, on the other hand, has a steeper cutoff characteristic, more effectively suppressing noise outside the passband.
[0074] In pantograph crack detection applications, a Gaussian design is typically chosen to achieve better time-frequency localization characteristics. The filter's transfer function is defined as a Gaussian curve with the dynamic center frequency as the peak value and the bandwidth parameter determining the full width at half maximum (FWHM). In practice, for each frequency point in the frequency domain, its distance from the center frequency is calculated, and then the Gaussian function is substituted to calculate the gain coefficient at that frequency point. The closer to the center frequency, the closer the gain is to 1, indicating that the frequency component is completely preserved; the farther the distance, the gain gradually decreases to 0, indicating that the frequency component is suppressed.
[0075] The constructed dynamic bandpass filter is stored as a two-dimensional matrix with the same size as the image spectrum, and each element corresponds to the filter gain at a specific frequency point. In the actual filtering operation, this filter matrix is multiplied element-wise with the image's spectral distribution to achieve frequency domain filtering. Through this dot-matrix operation, the crack feature frequency components within the passband are preserved, while low-frequency background noise and high-frequency random noise outside the passband are effectively suppressed, thereby achieving motion blur compensation and image quality improvement.
[0076] This embodiment solves the problem of unstable performance of traditional fixed-parameter filters in variable-speed scenarios by constructing an adaptive bandpass filter with a center frequency that dynamically adjusts with speed. The initial center frequency and bandwidth parameters, determined based on spectral statistical analysis of a large number of crack samples, ensure the scientific and targeted nature of the filter design, laying the foundation for accurate crack feature extraction. The mechanism of calculating the spectral shift factor based on real-time speed and converting it into a blue shift correction achieves a precise correlation between filter parameters and motion state, enabling the filter to automatically compensate for the spectral compression effect caused by motion blur. The strategy of dynamically adjusting the center frequency allows the filter to maintain optimal filtering characteristics under complex conditions such as train acceleration and deceleration, improving the robustness and adaptability of the system. The use of a Gaussian bandpass filter design effectively suppresses noise while avoiding the side effects of frequency domain truncation, ensuring the quality of the filtered image. The overall solution overcomes the image blurring problem in high-speed motion scenarios through frequency domain adaptive filtering technology, providing high-quality preprocessing results for subsequent crack feature extraction and recognition, and improving the real-time performance and accuracy of the detection system.
[0077] In one embodiment of this invention, the numerical magnitude of the high-dimensional feature space tensor is converted into amplitude, and the gradient direction of the high-dimensional feature space tensor is converted into phase, in order to construct a real-time object wave function, including the following steps: S410. Calculate the numerical value of each feature point in the high-dimensional feature space tensor as the amplitude component of the wave function; S420. Calculate the gradient vector of each feature point in the high-dimensional feature space tensor, and extract the direction angle of the gradient vector as the phase component of the wave function. S430: Combine the amplitude and phase components to construct a complex form of the real-time object wave function.
[0078] The sharpened feature map, after being processed by the convolutional neural network, is mapped to a high-dimensional feature space, forming a three-dimensional tensor structure with dimensions of height × width × number of channels. For each spatial location, there exists a multi-dimensional feature vector, whose components encode rich semantic information from the underlying texture to the mid-level shape.
[0079] When calculating the value of a feature point, first extract the complete feature vector corresponding to that location, then calculate the Euclidean norm of that vector, which is the square root of the sum of the squares of all its components. For example, if the feature vector of a certain feature point contains 256 channels, and the value of each channel is... The numerical value of that point is calculated as follows: This norm calculation method can comprehensively reflect the overall response intensity of all feature dimensions at a location. The larger the value, the more significant the feature at that location, and the higher the probability of the presence of target objects such as cracks.
[0080] To transform the numerical values into amplitude components suitable for wavefunction representation, normalization is required, mapping the numerical values of all feature points to the interval between 0 and 1. Normalization employs a max-min scaling method, finding the maximum and minimum values in the entire feature map, and then performing a linear transformation on the value of each feature point. The normalized amplitude components preserve the relative relationships of feature intensities at different locations while ensuring a uniform numerical range, facilitating subsequent wavefunction calculations.
[0081] Gradient information in the feature space reflects the spatial trend and directionality of the feature response, which is of great significance for identifying crack targets with directional features. When calculating the gradient vector, partial derivatives are taken with respect to the high-dimensional feature space tensor in both the horizontal and vertical spatial dimensions. Specifically, the Sobel operator or the central difference method is used. For each feature point, the differences between it and its neighboring pixels in each feature channel are calculated to obtain the horizontal and vertical gradients.
[0082] The combined horizontal gradient component is obtained by summing the horizontal gradients of all channels, and the combined vertical gradient component is obtained by summing the vertical gradients of all channels. These two components constitute the two-dimensional gradient vector of the feature point. The direction angle of the gradient vector is calculated using the arctangent function, specifically by dividing the vertical gradient component by the arctangent value of the horizontal gradient component; the result ranges from -π to π. To convert the direction angle into a phase component suitable for the wavefunction, it needs to be mapped to the interval 0 to 2π, achieved by adding π and taking the modulus.
[0083] Phase components encode the dominant direction of feature changes. For example, for a crack extending in a certain direction, its gradient direction remains relatively consistent along the crack path, and the corresponding phase components exhibit continuous changes. In contrast, the gradient direction in background noise regions is chaotic, and the phase components exhibit random distribution. This phase information provides important directional constraints for subsequent interference matching.
[0084] The amplitude and phase components are combined to construct a complex real-time object wavefunction, expressed using Euler's formula. For each spatial location (x, y) in the feature map, the corresponding wavefunction value is the amplitude component multiplied by the complex exponent of the phase, i.e. Where A(x,y) is the amplitude component at that position. is the phase component, and i is the imaginary unit.
[0085] This complex representation extends the characteristic information of the real number domain to the complex number domain, allowing for subsequent processing using the rich properties of complex number operations. The real-time object wavefunction is mathematically similar to the wave function in optics, with its amplitude corresponding to the intensity of the light wave and its phase corresponding to the phase of the light wave. Through this analogy, the pattern matching problem in the characteristic space can be transformed into an interference problem in wave optics.
[0086] The constructed real-time object wavefunction is stored in the form of a complex matrix. Each element of the matrix contains two components: a real part and an imaginary part. The real part is equal to the amplitude multiplied by the phase cosine, and the imaginary part is equal to the amplitude multiplied by the phase sine. This wavefunction representation not only preserves the intensity and direction information of the features but also establishes a bridge between the feature space and wave theory, creating conditions for efficient pattern matching using the principle of interferometry.
[0087] This embodiment transforms a high-dimensional feature space tensor into a real-time object wave function in complex form, combining deep learning feature representation with wave optics theory to provide a new method for image feature analysis and matching. By calculating the Euclidean norm of the feature vector as the amplitude component, the overall response intensity of multi-channel features is integrated, enabling accurate quantification of feature saliency. The direction angle of the gradient vector is extracted as the phase component, fully utilizing the spatial directionality information of the features and providing crucial evidence for identifying crack targets with directional characteristics. The use of a complex wave function representation unifies the encoding of amplitude and phase information, preserving the complete feature information and allowing for subsequent interferometric analysis using complex number operations. Compared to traditional real-domain features, this feature representation method has a richer mathematical structure and stronger expressive power, laying the foundation for efficient and accurate pattern matching and improving crack detection performance.
[0088] In one embodiment of this invention, the conjugate multiplication of the real-time object wavefunction and the standard reference wavefunction is performed in the characteristic space to calculate the interference intensity field, including the following steps: S510. Perform a conjugate operation on the standard reference wavefunction to obtain the conjugate reference wavefunction; S520. Add the real-time object wave function to the standard reference wave function using a complex number to obtain the superimposed wave function; S530. Calculate the square of the modulus of the superimposed wave functions to obtain the interference intensity field; S540. Multiply the real-time object wave function with the conjugate reference wave function by a complex number to obtain the phase matching degree distribution map; S550. Based on the phase matching degree distribution map and the numerical distribution of the interference intensity field, constructive interference regions and destructive interference regions are identified in the interference intensity field. Regions with interference intensity values higher than the average value are constructive interference regions, and regions with interference intensity values lower than the average value are destructive interference regions.
[0089] The standard reference wavefunction is obtained through deep learning training on a large number of labeled crack samples, representing the ideal wave pattern of typical crack features in the feature space. This reference wavefunction is also represented in complex form, containing two components: amplitude and phase. The amplitude distribution reflects the intensity pattern of the crack features, while the phase distribution reflects the morphology and orientation of the crack.
[0090] A conjugate operation is performed on the standard reference wavefunction to achieve interferometric matching. The conjugate operation is defined as keeping the real part of the complex number unchanged while inverting the imaginary part. For complex numbers... Its conjugate form is ,in For the amplitude of the reference wave function, The phase of the reference wave function.
[0091] From a physical perspective, conjugation is equivalent to reversing the propagation direction of the wavefunction, or inverting the phase rotation direction. In practical calculations, the conjugation operation is performed on each element of the reference wavefunction matrix. If the complex form of the element is... Then its conjugate is This conjugate operation provides the mathematical foundation for subsequent phase-matching analysis. When the real-time object wavefunction is multiplied by the conjugate reference wavefunction, if their phases are the same, the phase terms cancel each other out, resulting in a real number, indicating a perfect match; if their phases are different, the result contains an imaginary part, indicating a phase deviation. The conjugate reference wavefunction is stored in complex matrix form, having the same spatial dimension as the real-time object wavefunction, facilitating element-wise complex operations.
[0092] By complexly adding the real-time object wavefunction to the standard reference wavefunction, the physical process of superposition of two coherent waves in space was simulated. In wave optics, when two coherent light waves with the same frequency meet at a point in space, wave superposition occurs, and the result depends on the amplitude and phase relationship between the two waves.
[0093] For each location (x, y) in the feature space, the value of the real-time object wavefunction is The value of the standard reference wavefunction is The complex sum of the two yields the superimposed wave function. Expanding into complex form, the real part of the superposition wave function is... The imaginary part is .
[0094] This superposition operation is mathematically an element-by-element complex addition, but physically it reflects the principle of coherent superposition of waves. When the phase difference between two wave functions approaches zero or an integer multiple of 2π, that is... When two waves are in phase, their amplitude increases after superposition, a phenomenon known as constructive interference. When the phase difference is close to an odd multiple of π, i.e. When two waves are out of phase, their amplitudes decrease or even cancel each other out after superposition, a phenomenon known as destructive interference. The superimposed wavefunction completely preserves the phase information during the interference process, providing necessary intermediate results for subsequent calculations of the interference intensity field. In practical implementation, the superimposed wavefunction is also stored in the form of a complex matrix, where each element is obtained by complex addition of the corresponding real-time object wavefunction element and the reference wavefunction element.
[0095] The interference intensity field is a key physical quantity describing the wave interference result, reflecting the energy distribution of the superimposed waves. In optics, the light intensity is proportional to the square of the electric field amplitude, corresponding to the square of the modulus of the complex wave function. For the superimposed wave function S(x,y), the square of its modulus is calculated as follows: ,in yes . conjugate.
[0096] Expanding the calculation, the square of the modulus equals the square of the real part plus the square of the imaginary part, that is... Further simplification yields... This is the classic formula for the intensity of two-beam interference.
[0097] As can be seen from this formula, the interference intensity consists of three parts: the first term... It is the real-time intensity of the object wave itself, the second term. It is the intensity of the reference wave itself, the third term. The cosine term is the interference term, and its sign and magnitude depend on the phase difference between the two waves. When the phase difference is zero, the cosine value is 1, the interference term is positive and reaches its maximum value, the total intensity is the maximum, and constructive interference occurs; when the phase difference is π, the cosine value is -1, the interference term is negative and reaches its minimum value, the total intensity is the minimum, and destructive interference occurs.
[0098] The calculated interference intensity field is a real-number matrix, where each element represents the interference intensity at the corresponding spatial location. The spatial distribution of this intensity field exhibits a pattern of alternating bright and dark interference fringes. High-intensity regions correspond to locations where the real-time and reference features are highly matched, while low-intensity regions correspond to locations with significant feature differences. By analyzing the distribution pattern of the interference intensity field, suspected crack regions can be quickly located. Compared to traditional point-by-point feature comparison methods, the interferometry method can utilize both amplitude and phase information simultaneously, resulting in higher matching accuracy and better computational efficiency.
[0099] In addition to calculating the interferometric intensity field, a separate phase-matching degree distribution map is also needed to more precisely assess the phase consistency between the real-time and reference features. This is achieved by complex multiplying the real-time object wavefunction and its conjugate reference wavefunction. .
[0100] The result of this complex multiplication operation has an amplitude equal to the product of the amplitudes of the two wave functions. The phase is equal to the difference between the phases of the two wave functions. The phase difference directly reflects the degree of directional matching between real-time features and reference features. When the phase difference is close to zero, it indicates that the gradient directions of the two are highly consistent and the feature patterns are well matched; when the phase difference is close to π, it indicates that the gradient directions of the two are opposite and the feature patterns are mismatched.
[0101] Real part of the phase matching degree distribution diagram It reflects the strength of phase consistency; a larger real part value indicates a higher degree of matching; the imaginary part... This reflects phase orthogonality, with the imaginary part approaching zero indicating good phase alignment. By analyzing the phase matching degree distribution map, regions that match not only in intensity but also in directional characteristics can be identified; these regions are high-confidence candidates for real cracks. Phase matching degree analysis provides supplementary information to the interference intensity field because interference intensity alone cannot distinguish whether constructive interference is due to genuine characteristic matching or accidental phase coincidence. The phase matching degree distribution map, by explicitly calculating the phase difference, provides a more reliable basis for judgment.
[0102] Based on a comprehensive analysis of the interference intensity field and phase matching degree distribution map, constructive and destructive interference regions can be accurately identified. First, the global statistical characteristics of the interference intensity field are calculated, including the mean, standard deviation, maximum, and minimum values. The mean value represents the average interference intensity level across the entire feature map, serving as a benchmark threshold for distinguishing between constructive and destructive interference.
[0103] Iterate through each pixel in the interference intensity field and compare its intensity value with the average value. If the intensity value is significantly higher than the average value, for example, exceeding the average value plus one standard deviation, then the point is marked as a candidate point of the constructive interference region; if the intensity value is significantly lower than the average value, for example, lower than the average value minus one standard deviation, then it is marked as a candidate point of the destructive interference region.
[0104] However, the intensity of interference alone is insufficient for accurate judgment, as high intensity may be caused by accidental superposition of background noise, and low intensity may be due to weak features rather than true destructive interference. Therefore, secondary verification using a phase matching degree distribution map is necessary. For candidate points marked as constructive interference, their real part values in the phase matching degree distribution map are checked. If the real part value is also large, it indicates that the phase difference is close to zero, confirming it as true constructive interference; if the real part value is small or even negative, it may be a misjudgment, and its confidence level needs to be lowered. Similarly, for candidate points marked as destructive interference, their phase difference is checked to see if it is close to π; if so, it is confirmed as true destructive interference.
[0105] This dual verification mechanism effectively eliminates noise interference and improves identification accuracy. The identified constructive interference regions often exhibit a continuous spatial distribution, forming slender strip-like or curved structures. The geometric features of these structures highly match the morphology of cracks. Connectivity analysis of the constructive interference regions extracts continuous high-intensity areas, which are the candidate defect substrate regions with mid-level shape semantic features. Destructive interference regions correspond to background regions with feature mismatches and can be directly excluded without further analysis. Through joint analysis of the interference intensity field and phase matching degree, efficient pattern matching in the feature space is achieved, improving the detection accuracy and computational efficiency of crack candidate regions.
[0106] This embodiment achieves efficient and accurate crack candidate region localization by performing interference operations between the real-time object wavefunction and the standard reference wavefunction in the feature space, applying the interference principle of wave optics to pattern matching of deep learning features. Conjugate operations on the reference wavefunction establish a mathematical foundation for subsequent phase matching analysis, allowing direct extraction of phase difference information through complex number operations. The complex addition of wavefunctions simulates the coherent superposition process, fully preserving amplitude and phase information during interference, laying the foundation for accurate calculation of interference intensity. The square of the superimposed wavefunction modulus yields the interference intensity field, which intuitively reflects the matching degree between the real-time feature and the reference feature; high-intensity regions correspond to high-matching positions, providing an effective means for rapid candidate region selection. The complex multiplication of the real-time object wavefunction and the conjugate reference wavefunction yields a phase matching degree distribution map, providing a fine evaluation of feature directional matching and an important supplement to interference intensity analysis. The interference intensity field and phase matching degree distribution map are combined to identify constructive and destructive interference regions. A dual verification mechanism effectively eliminates noise interference, improving the accuracy of candidate region identification. The overall solution transforms the abstract feature matching problem into an intuitive wave interference problem, making full use of the mathematical advantages of complex number operations and the physical characteristics of interference phenomena. Compared with traditional feature matching methods, it has improved both computational efficiency and matching accuracy, providing high-quality candidate regions for subsequent high-level semantic analysis and ensuring the overall performance of crack detection.
[0107] In one embodiment of this invention, extracting regions with energy values above a preset energy threshold as candidate defect substrate regions with mid-level shape semantic features includes the following steps: S610. Traverse all pixels in the interference intensity field and obtain the interference intensity value of each pixel; S620. Compare the interference intensity value with the preset energy threshold and mark all pixels whose interference intensity value is higher than the preset energy threshold. S630. Perform connected component analysis on the marked pixels to extract continuous high-energy regions; S640. The extracted continuous high-energy regions are used as candidate defect substrate regions with mid-level shape semantic features.
[0108] The interferometric intensity field is stored as a two-dimensional matrix, with its row and column dimensions corresponding to the spatial resolution of the original image. Each matrix element represents the interferometric intensity value at its corresponding spatial location. The traversal operation employs a double-loop structure: the outer loop iterates through the row indices, and the inner loop iterates through the column indices, ensuring that every pixel in the interferometric intensity field is accessed. For a resolution of [missing information], [missing information]. The images require a total of [number] visits. Each pixel.
[0109] During the traversal, for each pixel coordinate (x, y), the corresponding value I(x, y) in the interference intensity field matrix is read. This value is the interference intensity value at that point. The interference intensity value is a real number, and its magnitude reflects the degree of matching between the real-time feature at that location and the reference crack feature. The larger the value, the higher the matching degree and the greater the probability of the existence of a crack.
[0110] To improve traversal efficiency, matrix vectorization is employed in the actual implementation. The two-dimensional matrix is flattened into a one-dimensional array, and all pixels are accessed in a single loop, avoiding the computational overhead of nested loops. Simultaneously, a mapping table between pixel coordinates and interference intensity values is established during traversal, facilitating subsequent threshold comparison and region extraction operations. The traversal operation also synchronously collects global statistics of the interference intensity field, including maximum, minimum, mean, and standard deviation. These statistics provide a reference for the dynamic adjustment of the preset energy threshold. By fully traversing the interference intensity field, precise intensity information for each pixel is obtained, laying the data foundation for subsequent threshold selection.
[0111] The preset energy threshold is a key parameter for distinguishing high-energy candidate regions from low-energy background regions, and its setting needs to strike a balance between detection sensitivity and false positive rate. Setting the threshold too low will cause a large amount of background noise to be misclassified as candidate regions, increasing the burden of subsequent processing and the false positive rate; setting the threshold too high may miss real micro-cracks and reduce the detection recall rate.
[0112] In practical applications, the preset energy threshold is usually determined dynamically based on the statistical characteristics of the interference intensity field using an adaptive strategy. A common method is to set the threshold as the mean of the interference intensity field plus a certain multiple of the standard deviation, for example... ,in The mean, Here, k represents the standard deviation, and k is the adjustment coefficient, typically ranging from 1.5 to 3. A larger k value increases the threshold, ensuring that only pixels significantly above the average level are selected, reducing false positives but potentially increasing false negatives; a smaller k value lowers the threshold, increasing detection sensitivity but potentially increasing false positives.
[0113] When traversing each pixel, its interference intensity value I(x,y) is compared with a preset energy threshold T. If I(x,y) > T, the pixel is marked as a high-energy point. The marking operation is implemented by creating a binary mask matrix of the same size as the interference intensity field. For pixels that meet the conditions, a value of 1 is assigned to the corresponding position in the mask matrix, indicating that the point is selected; for pixels that do not meet the conditions, a value of 0 is assigned, indicating that the point is excluded.
[0114] This binarization process transforms the continuous interference intensity field into a discrete candidate point distribution map, significantly simplifying subsequent region extraction operations. The marking process also counts the total number of selected pixels, which reflects the richness of suspected crack features in the current frame image and can serve as an auxiliary indicator for judging image quality and the likelihood of crack presence.
[0115] High-energy marked pixels often exhibit spatial clustering, forming several continuous regions. These regions correspond to local structures in the real-time image similar to the features of the reference crack. Connected component analysis is a classic image processing technique used to identify and extract interconnected sets of pixels in a binary image. In this application, connected component analysis is performed on the marked mask matrix to identify all interconnected high-energy pixel groups.
[0116] Connectivity determination uses the 8-neighborhood criterion, which considers a pixel to be connected to its eight neighboring pixels in eight directions: four orthogonal directions (up, down, left, right) and four diagonal directions. Connectivity analysis is typically implemented using a two-pass scanning algorithm or a region growing algorithm based on seed filling. The two-pass scanning algorithm assigns a temporary label to each foreground pixel in the first pass and records equivalent label pairs; in the second pass, labels are merged according to equivalence relationships to obtain the final connected component labels. The region growing algorithm starts from each unvisited foreground pixel and traverses all its connected neighboring pixels using breadth-first search or depth-first search, marking them as belonging to the same connected component.
[0117] The result of connected component analysis is a label matrix, where pixels with the same label value belong to the same connected component. The algorithm also outputs statistical information such as the number of connected components, the number of pixels contained in each connected component, and the bounding box coordinates of the connected components. For each connected component, the coordinates of all its contained pixels are extracted to form a continuous high-energy region.
[0118] These regions may exhibit various geometric shapes, including elongated strips, curved lines, or irregular patches. Elongated and curved regions are more consistent with the morphological characteristics of cracks, while circular or square patchy regions are more likely to be background noise or other interference. For further screening, geometric feature filtering can be applied to the extracted connected regions. For example, connected regions with excessively small areas (which may be isolated noise points) or aspect ratios close to 1 (which do not conform to the elongated characteristics of cracks) can be removed, retaining those connected regions with moderate areas and elongated shapes as valid candidate regions.
[0119] After connected component analysis and geometric feature filtering, the resulting continuous high-energy regions were formally identified as candidate defect substrate regions. These candidate regions have undergone sharpening processing of low-level texture semantic features and interference matching of mid-level shape semantic features, and have a high probability of being cracked, but still require subsequent high-level semantic category recognition for final confirmation.
[0120] Each candidate defect substrate region is stored in structured data format, including the region's bounding box coordinates, a list of contained pixels, region area, perimeter, centroid location, major axis direction, aspect ratio, and other geometric attributes. These attributes are used not only for subsequent topological skeleton extraction and semantic classification, but also for the visualization of candidate regions and manual review.
[0121] In practical applications, the number of candidate defect substrate regions is usually much smaller than the number of pixels in the original image. For example, for a 2048×2048 image, only a few dozen candidate regions may be extracted, each containing hundreds to thousands of pixels. This significant dimensionality reduction allows subsequent depth analysis to focus on a small number of key regions, avoiding the huge computational overhead of processing the entire image pixel by pixel.
[0122] The extraction of candidate defect substrate regions transforms dense pixel-level representations into sparse region-level representations, simplifying the problem of rapid image content screening from global search to local fine-grained analysis and improving processing efficiency. Simultaneously, through energy threshold screening of the interference intensity field and connected component analysis, candidate regions possess clear mid-level shape semantic features—that is, spatial distribution features highly similar to standard crack patterns in the feature space. This provides high-quality input for subsequent high-level semantic recognition, ensuring the accuracy and reliability of the overall detection process.
[0123] This embodiment achieves efficient transformation from dense interference intensity fields to sparse candidate defect substrate regions through a systematic energy threshold screening and connected component analysis process. By traversing the interference intensity field to obtain precise intensity information for each pixel, a complete data foundation is provided for subsequent threshold judgment, ensuring the comprehensiveness of the screening process. An adaptive energy threshold strategy is adopted, dynamically determining threshold parameters based on the statistical characteristics of the interference intensity field, achieving a good balance between detection sensitivity and false detection rate. This allows threshold screening to effectively eliminate low-energy background noise while retaining true crack features. Connected component analysis identifies spatially continuous high-energy pixel clusters, transforming discrete pixel markers into continuous region representations. This fully utilizes the spatial continuity and clustering of crack features, improving the completeness and accuracy of candidate regions. Combined with geometric feature filtering, interference regions that do not conform to crack morphology are further eliminated, enhancing the robustness of the screening. The extracted continuous high-energy regions are identified as candidate defect substrate regions with mid-level shape semantic features, achieving data dimensionality reduction from pixel-level to region-level, significantly reducing the amount of data requiring in-depth analysis and improving the computational efficiency of subsequent processing. The overall solution, through a multi-level screening and analysis mechanism, optimizes the allocation of computing resources while ensuring high detection accuracy, providing effective technical support for pantograph detection applications in high-speed railways with stringent real-time requirements.
[0124] In one embodiment of this example, for each candidate defect substrate region, a geometric topological skeleton is extracted, including the following steps: S710. Binarize the candidate defect substrate region; S720: Apply morphological thinning algorithm to gradually erode the boundary of the candidate defect substrate region after binarization until a skeleton structure with a single pixel width is obtained. S730, Extract the endpoints, branching points, and connection paths of the skeleton structure; S740. Based on endpoints, bifurcation points, and connection paths, construct the geometric topological skeleton of the candidate defect substrate region as a structural feature for high-level semantic classification.
[0125] When the candidate defect substrate region is extracted from the interference intensity field, the original grayscale intensity information is preserved. Each pixel has a different intensity value, reflecting the degree of matching between that location and the reference crack feature. To facilitate subsequent morphological processing and topological analysis, the candidate region needs to be converted into a binary image, containing only foreground and background pixel types.
[0126] The binarization process employs a local adaptive thresholding method, independently calculating the statistical characteristics of pixel intensity values within each candidate region. Specifically, it first calculates the mean and standard deviation of all pixel intensity values within the candidate region, then sets the binarization threshold to the mean minus 0.5 times the standard deviation. This setting ensures that high-intensity pixels at the region center are retained as foreground, while low-intensity pixels at the edges are removed as background.
[0127] For each pixel, if its intensity value is higher than a local threshold, it is marked as a foreground pixel in the binary image and assigned a value of 1; if the intensity value is lower than the threshold, it is marked as a background pixel and assigned a value of 0. This local adaptive strategy has better robustness than a globally fixed threshold and can adapt to differences in intensity distribution between different candidate regions. The binarized candidate regions exhibit clear foreground-background segmentation, with the foreground region corresponding to the core part of the suspected crack, and its shape and topology are preserved, providing clean input data for subsequent skeleton extraction.
[0128] Morphological thinning is a classic skeleton extraction method that gradually peels away the outer pixels of a region through iterative erosion operations while ensuring that the connectivity and topological properties of the region are not destroyed, ultimately shrinking the region into a central axis structure with a width of one pixel.
[0129] The refinement process employs an iterative algorithm based on structuring elements. In each iteration, all foreground pixels are traversed, and it is determined whether a pixel meets the deletion criteria. The deletion criteria include three constraints: first, the region remains connected after deleting the pixel and does not split into multiple independent parts; second, deleting the pixel does not change the region's topological characteristics, such as maintaining its Euler number; and finally, the pixel is located on the boundary of the region, not inside it. Pixels that meet all the conditions are marked as points to be deleted and are deleted uniformly after the current iteration ends.
[0130] Thinning algorithms typically employ two alternating sub-iterations. The first sub-iteration processes boundary points in a specific direction, while the second sub-iteration processes boundary points in another direction. This alternation strategy ensures the symmetry and centrality of the skeleton. The iterative process continues until no pixel meets the deletion criteria, at which point the region has shrunk to a skeleton structure with a single pixel width.
[0131] For slender crack regions, the skeleton typically appears as one or more continuous curves distributed along the crack's extension direction; for branching cracks, the skeleton forms intersection nodes at the branches. Geometrically, the skeleton structure preserves the main shape features of the original region; topologically, it retains the region's connectivity and branching structure, while significantly simplifying data representation and providing a compact and information-rich representation for subsequent topological analysis.
[0132] The topological features of the skeleton structure are mainly described by three types of key nodes: endpoints, bifurcation points, and connection paths. An endpoint is the termination position of the skeleton, corresponding to the start or end of a crack; in the skeleton image, it is represented by a skeleton point with only one neighboring pixel. When identifying endpoints, all skeleton pixels are traversed, and for each point, the number of skeleton pixels in its 8-neighborhood is counted. If the count is equal to 1, then that point is an endpoint.
[0133] A bifurcation point is the intersection of multiple skeleton branches, reflecting the branching or crossing characteristics of the crack. In a skeleton image, it is represented by the location of three or more neighboring skeleton points. When identifying bifurcation points, the number of neighboring skeleton points for each skeleton point is also counted. If the number is greater than or equal to 3, then the point is a bifurcation point.
[0134] A connection path is a skeleton segment between an endpoint and a bifurcation point, or between two bifurcation points, representing the direction and length of crack extension. When extracting connection paths, starting from each endpoint or bifurcation point, the path is traced along the skeleton until another endpoint or bifurcation point is reached. During the tracing process, the coordinates of all skeleton pixels traversed are recorded, forming a complete connection path. For each connection path, its geometric properties, such as length (number of pixels), average direction (obtained through principal component analysis), and curvature change (calculated through the direction changes of adjacent path segments), are calculated. By extracting endpoints, bifurcation points, and connection paths, the topological structure of the skeleton is fully described. These topological elements not only contain geometric information but also rich semantic information, providing key features for high-level semantic classification.
[0135] Based on the extracted endpoints, bifurcation points, and connecting paths, a geometric topological skeleton representation of the candidate defect substrate region is constructed. This representation adopts a graph structure, with endpoints and bifurcation points as nodes and connecting paths as edges. Each node stores attributes such as its spatial coordinates and node type (endpoint or bifurcation point); each edge stores attributes such as the two nodes it connects, path length, the sequence of pixel coordinates on the path, average direction, and curvature statistics. This graph structure representation not only preserves the complete topological information of the skeleton but also facilitates analysis using various graph theory algorithms.
[0136] Based on a geometric topological skeleton, various high-level semantic features can be calculated for classification. For example, the bifurcation rate feature is defined as the ratio of the number of bifurcation points to the total length of the skeleton, reflecting the complexity of the crack. Real cracks usually have a moderate bifurcation rate; too high a rate may be due to background texture interference, while too low a rate may indicate a simple linear structure. The curvature continuity feature is evaluated by analyzing the rate of curvature change of the connection paths. Crack propagation usually follows the material stress distribution law, with relatively smooth and continuous curvature changes, while background interference often presents irregular curvature jumps. The aspect ratio feature is obtained by calculating the aspect ratio of the minimum bounding rectangle of the skeleton. Cracks usually exhibit slender characteristics, with an aspect ratio much greater than 1.
[0137] In addition, topological statistics such as the total length of the skeleton, average path length, longest path length, and node degree distribution can be calculated. The constructed geometric topological skeleton serves as the structural feature representation of the candidate region and is input into a high-level semantic classifier for category determination. Based on these topological and geometric features, combined with pre-learned crack pattern knowledge, the classifier determines whether the candidate region is a real crack or background interference, thereby achieving accurate defect identification.
[0138] This embodiment utilizes a systematic geometric topological skeleton extraction process to transform candidate defect substrate regions from pixel-level representations to compact topological structure representations, providing high-quality structural features for high-level semantic classification. Binarization effectively separates the foreground and background of candidate regions, eliminating interference from intensity variations and providing clear input for subsequent morphological processing. The morphological thinning algorithm shrinks the region into a single-pixel skeleton while preserving topological characteristics, significantly simplifying data representation while retaining the region's main shape and connectivity features, achieving efficient information compression. By extracting topological elements such as endpoints, bifurcation points, and connection paths, the structural characteristics of the skeleton are fully described. These elements contain not only geometric information but also semantic information, providing crucial evidence for distinguishing cracks and interference. The constructed geometric topological skeleton is represented in graph form, facilitating the calculation of various topological and geometric features and supporting accurate judgment by high-level semantic classifiers. The overall scheme, through multi-level abstraction from pixels to skeleton to topological graph, achieves an effective conversion from low-level visual features to high-level semantic features, improving the accuracy and robustness of crack recognition and providing a technical foundation for intelligent pantograph defect detection.
[0139] In one embodiment of this invention, high-level semantic category recognition is performed based on a geometric topological skeleton, and an activation response value is determined based on the high-level semantic category recognition result, including the following steps: S801. Construct a high-level semantic classifier, which includes a crack bifurcation rate constraint operator, a curvature continuity constraint operator, and an aspect ratio constraint operator, used to classify candidate defect substrates into crack categories or non-crack categories. S802. Based on the geometric topological skeleton, calculate the ratio of the number of bifurcation points in the candidate defect substrate region to the total length, and use it as the actual bifurcation rate. S803. Calculate the matching degree between the actual bifurcation rate and the standard bifurcation rate range defined by the crack bifurcation rate constraint operator to obtain the bifurcation rate matching component. S804. Calculate the rate of curvature change of the candidate defect substrate region based on the geometric topological skeleton. S805. Calculate the degree of matching between the rate of change of curvature and the standard curvature continuity defined by the curvature continuity constraint operator to obtain the curvature fitting component. S806. Calculate the aspect ratio of the candidate defect substrate region based on the geometric topological skeleton; S807. Calculate the matching degree between the aspect ratio and the standard aspect ratio range defined by the aspect ratio constraint operator to obtain the aspect ratio matching component; S808. Perform high-level semantic category recognition on the bifurcation rate matching component, curvature matching component and aspect ratio matching component to obtain induced matching degree, wherein the induced matching degree is used to characterize the high-level semantic category recognition result. S809. Calculate the reduction in activation energy based on the induced fit, where the higher the induced fit, the greater the reduction in activation energy. S810. Subtract the activation energy reduction from the preset baseline activation energy threshold to obtain the actual activation energy. S811. Substitute the actual activation energy into the exponential decay function to calculate the reaction rate constant; S812, Output the reaction rate constant as the activation response value.
[0140] The high-level semantic classifier is constructed based on in-depth analysis of a large number of real crack samples and background interference samples, extracting key feature constraints that can effectively distinguish between the two types of samples. The classifier adopts a multi-constraint operator collaborative discrimination architecture, which includes three core constraint operators, each constraining an essential characteristic of the crack.
[0141] The crack bifurcation rate constraint operator is designed based on the physical mechanism of crack propagation. When a real crack propagates within a material, it extends along the stress concentration direction. When it encounters material defects or complex stress fields, it may bifurcate. However, the bifurcation frequency is constrained by material properties and stress distribution, and is usually kept within a specific range. Through statistical analysis of the bifurcation characteristics of hundreds of real crack samples, a standard bifurcation rate range of 0.5 to 2 bifurcation points per 100 pixels was determined. An excessively high bifurcation rate often corresponds to complex background texture interference, while an excessively low rate may indicate a simple linear structure rather than a crack.
[0142] The curvature continuity constraint operator is designed based on the continuity characteristics of the crack path. The crack extends along the direction of the maximum principal stress inside the material, and its path curvature change is usually smooth and continuous, without abrupt changes or violent oscillations. By analyzing the curvature distribution of the crack skeleton, the standard curvature continuity is defined as the rate of curvature change of adjacent skeleton segments being less than a preset threshold, typically set to no more than 15 degrees of curvature change within every 10 pixels.
[0143] The aspect ratio constraint operator is designed based on the geometric morphology of cracks. Cracks, as linear defects on the material surface, have a length much greater than their width, exhibiting a significant elongated characteristic. Through statistical analysis, the standard aspect ratio range is determined to be between 5 and 50. Regions with excessively small aspect ratios may be speckled noise, while those with excessively large aspect ratios may be image edges or other linear structures. Three constraint operators work together to comprehensively evaluate candidate regions from three dimensions: topological complexity, path continuity, and geometric morphology, forming the core discrimination mechanism of the high-level semantic classifier.
[0144] Based on the extracted geometric topological skeleton, the actual bifurcation rate characteristics of the candidate defect substrate region are calculated. First, the total number of bifurcation points in the skeleton is counted. Each bifurcation point corresponds to a node with a degree greater than or equal to 3 in the topological graph structure; each such node represents a branch or intersection of the crack path. Then, all nodes in the topological graph are traversed, and the number of nodes satisfying the degree condition is counted to obtain the total number of bifurcation points. .
[0145] Then, the total length of the skeleton is calculated by summing the lengths of all edges in the topology graph. The length of each edge represents the number of pixels contained in the corresponding connection path. The total length of the skeleton is obtained by summing the lengths of all edges. The actual bifurcation rate is defined as the ratio of the number of bifurcation points to the total length of the skeleton, i.e. This ratio reflects the average bifurcation density per unit length of skeleton.
[0146] To facilitate comparison with a standard bifurcation rate range, the actual bifurcation rate is typically normalized to the number of bifurcation points per 100 pixels. For example, if a candidate region's skeleton contains 3 bifurcation points and its total length is 250 pixels, then the actual bifurcation rate is 3 / 250 = 0.012, which, after normalization, is 1.2 bifurcation points per 100 pixels. This normalization process makes the bifurcation rates of candidate regions of different sizes comparable, facilitating a unified threshold judgment. The actual bifurcation rate, as a quantitative indicator of the topological complexity of the candidate region, provides basic data for subsequent matching degree calculations.
[0147] Calculate the matching degree between the actual bifurcation rate and the standard bifurcation rate range to evaluate whether the bifurcation characteristics of the candidate region conform to the statistical laws of real cracks. The standard bifurcation rate range is defined by the crack bifurcation rate constraint operator, and its lower bound is set as follows. The upper boundary is ,For example , (per 100 pixels).
[0148] The matching degree calculation uses a fuzzy membership function, when the actual bifurcation rate... When the value falls within the standard range, the matching degree is 1, indicating a perfect match; when the actual forking rate deviates from the standard range, the matching degree decreases with the degree of deviation. The specific calculation is divided into three cases: if... Then the matching degree ;like Then the matching degree ,in This is a tolerance parameter that controls the rate at which the matching degree decays; it is typically set to 1 / 3 of the standard range width. Then the matching degree .
[0149] This Gaussian decay function ensures that the matching degree is at its maximum within the standard range, and decreases smoothly with the deviation distance outside the range, avoiding the discontinuity of hard threshold judgment. The calculated matching degree value range is [0,1], which is output as the bifurcation rate matching component. This component quantifies the similarity between the candidate region and the real crack in the dimension of topological complexity. The higher the component value, the more the bifurcation characteristics match the crack mode.
[0150] The curvature change rate of candidate defect substrate regions is calculated based on a geometric topological skeleton to evaluate the continuity characteristics of crack paths. For each connection path in the skeleton, it is discretized into an ordered sequence of pixel coordinates. At each interior point At a certain point, the local curvature is calculated using a three-point circle fitting method, utilizing the point and its adjacent points before and after it. and If a circle is fitted, the reciprocal of the circle's radius of curvature is the curvature at that point. .
[0151] The sign of curvature reflects the direction of the path's bending; a positive value indicates bending to one side, and a negative value indicates bending to the other side. For all interior points on the path, the curvature difference between adjacent points is calculated. This difference reflects the magnitude of the curvature change. Summing all curvature differences and dividing by the total path length yields the average rate of curvature change for that path. ,in This represents the path length.
[0152] For a skeleton containing multiple connecting paths, calculate the weighted average of the curvature change rates of all paths, with the weights being the proportion of each path's length to the total length, to obtain the overall curvature change rate of the candidate region. ,in A smaller rate of curvature change indicates a smoother and more continuous path, consistent with the characteristic of cracks propagating steadily along the stress direction; a larger rate of curvature change indicates a drastic change in the path's direction, which is more likely to be due to background texture interference.
[0153] The degree of matching between the rate of curvature change and the standard curvature continuity is calculated to evaluate whether the path smoothness of the candidate region conforms to crack characteristics. The standard curvature continuity defined by the curvature continuity constraint operator is represented as a threshold. This represents the upper limit of the typical rate of curvature change of a real crack, and is usually set to no more than 0.05 radians per unit length.
[0154] The matching degree calculation also uses a fuzzy membership function, when the actual rate of curvature change... When the value is less than the standard threshold, it indicates a smooth and continuous path with a high degree of matching; when the rate of curvature change exceeds the threshold, the matching degree decreases with the degree of excess. The specific calculation formula is: If Then the matching degree ;like Then the matching degree ,in This is a tolerance parameter that controls the sensitivity of the matching degree to excessive curvature change rate; it is usually set to 1 / 2 of the standard threshold.
[0155] This unilateral constraint design reflects the essential characteristic of crack path continuity: the smaller the rate of curvature change, the better, with no lower bound constraint. The calculated matching degree range is [0,1], which is output as the curvature matching component. This component quantifies the similarity between the candidate region and the real crack in the path continuity dimension; a higher component value indicates a smoother path, which better conforms to the physical propagation law of the crack. The curvature matching component and the bifurcation rate matching component together constitute a multi-dimensional evaluation of the topological and geometric properties of the candidate region.
[0156] The aspect ratio of candidate defect substrate regions is calculated based on a geometric topological skeleton to evaluate whether their geometry exhibits the elongated characteristics of cracks. The aspect ratio calculation employs the minimum bounding rectangle method, which accurately captures the main extension direction and geometric dimensions of the region. First, the coordinates of all pixels in the skeleton are extracted to form a point set. Principal component analysis is performed on the point set to calculate the eigenvalues and eigenvectors of the covariance matrix. The eigenvector corresponding to the largest eigenvalue indicates the main extension direction of the skeleton, i.e., the major axis direction.
[0157] Rotate the coordinate system to the principal axis direction, making the major axis parallel to the horizontal axis of the new coordinate system. In the new coordinate system, calculate the projection range of all points along the horizontal and vertical axes. The span of the horizontal projection range is the length of the minimum bounding rectangle. The span of the vertical axis projection range is the width. Aspect ratio is defined as... For elongated crack regions, the aspect ratio is typically large, such as between 10 and 30; for circular or square interference regions, the aspect ratio is close to 1. It's important to note that extremely elongated regions (aspect ratio greater than 50) may correspond to image edges or other non-crack linear structures and should also be excluded. The aspect ratio, as a compact representation of the candidate region's geometry, provides crucial shape features for subsequent matching degree calculations.
[0158] Calculate the degree of matching between the aspect ratio and the standard aspect ratio range, and evaluate whether the geometry of the candidate region conforms to the slender characteristics of the crack. The standard aspect ratio range defined by the aspect ratio constraint operator is set to a lower bound of [value missing]. The upper boundary is ,For example , .
[0159] The matching degree calculation uses a two-sided constrained fuzzy membership function similar to the bifurcation rate. When the actual aspect ratio (AR) falls within the standard range, the matching degree is 1; when the aspect ratio deviates from the standard range, the matching degree decreases with the degree of deviation. The specific calculation is divided into three cases: if... Then the matching degree ;like Then the matching degree ;like Then the matching degree ,in This is a tolerance parameter, typically set to 1 / 4 of the standard range width.
[0160] This design ensures that candidate regions with a reasonable aspect ratio achieve a high matching degree, while regions that are too rounded or too elongated have a lower matching degree. The calculated matching degree value range is [0,1], which is output as the aspect ratio matching component. This component quantifies the similarity between the candidate region and the real crack in the geometric dimension, and together with the first two matching components, constitutes a comprehensive feature evaluation of the candidate region.
[0161] For bifurcation rate matching components Curvature fit component The proportions match the weight. A comprehensive analysis was conducted to obtain the induced fit score, which comprehensively characterizes the similarity between candidate regions and real cracks at a high-level semantic level. The induced fit score was calculated using a weighted geometric mean method, which has stronger constraints than the arithmetic mean. This method requires each fit component to reach a high level in order to obtain a high induced fit score, thus avoiding the situation where a high score in one dimension masks a low score in other dimensions.
[0162] The calculation formula is ,in The weighting coefficients for each fitting component reflect the importance of different feature dimensions in crack detection. In practical applications, based on statistical analysis and cross-validation of a large number of samples, the weighting allocation is determined as follows: The curvature fit component has a slightly higher weight because path continuity is the most essential physical characteristic of a crack.
[0163] The mathematical properties of geometric mean ensure that if any matching component is close to zero, the induced fit is also close to zero, achieving strict multi-dimensional constraints. The range of the induced fit is [0,1]. The closer the value is to 1, the more the candidate region closely matches the crack characteristics in the three dimensions of topological complexity, path continuity, and geometric morphology, and the greater the probability that it is a real crack. The closer the value is to 0, the more significant the deviation is in at least one dimension, and the more likely it is background interference. As the core result of high-level semantic category recognition, the induced fit provides a quantitative basis for the subsequent calculation of activation response values.
[0164] The reduction in activation energy is calculated based on the induced fit, establishing a mapping relationship between high-level semantic recognition results and chemical reaction kinetics models. In chemical reaction kinetics, reactant molecules need to overcome a certain energy barrier (activation energy) to transform into products. The role of a catalyst is to lower the activation energy, making the reaction easier to occur. Analogous to crack detection, the classification of candidate regions can be regarded as a reaction process. Candidate regions with high induced fit are equivalent to catalyzed reactants, making it easier for them to cross the discrimination threshold and be identified as cracks.
[0165] A positive correlation was established between the reduction in activation energy and the induced fit. The higher the induced fit, the more closely the candidate region's characteristics conform to the crack mode, which is equivalent to a stronger catalytic effect and a greater reduction in activation energy. The specific calculations employed a nonlinear mapping function to ensure that the reduction in activation energy accelerates with increasing induced fit, reflecting that high-fit candidate regions should achieve more significant catalytic advantages.
[0166] The calculation formula is ,in This represents the maximum reduction in activation energy, signifying the upper limit of activation energy reduction under perfect matching conditions, and is typically set to 70%–80% of the baseline activation energy. The kurtosis parameter controls the rate at which the activation energy reduction increases with induced fit, typically set to 3-5. This exponential mapping function is characterized by: slow activation energy reduction when induced fit is low; rapid reduction when induced fit exceeds a certain threshold (approximately 0.5); and near-maximum activation energy reduction when induced fit approaches 1. This non-linear characteristic aligns with practical needs, ensuring that candidate regions only achieve significant discriminative advantages when performing well across multiple feature dimensions.
[0167] Subtracting the activation energy reduction from the preset baseline activation energy threshold yields the actual activation energy for the current candidate region. (Baseline activation energy threshold) It is a global parameter that represents the energy barrier that a candidate region needs to overcome to be identified as a crack in the absence of any feature matching information. This value is determined through statistical analysis of a large number of samples and is set to a high level to ensure that only candidate regions that truly possess crack characteristics can pass the discrimination.
[0168] The formula for calculating the actual activation energy is as follows: This formula reflects the impact of feature matching on the difficulty of discrimination: the higher the induced fit, the greater the reduction in activation energy, and the lower the actual activation energy, indicating that the candidate region is more easily identified as a crack. The lower limit of the actual activation energy is... This corresponds to the ideal case where the induced fit is 1; the upper limit is... This corresponds to the worst-case scenario where the induced fit is 0. Through this dynamic adjustment mechanism, different candidate regions obtain differentiated activation energies based on their feature matching degree, achieving personalized discrimination criteria. The actual activation energy, as a physical quantity reflecting the reactivity of the candidate region, provides a key parameter for subsequent calculation of the reaction rate constant, establishing a quantitative bridge from high-level semantic features to the final discrimination result.
[0169] Substituting the actual activation energy into the exponential decay function, the reaction rate constant is calculated. This constant quantifies the reaction rate at which the candidate region is determined to be a crack. In chemical reaction kinetics, the reaction rate constant and activation energy follow the Arrhenius equation, which describes the effects of temperature and activation energy on the reaction rate. Drawing upon this theoretical framework, the formula for calculating the reaction rate constant is defined as follows: , where A is the pre-exponential factor, representing the maximum reaction rate under ideal conditions, and is usually set as a normalization constant of 1; R is the actual activation energy; R is the gas constant, which is used as a scale parameter in this application. Its value is adjusted according to the numerical range of the activation energy so that the value of the exponential term falls within a reasonable range; T is the temperature parameter, which is used as a sensitivity adjustment parameter in this application. A higher T value makes the reaction rate constant less sensitive to changes in the activation energy, while a lower T value makes the reaction rate constant highly sensitive to the activation energy.
[0170] By adjusting the R and T parameters, the numerical range of the reaction rate constant and its response to characteristic changes can be controlled. The mathematical properties of the exponential decay function ensure that the lower the activation energy, the larger the reaction rate constant, exhibiting an exponential growth relationship. For example, if the actual activation energy is reduced by 50%, the reaction rate constant may increase several times or even tens of times. This strong nonlinear characteristic allows high-fit candidate regions to obtain reaction rate constants significantly higher than those of low-fit regions, achieving effective discriminative power.
[0171] The calculated reaction rate constant is output as the activation response value, which comprehensively reflects the confidence and priority of candidate regions being classified as cracks at the high-level semantic level. The physical meaning of the activation response value is the rate at which a candidate region is converted into a crack classification result; a larger value indicates that the region should be preferentially identified as a crack. The range of the activation response value is... The lower bound corresponds to the case where the induced fit is 0, and the upper bound corresponds to the case where the induced fit is 1.
[0172] In practical applications, the activation response value is compared with a critical response threshold. If the activation response value exceeds the threshold, the candidate region is determined to be a real crack, and the current frame is a keyframe containing the crack. If the activation response value is below the threshold, it is determined to be background interference, and the current frame is a normal frame. The setting of the critical response threshold needs to strike a balance between detection sensitivity and false detection rate, and the optimal threshold point is usually determined through ROC curve analysis. The activation response value, as the final discrimination index, not only includes comprehensive information from low-level texture features (through sharpened feature maps), mid-level shape features (through interference intensity fields), and high-level semantic features (through topological skeleton analysis), but also, through the introduction of a chemical kinetic model, achieves a transformation from discrete feature matching to continuous confidence scoring, providing a more refined and reliable basis for subsequent decision-making.
[0173] This embodiment achieves accurate semantic recognition and quantitative evaluation of candidate defect substrate regions by constructing a high-level semantic classifier based on multiple constraint operators and combining it with a chemical reaction kinetics model to calculate activation response values. The high-level semantic classifier comprehensively evaluates the topological complexity, path continuity, and geometric morphological features of candidate regions through three dimensions: crack bifurcation rate constraint operator, curvature continuity constraint operator, and aspect ratio constraint operator, ensuring the comprehensiveness and accuracy of feature evaluation. By calculating the matching degree between actual feature values and standard ranges, a fuzzy membership function is used to achieve the transformation from hard threshold judgment to soft matching evaluation, improving the classifier's robustness and fault tolerance to feature changes. The induced fit degree is obtained by integrating multiple fit components using a weighted geometric average method, achieving effective fusion of multi-dimensional features and avoiding the influence of single feature dimension bias on the overall judgment. The activation energy theory of chemical reaction kinetics is introduced, mapping the induced fit degree to the activation energy reduction, and then calculating the actual activation energy and reaction rate constant, establishing a quantitative bridge from semantic features to discrimination confidence, making the classification results more refined and interpretable. By outputting the reaction rate constant as the activation response value, a continuous confidence score is provided for subsequent threshold judgment. Compared with traditional binary classification results, the activation response value can better reflect the crack probability of candidate regions, supporting more flexible decision-making strategies. The overall scheme realizes a complete transformation link from geometric topological skeleton to high-level semantic category recognition and then to activation response value through multi-level feature extraction, multi-dimensional matching evaluation, and quantitative calculation based on physical models. This improves the accuracy, robustness, and interpretability of crack detection, providing technical support for real-time and efficient pantograph defect detection.
[0174] This application embodiment also provides an image processing server, including: The memory is configured to store instructions; and The processor is configured to retrieve the instructions from the memory and, when executing the instructions, to implement the aforementioned method for rapid image content filtering based on multi-layer semantic modeling.
[0175] like Figure 4 This application also provides an image content rapid filtering system, including: High-speed camera; Speed sensor; The backend server is connected to both the high-speed camera and the speed sensor. An image processing server, connected to a backend server, is used to execute the operation steps described above.
[0176] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0177] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0178] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0179] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0180] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0181] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, like read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0182] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0183] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0184] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for rapid image content filtering based on multi-layer semantic modeling, characterized in that, This method, applied to a rapid image content filtering system, which includes a high-speed camera, a speed sensor, and a backend server, includes: The system acquires video streams from high-speed cameras and real-time velocity data from velocity sensors, maps each frame of the video stream to the frequency domain, and generates an initial spectral distribution map. Based on the real-time speed, the spectral offset factor of the background texture is calculated, and a dynamic bandpass filter with the center frequency adjusted for blue shift correction by the spectral offset factor is constructed. The initial spectral distribution map is filtered by the dynamic bandpass filter to remove low-frequency aliasing noise caused by motion blur, and then inversely transformed back to the spatial domain to obtain a sharpened feature map representing the semantic features of the underlying texture. Convolutional layers are used to transform the sharpened feature map into a high-dimensional feature space tensor, and the numerical value of the high-dimensional feature space tensor is transformed into the amplitude, and the gradient direction of the high-dimensional feature space tensor is transformed into the phase, in order to construct a real-time object wave function and load a pre-constructed standard reference wave function. Perform the conjugate multiplication operation of the real-time object wave function and the standard reference wave function in the feature space to calculate the interference intensity field and extract the region above the preset energy threshold as the candidate defect substrate region with mid-level shape semantic features; For each candidate defect substrate region, a geometric topological skeleton is extracted, high-level semantic category recognition is performed based on the geometric topological skeleton, and the activation response value is determined based on the high-level semantic category recognition result. If the activation response value of any candidate defect substrate region exceeds the critical response threshold, then the current frame is determined to be a key frame containing cracks based on the recognition results of the bottom layer texture semantic features, the middle layer shape semantic features and the high layer semantic category. If the current frame is a keyframe containing cracks, then output the current frame to the backend server.
2. The method according to claim 1, characterized in that, Based on the real-time speed, the spectral offset factor of the background texture is calculated, including: Obtain the image acquisition frame rate; Calculate the spectral compression coefficient caused by motion blur based on the real-time speed and image acquisition frame rate; Based on the spectral compression coefficient and the preset background texture reference frequency, the virtual Doppler frequency shift of the background texture in the frequency domain is calculated as the spectral offset factor.
3. The method according to claim 1, characterized in that, Constructing a dynamic bandpass filter with a center frequency that undergoes spectral blue shift correction based on a spectral shift factor includes: Determine the initial center frequency and bandwidth parameters of the reference bandpass filter; Calculate the blue shift correction amount for the center frequency based on the spectral shift factor; The initial center frequency is added to the blue shift correction to obtain the dynamically adjusted center frequency; A dynamic bandpass filter is constructed based on the dynamically adjusted center frequency and bandwidth parameters.
4. The method according to claim 1, characterized in that, The numerical magnitude of the high-dimensional feature space tensor is converted into amplitude, and the gradient direction of the high-dimensional feature space tensor is converted into phase, in order to construct the real-time object wave function, including: Calculate the numerical value of each feature point in the high-dimensional feature space tensor, and use it as the amplitude component of the wave function; Calculate the gradient vector of each feature point in the high-dimensional feature space tensor, and extract the direction angle of the gradient vector as the phase component of the wave function; By combining the amplitude and phase components, a real-time object wave function in complex form is constructed.
5. The method according to claim 1, characterized in that, Perform the conjugate multiplication of the real-time object wavefunction and the standard reference wavefunction in the characteristic space to calculate the interference intensity field, including: Perform a conjugate operation on the standard reference wavefunction to obtain the conjugate reference wavefunction; The real-time object wavefunction is added to the standard reference wavefunction by a complex number to obtain the superimposed wavefunction; The interference intensity field is obtained by calculating the square of the modulus of the superimposed wave functions; The phase matching degree distribution map is obtained by complex multiplying the real-time object wave function with the conjugate reference wave function. Based on the phase matching degree distribution map and the numerical distribution of the interference intensity field, constructive interference regions and destructive interference regions are identified in the interference intensity field. Regions with interference intensity values higher than the average value are constructive interference regions, while regions with interference intensity values lower than the average value are destructive interference regions.
6. The method according to claim 1, characterized in that, Regions with energy levels above a preset energy threshold are extracted as candidate defect substrate regions with mid-level shape semantic features, including: Iterate through all pixels in the interference intensity field and obtain the interference intensity value of each pixel; The interference intensity value is compared with a preset energy threshold, and all pixels with interference intensity values higher than the preset energy threshold are marked. Connectivity analysis is performed on the marked pixels to extract continuous high-energy regions; The extracted continuous high-energy regions are used as candidate defect substrate regions with mid-level shape semantic features.
7. The method according to claim 1, characterized in that, For each candidate defect substrate region, a geometric topological skeleton is extracted, including: Binarization is performed on the candidate defect substrate region; A morphological thinning algorithm is applied to gradually erode the boundary of the candidate defect substrate region after binarization until a skeleton structure with a single pixel width is obtained. Extract the endpoints, branching points, and connection paths of the skeleton structure; Based on endpoints, bifurcation points, and connection paths, a geometric topological skeleton of the candidate defect substrate region is constructed as a structural feature for high-level semantic classification.
8. The method according to claim 7, characterized in that, High-level semantic category recognition is performed based on a geometric topological skeleton, and activation response values are determined based on the high-level semantic category recognition results, including: A high-level semantic classifier is constructed, which includes a crack bifurcation rate constraint operator, a curvature continuity constraint operator, and an aspect ratio constraint operator, to classify candidate defect substrates into crack categories or non-crack categories. Based on the geometric topological skeleton, the ratio of the number of bifurcation points in the candidate defect substrate region to the total length is calculated as the actual bifurcation rate. Calculate the matching degree between the actual bifurcation rate and the standard bifurcation rate range defined by the crack bifurcation rate constraint operator to obtain the bifurcation rate matching component; Calculate the rate of curvature change of the candidate defect substrate region based on the geometric topological skeleton; Calculate the degree of matching between the rate of curvature change and the standard curvature continuity defined by the curvature continuity constraint operator to obtain the curvature fitting component; Calculate the aspect ratio of the candidate defect substrate region based on the geometric topological skeleton; Calculate the matching degree between the aspect ratio and the standard aspect ratio range defined by the aspect ratio constraint operator to obtain the aspect ratio matching component; High-level semantic category recognition is performed on the bifurcation rate matching component, curvature matching component and aspect ratio matching component to obtain induced matching degree, which is used to characterize the high-level semantic category recognition result. The reduction in activation energy is calculated based on the induced fit, where the higher the induced fit, the greater the reduction in activation energy. The actual activation energy is obtained by subtracting the reduction in activation energy from the preset baseline activation energy threshold. Substitute the actual activation energy into the exponential decay function to calculate the reaction rate constant; The reaction rate constant is output as the activation response value.
9. An image processing server, applied to the image content fast filtering method based on multi-layer semantic modeling as described in any one of claims 1-8, characterized in that, include: The memory is configured to store instructions; as well as A processor is configured to retrieve the instructions from the memory and, when executing the instructions, to implement the image content fast filtering method based on multi-layer semantic modeling according to any one of claims 1 to 8.
10. A rapid image content filtering system, characterized in that, include: High-speed camera; Speed sensor; The backend server is connected to both the high-speed camera and the speed sensor. The image processing server according to claim 9 is connected to a backend server and is used to perform the operation steps of the method described in any one of claims 1-8.