A thrust bearing surface damage large model detection method and system

By constructing a cross-modal target detection model and a segmentation model (SAM), and combining knowledge-guided and reinforcement learning, the problems of poor adaptability and low accuracy in thrust bearing surface damage detection are solved, achieving efficient and accurate damage identification and segmentation in complex backgrounds.

CN121616593BActive Publication Date: 2026-04-17XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
XI AN JIAOTONG UNIV
Filing Date
2026-02-02
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies have poor adaptability and weak anti-interference ability in the detection of surface damage of thrust bearings in complex backgrounds. Deep learning methods rely on a large number of labeled samples and have insufficient generalization ability, resulting in low damage detection accuracy, blurred damage area segmentation edges, and unreasonable distribution of confidence in multiple types of damage, leading to high false detection and false negative rates.

Method used

We construct a text-guided cross-modal target detection model and a segmentation model (SAM). By combining knowledge guidance and reinforcement learning, we generate multi-scale feature images and text features through feature encoding and fusion of grayscale and depth maps, perform cross-modal fusion, and optimize the damage confidence distribution through reinforcement learning to achieve fine segmentation of the damage region.

Benefits of technology

It significantly improves the accuracy of damage detection and segmentation quality in complex backgrounds, reduces the dependence on large-scale labeled data, enhances the model's generalization ability and engineering applicability, and achieves efficient and accurate identification and quantification of damaged areas.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121616593B_ABST
    Figure CN121616593B_ABST
Patent Text Reader

Abstract

The application discloses a kind of thrust bearing surface damage big model detection method and system, belong to machine wear state monitoring technical field.The method includes: using three-dimensional imaging camera obtains thrust bearing surface gray scale diagram and depth map, generates structure perception fusion feature by double branch coding and dynamic fusion;Hierarchical semantic knowledge base is constructed, text prompt vector is generated and is aligned with fusion feature cross modal;Joint Grounding DINO and SAM damage fine segmentation model, realize damage candidate prediction frame positioning and fine segmentation;Based on PPO framework, construct reinforcement learning fine-tuning mechanism, dynamically optimize multiclass damage confidence distribution.The application cooperates through multimodal fusion, knowledge guidance and reinforcement learning, improves the detection precision, generalization ability and segmentation fineness of multiple types of damage under complex background, reduces the dependence on large-scale labeled samples, and is suitable for intelligent damage detection of thrust bearing in complex service environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of machine wear condition monitoring technology, specifically relating to a method and system for detecting large-scale surface damage of thrust bearings. Background Technology

[0002] Thrust bearings, operating under complex conditions, exhibit various types of damage with significant structural differences on their surfaces. These damages are often accompanied by non-destructive interference such as machining marks, rendering existing detection methods inaccurate in identifying these diverse damage types under complex backgrounds. This hinders the need for refined assessment of thrust bearing operating conditions. Therefore, developing a detection method specifically for thrust bearing damage is of significant engineering importance for advancing intelligent detection technology for thrust bearing surface damage, aiming to improve the robustness and quantification accuracy of detection under complex operating conditions.

[0003] Damage detection intelligently identifies damage types based on surface morphology, revealing the degradation mechanism of thrust bearings. Thrust bearing surface damage types are diverse, with significant structural differences, and are often accompanied by non-damage interference, severely limiting the accuracy of damage detection. For example, existing image processing-based methods, such as low-rank decomposition and threshold segmentation, have limited adaptability to complex damage and are highly sensitive to noise and illumination changes. Deep learning methods have been widely researched and applied in recent years, but their damage detection accuracy heavily relies on a large number of labeled samples, limiting their widespread application in diverse real-world damage scenarios. Therefore, the large differences in damage structure and the abundance of interference signals remain key challenges restricting detection accuracy.

[0004] Using two-dimensional images or three-dimensional topography of the thrust bearing surface as input, damage detection algorithms are used to locate and characterize surface damage areas, facilitating the automatic extraction of damage geometry. The surface damage detection algorithm system mainly comprises two categories: one is classical algorithms, which have accumulated a solid theoretical foundation and practical experience over time; the other is algorithms based on convolutional neural networks, which, relying on the powerful capabilities of deep learning, demonstrate superior performance in complex image recognition tasks and have rapidly become the mainstream method in the field of surface damage detection.

[0005] Classical algorithms include structural methods, statistical methods, filtering methods, and model-based methods. These algorithms heavily rely on manually preset feature vector extraction; however, in complex and ever-changing background environments, their ability to capture damage features is often limited, making it difficult to guarantee recognition accuracy. Furthermore, when faced with subtle changes in image background or damage type, traditional algorithms typically require cumbersome parameter adjustments or algorithm reconstruction, demonstrating low flexibility and adaptability.

[0006] Deep learning-based damage detection methods mainly fall into two categories: Convolutional Neural Networks (CNNs) and large-scale models. CNNs can autonomously extract and represent damage features by learning from large-scale labeled images, demonstrating high accuracy and robustness in detecting complex backgrounds and various types of damage. Depending on the data annotation method, CNN applications can be categorized into four paradigms: fully supervised, semi-supervised, weakly supervised, and unsupervised learning. Among these, fully supervised learning, with its efficiency and accuracy, currently dominates research.

[0007] Large-scale model methods are characterized by their larger parameter scale, pre-trained features, and generalization ability with few or even zero samples. In recent years, large-scale models, represented by Transformers, Vision Transformers (ViT), and Generative Pre-trained Transformers (GPT), have demonstrated outstanding global modeling and pattern recognition capabilities in natural language processing and computer vision, effectively representing weak or irregular damage features in complex backgrounds. Currently, no research has directly applied large-scale models to surface damage detection in sliding thrust bearings. This research gap indicates the significant exploratory value of large-scale models in this field. Compared to CNNs, Transformers exhibit stronger robustness in complex textured backgrounds, small target detection, and imbalanced sample conditions. Their global modeling capabilities can effectively distinguish real damage from background noise or artifacts, which is particularly crucial for detecting interference from lubricating oil films, machining marks, etc., in thrust bearing images. Therefore, existing industrial defect detection models and their optimization strategies for small targets and complex backgrounds have the potential for direct transfer to the identification and segmentation of surface damage in thrust bearings.

[0008] In summary, deep learning-based algorithms demonstrate significant advantages in detecting complex backgrounds and diverse damage types due to their automatic feature extraction and learning capabilities. However, these methods still have certain limitations in practical applications: on the one hand, they are heavily reliant on large-scale labeled data and model optimization processes; on the other hand, feature extraction and recognition accuracy may be compromised in complex textured backgrounds or when dealing with small, irregular damage. In recent years, the development of large-scale model technology has provided new possibilities for solving these problems. Large-scale models possess stronger feature representation and cross-task transfer capabilities, achieving better detection performance under limited sample conditions and demonstrating advantages in handling small targets and complex backgrounds. Although there is currently a lack of mature research directly applied to thrust bearing surface damage detection, related industrial defect detection results indicate that large-scale models have significant exploratory potential in this field. In the future, combining deep learning with large-scale models, optimizing structural design, and improving model generalization and robustness are expected to further promote the development of thrust bearing surface damage detection methods. Summary of the Invention

[0009] The technical problem to be solved by this invention is to address the shortcomings of the prior art by providing a large-scale model detection method and system for thrust bearing surface damage. Based on text-guided cross-modal target detection models and the Segment Anything Model (SAM), a large-scale model of thrust bearing surface damage is constructed using knowledge-guided and reinforcement learning collaboration. This addresses the technical problems of existing image processing-based methods having poor adaptability and weak anti-interference ability in complex backgrounds; deep learning methods relying on a large number of labeled samples and having insufficient generalization ability; zero-shot segmentation models lacking domain knowledge and having low damage detection accuracy; blurred damage region segmentation edges that are difficult to meet quantization requirements; and unreasonable distribution of confidence scores for multiple damage categories, leading to high false positive and false negative rates.

[0010] The present invention adopts the following technical solution:

[0011] A method for detecting surface damage on a thrust bearing using a large model includes the following steps:

[0012] S1. Obtain the grayscale image and depth image of the thrust bearing surface, and perform feature encoding and fusion on the grayscale image and the depth image to obtain the structure-aware fusion feature;

[0013] S2. Construct a hierarchical semantic knowledge base for thrust bearing surface damage, generate text prompt vectors based on the semantic information of the hierarchical semantic knowledge base, perform cross-modal alignment between the text prompt vectors and the structure-aware fusion features obtained in step S1, and output the cross-modal aligned text features.

[0014] S3. Construct a large-scale model of thrust bearing surface damage. This model includes a multi-scale feature-enhanced damage feature image backbone network, a damage semantic text backbone network, a cross-modal fusion module, and a SAM damage fine segmentation model. Based on the structure-aware fusion features output in step S1, multi-scale image features of thrust bearing surface damage are generated through the damage feature image backbone network. Based on the text cue vector generated in step S2, text features with the same dimension as the multi-scale image features of the thrust bearing surface damage are generated through the damage semantic text backbone network. The multi-modal fusion module performs cross-modal fusion of the multi-scale image features and text features of the thrust bearing surface damage to obtain cross-modal fused image-text features. Then, a visual-text jointly represented candidate prediction box for thrust bearing surface damage is output. The candidate prediction box guides the SAM damage fine segmentation model for feature extraction, mask generation, and edge optimization, outputting the fine segmentation result of the damaged area.

[0015] S4. Based on the PPO framework, a reinforcement learning-driven fine-tuning mechanism is constructed. Using the structure-aware fusion features obtained in step S1, the cross-modal aligned text features obtained in step S2, the cross-modal fused image-text features obtained in step S3, the candidate prediction boxes for thrust bearing surface damage, and the fine segmentation results of the damage area as the basis, the image-text features obtained in step S3 are encoded as the state space. A composite reward mechanism is designed by combining the candidate prediction boxes for thrust bearing surface damage and the fine segmentation results of the damage area. The action form corresponding to the damage confidence distribution is redefined, and the large model of thrust bearing surface damage is iteratively fine-tuned to dynamically optimize the multi-class damage confidence distribution.

[0016] Preferably, in step S1, the process of feature encoding and fusion of the grayscale image and the depth image includes:

[0017] Construct a depth branch based on the depth map, encode and enhance the depth map to extract structural abrupt change features;

[0018] Construct a grayscale branch based on the grayscale image, encode the grayscale image, and extract texture and local edge features;

[0019] A feature dynamic fusion module combining the depth branch and the grayscale branch is constructed. A gating mechanism is used to dynamically fuse the features extracted by the two branches. A detection-guided weight adjustment mechanism is introduced. Based on the difference between the number of target boxes detected based on the depth map and the number of target boxes detected based on the fused image, the feature weights of the grayscale branch and the depth branch are dynamically adjusted to obtain the detection-guided optimized structure-aware fusion features.

[0020] Preferably, the fusion feature implemented by the gating mechanism is as follows:

[0021]

[0022] in, The initial fused features are the results of the dynamic fusion of the two branch features through a gating mechanism. Here, ⊙ represents the Sigmoid activation function, and ⊙ represents the Hadamard product. For trainable weights, , For branch encoding tensors, For deep branch attention weights, This is a bias term.

[0023] Preferably, structurally-aware fusion features for:

[0024]

[0025] in, Here, ⊙ represents the Sigmoid activation function, and ⊙ represents the Hadamard product. For trainable weights, , For branch encoding tensors, For grayscale branch weights, For depth branch weights, For adjustment coefficients, This is a bias term.

[0026] Preferably, in step S2, the process of generating text cue vectors and cross-modal alignment includes:

[0027] Text prompt vector ,in, For category encoding vectors, For the basic definition of vector, A morphological description vector. Embedded descriptions for synonyms or extended classes;

[0028] The text prompt vector is input into the text encoder of Grounding DINO. The semantic information is deeply modeled through the Transformer structure. The feature dimension is adjusted by linear projection so that the generated text features are consistent with the feature dimensions of the multi-scale image of thrust bearing surface damage in step S3, thus completing cross-modal alignment.

[0029] Preferably, in step S3, the process of cross-modal fusion of multi-scale image features and text features is as follows:

[0030] By using the Grounding DINO cross-modal attention module, a multi-head attention mechanism is employed to achieve dynamic semantic interaction between multi-scale image features and text features, resulting in cross-modal fused image and text features.

[0031] The cross-modal fused image and text features are input into the detection box encoder, and the output is a candidate prediction box for thrust bearing surface damage, which is jointly represented by vision and text.

[0032] Preferably, in step S3, the process of using the candidate prediction box for thrust bearing surface damage to guide the SAM damage fine segmentation model to segment the damage region is as follows:

[0033] The candidate prediction box is input into the SAM damage fine segmentation model. The feature extraction network of the SAM damage fine segmentation model is used to enhance the features of the damage area and generate an initial mask.

[0034] Edge optimization processing is performed on the initial mask to eliminate the blurring of the mask edges and output the fine segmentation result of the damaged area.

[0035] Preferably, in step S4, the definition of the state space and action form is as follows:

[0036] The state space is composed of the image-text feature encoding obtained from the cross-modal fusion in step S3, denoted as state. ;

[0037] The action mode is the surface of the thrust bearing The confidence distribution of damage class is expressed as follows: , Given a set of real numbers, the damage confidence distribution is modeled as the policy output in a continuous action space. ,in, Class 1, Class 2 to Class 3 Confidence response to damage type.

[0038] Preferably, in step S4, a composite reward mechanism is designed to redefine the action form corresponding to the damage confidence distribution, and to iteratively fine-tune the large-scale model of the thrust bearing surface damage, including:

[0039] Composite reward function for:

[0040]

[0041] in, Basic rewards; As a reward for attention; This is a sparse penalty term;

[0042] Construct a commentator network for cross-modal fused image-text features, and process the cross-modal fused image-text feature tensor. Perform global pooling to evaluate the value of the current policy state s;

[0043] Total loss function for:

[0044]

[0045] in, For strategic losses, This is due to the loss of testing and supervision from the original Grounding DINO. This is the balance coefficient;

[0046] The process involves a closed-loop iterative fine-tuning process: state-action generation, confidence assessment, reward feedback, and feature adjustment, until the large-scale model of thrust bearing surface damage meets the iteration termination condition.

[0047] Secondly, embodiments of the present invention provide a large-scale model detection system for surface damage of thrust bearings, comprising:

[0048] An extraction module is used to acquire grayscale and depth maps of the thrust bearing surface, and to perform feature encoding and fusion on the grayscale and depth maps to obtain structure-aware fusion features;

[0049] The semantic module is used to construct a hierarchical semantic knowledge base for thrust bearing surface damage. Based on the semantic information of the hierarchical semantic knowledge base, a text prompt vector is generated. The text prompt vector is then aligned with the obtained structure-aware fusion features across modalities, and the cross-modal aligned text features are output.

[0050] The segmentation module is used to construct a large-scale model of thrust bearing surface damage. This model includes a multi-scale feature-enhanced damage feature image backbone network, a damage semantic text backbone network, a cross-modal fusion module, and a SAM (Structured Aspect-Based Model) damage fine segmentation model. Based on the output structure-aware fusion features, the damage feature image backbone network generates multi-scale image features of thrust bearing surface damage. Based on the generated text cue vectors, the damage semantic text backbone network generates text features with dimensions consistent with the multi-scale image features of the thrust bearing surface damage. The cross-modal fusion module performs cross-modal fusion of the multi-scale image features and text features of the thrust bearing surface damage, obtaining cross-modal fused image-text features, and outputs a visual-text joint representation of the thrust bearing surface damage candidate prediction box. The candidate prediction box guides the SAM damage fine segmentation model to perform feature extraction, mask generation, and edge optimization, outputting the fine segmentation result of the damage region.

[0051] The output module is used to construct a reinforcement learning-driven fine-tuning mechanism based on the PPO framework. Based on the obtained structure-aware fusion features, cross-modal aligned text features, cross-modal fused image-text features, thrust bearing surface damage candidate prediction boxes, and fine segmentation results of damage regions, the cross-modal fused image-text features are encoded as the state space. A composite reward mechanism is designed by combining the thrust bearing surface damage candidate prediction boxes and fine segmentation results of damage regions. The action form corresponding to the damage confidence distribution is redefined, and the large-scale model of thrust bearing surface damage is iteratively fine-tuned to dynamically optimize the multi-class damage confidence distribution.

[0052] Thirdly, a computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method for detecting large-scale surface damage of thrust bearings.

[0053] Fourthly, embodiments of the present invention provide a computer-readable storage medium including a computer program, which, when executed by a processor, implements the steps of the above-described method for detecting large-scale surface damage of thrust bearings.

[0054] Fifthly, a chip includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method for detecting large-scale surface damage of thrust bearings.

[0055] In a sixth aspect, embodiments of the present invention provide an electronic device, including a computer program, which, when executed by the electronic device, implements the steps of the above-described method for detecting large-scale surface damage of thrust bearings.

[0056] Compared with the prior art, the present invention has at least the following beneficial effects:

[0057] A large-scale model detection method for thrust bearing surface damage is proposed. This method overcomes the limitation of single-modal approaches in fully characterizing damage features by fusing grayscale and depth maps as dual-modal inputs. It simultaneously captures texture edges and structural abrupt changes, providing comprehensive data support for subsequent detection. Secondly, it innovatively introduces a hierarchical semantic knowledge base and a reinforcement learning collaborative mechanism. This reduces the model's dependence on large-scale labeled samples through knowledge guidance, improving zero-shot generalization ability, and dynamically optimizes confidence distribution through the PPO framework to adapt to multiple damage scenarios. Thirdly, it combines Grounding DINO and SAM damage fine segmentation models to achieve a closed-loop processing of candidate prediction box localization and fine segmentation, balancing detection efficiency and segmentation accuracy. Finally, the steps are logically coherent, and the inputs and outputs are closely linked, forming a complete technical chain from feature extraction to model optimization. This ensures the feasibility and stability of the method, providing a systematic solution for thrust bearing damage detection under complex operating conditions.

[0058] Furthermore, a dual-branch coding architecture is adopted. The depth branch focuses on extracting structural abrupt changes, while the grayscale branch emphasizes capturing texture and local edge features. The two branches complement each other, effectively compensating for the shortcomings of single-modality feature representation in complex damage. A dynamic feature fusion module is designed, using a gating mechanism to achieve adaptive fusion of features from the two branches, rather than simple splicing, ensuring the relevance and effectiveness of the fused features. A detection-guided weight adjustment mechanism is introduced, dynamically adjusting the weights of the two branches based on the difference in the number of bounding boxes between the depth map and the fused image. This adaptively suppresses noise interference in grayscale features or compensates for missing texture information, solving the problem of poor adaptability of fixed-weight fusion in different damage scenarios. This significantly improves the quality of structure-aware fusion features, laying a highly robust feature foundation for cross-modal alignment and damage detection.

[0059] Furthermore, a sigmoid activation function σ is introduced to dynamically control the fusion weights of the two branches, thereby enhancing effective features and suppressing noisy features, improving the signal-to-noise ratio of the fused features. The Hadamard product ⊙ is used to perform element-wise matching and multiplication of deep branch features with attention weights and trainable weights, ensuring that the structural abrupt changes in the deep branch are accurately enhanced and avoiding feature information dilution. Simultaneously, the original features of the deep and gray-scale branches are integrated, preserving the integrity of the original features on the basis of weighted fusion, balancing feature specificity and comprehensiveness. Efficient integration of complementary information from both modalities generates initial fused features that highlight core damage features while retaining detailed information, providing a high-quality foundation for subsequent detection-guided weight optimization.

[0060] Furthermore, the weights are adjusted based on the difference in the number of bounding boxes between the depth map and the fused image, enabling the fusion process to have feedback adjustment capabilities and dynamically adapt to different damage scenarios according to the actual detection results. When there are too many bounding boxes in the fused image, the gray-level branch weights are suppressed to reduce noise, and when there are too few bounding boxes, the gray-level branch weights are increased to supplement information. Continuing the core design of the gating mechanism and Hadamard product, the scientific nature of feature fusion is maintained on the basis of dynamic weight updates, ensuring that the optimized features are both mathematically logical and practically applicable in engineering. This directly determines the effect of subsequent cross-modal alignment and damage detection. Its optimized design effectively improves the feature's ability to represent complex damage, especially in damage scenarios with blurred boundaries and significant structural abrupt changes, exhibiting stronger robustness.

[0061] Furthermore, the text prompt vectors encompass multi-dimensional semantic information including category, definition, morphology, and synonyms. Built upon a hierarchical semantic knowledge base, this ensures the comprehensiveness and professionalism of the semantic information, providing rich semantic support for cross-modal alignment. Leveraging Grounding DINO's text encoder and Transformer structure, deep modeling of semantic information is achieved. Linear projection adjusts feature dimensions, ensuring consistency between text and image features and eliminating dimensional barriers in cross-modal interaction. The cross-modal alignment process enables the model to accurately correlate text semantics with visual features, improving the interpretability of damage detection while reducing the risk of false detections due to non-damage interference. Especially in low-sample scenarios, semantic guidance allows for rapid adaptation to new types of damage, significantly enhancing the model's generalization ability.

[0062] Furthermore, the Grounding DINO cross-modal attention module is employed to achieve dynamic semantic interaction between multi-scale image features and text features through a multi-head attention mechanism, rather than static fusion. This makes feature interaction more targeted and can accurately capture the correspondence between visual details and semantic descriptions. The fused image and text features combine damage detail information from the visual end with semantic category information from the text end, significantly improving the localization accuracy and category determination reliability of candidate prediction boxes. This effectively solves the problem of difficulty in distinguishing between damage and non-damage interference in complex backgrounds. The fused features are input into the detection box encoder, which directly outputs candidate prediction boxes with joint visual-text representations, simplifying the detection process and balancing detection efficiency and accuracy. This provides a precise localization foundation for the subsequent fine segmentation of the SAM damage fine segmentation model, ensuring efficient connection between the detection and segmentation stages.

[0063] Furthermore, guided by candidate bounding boxes, the SAM damage fine segmentation model can quickly focus on the damage region, avoiding indiscriminate feature extraction from the entire image and improving segmentation efficiency and specificity. This is particularly advantageous in scenarios with multiple damages and small targets. The feature extraction network of the SAM damage fine segmentation model enhances the features of the damage region, generating an initial mask with good regional integrity. Edge optimization further eliminates mask edge blurring, improving the contour accuracy of the damage region. This achieves seamless integration of localization and segmentation. Candidate bounding boxes provide precise range guidance for segmentation, and the segmentation results, in turn, verify the accuracy of localization, forming a closed-loop optimization. This ensures that the segmentation results both conform to the actual damage morphology and remain consistent with the detection and localization, providing reliable regional data for damage quantification analysis.

[0064] Furthermore, the image and text feature encoding after cross-modal fusion is defined as the state space s, which combines visual and semantic information and can comprehensively represent the state of the current detection scene, providing an accurate basis for policy evaluation. The action form is defined as the confidence distribution of N types of damage, which is directly related to the core output of the model, so that the optimization goal of reinforcement learning is highly consistent with the detection task, avoiding the optimization direction from deviating from the actual needs. The policy output expression establishes the mapping relationship between state and action, ensuring that the model can dynamically adjust the confidence distribution according to the current scene and adapt to the representation needs of different damage types. Especially in the scenario where multiple types of damage coexist and the confidence distribution is unreasonable, it effectively improves the accuracy of category determination.

[0065] Furthermore, the composite reward function R integrates basic rewards, attention rewards, and sparse penalty terms to evaluate the model's detection performance from multiple dimensions. Basic rewards ensure detection accuracy, attention rewards strengthen the weight of target regions, and sparse penalty terms suppress continuous missed detections, making reward feedback more comprehensive and accurate. The total loss function L balances the PPO strategy loss and detection supervision loss, optimizing the confidence distribution while avoiding disrupting the model's original detection architecture and ensuring the stability of the fine-tuning process. Through a closed-loop iteration of state-action generation, confidence evaluation, reward feedback, and feature adjustment, the model can adaptively and dynamically optimize the confidence distribution of multi-class damages. In complex scenarios with insufficient samples and varied damage morphologies, it continuously improves detection robustness and accuracy, solving the problem of fixed performance after traditional model training and difficulty in adapting to dynamic scenarios.

[0066] It is understood that the beneficial effects of the second to sixth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here.

[0067] In summary, this invention, through technological innovations in four aspects—adaptive fusion of multimodal features, cross-modal alignment guided by domain knowledge, joint reasoning for detection and segmentation, and self-optimization through reinforcement learning—forms a complete, robust, and scalable intelligent detection system for thrust bearing damage. It significantly improves the recognition accuracy and segmentation quality for various types of small-target, weakly visible damage in complex industrial scenarios, while reducing reliance on large-scale labeled data and enhancing the model's generalization ability and engineering practicality.

[0068] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0069] Figure 1 This is the overall flowchart of the present invention;

[0070] Figure 2 A framework for damage detection in complex structures;

[0071] Figure 3 Structure diagram of the complex damage feature extraction module;

[0072] Figure 4 The images show the fusion results of cavitation damage, where (a) is the depth map, (b) is the grayscale image, and (c) is the fused image.

[0073] Figure 5 Fine-tune the structure diagram for the strategy;

[0074] Figure 6 A schematic diagram of a computer device provided in an embodiment of the present invention;

[0075] Figure 7 This is a block diagram of a chip according to an embodiment of the present invention;

[0076] Figure 8 Typical damage morphology diagrams are shown, where (a) is abrasive wear morphology, (b) is cavitation morphology, (c) is scratch damage morphology, (d) is pitting morphology, (e) is corrosion wear morphology, (f) is fatigue spalling morphology, (g) is crack damage morphology, (h) is adhesive wear morphology, and (i) is plastic deformation morphology.

[0077] Figure 9 The images show typical damage detection results, where (a) is a comparison of the detection results of abrasive wear, (b) is a comparison of the detection results of cavitation damage, (c) is a comparison of the detection results of pitting damage, and (d) is a comparison of the detection results of fatigue spalling. #1 is the original input image, #2 is the detection / segmentation result of CLIPSeg, #3 is the detection / segmentation result of Mask2Former, and #4 is the final detection and segmentation result of the method of the present invention.

[0078] Figure 10 The following are example images of the damage detection effects of each module: (a) before optimization, (b) with only knowledge guidance added, (c) with only feature fusion added, (d) feature fusion + knowledge guidance, (e) feature fusion + reinforcement learning, and (f) all modules.

[0079] Among them, 60. Computer equipment; 61. Processor; 62. Memory; 63. Computer program; 600. Electronic device; 610. Processing unit; 620. Storage unit; 6201. Random access memory unit; 6202. Cache memory unit; 6203. Read-only memory unit; 6204. Program / utility; 6205. Program module; 630. Bus; 640. Display unit; 650. Input / output interface; 660. Network adapter; 700. External device. Detailed Implementation

[0080] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0081] In the description of this invention, it should be understood that the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0082] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0083] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Additionally, the character " / " in this invention generally indicates that the preceding and following objects have an "or" relationship.

[0084] It should be understood that although terms such as first, second, third, etc., may be used in the embodiments of the present invention to describe the preset range, these preset ranges should not be limited to these terms. These terms are only used to distinguish the preset ranges from one another. For example, without departing from the scope of the embodiments of the present invention, the first preset range may also be referred to as the second preset range, and similarly, the second preset range may also be referred to as the first preset range.

[0085] Depending on the context, the word "if" as used here can be interpreted as "when," "when," "in response to determination," or "in response to detection." Similarly, depending on the context, the phrase "if determination" or "if detection (of the stated condition or event)" can be interpreted as "when determination," "in response to determination," "when detection (of the stated condition or event)," or "in response to detection (of the stated condition or event)."

[0086] The accompanying drawings illustrate various structural schematic diagrams according to embodiments disclosed in this invention. These drawings are not to scale, and some details have been enlarged for clarity, and some details may have been omitted. The shapes of the various regions and layers shown in the drawings, as well as their relative sizes and positional relationships, are merely exemplary and may deviate from reality due to manufacturing tolerances or technical limitations. Furthermore, those skilled in the art can design regions / layers with different shapes, sizes, and relative positions as needed.

[0087] This invention provides a large-scale model detection method for thrust bearing surface damage. Using grayscale and depth maps of the thrust bearing surface morphology acquired by a structured light camera as the target, a structure-aware extraction and fusion module for thrust bearing surface damage features is constructed through collaborative dual-branch coding, gated dynamic fusion, and detection-guided weight adjustment. This module accurately captures the surface damage features of the thrust bearing. A hierarchical semantic knowledge base for thrust bearing surface damage is built. Based on its semantic information, knowledge guidance is applied to generate a text encoder that embeds text prompt vectors into a text-guided cross-modal target detection model (Grounding DINO), achieving cross-modal alignment of text and image of thrust bearing surface damage. A large-scale model detection method for thrust bearing surface damage is constructed by combining Grounding DINO and SAM. A cross-modal collaborative mechanism is used to output candidate prediction boxes for thrust bearing surface damage, and a SAM damage fine segmentation model guided by these candidate prediction boxes is constructed to achieve fine segmentation of the damage region. Finally, a proximal policy optimization method is implemented. This invention constructs a reinforcement learning-driven fine-tuning mechanism for thrust bearing surface damage detection using the Optimization (PPO) framework. It redefines the state space, action forms, and composite reward mechanism, guiding a large-scale thrust bearing surface damage model to dynamically optimize the confidence distribution under multi-class damage scenarios, thereby improving the accuracy and robustness of multi-class damage detection. This invention enhances the representation of damage boundaries and contours in thrust bearing surface damage detection, alleviating depth ambiguity; and improves the generalization ability of the large-scale thrust bearing surface damage model in damage identification, significantly improving detection accuracy and adaptability.

[0088] Please see Figure 1 The present invention provides a large-scale model detection method for surface damage of thrust bearings, comprising the following steps:

[0089] S1. Using the grayscale image and depth image of the thrust bearing surface morphology acquired by the structured light camera as the object, a structure perception extraction and fusion module for the surface damage features of the thrust bearing is constructed by coordinating dual-branch coding, gated dynamic fusion and detection-guided weight adjustment, so as to accurately capture the surface damage features of the thrust bearing.

[0090] S101. Construct a depth branch based on the depth map of the thrust bearing surface, and encode the depth map into a tensor. , The number of feature channels, For feature map height, The feature map width is used, and gradient enhancement and attention mechanisms are applied to generate the feature map for accurate extraction of abrupt structural features on the surface of the thrust bearing.

[0091] S102. Construct a grayscale branch based on the grayscale image of the thrust bearing surface, and encode the grayscale image into a tensor. It is used to efficiently capture the surface texture and local edge features of thrust bearings;

[0092] S103. Construct a dynamic fusion module for thrust bearing surface features that combines deep branch and grayscale branch features. Use a gating mechanism to dynamically fuse the features extracted from the two branches, as shown in Equation (1).

[0093] (1)

[0094] in, is the Sigmoid activation function, and ⊙ is the Hadamard product. For trainable weights, , For branch encoding tensors.

[0095] A detection-guided weight adjustment mechanism is introduced to dynamically adjust the feature weights of the grayscale branch and the depth branch based on the difference in the number of target boxes of the two types. Let the number of target boxes detected by the depth map be... The number of bounding boxes detected by the fused image is .like Significantly greater than This indicates that the grayscale feature has introduced too much noise, and in this case, the grayscale branch weight should be reduced. Conversely, if Less than If the proportion of grayscale features is appropriately increased to compensate for the lack of edge and texture information, the weight dynamic update rules are shown in Equation (2) and Formula (3), and the final fusion features after detection-guided optimization are shown in Equation (4).

[0096] (2)

[0097] (3)

[0098] in, This represents the number of bounding boxes detected by the depth map. To fuse the number of target boxes detected in the image, For grayscale branch weights, For depth branch weights, This is the adjustment coefficient.

[0099] (4)

[0100] in, Here, ⊙ represents the Sigmoid activation function, and ⊙ represents the Hadamard product. For branch weights, , For branch encoding tensors.

[0101] Please see Figure 3The structure-aware extraction and fusion module for thrust bearing surface damage features adopts a "parallel branching, co-coding" architecture, using both depth and grayscale images of the thrust bearing surface as dual inputs. The depth-map-based extraction branch focuses on capturing 3D structural abrupt changes in the damaged area, while the grayscale-map-based extraction branch concentrates on extracting the texture features of the damage. Both branches are based on convolutional neural networks (CNNs). Figure 2 Feature encoding is implemented in the multimodal fusion module, which gradually transforms the original image into a high-dimensional feature tensor. Finally, multimodal features containing three-dimensional structural information and two-dimensional texture information are output, providing complementary information support for subsequent fusion stages and solving the problem that a single modality is difficult to fully characterize complex damage morphology.

[0102] Please see Figure 4 The images show the depth map, grayscale map, and fusion map of typical cavitation damage, clearly demonstrating the enhancement effect of this module on complex boundaries and blurred regions. Through the above multi-level processing, the fused features exhibit stronger robustness in dealing with the blurred boundaries and structural abrupt changes of complex damage, significantly improving the ability to perceive complex structural damage and providing robust input representation for subsequent semantic guidance and segmentation accuracy.

[0103] S2. Construct a hierarchical semantic knowledge base for thrust bearing surface damage, guide knowledge based on its semantic information, generate text prompt vectors and embed them into a text encoder of Grounding DINO to achieve cross-modal alignment of thrust bearing surface damage text and image.

[0104] S201. Based on the failure mechanism of thrust bearings, engineering experience and literature, construct a hierarchical knowledge base for surface damage of thrust bearings.

[0105] The thrust bearing surface damage linguistic knowledge base is built based on thrust bearing failure mechanisms, engineering experience, and literature. Its structure is shown in the table below, employing a hierarchical design and containing the following five core attributes:

[0106] (1) Damage types, including six common types of thrust bearing surface damage such as abrasive wear, cavitation, pitting, corrosive wear, fatigue spalling, and plastic deformation;

[0107] (2) Basic definition: the physical causes, formation process and common inducing conditions of each type of injury;

[0108] (3) Morphological characteristics: such as "cavitation presents as pinhole-like, honeycomb-like or fish-scale-like pits with sharp and raised edges", etc.

[0109] (4) Key prompts: Match each type of injury with a descriptive statement that can be recognized by the visual language model;

[0110] (5) Synonyms: To support different industrial terms, add semantic mappings such as “micro-pit” and “dentation”.

[0111]

[0112] S202. Organize the semantic information of the semantic knowledge base of thrust bearing surface damage in a structured manner to ensure compatibility with prompt generation, label mapping and model training processes;

[0113] S203. Generate a text prompt vector for thrust bearing surface damage based on a semantic knowledge base of thrust bearing surface damage. This transforms semantic information into a text signal recognizable by a text-guided cross-modal target detection model. The text prompt vector... As shown in equation (5);

[0114] (5)

[0115] in, For category encoding vectors, For the basic definition of vector, A morphological description vector. Embedded descriptions for synonyms or extended classes.

[0116] S204. Construct a text encoding strategy for thrust bearing surface damage. Input the thrust bearing surface damage prompt vector into the text encoder of Grounding DINO for encoding. This achieves cross-modal alignment of thrust bearing surface damage text and image, improving the accuracy, interpretability, and zero-sample capability of the large-model detection method for thrust bearing surface damage.

[0117] S3. A large-scale model detection method for thrust bearing surface damage is constructed by combining Grounding DINO and SAM. A cross-modal collaborative mechanism is used to output candidate prediction boxes for thrust bearing surface damage. A SAM damage fine segmentation model guided by the prediction boxes is constructed to achieve fine segmentation of the damage area.

[0118] S301. Using the structure-aware fusion features output in step S1 as the object, construct a damage feature image backbone network based on multi-scale feature enhancement to generate multi-scale image features of thrust bearing surface damage.

[0119] The input structure-aware fusion features have been processed through dual-branch encoding, gated modal fusion, and detection-guided weight adjustment in step S1, while retaining grayscale texture and deep structural information. After being input into the image backbone network, they are transformed into multimodal image features containing multidimensional visual information through multi-scale feature extraction, providing a stable visual foundation for subsequent cross-modal alignment of text and images.

[0120] S302. Using the multi-dimensional text prompt vector generated by the semantic knowledge base of thrust bearing surface damage in step S2 as the object, construct a damage semantic text backbone network to generate text features consistent with the image feature dimensions.

[0121] By using the Transformer structure to perform deep modeling of semantic information, and then adjusting the feature dimensions through linear projection, we can ensure that the dimensions of the generated text features are consistent with the dimensions of the output image features, thereby achieving dimensional matching of visual and semantic features and eliminating dimensional barriers for subsequent cross-modal interactions.

[0122] S303. Construct a cross-modal attention module based on Grounding DINO to fuse the image and text features of the thrust bearing surface damage across modalities, output the fused image and text features, and then output the candidate prediction boxes of the thrust bearing surface damage jointly represented by the detection box encoder.

[0123] To address the challenge of distinguishing between complex damage and non-damage interference in a single modality, the Grounding DINO cross-modal attention module employs a multi-head attention mechanism to enable dynamic semantic interaction between image and text features. The resulting cross-modal fused image and text features combine visual details with semantic information, significantly reducing the risk of false detections caused by non-damage interference.

[0124] S304. Establish a SAM damage fine segmentation model guided by the prediction box. After feature extraction, mask generation and edge optimization of SAM, output the corresponding thrust bearing surface damage fine segmentation results.

[0125] S4. Based on the PPO framework, a reinforcement learning-driven fine-tuning mechanism for thrust bearing surface damage detection is constructed. The state space, action form, and composite reward mechanism are redefined to guide the model to dynamically optimize the confidence distribution in multi-class damage scenarios, thereby improving the accuracy and robustness of multi-class damage detection.

[0126] S401. Define the image and text feature encoding after cross-modal fusion as the current state s. For N types of damage on the surface of the thrust bearing, model the damage confidence distribution (as in Equation (6)) as the strategy output in the continuous action space (as in Equation (7)).

[0127] (6)

[0128] (7)

[0129] in, The current image state is represented by fused image feature encoding; For the first Confidence response to damage type for 3D real space.

[0130] S402. Construct a commentator network for the cross-modal fused image-text features, and process the cross-modal fused image-text feature tensor. Perform global pooling. To fuse feature map height, To fuse feature map widths, As a feature dimension, evaluate the value of the current policy state s. , The value assessment mapping function for the critic network is shown in Equation (8);

[0131] (8)

[0132] S403. Design a composite reward function based on the accuracy and semantic rationality of thrust bearing surface damage detection, as shown in Equation (9), to provide feedback for the optimization of damage confidence distribution strategy and guide the model to adapt to the confidence expression requirements of different damage forms.

[0133] (9)

[0134]

[0135]

[0136]

[0137] in, The basic reward is allocated based on whether the prediction is correct. For attention rewards, weighting is applied based on the salience of the target region; As a sparse penalty item, a penalty is imposed for 10 consecutive missed detections.

[0138] S404. Construct a reinforcement learning-guided cross-modal fusion post-image feature adjustment module, and use the policy loss as the total loss function (as shown in Equation (10)) to achieve parameter backpropagation, and adjust multi-scale image features locally without destroying the original detection architecture.

[0139] (10)

[0140] in, For strategic losses, This is due to the loss of testing and supervision from the original Grounding DINO. This is the balance coefficient.

[0141] S405. Design an iterative execution fine-tuning strategy based on PPO, following a closed-loop iteration of "state-action generation—confidence assessment—reward feedback—feature adjustment" to dynamically adjust the confidence distribution of multiple types of damage, stabilize the confidence output, and complete the fine-tuning until the model meets the iteration termination condition.

[0142] Please see Figure 5 Through strategy fine-tuning, the model can adaptively adjust the confidence distribution of multiple types of damage, and can effectively enhance the robustness and adaptability of the detector, especially under conditions of insufficient samples.

[0143] In another embodiment of the present invention, a large-scale model detection system for thrust bearing surface damage is provided. This system can be used to implement the above-mentioned large-scale model detection method for thrust bearing surface damage. Specifically, the large-scale model detection system for thrust bearing surface damage includes an extraction module, a semantic module, a segmentation module, and an output module.

[0144] The extraction module is used to acquire grayscale and depth maps of the thrust bearing surface, and to perform feature encoding and fusion on the grayscale and depth maps to obtain structure-aware fusion features.

[0145] The semantic module is used to construct a hierarchical semantic knowledge base for thrust bearing surface damage. Based on the semantic information of the hierarchical semantic knowledge base, a text prompt vector is generated. The text prompt vector is then aligned with the obtained structure-aware fusion features across modalities, and the cross-modal aligned text features are output.

[0146] The segmentation module is used to construct a large-scale model of thrust bearing surface damage. This large model includes a damage feature image backbone network with enhanced multi-scale features of thrust bearing surface damage, a damage semantic text backbone network, a cross-modal fusion module, and a SAM (Structured Aspect-Based Model) damage fine segmentation model. Based on the output structure-aware fusion features, the damage feature image backbone network generates multi-scale image features of thrust bearing surface damage. Based on the generated text cue vectors, the damage semantic text backbone network generates text features with dimensions consistent with the multi-scale image features of the thrust bearing surface damage. The cross-modal fusion module performs cross-modal fusion of the multi-scale image features and text features of the thrust bearing surface damage, obtaining cross-modal fused image-text features, and then outputs a visual-text joint representation of the thrust bearing surface damage candidate prediction box. The thrust bearing surface damage candidate prediction box guides the SAM damage fine segmentation model to perform feature extraction, mask generation, and edge optimization, outputting the fine segmentation result of the damaged region.

[0147] The output module is used to construct a reinforcement learning-driven fine-tuning mechanism based on the PPO framework. Based on the obtained structure-aware fusion features, cross-modal aligned text features, cross-modal fused image-text features, thrust bearing surface damage candidate prediction boxes, and fine segmentation results of damage regions, the cross-modal fused image-text features are encoded as the state space. A composite reward mechanism is designed by combining the thrust bearing surface damage candidate prediction boxes and fine segmentation results of damage regions. The action form corresponding to the damage confidence distribution is redefined, and the large-scale model of thrust bearing surface damage is iteratively fine-tuned to dynamically optimize the multi-class damage confidence distribution.

[0148] This invention provides a terminal device comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to achieve a corresponding method flow or corresponding function. The processor described in this embodiment can be used in the operation of a large-scale model detection method for thrust bearing surface damage, including:

[0149] A grayscale image and depth map of the thrust bearing surface are acquired. Feature encoding and fusion are performed on the grayscale image and depth map to obtain structure-aware fusion features. A hierarchical semantic knowledge base for thrust bearing surface damage is constructed. Text prompt vectors are generated based on the semantic information of the hierarchical semantic knowledge base. The text prompt vectors are then cross-modal aligned with the obtained structure-aware fusion features to output the cross-modal aligned text features. A large-scale model of thrust bearing surface damage is constructed, comprising a multi-scale feature-enhanced damage feature image backbone network, a damage semantic text backbone network, a cross-modal fusion module, and a SAM damage fine segmentation model. Based on the output structure-aware fusion features, multi-scale image features of thrust bearing surface damage are generated through the damage feature image backbone network. Based on the generated text prompt vectors, text features with the same dimension as the multi-scale image features of the thrust bearing surface damage are generated through the damage semantic text backbone network. The cross-modal fusion module is used to process the... The multi-scale image features and text features of thrust bearing surface damage are fused across modally to obtain cross-modal fused image-text features. Then, a visual-text joint representation of the thrust bearing surface damage candidate prediction box is output. The thrust bearing surface damage candidate prediction box guides the SAM damage fine segmentation model to perform feature extraction, mask generation, and edge optimization, outputting the damage region fine segmentation result. Based on the PPO framework, a reinforcement learning-driven fine-tuning mechanism is constructed. Based on the obtained structure-aware fusion features, cross-modal aligned text features, cross-modal fused image-text features, thrust bearing surface damage candidate prediction boxes, and damage region fine segmentation results, the cross-modal fused image-text features are encoded as the state space. A composite reward mechanism is designed by combining the thrust bearing surface damage candidate prediction boxes and damage region fine segmentation results to redefine the action form corresponding to the damage confidence distribution. The large-scale model of thrust bearing surface damage is iteratively fine-tuned to dynamically optimize the multi-class damage confidence distribution.

[0150] Please see Figure 6 The terminal device is a computer device. In this embodiment, the computer device 60 includes a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When executed by the processor 61, the computer program 63 implements the large-scale model detection method for thrust bearing surface damage in this embodiment. To avoid repetition, these details are not elaborated here. Alternatively, when executed by the processor 61, the computer program 63 implements the functions of each model / unit in the large-scale model detection system for thrust bearing surface damage in this embodiment. To avoid repetition, these details are not elaborated here.

[0151] Computer device 60 can be a desktop computer, laptop, handheld computer, cloud server, or other computing device. Computer device 60 may include, but is not limited to, a processor 61 and a memory 62. Those skilled in the art will understand that... Figure 6 This is merely an example of computer device 60 and does not constitute a limitation on computer device 60. It may include more or fewer components than shown, or combine certain components, or different components. For example, computer device may also include input / output devices, network access devices, buses, etc.

[0152] The processor 61 may be a Central Processing Unit (CPU), or other general-purpose processors, graphics processing units (GPUs), tensor processing units (TPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0153] The memory 62 can be an internal storage unit of the computer device 60, such as a hard disk or memory of the computer device 60. The memory 62 can also be an external storage device of the computer device 60, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on the computer device 60.

[0154] Furthermore, memory 62 may include both internal storage units and external storage devices of the computer device 60. Memory 62 is used to store computer programs and other programs and data required by the computer device. Memory 62 can also be used to temporarily store data that has been output or will be output.

[0155] Please see Figure 7 The terminal device is an electronic device 600, which is manifested in the form of a general-purpose computing device. The components of the electronic device may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different platform components (including storage unit 620 and processing unit 610), a display unit 640, etc.

[0156] The storage unit stores program code, which can be executed by the processing unit 610 to perform the steps described in the method section of this specification according to various exemplary embodiments of the present invention. For example, the processing unit 610 can perform actions such as... Figure 1 The steps are shown in the figure.

[0157] Storage unit 620 may include a readable medium in the form of a volatile storage unit, such as random access memory (RAM) 6201 and / or cache memory 6202, and may further include a read-only memory (ROM) 6203.

[0158] Storage unit 620 may also include a program / utility 6204 having a set (at least one) program module 6205, such program module 6205 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0159] Bus 630 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the multiple bus structures.

[0160] Electronic device 600 can also communicate with one or more external devices 700 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 600, and / or with any device that enables electronic device 600 to communicate with one or more other computing devices (e.g., router, modem). This communication can be performed via input / output interface 650. Furthermore, electronic device 600 can also communicate with one or more networks (e.g., local area network, wide area network, and / or public network, such as the Internet) via network adapter 660. Network adapter 660 can communicate with other modules of electronic device 600 via bus 630. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage platforms.

[0161] Example 4

[0162] This invention also provides a storage medium, specifically a computer-readable storage medium, which is a memory device in a terminal device for storing programs and data. It is understood that the computer-readable storage medium here can include both built-in storage media in the terminal device and extended storage media supported by the terminal device; it can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor, which can be one or more computer programs (including program code). More specific examples of the computer-readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, random access memory, read-only memory, erasable programmable read-only memory, optical fiber, portable compact disk read-only memory, optical storage device, magnetic storage device, or any suitable combination thereof.

[0163] Computer-readable storage media also include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable storage medium can also be any readable medium other than a readable storage medium that can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, radio frequency, etc., or any suitable combination thereof.

[0164] Program code for performing the operations of this invention can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0165] One or more instructions stored in a computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the large-scale model detection method for thrust bearing surface damage in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by the processor in the following steps:

[0166] A grayscale image and depth map of the thrust bearing surface are acquired. Feature encoding and fusion are performed on the grayscale image and depth map to obtain structure-aware fusion features. A hierarchical semantic knowledge base for thrust bearing surface damage is constructed. Text prompt vectors are generated based on the semantic information of the hierarchical semantic knowledge base. The text prompt vectors are then cross-modal aligned with the obtained structure-aware fusion features to output the cross-modal aligned text features. A large-scale model of thrust bearing surface damage is constructed, comprising a multi-scale feature-enhanced damage feature image backbone network, a damage semantic text backbone network, a cross-modal fusion module, and a SAM damage fine segmentation model. Based on the output structure-aware fusion features, multi-scale image features of thrust bearing surface damage are generated through the damage feature image backbone network. Based on the generated text prompt vectors, text features with the same dimension as the multi-scale image features of the thrust bearing surface damage are generated through the damage semantic text backbone network. The cross-modal fusion module is used to process the... The multi-scale image features and text features of thrust bearing surface damage are fused across modally to obtain cross-modal fused image-text features. Then, a visual-text joint representation of the thrust bearing surface damage candidate prediction box is output. The thrust bearing surface damage candidate prediction box guides the SAM damage fine segmentation model to perform feature extraction, mask generation, and edge optimization, outputting the damage region fine segmentation result. Based on the PPO framework, a reinforcement learning-driven fine-tuning mechanism is constructed. Based on the obtained structure-aware fusion features, cross-modal aligned text features, cross-modal fused image-text features, thrust bearing surface damage candidate prediction boxes, and damage region fine segmentation results, the cross-modal fused image-text features are encoded as the state space. A composite reward mechanism is designed by combining the thrust bearing surface damage candidate prediction boxes and damage region fine segmentation results to redefine the action form corresponding to the damage confidence distribution. The large-scale model of thrust bearing surface damage is iteratively fine-tuned to dynamically optimize the multi-class damage confidence distribution.

[0167] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0168] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0169] To verify the effectiveness of the proposed method in detecting complex surface damage in sliding thrust bearings, this paper randomly allocated 20% of the image dataset (93 images in total) as a test set to evaluate the model's recognition accuracy. The tested damage types covered typical fatigue, mechanical wear, cavitation, etc., such as... Figure 8 As shown.

[0170] In the comparative experiments, the performance of the method of this invention was compared with that of two representative methods: the Masked Two-Stage Transformer for Universal Image Segmentation (Mask2Former) and the CLIP-based Segmentation Model (CLIPSeg).

[0171] Please see Figure 9 The method of the present invention ( Figure 9 (d) effectively detected various complex damages. It should be noted that Mask2Former ( Figure 9 (b) and CLIPSeg ( Figure 9 The zero-sample segmentation large model (a) is trained based on natural scene data and lacks domain knowledge. It is not adaptable to the thrust bearing surface scene with variable damage scale, blurred boundaries and strong regional noise. The method of the present invention is significantly better than the comparative method in all indicators, which fully verifies the effectiveness of the method of the present invention in industrial surface damage detection.

[0172] Please see Figure 10 To further analyze the contribution of each module in the method of this invention to the detection performance, an ablation experiment was designed, and four types of combination models were set up.

[0173] Combination 1: Feature fusion module, which performs detection only by fusing depth and grayscale features.

[0174] Combination 2: Feature fusion + knowledge guidance module, which introduces a damage prior knowledge enhancement strategy based on Combination 1.

[0175] Combination 3: Feature fusion + reinforcement learning module, which introduces a confidence dynamic optimization strategy based on Combination 1.

[0176] Combination 4: A complete model that integrates feature fusion, knowledge guidance, and reinforcement learning modules.

[0177] Overall, the complete model achieved optimal performance across all four metrics, significantly outperforming other combinations, indicating that the synergistic effect of each module can bring stable and considerable performance gains.

[0178] In summary, the present invention provides a large-scale model detection method and system for thrust bearing surface damage, which has the following characteristics:

[0179] (1) This invention takes the grayscale image and depth image of the surface morphology of the thrust bearing acquired by the structured light camera as the object, and coordinates the dual-branch coding, gated dynamic fusion and detection-guided weight adjustment to construct a structure perception extraction and fusion module for the surface damage features of the thrust bearing, so as to accurately capture the surface damage features of the thrust bearing.

[0180] (2) This invention constructs a hierarchical semantic knowledge base for thrust bearing surface damage, uses its semantic information for knowledge guidance, generates text prompt vectors and embeds them into a text encoder of Grounding DINO, and achieves cross-modal alignment of thrust bearing surface damage “text-image”.

[0181] (3) The present invention combines Grounding DINO and SAM to construct a large model detection method for surface damage of thrust bearings. By using a cross-modal collaborative mechanism, the method outputs candidate prediction boxes for surface damage of thrust bearings and constructs a SAM damage fine segmentation model guided by the prediction boxes to achieve fine segmentation of the damage area.

[0182] (4) Based on the PPO framework, this invention constructs a reinforcement learning-driven fine-tuning mechanism for thrust bearing surface damage detection, redefines the state space, action form and compound reward mechanism, guides the model to dynamically optimize the confidence distribution in multi-class damage scenarios, and improves the accuracy and robustness of multi-class damage detection.

[0183] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0184] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0185] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed in this invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0186] In the embodiments provided by this invention, it should be understood that the disclosed devices / terminals and methods can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0187] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0188] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0189] If the integrated module / unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random-access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately added or removed according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable media do not include electrical carrier signals and telecommunication signals.

[0190] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus, and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0191] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0192] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0193] The above content is only for illustrating the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solution based on the technical concept proposed in this invention shall fall within the scope of protection of the claims of this invention.

Claims

1. A method of detecting a large model of damage to a thrust bearing surface, characterized by, Includes the following steps: S1. Obtain the grayscale image and depth image of the thrust bearing surface, and perform feature encoding and fusion on the grayscale image and the depth image to obtain structure-aware fusion features, specifically including: Construct a depth branch based on the depth map, encode and enhance the depth map to extract structural abrupt change features; Construct a grayscale branch based on the grayscale image, encode the grayscale image, and extract texture and local edge features; A feature dynamic fusion module combining the depth branch and the grayscale branch is constructed. A gating mechanism is used to dynamically fuse the features extracted by the two branches. A detection-guided weight adjustment mechanism is introduced. Based on the difference between the number of target boxes detected based on the depth map and the number of target boxes detected based on the fused image, the feature weights of the grayscale branch and the depth branch are dynamically adjusted to obtain the detection-guided optimized structure-aware fusion features. S2. Construct a hierarchical semantic knowledge base for thrust bearing surface damage, generate text prompt vectors based on the semantic information of the hierarchical semantic knowledge base, perform cross-modal alignment between the text prompt vectors and the structure-aware fusion features obtained in step S1, and output the cross-modal aligned text features. S3. Construct a large-scale model of thrust bearing surface damage. This model includes a multi-scale feature-enhanced damage feature image backbone network, a damage semantic text backbone network, a cross-modal fusion module, and a SAM damage fine segmentation model. Based on the structure-aware fusion features output in step S1, multi-scale image features of thrust bearing surface damage are generated through the damage feature image backbone network. Based on the text cue vector generated in step S2, text features with the same dimension as the multi-scale image features of the thrust bearing surface damage are generated through the damage semantic text backbone network. The multi-modal fusion module performs cross-modal fusion of the multi-scale image features and text features of the thrust bearing surface damage to obtain cross-modal fused image-text features. Then, a visual-text jointly represented candidate prediction box for thrust bearing surface damage is output. The candidate prediction box guides the SAM damage fine segmentation model for feature extraction, mask generation, and edge optimization, outputting the fine segmentation result of the damaged area. S4. Based on the PPO framework, a reinforcement learning-driven fine-tuning mechanism is constructed. Using the structure-aware fusion features obtained in step S1, the cross-modal aligned text features obtained in step S2, the cross-modal fused image-text features obtained in step S3, the candidate prediction boxes for thrust bearing surface damage, and the fine segmentation results of the damage area as the basis, the image-text features obtained in step S3 are encoded as the state space. A composite reward mechanism is designed by combining the candidate prediction boxes for thrust bearing surface damage and the fine segmentation results of the damage area. The action form corresponding to the damage confidence distribution is redefined, and the large model of thrust bearing surface damage is iteratively fine-tuned to dynamically optimize the multi-class damage confidence distribution.

2. The thrust bearing surface damage large model detection method according to claim 1, characterized by, The fusion features achieved by the gating mechanism are as follows: in, The initial fused features are the results of the dynamic fusion of the two branch features through a gating mechanism. Here, ⊙ represents the Sigmoid activation function, and ⊙ represents the Hadamard product. For trainable weights, , For branch encoding tensors, For deep branch attention weights, This is a bias term.

3. The thrust bearing surface damage large model detection method according to claim 1, characterized by, Structure-aware fused features To: wherein, is a Sigmoid activation function, and is a Hadamard product, is a trainable weight, , is a branch encoding tensor, is a grayscale branch weight, is a depth branch weight, is a bias term.

4. The thrust bearing surface damage large model detection method according to claim 1, characterized by, Step S2, the process of generating text cue vectors and cross-modal alignment includes: textual cue vector wherein, is a category encoding vector, is a base definition vector, is a form description vector, is a synonym or expanded class description embedding; The text prompt vector is input into the text encoder of the text-guided cross-modal target detection model. The semantic information is deeply modeled through the Transformer structure, and the feature dimension is adjusted by linear projection so that the generated text features are consistent with the feature dimensions of the multi-scale image of thrust bearing surface damage in step S3, thus completing the cross-modal alignment.

5. The thrust bearing surface damage large model detection method according to claim 1, characterized by, In step S3, the process of cross-modal fusion of multi-scale image features and text features is as follows: The cross-modal attention module of the text-guided cross-modal object detection model uses a multi-head attention mechanism to achieve dynamic semantic interaction between multi-scale image features and text features, resulting in cross-modal fused image and text features. The cross-modal fused image and text features are input into the detection box encoder, and the output is a candidate prediction box for thrust bearing surface damage, which is jointly represented by vision and text.

6. The method for detecting large-scale surface damage of thrust bearings according to claim 1, characterized in that, In step S3, the process of using the candidate prediction box for thrust bearing surface damage to guide the SAM damage fine segmentation model to segment the damage region is as follows: The candidate prediction box is input into the SAM damage fine segmentation model. The feature extraction network of the SAM damage fine segmentation model is used to enhance the features of the damage area and generate an initial mask. Edge optimization processing is performed on the initial mask to eliminate the blurring of the mask edges and output the fine segmentation result of the damaged area.

7. The thrust bearing surface damage large model detection method according to claim 1, characterized by, In step S4, the definition of the state space and action form is as follows: The state space is constituted by the cross-modal fused image-text feature encoding of step S3, denoted as state ; Action Form For the surface of the thrust bearing The confidence distribution of damage class is expressed as follows: , Given a set of real numbers, the damage confidence distribution is modeled as the policy output in a continuous action space. ,in, Class 1, Class 2 to Class 3 Confidence response to damage type.

8. The thrust bearing surface damage large model detection method of claim 1, wherein, In step S4, a composite reward mechanism is designed, the action form corresponding to the damage confidence distribution is redefined, and the large-scale model of thrust bearing surface damage is iteratively fine-tuned, including: Composite reward function is: wherein, is a base reward; is an attention reward; is a sparsity penalty term; Construct a commentator network for cross-modal fused image-text features, and utilize the commentator network to process the cross-modal fused image-text feature tensor. Perform global pooling to evaluate the current state of the policy. s Value; Total loss function is: in, For strategic losses, This is due to the loss of testing and supervision from the original Grounding DINO. This is the balance coefficient; The process involves closed-loop iterative fine-tuning of state-action generation, confidence assessment, reward feedback, and feature adjustment until the large model of thrust bearing surface damage meets the iteration termination condition.

9. A thrust bearing surface damage large model detection system, characterized by, include: The extraction module is used to acquire grayscale and depth maps of the thrust bearing surface, and to perform feature encoding and fusion on the grayscale and depth maps to obtain structure-aware fusion features, specifically including: Construct a depth branch based on the depth map, encode and enhance the depth map to extract structural abrupt change features; Construct a grayscale branch based on the grayscale image, encode the grayscale image, and extract texture and local edge features; A feature dynamic fusion module combining the depth branch and the grayscale branch is constructed. A gating mechanism is used to dynamically fuse the features extracted by the two branches. A detection-guided weight adjustment mechanism is introduced. Based on the difference between the number of target boxes detected based on the depth map and the number of target boxes detected based on the fused image, the feature weights of the grayscale branch and the depth branch are dynamically adjusted to obtain the detection-guided optimized structure-aware fusion features. The semantic module is used to construct a hierarchical semantic knowledge base for thrust bearing surface damage. Based on the semantic information of the hierarchical semantic knowledge base, a text prompt vector is generated. The text prompt vector is then aligned with the obtained structure-aware fusion features across modalities, and the cross-modal aligned text features are output. The segmentation module is used to construct a large-scale model of thrust bearing surface damage. This model includes a multi-scale feature-enhanced damage feature image backbone network, a damage semantic text backbone network, a cross-modal fusion module, and a SAM (Structured Aspect-Based Model) damage fine segmentation model. Based on the output structure-aware fusion features, the damage feature image backbone network generates multi-scale image features of thrust bearing surface damage. Based on the generated text cue vectors, the damage semantic text backbone network generates text features with dimensions consistent with the multi-scale image features of the thrust bearing surface damage. The cross-modal fusion module performs cross-modal fusion of the multi-scale image features and text features of the thrust bearing surface damage, obtaining cross-modal fused image-text features, and then outputs a visual-text joint representation of the thrust bearing surface damage candidate prediction box. The thrust bearing surface damage candidate prediction box guides the SAM damage fine segmentation model to perform feature extraction, mask generation, and edge optimization, outputting the fine segmentation result of the damaged region. The output module is used to construct a reinforcement learning-driven fine-tuning mechanism based on the PPO framework. Based on the obtained structure-aware fusion features, cross-modal aligned text features, cross-modal fused image-text features, thrust bearing surface damage candidate prediction boxes, and fine segmentation results of damage regions, the cross-modal fused image-text features are encoded as the state space. A composite reward mechanism is designed by combining the thrust bearing surface damage candidate prediction boxes and fine segmentation results of damage regions. The action form corresponding to the damage confidence distribution is redefined, and the large-scale model of thrust bearing surface damage is iteratively fine-tuned to dynamically optimize the multi-class damage confidence distribution.

Citation Information

Patent Citations

  • Methods and systems for creating virtual and augmented reality

    CN106937531A

  • Bearing defect detection system and method based on multi-combination lightsource and multi-view multi-channel YOLO network structure detection algorithm

    WO2025199743A1