Gear shaft automated assembly system and method based on computer vision

Through computer vision technology based on deep learning, semantic feature matching of gear shaft assembly images, the problem of low traditional manual detection efficiency is solved, automatic and efficient detection of gear shaft assembly quality is realized, and detection accuracy and production efficiency are improved.

CN118735859BActive Publication Date: 2025-08-29TAIZHOU WUBIAO MASCH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410722107.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-05
Publication Date
2025-08-29
Estimated Expiration
2044-06-05

AI Technical Summary

Technical Problem

Traditional gear shaft assembly inspection relies on manual operation, with low efficiency, strong subjectivity and large errors, making it difficult to meet the large-scale and high-efficiency production needs of modern machinery manufacturing.

Method used

Computer vision technology based on deep learning is used to measure semantic feature matching measurements on gear shaft assembly images and reference images, generate early warning prompts, and realize automated and efficient detection.

Benefits of technology

It improves the accuracy and reliability of gear shaft assembly quality inspection, reduces the cost and error of manual inspection, and improves production efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118735859B_ABST
    Figure CN118735859B_ABST
Patent Text Reader

Abstract

The present application relates to the field of image processing technology, and specifically discloses a computer vision-based gear shaft automated assembly system and method. The system uses deep learning-based computer vision technology to compare and analyze the gear shaft assembly image to be inspected and the qualified gear shaft assembly reference image, respectively extracting the image semantic features of the gear shaft assembly image to be inspected and the qualified gear shaft assembly reference image. By performing semantic feature matching measurement on the two, the system intelligently determines whether the assembly quality of the gear shaft to be inspected is unqualified, and then generates an early warning prompt. In this way, the automated and efficient detection of the gear shaft assembly quality can be achieved to improve production efficiency, while reducing the cost and error of manual detection and improving the accuracy and reliability of assembly quality detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and more specifically, to a computer vision-based gear shaft automated assembly system and method. Background Art

[0002] Gear assembly is a complex and critical process in mechanical manufacturing, involving the accurate assembly of multiple precision components to ensure the efficient and smooth operation of the gear system. The quality of gear assembly directly impacts the performance and reliability of mechanical equipment. Therefore, after assembly, comprehensive testing is typically required to ensure that the assembly meets design requirements.

[0003] Traditional gear shaft assembly inspection methods rely primarily on manual labor, inspecting each component individually through visual inspection or measurement tools. However, this method suffers from low inspection efficiency, high subjectivity, and large errors. Furthermore, due to the complex structure of gear shafts and the difficulty of inspection, traditional manual inspection methods are no longer able to meet the large-scale, high-efficiency production requirements of modern machinery manufacturing.

[0004] With the rapid development of computer vision technology, the application of computer vision technology in the automated detection of gear shaft assembly processes has gradually become a research hotspot in the field of mechanical manufacturing. Therefore, a computer vision-based automated gear shaft assembly system and method is expected. Summary of the Invention

[0005] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a computer vision-based gear shaft automated assembly system and method, which uses deep learning-based computer vision technology to compare and analyze the gear shaft assembly image to be inspected and the qualified gear shaft assembly reference image, respectively extract the image semantic features of the gear shaft assembly image to be inspected and the qualified gear shaft assembly reference image, and perform semantic feature matching measurement on the two, so as to intelligently judge whether the assembly quality of the gear shaft to be inspected is unqualified, and then generate an early warning prompt. In this way, the automated and efficient detection of the gear shaft assembly quality can be achieved to improve production efficiency, while reducing the cost and error of manual detection and improving the accuracy and reliability of assembly quality detection.

[0006] Accordingly, according to one aspect of the present application, a computer vision-based automated gear shaft assembly system is provided, comprising:

[0007] The gear shaft assembly image acquisition module is used to acquire the gear shaft assembly image captured by the camera and extract a set of gear shaft assembly reference images marked as qualified from the background database;

[0008] a reference image semantic feature extraction module, configured to perform assembly state feature extraction and semantic distillation on each of the qualified gear shaft assembly reference images in the set of qualified gear shaft assembly reference images to obtain a set of gear shaft assembly reference semantic distillation feature vectors;

[0009] A detection image semantic feature extraction module is used to perform assembly state feature extraction and semantic distillation on the gear shaft assembly image to obtain a gear shaft assembly detection semantic distillation feature vector;

[0010] a semantic matching measurement module, configured to perform semantic matching measurement on each gear shaft assembly reference semantic distillation feature vector in the set of the gear shaft assembly reference semantic distillation feature vectors and the gear shaft assembly detection semantic distillation feature vector to obtain a set of semantic matching degrees;

[0011] The unqualified warning module is used to generate a warning prompt of unqualified gear shaft assembly in response to at least one semantic matching degree in the set of semantic matching degrees being less than or equal to a predetermined threshold.

[0012] In the above-mentioned computer vision-based gear shaft automated assembly system, the reference image semantic feature extraction module includes: an assembly state feature extraction unit, which is used to pass each of the qualified gear shaft assembly reference images in the set of qualified gear shaft assembly reference images through an assembly state feature extractor based on a DenseNet model to obtain a set of gear shaft assembly reference feature maps; and a multi-scale semantic distillation unit, which is used to pass each of the gear shaft assembly reference feature maps in the set of gear shaft assembly reference feature maps through a multi-scale semantic distillation module to obtain a set of gear shaft assembly reference semantic distillation feature vectors.

[0013] In the above-mentioned computer vision-based gear shaft automated assembly system, the multi-scale semantic distillation unit includes: a multi-scale semantic feature extraction sub-unit, used to respectively extract the global semantic features and local semantic features of the gear shaft assembly reference feature map to obtain a gear shaft assembly reference global semantic feature vector and a gear shaft assembly reference local semantic feature vector; a multi-scale semantic feature fusion sub-unit, used to fuse the gear shaft assembly reference global semantic feature vector and the gear shaft assembly reference local semantic feature vector to obtain the gear shaft assembly reference semantic distillation feature vector.

[0014] In the above-mentioned computer vision-based gear shaft automated assembly system, the multi-scale semantic feature extraction subunit is used to: upsample the gear shaft assembly reference feature map to obtain an upsampled gear shaft assembly reference feature map; pass the upsampled gear shaft assembly reference feature map through a global semantic feature extraction module to obtain the gear shaft assembly reference global semantic feature vector; pass the upsampled gear shaft assembly reference feature map through a local semantic feature extraction module to obtain the gear shaft assembly reference local semantic feature vector.

[0015] In the above-mentioned computer vision-based gear shaft automated assembly system, the global semantic feature extraction module includes a global average pooling layer, a first point convolution layer, a first batch normalization layer and a first activation layer, and the local semantic feature extraction module includes a second point convolution layer, a second batch normalization layer and a second activation layer.

[0016] In the above-mentioned computer vision-based gear shaft automated assembly system, the detection image semantic feature extraction module is used to: pass the gear shaft assembly image through the assembly state feature extractor based on the DenseNet model and the multi-scale semantic distillation module to obtain the gear shaft assembly detection semantic distillation feature vector.

[0017] In the above-mentioned computer vision-based gear shaft automated assembly system, the semantic matching measurement module includes: a feature distribution optimization unit, which is used to perform feature distribution optimization on the set of gear shaft assembly reference semantic distillation feature vectors to obtain a set of optimized gear shaft assembly reference semantic distillation feature vectors; a semantic matching degree calculation unit, which is used to calculate the semantic matching degree between the gear shaft assembly detection semantic distillation feature vector and each optimized gear shaft assembly reference semantic distillation feature vector in the set of optimized gear shaft assembly reference semantic distillation feature vectors to obtain the set of semantic matching degrees.

[0018] In the above-mentioned computer vision-based gear shaft automated assembly system, the semantic matching degree calculation unit is used to calculate the semantic matching degree between the gear shaft assembly detection semantic distillation feature vector and the optimized gear shaft assembly reference semantic distillation feature vector using the following semantic matching formula, wherein the semantic matching formula is:

[0019] r=sigmoid(V w1 *V1+V w2 *V2)

[0020] Among them, V1 represents the semantic distillation feature vector of the gear shaft assembly detection, V2 represents the semantic distillation feature vector of the optimized gear shaft assembly reference, V w1 represents the first weight vector, V w2represents the second weight vector, * represents vector multiplication, + represents addition, sigmoid represents the activation function, and r represents the semantic matching degree.

[0021] According to another aspect of the present application, a computer vision-based automated assembly method for a gear shaft is provided, comprising:

[0022] Acquire the gear shaft assembly image captured by the camera, and extract a set of gear shaft assembly reference images marked as qualified from the background database;

[0023] performing assembly state feature extraction and semantic distillation on each qualified gear shaft assembly reference image in the set of qualified gear shaft assembly reference images to obtain a set of gear shaft assembly reference semantic distillation feature vectors;

[0024] Performing assembly state feature extraction and semantic distillation on the gear shaft assembly image to obtain a gear shaft assembly detection semantic distillation feature vector;

[0025] Performing semantic matching measurement on each gear shaft assembly reference semantic distillation feature vector in the set of the gear shaft assembly reference semantic distillation feature vectors and the gear shaft assembly detection semantic distillation feature vector to obtain a set of semantic matching degrees;

[0026] In response to at least one semantic matching degree in the set of semantic matching degrees being less than or equal to a predetermined threshold, an early warning prompt indicating that the gear shaft assembly is unqualified is generated.

[0027] Compared to the prior art, the computer vision-based automated gear shaft assembly system and method provided by this application utilizes deep learning-based computer vision technology to perform comparative analysis between an image of the gear shaft assembly to be inspected and a reference image of a qualified gear shaft assembly. The system extracts semantic features from the image of the gear shaft assembly to be inspected and the reference image of a qualified gear shaft assembly, respectively. By performing semantic feature matching measurement on the two, the system intelligently determines whether the assembly quality of the gear shaft to be inspected is unqualified, thereby generating an early warning. This allows for automated and efficient inspection of gear shaft assembly quality, thereby improving production efficiency while reducing the cost and error of manual inspection and improving the accuracy and reliability of assembly quality inspection. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0029] Figure 1 4 is a block diagram of a computer vision-based gear shaft automated assembly system according to an embodiment of the present application.

[0030] Figure 2 Schematic diagram of the architecture of a computer vision-based gear shaft automated assembly system according to an embodiment of the present application.

[0031] Figure 3 4 is a block diagram of a reference image semantic feature extraction module in a computer vision-based gear shaft automated assembly system according to an embodiment of the present application.

[0032] Figure 4 4 is a block diagram of a multi-scale semantic distillation unit in a computer vision-based gear shaft automated assembly system according to an embodiment of the present application.

[0033] Figure 5 4 is a block diagram of a semantic matching measurement module in a computer vision-based gear shaft automated assembly system according to an embodiment of the present application.

[0034] Figure 6 The figure is a flowchart of a method for automatic assembly of a gear shaft based on computer vision according to an embodiment of the present application. DETAILED DESCRIPTION

[0035] Below, the embodiments of the present application will be described in more detail with reference to the accompanying drawings, and the above-mentioned and other purposes, features, and advantages of the present application will become more apparent. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments of the present application, and it should be understood that the present application is not limited to the example embodiments described herein.

[0036] As mentioned in the background art above, traditional gear shaft assembly inspection methods primarily rely on manual operation, namely, visually inspecting the assembled components one by one or using measuring tools. However, this method has significant shortcomings, such as low inspection efficiency, excessive subjective judgment factors, and high error rates. In addition, given the complexity of gear shaft structures and the increasing difficulty of inspection, traditional manual inspection methods are no longer able to meet the urgent requirements of the modern machinery manufacturing industry for large-scale, high-efficiency production. To address the above technical issues, the technical concept of this application is to use deep learning-based computer vision technology to compare and analyze the gear shaft assembly image to be inspected and the reference image of a qualified gear shaft assembly, respectively extract the image semantic features of the gear shaft assembly image to be inspected and the qualified gear shaft assembly reference image, and then intelligently determine whether the gear shaft to be inspected has unqualified assembly quality by performing semantic feature matching measurement on the two images, thereby generating an early warning prompt. In this way, automated and efficient inspection of gear shaft assembly quality can be achieved to improve production efficiency, while reducing the cost and error of manual inspection and improving the accuracy and reliability of assembly quality inspection.

[0037] Figure 1 4 is a block diagram of a computer vision-based gear shaft automated assembly system according to an embodiment of the present application. Figure 2 FIG. 1 is a schematic diagram of the architecture of a gear shaft automatic assembly system based on computer vision according to an embodiment of the present application. Figure 1 and Figure 2 As shown, according to an embodiment of the present application, a computer vision-based gear shaft automated assembly system 100 includes: a gear shaft assembly image acquisition module 110, which is used to acquire a gear shaft assembly image captured by a camera and extract a set of gear shaft assembly reference images marked as qualified from a background database; a reference image semantic feature extraction module 120, which is used to perform assembly state feature extraction and semantic distillation on each qualified gear shaft assembly reference image in the set of qualified gear shaft assembly reference images to obtain a set of gear shaft assembly reference semantic distillation feature vectors; a detection image semantic feature extraction module 130, which is used to perform assembly state feature extraction and semantic distillation on the gear shaft assembly image to obtain a gear shaft assembly detection semantic distillation feature vector; a semantic matching measurement module 140, which is used to perform semantic matching measurement on each gear shaft assembly reference semantic distillation feature vector in the set of gear shaft assembly reference semantic distillation feature vectors and the gear shaft assembly detection semantic distillation feature vector to obtain a set of semantic matching degrees; and an unqualified warning module 150, which is used to generate an early warning prompt of unqualified gear shaft assembly in response to the presence of at least one semantic matching degree less than or equal to a predetermined threshold in the set of semantic matching degrees.

[0038] In the aforementioned computer vision-based automated gear shaft assembly system 100, the gear shaft assembly image acquisition module 110 is configured to acquire gear shaft assembly images captured by a camera and extract a set of gear shaft assembly reference images labeled as qualified from a backend database. It should be understood that capturing gear shaft assembly images via a camera captures the visual characteristics of the gear shaft's assembly state, providing a data foundation for subsequent assembly quality assessment. Furthermore, to enable more intuitive and effective assembly quality assessment, a set of gear shaft assembly reference images labeled as qualified is further extracted from the backend database as a comparison benchmark. Each gear shaft assembly reference image in this set represents the correct state of the gear shaft assembly and provides a clear acceptance standard. By performing image semantic feature matching analysis between the gear shaft assembly image to be inspected and the qualified reference images, the gear shaft assembly quality can be more accurately determined. Furthermore, extracting this set of gear shaft assembly reference images overcomes the limitations of a single reference image, enabling a more comprehensive and accurate assessment of assembly quality. Furthermore, in a specific implementation, different types of qualified gear shaft assembly reference images can be added or updated in the background database to accommodate different types of gear shafts or assembly requirements, thereby enhancing the flexibility and scalability of the system.

[0039] In the above-mentioned computer vision-based gear shaft automated assembly system 100, the reference image semantic feature extraction module 120 is used to perform assembly state feature extraction and semantic distillation on each qualified gear shaft assembly reference image in the set of qualified gear shaft assembly reference images to obtain a set of gear shaft assembly reference semantic distillation feature vectors. Figure 3 FIG. 1 is a block diagram of a reference image semantic feature extraction module in a computer vision-based gear shaft automated assembly system according to an embodiment of the present application. Figure 3 As shown, the reference image semantic feature extraction module 120 includes: an assembly state feature extraction unit 121, which is used to pass each of the qualified gear shaft assembly reference images in the set of qualified gear shaft assembly reference images through an assembly state feature extractor based on a DenseNet model to obtain a set of gear shaft assembly reference feature maps; a multi-scale semantic distillation unit 122, which is used to pass each of the gear shaft assembly reference feature maps in the set of gear shaft assembly reference feature maps through a multi-scale semantic distillation module to obtain a set of gear shaft assembly reference semantic distillation feature vectors.

[0040] Specifically, the assembly state feature extraction unit 121 is used to pass each of the assembled gear shaft assembly reference images in the set of assembled gear shaft assembly reference images through an assembly state feature extractor based on the DenseNet model to obtain a set of gear shaft assembly reference feature maps. That is, in order to extract the assembled gear shaft state features from each gear shaft assembly reference image, in the technical solution of the present application, an assembly state feature extractor based on the DenseNet model is used to process each of the assembled gear shaft assembly reference images in the set of assembled gear shaft assembly reference images. Those skilled in the art should know that DenseNet (Densely Connected Convolutional Networks) is a deep learning model architecture that uses a dense connectivity structure so that the output of each layer is connected to the outputs of all previous layers, which facilitates information transfer and helps alleviate the gradient vanishing problem. Here, the assembly status feature extractor based on the DenseNet model performs deep convolution operations on each gear shaft assembly reference image through its unique dense connection architecture. It can effectively extract key visual features such as component shape, texture, and color in the image, fully express the qualified assembly status of the gear shaft, and provide effective data support for subsequent comparative analysis.

[0041] Specifically, the multi-scale semantic distillation unit 122 is configured to process each gear shaft assembly reference feature map in the set of gear shaft assembly reference feature maps through a multi-scale semantic distillation module to obtain a set of gear shaft assembly reference semantic distillation feature vectors. It should be understood that, given that the gear shaft assembly reference images typically contain assembly structures at multiple scales, such as local component connection structures and overall assembly layout structures, in order to more comprehensively and accurately represent the qualified state of the gear shaft assembly, the technical solution of the present application utilizes a multi-scale semantic distillation module to extract and fuse multi-scale semantic features from each gear shaft assembly reference feature map. Specifically, the multi-scale semantic distillation module mines global and local semantic features from the gear shaft assembly reference feature maps using a global semantic feature extraction module and a local semantic feature extraction module, respectively, to capture local assembly detail information and overall structural context information in the gear shaft assembly reference images. In particular, a global average pooling layer is provided in the global semantic feature extraction module to globally condense the gear shaft assembly reference feature maps to learn their overall distribution pattern. In this way, by fusing the global and local assembly state semantic features to obtain the gear shaft assembly reference semantic distillation feature vector, the assembly state of the gear shaft can be described more accurately, thereby improving the understanding of the gear shaft assembly reference image.

[0042] Figure 4 FIG is a block diagram of a multi-scale semantic distillation unit in a computer vision-based gear shaft automated assembly system according to an embodiment of the present application. Figure 4 As shown, the multi-scale semantic distillation unit 122 includes: a multi-scale semantic feature extraction sub-unit 1221, which is used to extract the global semantic features and local semantic features of the gear shaft assembly reference feature map respectively to obtain the gear shaft assembly reference global semantic feature vector and the gear shaft assembly reference local semantic feature vector; a multi-scale semantic feature fusion sub-unit 1222, which is used to fuse the gear shaft assembly reference global semantic feature vector and the gear shaft assembly reference local semantic feature vector to obtain the gear shaft assembly reference semantic distillation feature vector.

[0043] More specifically, the multi-scale semantic feature extraction subunit 1221 is configured to: upsample the gear shaft assembly reference feature map to obtain an upsampled gear shaft assembly reference feature map; pass the upsampled gear shaft assembly reference feature map through a global semantic feature extraction module to obtain a gear shaft assembly reference global semantic feature vector; and pass the upsampled gear shaft assembly reference feature map through a local semantic feature extraction module to obtain a gear shaft assembly reference local semantic feature vector. The global semantic feature extraction module includes a global average pooling layer, a first point convolution layer, a first batch normalization layer, and a first activation layer, and the local semantic feature extraction module includes a second point convolution layer, a second batch normalization layer, and a second activation layer.

[0044] In the above-mentioned computer vision-based gear shaft automated assembly system 100, the detection image semantic feature extraction module 130 is used to perform assembly state feature extraction and semantic distillation on the gear shaft assembly image to obtain a gear shaft assembly detection semantic distillation feature vector. In a specific example of the present application, the processing method for performing assembly state feature extraction and semantic distillation on the gear shaft assembly image is to pass the gear shaft assembly image through the assembly state feature extractor based on the DenseNet model and the multi-scale semantic distillation module to obtain the gear shaft assembly detection semantic distillation feature vector. That is, for the gear shaft assembly image to be detected, its processing flow is similar to the processing flow of the gear shaft assembly reference image. Similarly, the gear shaft assembly image is processed using the DenseNet model-based assembly state feature extractor and the multi-scale semantic distillation module to capture key visual features such as component shape, texture, and color in the gear shaft assembly image, and to mine multi-scale assembly structure semantic information, thereby obtaining a gear shaft assembly detection semantic distillation feature vector. This allows for a comprehensive and accurate expression of the assembly state of the gear shaft to be inspected, providing data support for subsequent comparative analysis with the gear shaft assembly reference semantic distillation feature vector. Furthermore, through this same image processing approach, the semantic feature space consistency and comparability between the gear shaft assembly reference semantic distillation feature vector and the gear shaft assembly detection semantic distillation feature vector can be ensured, thereby improving the accuracy and reliability of assembly quality assessment.

[0045] In the above-mentioned computer vision-based gear shaft automated assembly system 100, the semantic matching metric module 140 is used to perform semantic matching metric on each gear shaft assembly reference semantic distillation feature vector in the set of gear shaft assembly reference semantic distillation feature vectors and the gear shaft assembly detection semantic distillation feature vector to obtain a set of semantic matching degrees. Figure 5 FIG. 1 is a block diagram of a semantic matching metric module in a computer vision-based gear shaft automated assembly system according to an embodiment of the present application. Figure 5 As shown, the semantic matching metric module 140 includes: a feature distribution optimization unit 141, which is used to perform feature distribution optimization on the set of the gear shaft assembly reference semantic distillation feature vectors to obtain a set of optimized gear shaft assembly reference semantic distillation feature vectors; a semantic matching degree calculation unit 142, which is used to calculate the semantic matching degree between the gear shaft assembly detection semantic distillation feature vector and each optimized gear shaft assembly reference semantic distillation feature vector in the set of optimized gear shaft assembly reference semantic distillation feature vectors to obtain the set of semantic matching degrees.

[0046] Specifically, the feature distribution optimization unit 141 is used to perform feature distribution optimization on the set of the gear shaft assembly reference semantic distillation feature vectors to obtain a set of optimized gear shaft assembly reference semantic distillation feature vectors. In particular, considering the differences in the image semantic feature expressions of the source image semantic differences of each qualified gear shaft assembly reference image in the set of qualified gear shaft assembly reference images after image semantic feature extraction, each gear shaft assembly reference semantic distillation feature vector will also have feature deviation differences relative to the image semantic features of the gear shaft assembly detection semantic distillation feature vector, thereby affecting the mapping regression constraints of each gear shaft assembly reference semantic distillation feature vector to the gear shaft assembly detection semantic distillation feature vector as the semantic feature deviation reference source under the semantic matching domain. Based on this, in the technical solution of the present application, feature distribution optimization is further performed on the set of gear shaft assembly reference semantic distillation feature vectors.

[0047] Preferably, the feature distribution optimization of the set of the gear shaft assembly reference semantic distillation feature vectors to obtain the set of optimized gear shaft assembly reference semantic distillation feature vectors includes the following steps: cascading the individual gear shaft assembly reference semantic distillation feature vectors in the set of the gear shaft assembly reference semantic distillation feature vectors to obtain a gear shaft assembly reference semantic distillation joint feature vector; calculating the sum of the point sum of the gear shaft assembly reference semantic distillation joint feature vector and the square root of its length and the reciprocal of the square root of its second norm to obtain a first gear shaft assembly reference semantic distillation joint intermediate feature vector; calculating the exponential of the first gear shaft assembly reference semantic distillation joint intermediate feature vector with a natural constant as the base function to obtain the second gear shaft assembly reference semantic distillation joint intermediate feature vector; calculate the dot product of the gear shaft assembly reference semantic distillation joint feature vector and its norm and weight hyperparameter to obtain the third gear shaft assembly reference semantic distillation joint intermediate feature vector; calculate the dot sum of the second gear shaft assembly reference semantic distillation joint intermediate feature vector and the third gear shaft assembly reference semantic distillation joint intermediate feature vector to obtain the optimized gear shaft assembly reference semantic distillation joint feature vector; convert the optimized gear shaft assembly reference semantic distillation joint based on the cascade of the individual gear shaft assembly reference semantic distillation feature vectors into a set of optimized gear shaft assembly reference semantic distillation feature vectors.

[0048] In the above preferred example, the structured norm of the gear shaft assembly reference semantic distillation joint feature vector, which is the joint semantic representation of each gear shaft assembly reference semantic distillation feature vector in the set of the gear shaft assembly reference semantic distillation feature vector, is used as the local canonical coordinate for each eigenvalue of each gear shaft assembly reference semantic distillation feature vector to determine the overall distribution of features of each gear shaft assembly reference semantic distillation feature vector, representing the rotational offset of the eigenvalue relative to the eigenvalue for the offset prediction direction of each eigenvalue of the each gear shaft assembly reference semantic distillation feature vector as the center, and the bounding box of the eigenvalue distribution of each gear shaft assembly reference semantic distillation feature vector is used to perform eigenvalue constraints to improve the mapping regression constraints of each gear shaft assembly reference semantic distillation feature vector to the gear shaft assembly detection semantic distillation feature vector as the semantic feature deviation reference source under the semantic matching domain, thereby improving the training speed of the model and the consistency of the calculation results of each semantic matching degree in the set of semantic matching degrees.

[0049] Specifically, the semantic matching degree calculation unit 142 is configured to calculate the semantic matching degree between the gear shaft assembly detection semantic distilled feature vector and each optimized gear shaft assembly reference semantic distilled feature vector in the set of optimized gear shaft assembly reference semantic distilled feature vectors, thereby obtaining the set of semantic matching degrees. In other words, by calculating the semantic matching degree between the gear shaft assembly detection semantic distilled feature vector and each optimized gear shaft assembly reference semantic distilled feature vector in the set of optimized gear shaft assembly reference semantic distilled feature vectors, the degree of similarity between the gear shaft assembly image and each gear shaft assembly reference image is quantitatively described. The magnitude of the semantic matching degree directly reflects the assembly quality of the gear shaft being inspected. A higher semantic matching degree indicates a closer match to a qualified gear shaft assembly state and higher assembly quality. Conversely, a lower semantic matching degree may indicate assembly quality issues. This matching metric achieves a quantitative assessment of gear shaft assembly quality, improving the accuracy and objectivity of the assessment.

[0050] In specific examples of the present application, the semantic matching degree calculation unit 142 is used to calculate the semantic matching degree between the gear shaft assembly detection semantic distillation feature vector and the optimized gear shaft assembly reference semantic distillation feature vector using the following semantic matching formula, wherein the semantic matching formula is:

[0051] r=sigmoid(V w1 *V1+V w2 *V2)

[0052] Among them, V1 represents the semantic distillation feature vector of the gear shaft assembly detection, V2 represents the semantic distillation feature vector of the optimized gear shaft assembly reference, V w1 represents the first weight vector, V w2 represents the second weight vector, * represents vector multiplication, + represents addition, sigmoid represents the activation function, and r represents the semantic matching degree.

[0053] In the above-mentioned computer vision-based gear shaft automated assembly system 100, the unqualified warning module 150 is used to generate a warning prompt of unqualified gear shaft assembly in response to the presence of at least one semantic matching degree in the set of semantic matching degrees being less than or equal to a predetermined threshold. That is, the set of semantic matching degrees is screened and judged by setting a predetermined threshold to determine whether the assembly quality of the gear shaft to be tested meets the preset requirements. When at least one semantic matching degree in the set of semantic matching degrees is less than or equal to the predetermined threshold, it means that the similarity between the gear shaft assembly image to be tested and at least one qualified gear shaft assembly reference image is low, does not meet the expected qualification standard, and may have an assembly quality problem. At this time, the system automatically generates a warning prompt of unqualified gear shaft assembly to remind the operator to conduct further inspection and processing. In this way, potential assembly problems can be discovered in a timely manner, avoiding equipment failure or performance degradation caused by improper assembly, thereby improving the reliability and stability of the equipment.

[0054] In summary, according to the embodiment of the present application, a computer vision-based automated gear shaft assembly system is described. It uses deep learning-based computer vision technology to compare and analyze the gear shaft assembly image to be inspected and the reference image of a qualified gear shaft assembly. It extracts the image semantic features of the gear shaft assembly image to be inspected and the qualified gear shaft assembly reference image, respectively. By performing semantic feature matching measurement on the two, it intelligently determines whether the assembly quality of the gear shaft to be inspected is unqualified, and then generates an early warning prompt. In this way, automated and efficient inspection of gear shaft assembly quality can be achieved to improve production efficiency, while reducing the cost and error of manual inspection and improving the accuracy and reliability of assembly quality inspection.

[0055] Figure 6 Flowchart of the automatic assembly method of gear shaft based on computer vision according to the embodiment of the present application. Figure 6As shown, the computer vision-based automatic assembly method of a gear shaft according to an embodiment of the present application includes the following steps: S1, obtaining a gear shaft assembly image captured by a camera, and extracting a set of gear shaft assembly reference images marked as qualified assembly from a background database; S2, performing assembly state feature extraction and semantic distillation on each qualified gear shaft assembly reference image in the set of qualified gear shaft assembly reference images to obtain a set of gear shaft assembly reference semantic distillation feature vectors; S3, performing assembly state feature extraction and semantic distillation on the gear shaft assembly image to obtain a gear shaft assembly detection semantic distillation feature vector; S4, performing semantic matching measurement on each gear shaft assembly reference semantic distillation feature vector in the set of gear shaft assembly reference semantic distillation feature vectors and the gear shaft assembly detection semantic distillation feature vector to obtain a set of semantic matching degrees; S5, generating an early warning prompt of unqualified gear shaft assembly in response to the presence of at least one semantic matching degree less than or equal to a predetermined threshold in the set of semantic matching degrees.

[0056] Here, those skilled in the art will understand that the specific operations of each step in the above-mentioned computer vision-based gear shaft automated assembly method have been described in detail above. Figures 1 to 5 The description of the computer vision-based gear shaft automated assembly system has been introduced in detail, and therefore, its repeated description will be omitted.

[0057] The basic principles of the present invention have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in the present invention are merely illustrative and non-limiting, and should not be construed as necessarily possessed by each embodiment of the present invention. Furthermore, the specific details of the above embodiments are provided for illustrative purposes and to facilitate understanding, and are not intended to be limiting. These details do not necessarily limit the present invention to being implemented using these specific details.

[0058] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, please refer to the relevant description of other embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiment described above is only schematic. For example, the unit division is only a logical function division, and there may be other division methods in actual implementation. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical subunits, that is, they may be located in one place, or they may be distributed on multiple network subunits. Some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0059] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be encompassed therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.

[0060] In addition, it is obvious that the word "comprising" does not exclude other subunits or steps, and the singular does not exclude the plural. Multiple subunits stated in the system claims can also be implemented by one subunit through software or hardware.

[0061] Finally, it should be noted that the above description has been provided for purposes of illustration and description. Furthermore, the above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to be limiting. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art will appreciate that the technical solutions of the present invention may be modified or replaced with equivalents without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A computer vision-based gear shaft automated assembly system, characterized in that: include: The gear shaft assembly image acquisition module is used to acquire the gear shaft assembly image captured by the camera and extract a set of gear shaft assembly reference images marked as qualified from the background database; a reference image semantic feature extraction module, configured to perform assembly state feature extraction and semantic distillation on each of the qualified gear shaft assembly reference images in the set of qualified gear shaft assembly reference images to obtain a set of gear shaft assembly reference semantic distillation feature vectors; A detection image semantic feature extraction module is used to perform assembly state feature extraction and semantic distillation on the gear shaft assembly image to obtain a gear shaft assembly detection semantic distillation feature vector; a semantic matching measurement module, configured to perform semantic matching measurement on each gear shaft assembly reference semantic distillation feature vector in the set of the gear shaft assembly reference semantic distillation feature vectors and the gear shaft assembly detection semantic distillation feature vector to obtain a set of semantic matching degrees; a non-conformity warning module, configured to generate a warning prompt indicating that the gear shaft assembly is non-conformity in response to at least one semantic matching degree in the set of semantic matching degrees being less than or equal to a predetermined threshold; The semantic matching metric module includes: a feature distribution optimization unit, configured to perform feature distribution optimization on the set of the gear shaft assembly reference semantic distillation feature vectors to obtain a set of optimized gear shaft assembly reference semantic distillation feature vectors; a semantic matching degree calculation unit, configured to calculate a semantic matching degree between the gear shaft assembly detection semantic distillation feature vector and each optimized gear shaft assembly reference semantic distillation feature vector in the set of optimized gear shaft assembly reference semantic distillation feature vectors to obtain the set of semantic matching degrees; Among them, the feature distribution optimization of the set of the gear shaft assembly reference semantic distillation feature vectors to obtain the set of optimized gear shaft assembly reference semantic distillation feature vectors includes the following steps: cascading the individual gear shaft assembly reference semantic distillation feature vectors in the set of the gear shaft assembly reference semantic distillation feature vectors to obtain a gear shaft assembly reference semantic distillation joint feature vector; calculating the sum of the point sum of the gear shaft assembly reference semantic distillation joint feature vector and the square root of its length and the reciprocal of the square root of its second norm to obtain a first gear shaft assembly reference semantic distillation joint intermediate feature vector; calculating the exponential function of the first gear shaft assembly reference semantic distillation joint intermediate feature vector with a natural constant as the base to obtain Obtain a second gear shaft assembly reference semantic distillation joint intermediate feature vector; calculate the dot product of the gear shaft assembly reference semantic distillation joint feature vector and its norm and weight hyperparameter to obtain a third gear shaft assembly reference semantic distillation joint intermediate feature vector; calculate the dot sum of the second gear shaft assembly reference semantic distillation joint intermediate feature vector and the third gear shaft assembly reference semantic distillation joint intermediate feature vector to obtain an optimized gear shaft assembly reference semantic distillation joint feature vector; convert the optimized gear shaft assembly reference semantic distillation joint feature vector into a set of optimized gear shaft assembly reference semantic distillation feature vectors based on the cascade of the individual gear shaft assembly reference semantic distillation feature vectors.

2. The computer vision-based gear shaft automated assembly system according to claim 1, characterized in that: The reference image semantic feature extraction module includes: an assembly state feature extraction unit, configured to pass each qualified gear shaft assembly reference image in the set of qualified gear shaft assembly reference images through an assembly state feature extractor based on a DenseNet model to obtain a set of gear shaft assembly reference feature maps; The multi-scale semantic distillation unit is used to pass each gear shaft assembly reference feature map in the set of gear shaft assembly reference feature maps through a multi-scale semantic distillation module to obtain a set of gear shaft assembly reference semantic distillation feature vectors.

3. The computer vision-based gear shaft automated assembly system according to claim 2, characterized in that: The multi-scale semantic distillation unit includes: a multi-scale semantic feature extraction subunit, configured to respectively extract global semantic features and local semantic features of the gear shaft assembly reference feature graph to obtain a gear shaft assembly reference global semantic feature vector and a gear shaft assembly reference local semantic feature vector; The multi-scale semantic feature fusion subunit is used to fuse the gear shaft assembly reference global semantic feature vector and the gear shaft assembly reference local semantic feature vector to obtain the gear shaft assembly reference semantic distillation feature vector.

4. The computer vision-based gear shaft automated assembly system according to claim 3, characterized in that: The multi-scale semantic feature extraction subunit is used to: Upsampling the gear shaft assembly reference feature map to obtain an upsampled gear shaft assembly reference feature map; Passing the upsampled gear shaft assembly reference feature map through a global semantic feature extraction module to obtain the gear shaft assembly reference global semantic feature vector; The upsampled gear shaft assembly reference feature map is passed through a local semantic feature extraction module to obtain the gear shaft assembly reference local semantic feature vector.

5. The computer vision-based gear shaft automated assembly system according to claim 4, characterized in that: The global semantic feature extraction module includes a global average pooling layer, a first point convolution layer, a first batch normalization layer and a first activation layer, and the local semantic feature extraction module includes a second point convolution layer, a second batch normalization layer and a second activation layer.

6. The computer vision-based gear shaft automated assembly system according to claim 5, characterized in that: The detection image semantic feature extraction module is used to: The gear shaft assembly image is passed through the assembly state feature extractor based on the DenseNet model and the multi-scale semantic distillation module to obtain the gear shaft assembly detection semantic distillation feature vector.

7. A computer vision-based gear shaft automated assembly method, using the computer vision-based gear shaft automated assembly system according to claim 1, characterized in that: include: Acquire the gear shaft assembly image captured by the camera, and extract a set of gear shaft assembly reference images marked as qualified from the background database; performing assembly state feature extraction and semantic distillation on each qualified gear shaft assembly reference image in the set of qualified gear shaft assembly reference images to obtain a set of gear shaft assembly reference semantic distillation feature vectors; Performing assembly state feature extraction and semantic distillation on the gear shaft assembly image to obtain a gear shaft assembly detection semantic distillation feature vector; Performing semantic matching measurement on each gear shaft assembly reference semantic distillation feature vector in the set of the gear shaft assembly reference semantic distillation feature vectors and the gear shaft assembly detection semantic distillation feature vector to obtain a set of semantic matching degrees; In response to at least one semantic matching degree in the set of semantic matching degrees being less than or equal to a predetermined threshold, an early warning prompt indicating that the gear shaft assembly is unqualified is generated.

Citation Information

Patent Citations

  • Water cooling control system and method for online quenching

    CN117127005A

  • Real-time monitoring system and method for plush toy production process

    CN118052793A