Spherical lens multi-mode optical detection sorting method and system

Through the multimodal optical detection and sorting method of spherical lenses, the quality score is generated using deep learning models, which solves the problem that the existing technology cannot comprehensively evaluate the quality of spherical lenses, and achieves efficient and accurate sorting effects.

CN120190143AActive Publication Date: 2025-06-24ZHONGKE BOCHUANG (GUANGDONG) TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510691170.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-24
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to comprehensively evaluate the various quality factors of spherical lenses, resulting in poor sorting results and the inability to effectively distinguish qualified products from unqualified products.

Method used

The multimodal optical detection and sorting method of spherical lenses is adopted to obtain curvature, surface defects and coating uniformity information, and combine with the deep learning model to generate objective quality scores to achieve a comprehensive and accurate evaluation of the quality of spherical lenses.

Benefits of technology

It realizes a comprehensive and accurate evaluation of the quality of spherical lenses, improves the efficiency and accuracy of sorting, and can effectively distinguish qualified products from unqualified products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120190143A_ABST
    Figure CN120190143A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of visual inspection, and particularly discloses a spherical lens multi-mode optical detection sorting method and system.The method comprises the steps that appearance data information of to-be-sorted lenses is obtained, and the appearance data information comprises curvature information, surface defect information and coating uniformity information; the appearance data information is put into a pre-trained quality scoring model to generate a quality score, so that the spherical lens sorting system sorts the to-be-sorted lens according to the quality score; according to the method, by integrating curvature, surface, coating and other multi-modal information and utilizing a deep learning model for comprehensive evaluation, the limitation that a traditional method only pays attention to a single index can be overcome, multiple quality factors of the spherical lens can be comprehensively considered, objective quality scores are generated, comprehensive and accurate evaluation of the quality of the spherical lens is achieved, and the method is suitable for large-scale popularization and application. And the sorting efficiency and precision are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of vision detection technology. Specifically, it relates to a multi-modal optical detection and sorting method and system for spherical lenses. Background Art

[0002] In the field of optical manufacturing, as a key component, the quality of spherical lenses directly affects the performance and stability of optical systems. Therefore, the quality inspection and sorting links in the production process of spherical lenses are crucial.

[0003] Currently, there are mainly two methods for quality inspection and sorting on the spherical lens production line, namely manual inspection and traditional machine vision inspection. In the manual inspection mode, workers need to check the lenses one by one with the naked eye and experience. Not only is the efficiency extremely low and it is difficult to meet the needs of large-scale production, but also the inspection results highly depend on the professional level and subjective judgment of workers, and it is very easy to miss inspections, misjudgments, etc., resulting in uneven product quality.

[0004] Although the existing machine vision detection technology has improved the detection efficiency to a certain extent, its functions are relatively single, and most of them only focus on the thickness detection of lenses. However, the quality problems of spherical lenses are not only reflected in the thickness deviation, but also the surface defects (such as scratches, stains, bubbles, etc.) and the deviation of curvature parameters will also seriously affect their optical performance. Since the traditional machine vision detection cannot comprehensively cover these key quality indicators, it is difficult to accurately evaluate the overall quality of the lenses, resulting in poor sorting effects, unable to effectively distinguish qualified products from unqualified products, and difficult to ensure the high-quality output of spherical lens products, restricting the further development of the industry.

[0005] In response to the above problems, there is currently no effective technical solution. Summary of the Invention

[0006] The purpose of the present application is to provide a multi-modal optical detection and sorting method and system for spherical lenses to comprehensively consider various quality factors of spherical lenses and generate an objective quality score to drive the sorting behavior.

[0007] In a first aspect, the present application provides a multi-modal optical detection and sorting method for spherical lenses, which is applied in a spherical lens sorting system. The method includes the following steps: S1. Obtain the appearance data information of the lens to be sorted, where the appearance data information includes curvature information, surface defect information, and coating uniformity information; S2. Place the appearance data information into a pre-trained quality scoring model to generate a quality score, so that the spherical lens sorting system sorts the lens to be sorted according to the quality score; The process of the quality scoring model generating the quality score includes: A1. Convert the curvature information, the surface defect information, and the coating uniformity information into curvature features, defect features, and uniformity features of the same dimension; A2. Use the uniformity feature as the query subject, the curvature feature and the defect feature as the objects to be queried, and establish a cross-modal correlation relationship by combining the attention mechanism to generate cross-modal features; A3. Fuse the curvature feature, defect feature, uniformity feature, and cross-modal feature to generate a fused feature; A4. Generate the quality score based on the fused feature according to the dual-task cooperation mechanism.

[0008] The multi-modal optical inspection and sorting method for spherical lenses of this application can overcome the limitations of traditional methods that only focus on a single index by integrating multi-modal information such as curvature, surface, and coating, and using a deep learning model for comprehensive evaluation. This method can comprehensively consider various quality factors of spherical lenses, generate an objective quality score, achieve a comprehensive and accurate evaluation of the quality of spherical lenses, and improve the efficiency and accuracy of sorting.

[0009] In the multi-modal optical inspection and sorting method for spherical lenses, step A1 is based on a heterogeneous modal encoder for feature conversion, and the heterogeneous modal encoder includes a curvature encoder, a surface defect encoder, and a coating encoder.

[0010] This design enables each encoder to be optimized according to the characteristics of its corresponding modal data, for example, by adopting a network structure or processing method suitable for this data type.

[0011] In the multi-modal optical inspection and sorting method for spherical lenses, the curvature encoder includes a 2D CNN layer and a Bi-GRU network. The Bi-GRU network is used to capture the forward and backward path dependencies under the action of the 2D CNN layer extracting local curvature mutation features from the curvature information to convert and output the curvature information into a 128-dimensional curvature feature.

[0012] In the multi-modal optical inspection and sorting method for spherical lenses, the surface defect encoder includes a ResNet-18 improvement module with the first layer of convolution replaced by a 7×7 kernel and an SPP layer. The SPP layer is used to adapt to defects of different sizes under the action of the ResNet-18 improvement module extracting the macroscopic morphology of defects to convert and output the surface defect information into a 128-dimensional defect feature.

[0013] The described spherical lens multi-modal optical detection and sorting method, wherein the coating encoder includes a graph convolutional network and three-layer GraphSAGE network. The three-layer GraphSAGE network is used to aggregate neighborhood information to convert and output the coating uniformity information into 128-dimensional uniformity features when the graph convolutional network constructs a topological graph from thickness point clouds and defines edge features as thickness gradients.

[0014] The described spherical lens multi-modal optical detection and sorting method, wherein step A3 includes: A31. Generate weight coefficients for the curvature features, defect features, and uniformity features based on the reliability of each modal data; A32. After weighting the curvature features, defect features, and uniformity features based on the weight coefficients, fuse them with cross-modal features to generate fused features.

[0015] The described spherical lens multi-modal optical detection and sorting method, wherein step A31 includes: A311. Predict the variance information for characterizing reliability corresponding to the curvature features, defect features, and uniformity features; A312. Generate the weight coefficients based on the inverse proportional weighting of the variance information.

[0016] The described spherical lens multi-modal optical detection and sorting method, wherein in step A4, the dual-task cooperation mechanism includes a main branch mechanism for the regression task and an auxiliary branch mechanism for the interpretation task that are executed synchronously. The main branch mechanism is used to output the quality score based on the fused features, and the auxiliary branch mechanism is used to output an interpretable modal importance score and a defect heat map based on the fused features and the curvature features, defect features, and uniformity features.

[0017] In a second aspect, the present application also provides a spherical lens sorting system, including a controller and an execution module. The controller is used to control the execution module to sort the lenses to be sorted according to the spherical lens multi-modal optical detection and sorting method provided in the first aspect.

[0018] The spherical lens sorting system of the present application controls the execution module to classify and sort the lenses to be sorted based on the quality score generated by the spherical lens multi-modal optical detection and sorting method provided in the first aspect. The scoring process integrates multi-modal information such as curvature, surface, and coating, and uses a deep learning model for comprehensive evaluation. It can comprehensively consider various quality factors of spherical lenses, generate an objective quality score, and achieve a comprehensive and accurate evaluation of the quality of spherical lenses, effectively improving the sorting accuracy of the execution module.

[0019] The spherical lens sorting system described above, wherein the spherical lens sorting system further includes a spectral component, the spectral component includes a spectroscopic ellipsometer, and the controller is further configured to collect the polarization light phase change of the lens to be sorted by using the spectral component, so as to obtain the coating uniformity information according to the polarization light phase change.

[0020] As can be seen from the above, the present application provides a method and system for multi-modal optical detection and sorting of spherical lenses. Among them, the multi-modal optical detection and sorting method of spherical lenses in the present application integrates multi-modal information such as curvature, surface, and coating, and uses a deep learning model for comprehensive evaluation. This method can overcome the limitations of traditional methods that only focus on a single index, can comprehensively consider various quality factors of spherical lenses, generate an objective quality score, and achieve a comprehensive and accurate evaluation of the quality of spherical lenses, improving the efficiency and accuracy of sorting. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a flowchart of the multi-modal optical detection and sorting method of spherical lenses provided by the embodiment of the present application.

[0022] Figure 2 It is a flowchart of generating a quality score by a quality scoring model.

[0023] Figure 3 It is a schematic diagram of the electrical control structure of the spherical lens sorting system provided by the embodiment of the present application.

[0024] Reference numerals: 201, controller; 202, execution module; 203, interference component; 204, camera component; 205, spectral component. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present application provided in the drawings below is not intended to limit the scope of the present application claimed, but only represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the scope of protection of the present application.

[0026] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0027] In a first aspect, please refer to Figure 1 and Figure 2 , some embodiments of the present application provide a multi-modal optical detection and sorting method for spherical lenses, which is applied in a spherical lens sorting system. The method includes the following steps: S1. Obtain the appearance data information of the lens to be sorted, where the appearance data information includes curvature information, surface defect information, and coating uniformity information; S2. Input the appearance data information into a pre-trained quality scoring model to generate a quality score, so that the spherical lens sorting system sorts the lens to be sorted according to the quality score; The process of the quality scoring model generating the quality score includes: A1. Convert the curvature information, surface defect information, and coating uniformity information into curvature features, defect features, and uniformity features of the same dimension; A2. Use the uniformity feature as the query subject, and the curvature feature and defect feature as the query objects, and establish a cross-modal correlation relationship by combining the attention mechanism to generate cross-modal features; A3. Fuse the curvature feature, defect feature, uniformity feature, and cross-modal feature to generate a fused feature; A4. Generate a quality score according to the fused feature.

[0028] Specifically, this method aims to solve the problem that traditional detection methods cannot comprehensively evaluate the quality of spherical lenses and achieve efficient and automated sorting. Among them, step S1 collects the curvature information, surface defect information, and coating uniformity information of the lens to be sorted, thereby collecting multi-modal data reflecting different quality dimensions of the lens. These information are different data types for evaluating the quality of the lens, overcoming the limitation of traditional methods that only focus on a single parameter. Step S2 inputs these multi-modal data into a pre-trained quality scoring model, and the model generates a quality score. The spherical lens sorting system sorts according to this score, realizing automated sorting based on comprehensive quality evaluation.

[0029] More specifically, the process of generating the quality score by the quality scoring model is broken down into several sub-steps. Among them, in step A1, a heterogeneous modality encoder can be used to process different types of appearance data information, and convert the curvature information, surface defect information, and coating uniformity information into curvature features, defect features, and uniformity features of the same dimension respectively. This conversion unifies different forms of raw data into a processable feature representation, enabling subsequent unified processing of data of different modalities. These features exist in the form of feature vectors; step A2 analyzes and captures the interactions between different quality aspects, improving the model's ability to identify complex quality problems; step A3 combines all the extracted individual features and the relationship features between them to form a fused feature with a comprehensive representation, laying a foundation for comprehensively evaluating the lens quality; step A4 converts the fused feature into a numerical value, which serves as a quantitative representation of the overall quality of the lens, directly guiding the sorting operation. The entire process realizes the comprehensive processing and evaluation of multi-modal quality information of spherical lenses, generates a quality score for sorting, and solves the problem that traditional methods cannot comprehensively evaluate the overall quality of lenses.

[0030] More specifically, after obtaining the curvature information, surface defect information, and coating uniformity information of the lens to be sorted, these appearance data information are placed into a pre-trained quality scoring model. The model first performs feature conversion on different types of appearance data information through a heterogeneous modality encoder, encoding the curvature information, surface defect information, and coating uniformity information into curvature features, defect features, and uniformity features of the same dimension respectively. Thus, data of different modalities are uniformly represented. Further, the model uses the uniformity feature as the query subject, and the curvature feature and defect feature as the query objects, and establishes a cross-modal association relationship by combining the attention mechanism to generate cross-modal features. This process captures the mutual influence between different quality dimensions. Subsequently, the curvature feature, defect feature, uniformity feature, and cross-modal feature are fused to generate a comprehensive fused feature. The fusion can be achieved in various ways, such as simple concatenation or weighted summation. This fused feature synthesizes the performance of the lens in each quality dimension and the relationship between these dimensions, forming a comprehensive description of the overall quality state of the lens. Finally, the model generates a quality score according to the fused feature. This quality score is a quantitative value representing the overall quality level of the lens. After receiving this quality score, the spherical lens sorting system sorts the lens to be sorted into different categories (such as qualified products, repairable non-conforming products, non-repairable non-conforming products, etc.) according to the preset sorting rules.

[0031] The spherical lens multi-modal optical detection and sorting method according to the embodiments of the present application integrates multi-modal information such as curvature, surface, and coating, and uses a deep learning model for comprehensive evaluation. This method can overcome the limitations of traditional methods that only focus on a single index, can comprehensively consider various quality factors of spherical lenses, generate objective quality scores, achieve a comprehensive and accurate evaluation of the quality of spherical lenses, and improve the efficiency and accuracy of sorting.

[0032] In some preferred embodiments, step A1 performs feature transformation based on a heterogeneous modal encoder, and the heterogeneous modal encoder includes a curvature encoder, a surface defect encoder, and a coating encoder.

[0033] Specifically, the heterogeneous modal encoder is composed of independent encoders designed for different modal data. Among them, the curvature encoder is specifically used to process curvature information, the surface defect encoder is specifically used to process surface defect information, and the coating encoder is specifically used to process coating uniformity information. This design enables each encoder to be optimized according to the characteristics of its corresponding modal data. For example, a network structure or processing method suitable for this data type can be adopted. By equipping each modal data with a dedicated encoder, this solution can more effectively extract representative features from heterogeneous raw data. Each encoder converts the data of its respective modality into feature vectors of the same dimension (curvature features, defect features, uniformity features). This conversion process retains the key information of each modality while unifying them into a comparable and fusible representation space, solving the problem of difficultly effectively extracting and unifying features when processing heterogeneous data in step A1, laying a foundation for subsequent cross-modal association and feature fusion, thereby improving the accuracy of quality scores.

[0034] In some preferred embodiments, the curvature encoder includes a 2D CNN layer and a Bi-GRU network. The Bi-GRU network is used to capture the forward and backward path dependencies under the action of the 2D CNN layer extracting local curvature mutation features according to curvature information to convert and output curvature features of 128 dimensions.

[0035] Specifically, the 2D CNN layer is responsible for processing the input curvature information, and extracts local features in the data through convolution operations, especially the mutation features of curvature. Preferably, the convolution kernel width of the 2D CNN layer is 5 and the stride is 2. It can scan the curvature data with a specific receptive field and sampling method to identify local anomalies or changes. The Bi-GRU network receives the output of the 2D CNN layer and further processes these local features. The Bi-GRU is a bidirectional recurrent neural network that can capture the forward and backward dependencies in the data sequence. The curvature information has a sequential nature in space or time, and the Bi-GRU network can understand this sequential structure, associate local features at different positions, and thus obtain a more comprehensive representation of curvature errors.

[0036] More specifically, in this embodiment, first, the curvature information is input into the 2D CNN layer. The 2D CNN layer scans and processes the curvature information to extract features reflecting the local changes in curvature, such as sudden elevation or decrease in curvature and other mutation features, thereby achieving effective identification of local curvature mutation features. Subsequently, the output of the 2D CNN layer is input into the Bi-GRU network. The Bi-GRU network can process sequence information in both forward and backward directions simultaneously, capturing the dependencies between curvature information at different positions, such as the continuity or periodicity of the curvature change trend. The curvature encoder combines the local features extracted by the 2D CNN and the sequence dependencies captured by the Bi-GRU to convert the original curvature information into a 128-dimensional curvature feature vector, achieving effective encoding of the curvature information. This feature vector comprehensively reflects the local characteristics and overall structure of the curvature error, providing accurate input for the subsequent quality scoring model and improving the accuracy of quality scoring.

[0037] In some preferred embodiments, the surface defect encoder includes a ResNet-18 improvement module with the first layer convolution replaced by a 7×7 kernel and an SPP layer. The SPP layer is used to adapt to defects of different sizes under the action of the ResNet-18 improvement module extracting the macroscopic morphology of the defects to convert and output the surface defect information into a 128-dimensional defect feature.

[0038] Specifically, the ResNet-18 improvement module is obtained by replacing the 3×3 convolution in the first layer of ResNet-18 with a 7×7 convolution. Such a large-size convolution kernel can capture a larger range of initial features in the input image. The SPP layer generates a fixed-length feature vector by pooling feature maps of different scales. Finally, the features processed by the SPP layer are converted into a 128-dimensional defect feature vector and output. By combining the improved ResNet-18 module to extract the macroscopic morphology and the SPP layer to adapt to different sizes, the surface defect encoder can more comprehensively and accurately convert the original surface defect information into a defect feature vector, thereby improving the accuracy of the subsequent quality scoring model.

[0039] In some preferred embodiments, the coating encoder includes a graph convolutional network and a three-layer GraphSAGE network. The three-layer GraphSAGE network is used to aggregate neighborhood information to convert and output the coating uniformity information into a 128-dimensional uniformity feature under the condition that the graph convolutional network constructs a topological graph from the thickness point cloud and defines the edge feature as the thickness gradient.

[0040] Specifically, the coating uniformity information usually exists in the form of thickness point clouds, which has a spatial topological structure. The graph convolutional network is used to convert the thickness point cloud data into graph-structured data, where the points serve as nodes, the connections between the points form edges, and the thickness difference or gradient between the points is defined as the edge feature. A three-layer GraphSAGE network is used to process the constructed graph-structured data, and the representation of each node is learned by aggregating the neighborhood information of each node layer by layer. The GraphSAGE network captures the local and broader features of the coating uniformity through the aggregation operation. Finally, the GraphSAGE network outputs a feature vector with a fixed dimension, that is, the uniformity feature. The graph convolutional network is responsible for converting the original point cloud data into a graph structure suitable for processing by the GraphSAGE network, and the GraphSAGE network is responsible for feature learning and aggregation on the graph structure, converting the variable-length graph data into a fixed-length feature vector.

[0041] More specifically, the graph convolutional network is used to process the thickness point cloud, construct the points in the point cloud as the nodes of the graph, establish edges according to the spatial relationship or distance between the points, and at the same time calculate the thickness gradient between adjacent points as the feature of the edge. Thus, the original thickness point cloud data is converted into a topological graph with node and edge features. Then, the constructed three-layer GraphSAGE network processes this topological graph. The GraphSAGE network aggregates information from the neighborhood nodes of each node by defining an aggregation function and updates the representation of the node by combining the features of the node itself. The first layer of the GraphSAGE network aggregates the information of the direct neighbors, the second layer aggregates the information of the two-hop neighbors, and the third layer aggregates the information of the three-hop neighbors, thereby gradually expanding the receptive field and capturing more comprehensive local and global structure information. After three-layer aggregation, each node obtains a feature representation containing its neighborhood information. Finally, the feature representations of these nodes are integrated into a vector with a fixed dimension as the output uniformity feature, and its dimension is preferably set to 128. This processing method effectively utilizes the spatial structure and thickness gradient information of the point cloud data and converts the variable-length graph-structured data into a fixed-length feature vector, providing a standardized input for subsequent feature fusion.

[0042] In some preferred embodiments, step A2 establishes a cross-modal association relationship based on the multi-head attention mechanism, where the multi-head attention mechanism satisfies: (1) where head j is the calculation result of the j-th attention head, that is, the j-th cross-modal association relationship, Q is the query subject, and K, V are the objects to be queried, that is, the concatenated vectors of the curvature feature and the defect feature, is the learnable projection matrix of the j-th attention head with respect to Q, is the learnable projection matrix of the j-th attention head with respect to K, is the learnable projection matrix of the j-th attention head with respect to V, and d k is the projection dimension of K, T is the transpose matrix mark. In the embodiments of the present application, the multi-head attention mechanism is a four-head attention mechanism.

[0043] Specifically, this technical solution improves the process of establishing cross-modal association relationships by introducing a multi-head attention mechanism. When establishing the association between the uniformity feature, the curvature feature, and the defect feature, the multi-head attention mechanism is adopted, and the query subject Q and the queried objects K, V can be projected into multiple different low-dimensional subspaces through different linear transformations respectively. The attention calculation is independently and parallelly performed in each subspace to obtain the outputs head j of multiple attention heads. Each attention head can learn and focus on different association patterns or information aspects in the input features. For example, one head may focus on the association between uniformity and the magnitude of curvature error, and another head may focus on the association between uniformity and defect distribution. This technical solution specifically adopts a four-head attention mechanism, indicating that four independent attention calculations are performed in parallel. Finally, the outputs of these four attention heads are concatenated and then passed through a linear transformation to generate the final cross-modal feature.

[0044] More specifically, the above processing method aims to solve the problem that a single attention mechanism may not be able to fully capture cross-modal association information in different dimensions or different aspects, which limits the expression ability of cross-modal features. By concatenating the outputs of multiple attention heads, a more comprehensive and richer cross-modal association representation can be obtained. The features after attention interaction already contain semantic associations between modalities, enabling the generated cross-modal features to more precisely reflect the complex interactions between the uniformity feature, the curvature feature, and the defect feature, providing more discriminative information for the generation of subsequent fusion features, and thus improving the accuracy of quality scoring.

[0045] In some preferred embodiments, the multi-head attention mechanism has learnable gating weights, which satisfy: (2) where g is the gating value, satisfying g ∈ [0, 1], and the final attention output of the multi-head attention mechanism is g·concat(head1,...,head n ), where n is the total number of attention heads, σ is the activation function, and MLP is the multi-layer perceptron.

[0046] Specifically, the gating weight processes the concatenation vector of the query subject and the queried object through a multi-layer perceptron. The output of the multi-layer perceptron is mapped to a range of 0 to 1 through an activation function, thereby generating a gating value. The gating value is used to control the amount of information flow. Specifically, the final output of the multi-head attention mechanism is obtained by multiplying the gating value with the standard output of the multi-head attention mechanism. When the gating value is close to 0, the amount of information flow decreases; when the gating value is close to 1, the amount of information flow increases. Thus, the amount of information flow is dynamically adjusted according to the input features.

[0047] More specifically, in this embodiment, after the multi-head attention mechanism calculates the results of each attention head and splices them, they are not directly used as the final cross-modal features. The gating value is multiplied by the spliced ​​output of the multi-head attention mechanism. Thus, the gating value dynamically adjusts the intensity of the attention output according to the specific content of the query subject and the queried object. When the input feature combination indicates that the cross-modal association is valid, the gating value is close to 1, and the information is almost completely passed. When the input feature combination indicates that the cross-modal association may contain noise, the gating value is close to 0, and the information is greatly attenuated. This enables the generated cross-modal features to better reflect the effective association between modalities, improves the quality of feature fusion, and thus improves the accuracy of the quality score.

[0048] In some preferred embodiments, step A3 comprises: A31. Generate weight coefficients for curvature features, defect features, and uniformity features based on the reliability of each modal data; A32. After weighting the curvature features, defect features, and uniformity features based on the weight coefficients, they are fused with the cross-modal features to generate fused features.

[0049] Specifically, considering the differences in the acquisition methods of different modal data, environmental factors, or sensor performance, their reliability may vary. For example, surface defect detection may be easily affected by changes in lighting, while curvature measurement may be more sensitive to vibrations. Simply averaging or splicing these features may not fully utilize the data with high reliability and may be negatively affected by the data with low reliability. Therefore, a weighted mechanism based on data reliability is introduced in the fusion process of this solution. First, according to the reliability of each modal data (corresponding to curvature features, defect features, and uniformity features), the corresponding weight coefficients are calculated and generated. The weight coefficient corresponding to the data with high reliability is larger, and the weight coefficient corresponding to the data with low reliability is smaller. Then, these weight coefficients are used to fuse the curvature features, defect features, uniformity features, and cross-modal features. Through this weighted method, the fused features can focus more on the information from the modalities with higher reliability, thus more accurately characterizing the overall quality of the lens. The generated fused features are then used to generate the final quality score based on the dual-task cooperation mechanism to guide the sorting system to perform sorting operations. Thus, by considering data reliability for weighted fusion, the quality of the fused features is improved, and further the accuracy of the quality score and the reliability of sorting are enhanced.

[0050] In some preferred embodiments, step A31 includes: A311. Predict the variance information used to characterize the reliability corresponding to the curvature features, defect features, and uniformity features; A312. Generate the weight coefficients based on the inverse proportional weighting of the variance information.

[0051] Specifically, the variance characterizes the reliability of the data. The smaller the variance, the higher the reliability, and the larger its weight coefficient. Thus, the reliable data affects the fusion result and improves the accuracy of the fused features. This processing method provides a mechanism for dynamically adjusting the weights of each modal feature based on the uncertainty of the data itself, making the generation of the weight coefficients data-driven, overcoming the problems that may be brought by simple averaging or fixed weights, and improving the effectiveness of multi-modal feature fusion.

[0052] In some preferred embodiments, in step A311, the variance is generated based on a preset feedforward neural network model, which satisfies: (3) Wherein, is the variance information of the i-th modal data, and MLP i is the feedforward neural network model adopted by the i-th modal data, and h i is the i-th modal data.

[0053] Specifically, the above processing method specifies using a preset feedforward neural network model to generate the variance information of the i-th modal data. For each type of modal data (curvature feature, defect feature, uniformity feature), there is a corresponding feedforward neural network model. The feedforward neural network model MLP i is introduced, making the prediction of variance no longer an arbitrary or undefined process, but achieved through a learnable model. This model can learn how to generate variance information reflecting its reliability based on the input feature data during the training process.

[0054] More specifically, the above processing method provides a structured, model-based variance prediction method. By training the feedforward neural network model, it can learn from the input modal features and output the variance information characterizing its reliability. Thus, the variance prediction process becomes controllable and stable, providing reliable input for the subsequent weight calculation based on variance information, thereby improving the accuracy of feature fusion and ultimately enhancing the generation effect of the quality score.

[0055] In some preferred embodiments, in step A312, the weight coefficient is calculated and generated according to the following formula: (4) where r is a hyperparameter, and w i is the weight coefficient of the i-th modal data.

[0056] Specifically, this method receives the variance information used to characterize the reliability of each modal data as input. In the formula the variance is converted into a numerical value, which has a positive correlation with reliability, that is, the smaller the variance (the higher the reliability), the larger this numerical value. The hyperparameter r is set according to the usage requirements, and it is used to adjust the influence degree of variance on the final weight. The larger the r value, the more significant the influence of variance on the weight. is to sum the exponential terms of all modalities, used to normalize the calculated weights to ensure that the sum of the weight coefficients of all modalities is 1. Thus, formula (4) converts the variance information representing unreliability into the weight coefficient for feature fusion. Data with high reliability obtains a higher weight, and data with low reliability obtains a lower weight, providing a quantitative basis based on data reliability for the subsequent feature weighted fusion and improving the effectiveness of the fused features.

[0057] In some preferred embodiments, in step A32, the fused feature satisfies: (5) where y is the fused feature, MLP is the multi-layer perceptron, and concat(h1, h2, h3) is the cross-modal feature.

[0058] Specifically, the fused feature y consists of two parts. The first part is , which represents the weighted sum of the curvature feature, defect feature, and uniformity feature (h_i), and the weight coefficient w i reflects the reliability of the corresponding modal data. The second part is MLP(concat(h1, h2, h3)), which represents the cross-modal feature generated in step A2. This fusion method combines the weighted original features based on reliability and the processing of cross-modal features to generate a fused feature. This processing method utilizes the cross-modal associations generated in step A2, combines the reliability of each modality itself, and dynamically balances the contributions of attention features and original features for feature fusion. That is, it comprehensively combines the method of establishing associations between modalities at the semantic level and realizing quantitative weighting at the feature level to construct the fused feature, jointly achieving high-precision, interpretable, and anti-interference multi-modal data fusion. The generated fused feature can more comprehensively and accurately represent the appearance data information of the lens to be sorted, overcoming the limitations of simple weighted summation in capturing complex associations between different modal information, and providing a basis for the accuracy of the subsequent quality scoring model.

[0059] In some preferred embodiments, step A4 generates a quality score based on a dual-task cooperation mechanism. Among them, the dual-task cooperation mechanism includes a main branch mechanism for the regression task and an auxiliary branch mechanism for the interpretation task that are executed synchronously. The main branch mechanism is used to output a quality score according to the fused feature, and the auxiliary branch mechanism is used to output an interpretable modal importance score and a defect heat map according to the fused feature, curvature feature, defect feature, and uniformity feature.

[0060] Specifically, through this dual-task cooperation mechanism, an interpretable index (modal importance score and defect heat map) is generated while outputting the product quality score (main branch), realizing the synchronous optimization of prediction accuracy and decision credibility.

[0061] More specifically, the design principle of this step is to achieve the unity of accuracy and interpretability through dual-branch cooperative optimization. Its core is to synchronously complete quality prediction (regression task) and decision basis analysis (interpretation task) through shared feature representation. Among them, the main branch mechanism uses a 3-layer MLP to output a quality score (0-100) according to the fused feature, and uses the Huber loss function to resist outliers. The Huber loss function is used to linearly penalize abnormal samples that exceed the threshold δ (such as outliers caused by sensor failures), avoiding gradient explosion.

[0062] Based on the foregoing, it can be seen that the quality scoring model is a multi-modal fusion deep learning model based on the attention mechanism, and its core architecture combines heterogeneous modal encoders, cross-modal association modules, and dynamic feature fusion mechanisms. The quality scoring model encodes multi-modal data such as curvature, defects, and coating uniformity into 128-dimensional feature vectors respectively through a curvature encoder (2D CNN + Bi-GRU), a surface defect encoder (improved ResNet-18 + SPP), and a coating encoder (GCN + GraphSAGE). Subsequently, the quality scoring model uses a four-head attention mechanism to establish cross-modal associations between the uniformity feature and the curvature feature and the defect feature, generates interaction features, and introduces a gating weight (dynamically adjusted by MLP) to optimize the information flow. Finally, the quality scoring model realizes a comprehensive evaluation of the quality of spherical lenses through weighted fusion based on variance prediction and a dual-task cooperation mechanism (regression scoring and interpretability analysis), with both high precision and interpretability.

[0063] It should be noted that the quality scoring model can be obtained through a training strategy that combines supervised learning and multi-task cooperation training. Its training process includes data preparation, loss function setting, dynamic optimization, and an end-to-end training process. These processes are common training processes for deep learning models and will not be elaborated here.

[0064] It should be noted that the scoring value output by the quality scoring model is from 0 to 100 points. Among them, 90 - 100 points, 70 - 89 points, and 0 - 69 points correspond to the qualified products, repairable non-conforming products, and non-repairable non-conforming products mentioned above respectively.

[0065] In a second aspect, please refer to Figure 3 , some embodiments of the present application further provide a spherical lens sorting system, including a controller 201 and an execution module 202. The controller 201 is used to control the execution module 202 to sort the lenses to be sorted according to the spherical lens multi-modal optical detection and sorting method provided in the first aspect.

[0066] Specifically, the controller 201 receives and processes the quality scores generated by the spherical lens multi-modal optical inspection and sorting method provided by the first aspect. The controller 201 can also be configured to execute the spherical lens multi-modal optical inspection and sorting method provided by the first aspect to generate lens quality scores, and then determine the destination of the lenses to be sorted according to the quality scores. The execution module 202, according to the instructions of the controller 201, performs physical operations such as grasping, moving, or pushing the lenses, and places them into the corresponding sorting areas or containers. The controller 201 is the core of the system, responsible for converting the abstract quality scores into specific physical control signals. Its technical contribution lies in realizing the connection between the method logic and physical execution. The execution module 202 is the execution mechanism of the system, responsible for completing the actual physical sorting actions. Its technical contribution lies in providing a physical means to achieve automated sorting. Through the collaborative work of the controller 201 and the execution module 202, the system converts the scoring results of the inspection and sorting method into actual physical sorting behaviors, realizing the automated inspection and sorting of spherical lenses.

[0067] More specifically, the controller 201 can control the execution module 202 to classify and sort the lenses to be sorted according to a preset score range, where, for example, they are classified into excellent products, good excellent products, qualified excellent products, and unqualified excellent products.

[0068] The spherical lens sorting system according to the embodiment of the present application controls the execution module 202 to classify and sort the lenses to be sorted based on the quality scores generated by the spherical lens multi-modal optical inspection and sorting method provided by the first aspect. The scoring process integrates multi-modal information such as curvature, surface, and coating, and uses a deep learning model for comprehensive evaluation. It can comprehensively consider various quality factors of spherical lenses, generate objective quality scores, achieve a comprehensive and accurate evaluation of the quality of spherical lenses, and effectively improve the sorting accuracy of the execution module 202.

[0069] In some preferred embodiments, the spherical lens sorting system further includes an interference component 203, and the interference component 203 is a laser interferometer or a white light interferometer; The controller 201 is used to use the interference component 203 to scan the surface height data of the lenses to be sorted to obtain curvature information.

[0070] Specifically, the interference component 203 specifically adopts a laser interferometer or a white light interferometer, and these devices can measure the height data of the optical surface in a non-contact manner. The controller 201 receives the surface height data obtained by the interference component 203. The controller 201 performs calculations based on the received surface height data, thereby obtaining the curvature information of the lenses to be sorted. The controller 201 uses this component to scan and perform calculations, completing the conversion from the original measurement data to curvature information. By adding the interference component 203 and its functions, the system obtains the ability to obtain curvature information.

[0071] In some preferred embodiments, the spherical lens sorting system further includes a camera assembly 204, and the controller 201 is further configured to use the camera assembly 204 to collect images of the lenses to be sorted, so as to obtain surface defect information based on the images.

[0072] Specifically, the camera assembly 204 may be an industrial camera. The controller 201 is electrically connected to the camera assembly 204 and controls the camera assembly 204 to perform image acquisition operations. The controller 201 receives the collected image data, processes and analyzes the image data, and identifies and extracts information related to surface defects therefrom. Thus, necessary surface defect data is provided for subsequent sorting methods.

[0073] In some preferred embodiments, the spherical lens sorting system further includes a spectroscopic assembly 205. The spectroscopic assembly 205 includes a spectroscopic ellipsometer. The controller 201 is further configured to use the spectroscopic assembly 205 to collect the polarization light phase changes of the lenses to be sorted, so as to obtain coating uniformity information based on the polarization light phase changes.

[0074] Specifically, the controller 201 is electrically connected to the spectroscopic assembly 205 and is configured to control the spectroscopic ellipsometer to perform measurement tasks. The spectroscopic ellipsometer is configured to perform optical measurements on the lenses to be sorted and collect polarization light phase change data caused by the coatings on the lens surfaces. The controller 201 receives the polarization light phase change data output by the spectroscopic ellipsometer and executes data processing algorithms to extract information characterizing the coating uniformity from these data. Thus, the spectroscopic assembly 205 provides a physical measurement means, and the controller 201 provides data acquisition and information processing capabilities, jointly realizing the acquisition of coating uniformity information, which provides necessary data input for subsequent quality scoring and sorting.

[0075] In addition, the units described as separation components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0076] Furthermore, in each embodiment of the present application, the various functional modules may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.

[0077] In this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.

[0078] The above are only embodiments of the present application and are not intended to limit the protection scope of the present application. For those skilled in the art, various modifications and variations can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A multimodal optical detection and sorting method for spherical lenses, which is applied in a spherical lens sorting system, is characterized in that The method includes the following steps: S1. Obtain the appearance data information of the lens to be sorted, where the appearance data information includes curvature information, surface defect information, and coating uniformity information; S2. Input the appearance data information into a pre-trained quality scoring model to generate a quality score, so that the spherical lens sorting system sorts the lens to be sorted according to the quality score; The process of the quality scoring model generating the quality score includes: A1. Convert the curvature information, the surface defect information, and the coating uniformity information into curvature features, defect features, and uniformity features of the same dimension; A2. Use the uniformity feature as the query subject, and the curvature feature and the defect feature as the objects to be queried, and establish a cross-modal correlation relationship by combining the attention mechanism to generate cross-modal features; A3. Fuse the curvature feature, defect feature, uniformity feature, and cross-modal feature to generate a fused feature; A4. Generate the quality score based on the fused feature according to the dual-task cooperation mechanism.

2. The spherical lens multi-modal optical detection and sorting method according to claim 1, wherein Step A1 is based on a heterogeneous modal encoder for feature conversion, and the heterogeneous modal encoder includes a curvature encoder, a surface defect encoder, and a coating encoder.

3. The spherical lens multi-modal optical detection and sorting method according to claim 2, characterized in that The curvature encoder includes a 2D CNN layer and a Bi-GRU network. The Bi-GRU network is used to capture the forward and backward path dependencies under the action of the 2D CNN layer extracting local curvature mutation features according to the curvature information, so as to convert and output the curvature information into curvature features of 128 dimensions.

4. The spherical lens multi-modal optical detection and sorting method according to claim 2, characterized in that, The surface defect encoder includes a ResNet-18 improvement module with the first layer convolution replaced by a 7×7 kernel and an SPP layer. The SPP layer is used to adapt to defects of different sizes under the action of the ResNet-18 improvement module extracting the macroscopic morphology of the defects, so as to convert and output the surface defect information into defect features of 128 dimensions.

5. The spherical lens multi-modal optical detection and sorting method according to claim 2, wherein The coating encoder includes a graph convolutional network and a three-layer GraphSAGE network. The three-layer GraphSAGE network is used to aggregate neighborhood information when the graph convolutional network constructs a topological graph from thickness point clouds and defines edge features as thickness gradients, so as to convert and output the coating uniformity information into uniformity features of 128 dimensions.

6. The spherical lens multi-modal optical detection and sorting method according to claim 1, wherein, Step A3 includes: A31. Generate weight coefficients for the curvature feature, defect feature, and uniformity feature based on the reliability of each modal data; A32. After weighting the curvature feature, defect feature, and uniformity feature based on the weight coefficients, fuse them with the cross-modal feature to generate a fused feature.

7. The spherical lens multi-modal optical detection and sorting method according to claim 6, characterized in that Step A31 includes: A311. Predict the variance information for characterizing reliability corresponding to the curvature feature, defect feature, and uniformity feature; A312. Generate the weight coefficients based on the inverse proportional weighting of the variance information.

8. The multi-modal optical inspection and sorting method for spherical lenses according to claim 1, wherein In step A4, the dual-task cooperation mechanism includes a main branch mechanism for the regression task and an auxiliary branch mechanism for the interpretation task that are executed synchronously. The main branch mechanism is used to output the quality score according to the fused feature, and the auxiliary branch mechanism is used to output an interpretable modal importance score and a defect heat map according to the fused feature, the curvature feature, the defect feature, and the uniformity feature.

9. A spherical lens sorting system, characterized in that, It includes a controller and an execution module. The controller is used to control the execution module to sort the lenses to be sorted according to the spherical lens multi-modal optical detection and sorting method described in any one of claims 1-8.

10. The spherical lens sorting system according to claim 9, wherein, The spherical lens sorting system further includes a spectral component. The spectral component includes a spectroscopic ellipsometer. The controller is further used to collect the polarization light phase change of the lens to be sorted by using the spectral component, so as to obtain the coating uniformity information according to the polarization light phase change.

Citation Information

Patent Citations

  • Multi-dimensional multi-modal group biological feature recognition system and method

    CN112507781A

  • Defect detection method and system for microscope lens production

    CN118348009A

  • Lens production quality management system based on optical characteristics

    CN119047695A

  • Lens-evaluating method and lens-evaluating apparatus

    US20030048436A1

  • Three-dimensional target detection method based on multimodal fusion and depth attention mechanism

    US20250037299A1