Product monitoring method and system based on machine vision and convolutional neural network
By fusing two-dimensional images and three-dimensional point cloud data through a multi-task convolutional neural network, decision thresholds and uncertainty quantification are dynamically generated, solving the problems of inspection efficiency and reliability in automotive bearing production. This enables adaptive production cycle time and reliability assessment, improving inspection efficiency and decision credibility.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-13
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies in the production of precision automotive bearing components suffer from insufficient testing efficiency and a lack of quantitative assessment of the reliability of unknown defects. In particular, it is difficult to achieve a dynamic balance between system efficiency and reliability under varying production line throughput rates.
A multi-task convolutional neural network is used to combine two-dimensional images and three-dimensional point cloud data. The fused features are extracted through a shared feature encoder to generate normal scores and depth feature vectors for automobile bearings. The decision threshold is dynamically generated by combining the real-time throughput rate of the production line and historical quality statistics. The uncertainty is quantified by Monte Carlo discarding operations to generate the final monitoring results.
It improves detection efficiency and feature reusability, constructs an intelligent diversion mechanism that adapts to production cycle time, resolves the contradiction between fixed computing resources and changing production needs, establishes an objective evaluation system for model decision reliability, and outputs a comprehensive report that includes specific defect diagnosis and detection confidence status.
Smart Images

Figure CN121837249A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent manufacturing, in particular to a product monitoring method and system based on machine vision and convolutional neural network. BACKGROUND
[0002] In the production of precision parts of automobile bearings, the dimensional accuracy directly affects the reliability and safety of the product. Although the traditional contact type measurement method has high precision, it has limitations such as low efficiency and inability to detect all. In recent years, the machine vision technology based on convolutional neural network has brought innovation to the field. The existing technical solution uses an industrial camera to collect a two-dimensional image of the product, and realizes automatic classification and recognition of defects through training of a deep convolutional neural network. The powerful feature extraction capability of CNN improves the automation level and consistency of detection, and has become the mainstream technical direction in the field.
[0003] In response to complex industrial scenarios with high precision, the existing technology still faces challenges. In terms of real-time performance, the complex neural network model designed to ensure accuracy takes a long time to infer, making it difficult to adapt to changing throughput rates in continuous production. When the production line speeds up, it can cause processing bottlenecks, and when it slows down, it can lead to idle computing power. In terms of decision reliability, conventional CNN models lack the ability to quantitatively evaluate the degree of confidence in their own decisions. For unknown defects or boundary samples outside the training data, they may still output incorrect judgments with high confidence, which poses a risk in precision manufacturing that emphasizes zero defects. The existing technology fails to effectively combine production rhythm information with model uncertainty evaluation to achieve a dynamic balance between system efficiency and reliability. SUMMARY
[0004] In view of the above existing problems, the present application is proposed.
[0005] Therefore, the present application provides a product monitoring method based on machine vision and convolutional neural network, which solves the problems of insufficient detection efficiency and lack of quantitative evaluation of unknown defect judgment reliability in the existing technology under changing production line throughput rates.
[0006] To solve the above technical problems, the present application provides the following technical solutions: In a first aspect, the present application provides a product monitoring method based on machine vision and convolutional neural network, which includes image acquisition and preprocessing of automobile bearings to obtain preprocessed two-dimensional image features and three-dimensional point cloud data of automobile bearings; The preprocessed two-dimensional image features and three-dimensional point cloud data of automobile bearings are input into a pre-trained multi-task convolutional neural network, and the shared feature encoder extracts the fused features. The first output head generates a normal score for the automobile bearing, and the second output head generates a preliminary defect judgment and a deep feature vector of high-dimensional features. According to the current production line real-time throughput rate and historical quality statistical information, a first-level decision threshold is generated, the automobile bearing normal score is compared with the first-level decision threshold, and the qualified and end process and deep check state are output respectively; Based on the deep feature vector, the second output head with the Monte Carlo discard operation enabled is forward propagated to obtain a cognitive uncertainty quantization value; According to the cognitive uncertainty quantization value and the real-time constraint, a final uncertainty decision threshold is determined, and a comprehensive report of the final monitoring result of the automobile bearing is generated by combining the preliminary diagnosis opinion of the automobile bearing.
[0007] As a preferred scheme of the product monitoring method based on machine vision and convolutional neural network, the method comprises the following steps: An industrial area array camera is used to collect the automobile bearing to obtain a two-dimensional gray image of the automobile bearing, and then Canny edge detection is performed to extract the preprocessed two-dimensional image features of the automobile bearing; A line laser three-dimensional scanner is used to scan the same automobile bearing to obtain three-dimensional point cloud of the automobile bearing, and a statistical-based outlier removal operation is performed to obtain noise-removed three-dimensional point cloud of the automobile bearing; The noise-removed three-dimensional point cloud of the automobile bearing is analyzed, and the orientation information of the surface of each point in the point cloud is estimated, and the preprocessed three-dimensional point cloud data of the automobile bearing is obtained.
[0008] As a preferred scheme of the product monitoring method based on machine vision and convolutional neural network, the method comprises the following steps: Based on the preprocessed two-dimensional image features of the automobile bearing and the preprocessed three-dimensional point cloud data of the automobile bearing, a training set is constructed, and the multi-task convolutional neural network is trained through the training set to obtain the pre-trained multi-task convolutional neural network; The preprocessed two-dimensional image features of the automobile bearing are input into the two-dimensional convolution branch of the pre-trained multi-task convolutional neural network to generate a two-dimensional feature map; The preprocessed three-dimensional point cloud data of the automobile bearing are input into the three-dimensional convolution branch of the pre-trained multi-task convolutional neural network to generate a three-dimensional feature vector.
[0009] As a preferred scheme of the product monitoring method based on machine vision and convolutional neural network, wherein: the fusion features are extracted by the shared feature encoder, the first output head generates the automobile bearing normal score, and the second output head generates the defect preliminary judgment and the deep feature vector of the high-dimensional feature. The two-dimensional feature map and the three-dimensional feature vector are spliced in the fusion layer of the shared feature encoder to obtain the fusion features. The fusion features are input into the first output head of the pre-trained multi-task convolutional neural network, and the first output head is configured to perform a simplified binary classification task, extract global patterns for judging whether the automobile bearing is qualified as a whole from the fusion features, and output the automobile bearing normal score. The fusion features are input into the second output head of the pre-trained multi-task convolutional neural network, and the second output head is configured to perform fine-grained multi-task learning, analyze specific semantic information for defect classification and geometric features for size regression from the fusion features, and output the deep feature vector of the high-dimensional feature of the defect preliminary judgment.
[0010] As a preferred scheme of the product monitoring method based on machine vision and convolutional neural network, wherein: the first-level decision threshold is generated according to the current production line real-time throughput rate and historical quality statistical information, including the following steps: The current production line real-time throughput rate is read from the data interface of the production line controller. The historical database is queried to obtain the historical quality statistical information of the automobile bearing normal score recorded in the historical batch matched with the current time period. Based on the current production line real-time throughput rate and the historical quality statistical information of the automobile bearing normal score, the first-level decision threshold is calculated by a pre-defined decision function.
[0011] As a preferred scheme of the product monitoring method based on machine vision and convolutional neural network, wherein: the automobile bearing normal score is compared with the first-level decision threshold, and the qualified and the end process and the deep review state are output respectively, including the following steps: The automobile bearing normal score is compared with the first-level decision threshold.
[0012] When the automobile bearing normal score is higher than the first-level decision threshold, the automobile bearing is determined to be qualified and the monitoring process is ended; When the automobile bearing normal score is not higher than the first-level decision threshold, the state of the monitoring process is marked as the deep review state.
[0013] As a preferred scheme of the product monitoring method based on machine vision and convolutional neural network, wherein: based on the deep feature vector, the second output head enabled with the Monte Carlo dropout operation is forward propagated to obtain a cognitive uncertainty quantization value, including the following steps: The deep feature vector is input into the second output head enabled with the Monte Carlo dropout operation, and the second output head enabled with the Monte Carlo dropout operation is forward propagated to obtain a plurality of sets of probability distributions of the second output head output generated by a plurality of times of forward propagation; The plurality of sets of probability distributions of the second output head output are statistically calculated to obtain the cognitive uncertainty quantization value.
[0014] As a preferred scheme of the product monitoring method based on machine vision and convolutional neural network, wherein: according to the cognitive uncertainty quantization value and the real-time constraint, a final uncertainty decision threshold is determined, including the following steps: The remaining allowed decision time is calculated as the real-time constraint based on the preset total processing time limit and the current process consumed time; According to the cognitive uncertainty quantization value and the real-time constraint, the final uncertainty decision threshold is dynamically determined through a preset trade-off strategy.
[0015] As a preferred scheme of the product monitoring method based on machine vision and convolutional neural network, wherein: a comprehensive report of the final monitoring result of the automobile bearing is generated by combining the preliminary diagnosis opinion of the automobile bearing, including the following steps: The cognitive uncertainty quantization value is compared with the final uncertainty decision threshold; When the cognitive uncertainty quantization value is lower than the final uncertainty decision threshold, the preliminary diagnosis opinion of the automobile bearing is adopted as the final diagnosis opinion of the automobile bearing; When the cognitive uncertainty quantization value is not lower than the final uncertainty decision threshold, the final diagnosis opinion of the automobile bearing is marked as a high-uncertainty state requiring manual review; The final diagnosis opinion of the automobile bearing, the cognitive uncertainty quantization value, and the specific information in the preliminary diagnosis opinion of the automobile bearing are integrated to generate a comprehensive report of the final monitoring result of the automobile bearing.
[0016] In a second aspect, the present application provides a product monitoring system based on machine vision and convolutional neural network, including a data processing module, which performs image acquisition and preprocessing on automobile bearings to obtain preprocessed automobile bearing two-dimensional image features and automobile bearing three-dimensional point cloud data; The perception module inputs the preprocessed two-dimensional image features of the automobile bearing and the preprocessed three-dimensional point cloud data of the automobile bearing into a pre-trained multi-task convolutional neural network, extracts fusion features through a shared feature encoder, a first output head generates an automobile bearing normal score, and a second output head generates a defect preliminary judgment and a deep feature vector of a high-dimensional feature; The decision module generates a first-level decision threshold according to the current production line real-time throughput rate and historical quality statistical information, compares the automobile bearing normal score with the first-level decision threshold, and respectively outputs qualified and ends the process and a deep verification state; The evaluation module performs forward propagation on the second output head with the Monte Carlo dropout operation enabled based on the deep feature vector to obtain a cognitive uncertainty quantization value. The decision module determines a final uncertainty decision threshold according to the cognitive uncertainty quantization value and real-time constraints, and generates a comprehensive report of the final monitoring result of the automobile bearing by combining the preliminary diagnosis opinion of the automobile bearing.
[0017] The present application has the following advantages: through multi-modal data fusion and multi-task collaborative processing architecture, a shared feature encoder is used to synchronously generate an automobile bearing normal score and a deep feature vector, detection efficiency and feature reusability are doubled, a first-level decision threshold is dynamically generated according to the production line real-time throughput rate and historical quality statistical information, an intelligent shunting mechanism with adaptive production rhythm is constructed, the contradiction between fixed computing resources and changing production needs is solved, a cognitive uncertainty quantization value is obtained by performing forward propagation on the Monte Carlo dropout network based on the deep feature vector, an objective evaluation system of model decision reliability is established, and a dual protection mechanism considering detection efficiency and result reliability is formed through the final uncertainty decision threshold linked with real-time constraints, and the comprehensive report output contains specific defect diagnosis and labeled detection confidence state, providing complete decision basis for quality control. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0019] Fig. 1 Flowchart of product monitoring method based on machine vision and convolutional neural network.
[0020] Fig. 2 Schematic diagram of product monitoring system based on machine vision and convolutional neural network. DETAILED DESCRIPTION
[0021] In order to make the above objectives, characteristics and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0022] In the following description, a large number of specific details are set forth in order to facilitate a thorough understanding of the present application, but the present application can also be implemented in other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the concept of the present application, therefore the present application is not limited to the specific embodiments disclosed below.
[0023] Secondly, one embodiment or embodiments referred to herein can include specific features, structures or characteristics that can be included in at least one implementation of the present application. In different places in the specification, one embodiment does not refer to the same embodiment, nor is it an independent or alternative embodiment that excludes other embodiments.
[0024] Reference Figs. 1-2 For one embodiment of the present application, the embodiment provides a product monitoring method based on machine vision and convolutional neural network, comprising the following steps: S1, image acquisition and preprocessing of automobile bearings are performed to obtain preprocessed automobile bearing two-dimensional image features and automobile bearing three-dimensional point cloud data.
[0025] S1.1, an industrial area array camera is used to collect automobile bearings to obtain two-dimensional gray scale images of automobile bearings, and then Canny edge detection is performed to extract preprocessed automobile bearing two-dimensional image features.
[0026] Furthermore, the two-dimensional gray scale images of automobile bearings collected by the industrial area array camera under uniform illumination conditions can completely retain the texture details and macroscopic topographic features of the bearing surface, and the operation of performing Canny edge detection on the two-dimensional gray scale images of automobile bearings can effectively identify the boundary lines of the inner and outer ring profiles, roller edges and possible surface cracks, pits and other defects of automobile bearings by calculating the image gradient amplitude and direction and applying double threshold detection and edge connection. The representation method based on edge features can convert the high-dimensional two-dimensional gray scale images of automobile bearings into more structural and discriminative preprocessed automobile bearing two-dimensional image features, which greatly reduces the data dimension of subsequent neural network processing and highlights the geometric and topological information most relevant to bearing size measurement and defect identification.
[0027] Specifically, the Canny edge detection is implemented to extract the preprocessed two-dimensional image features of the automobile bearing, the original pixel information is converted into higher-level features directly related to the bearing monitoring task, the conversion not only reduces the calculation burden, but also enhances the perception ability of the neural network to the key geometric features and abnormal areas of the bearing by highlighting the edge and contour information, avoids overfitting on irrelevant textures or noise, and the Canny edge detection is implemented on the two-dimensional gray image of the automobile bearing to extract the preprocessed two-dimensional image features of the automobile bearing.
[0028] S1.2, the same automobile bearing is scanned using a line laser three-dimensional scanner to obtain a three-dimensional point cloud of the automobile bearing, and a statistical-based outlier removal operation is performed to obtain a noise-removed three-dimensional point cloud of the automobile bearing.
[0029] Further, the three-dimensional point cloud of the automobile bearing obtained by the line laser three-dimensional scanner through the laser triangulation principle accurately records the three-dimensional space coordinates of each sampling point on the bearing surface, provides depth and three-dimensional geometric shape information that the two-dimensional gray image of the automobile bearing does not have, and is crucial for accurately measuring the size of the bearing such as the inner diameter, the outer diameter, the channel curvature and detecting three-dimensional topographic defects. Due to environmental interference, bearing surface reflection or dust and other factors during the scanning process, the three-dimensional point cloud of the automobile bearing inevitably contains some discrete noise points deviating from the main point group, i.e. outliers.
[0030] Specifically, the statistical-based outlier removal operation is performed to extract the statistical distribution of the average distance of each point from its neighboring points, identify and remove those points whose distance from the mean value exceeds several times the standard deviation, effectively filter out the scanning noise, purify the three-dimensional point cloud of the automobile bearing, and automatically determine the outlier judgment standard according to the density distribution of the local neighborhood of the point cloud. It can robustly handle the possible point cloud density changes in different parts of the bearing, and the noise-removed three-dimensional point cloud of the automobile bearing is obtained by performing the statistical-based outlier removal operation.
[0031] S1.3, the noise-removed three-dimensional point cloud of the automobile bearing is analyzed, the orientation information of the surface of each point in the point cloud is estimated, and the preprocessed three-dimensional point cloud data of the automobile bearing is obtained.
[0032] Further, the noise-removed automobile bearing three-dimensional point cloud contains accurate three-dimensional coordinates, but lacks the description of the local surface geometric properties. The surface orientation information of each point in the point cloud is estimated, and the normal vector estimation is usually based on the principal component analysis method. The characteristic value and characteristic vector of the covariance matrix of the local neighborhood composed of each point and its adjacent points are analyzed, and the characteristic vector corresponding to the minimum characteristic value is approximately the surface normal direction at the point. The estimated surface orientation information enriches the data connotation of the noise-removed automobile bearing three-dimensional point cloud, and upgrades the simple geometric point set to a three-dimensional feature set containing local surface direction description, i.e. the preprocessed automobile bearing three-dimensional point cloud data. The normal vector information in the preprocessed automobile bearing three-dimensional point cloud data has multiple values for subsequent three-dimensional feature extraction.
[0033] Specifically, the normal vector is used to obtain curvature and other local shape descriptors, which helps to identify bearing raceway, chamfer, and specific geometric feature areas. The consistency of the normal vector can be used to detect surface abnormalities, such as sudden direction changes that may indicate scratches or pits. In neural network processing, the normal vector can be used as an additional input channel to provide clues for local surface orientation and enhance the network's understanding of three-dimensional geometric structures. The normal vector estimation method based on principal component analysis is used to obtain the preprocessed automobile bearing three-dimensional point cloud data containing point coordinates and normal vectors.
[0034] S2, input the preprocessed automobile bearing two-dimensional image features and the preprocessed automobile bearing three-dimensional point cloud data into the pre-trained multi-task convolutional neural network.
[0035] S2.1, based on the preprocessed automobile bearing two-dimensional image features and the preprocessed automobile bearing three-dimensional point cloud data, construct a training set, train the multi-task convolutional neural network through the training set, and obtain the pre-trained multi-task convolutional neural network.
[0036] Further, the construction of the training set takes each pair of strictly aligned pre-processed 2D image features of the automobile bearing and pre-processed 3D point cloud data of the automobile bearing as a training sample, and is equipped with the automobile bearing normal score label and the preliminary diagnosis opinion label of the automobile bearing obtained from the real measurement and manual labeling of the bearing, and the process of training the multi-task convolutional neural network adopts a supervised joint learning paradigm, wherein the training target contains two items: one is to minimize the difference between the automobile bearing normal score predicted by the first output head of the multi-task convolutional neural network and the automobile bearing normal score label, and the other is to minimize the difference between the preliminary diagnosis opinion of the automobile bearing predicted by the second output head of the multi-task convolutional neural network and the preliminary diagnosis opinion label of the automobile bearing, and the network parameters of the shared feature encoder, the first output head and the second output head are iteratively updated by the back propagation algorithm and the optimizer, so that the total loss function of the network on the training set is continuously reduced until convergence, and the shared feature encoder is forced to learn a general and powerful feature representation that can simultaneously support rapid qualification and fine defect diagnosis.
[0037] Specifically, the first output head focuses on extracting global patterns highly related to overall qualification from the fused features, and the second output head focuses on analyzing fine-grained patterns related to specific defect types and size deviations, avoiding feature learning redundancy and waste of computing resources caused by training independent networks for the two tasks, and also enabling high-quality outputs for different decision stages to be obtained simultaneously through a single forward propagation in subsequent inference. By training the multi-task convolutional neural network with the training set, a pre-trained multi-task convolutional neural network with optimized network parameters and the ability to simultaneously output automobile bearing normal scores and preliminary defect judgments is ultimately obtained.
[0038] S2.2, input the pre-processed 2D image features of the automobile bearing into the 2D convolution branch of the pre-trained multi-task convolutional neural network to generate a 2D feature map.
[0039] Further, the pre-processed 2D image features of the automobile bearing are structured images highlighting the bearing outline and defect boundaries extracted by Canny edge detection, and the 2D convolution branch of the pre-trained multi-task convolutional neural network is usually composed of multiple convolution layers, activation function layers and pooling layers stacked together. The layers have learned the weight parameters for extracting multi-level visual features from the bearing edge image during the training process. When the pre-processed 2D image features of the automobile bearing are input into the 2D convolution branch, the shallow convolution kernels are responsible for detecting low-level features such as basic edge directions and corner points. As the multi-task convolutional neural network deepens, the subsequent convolution layers gradually abstract higher-level semantic features such as specific shape outline fragments, texture patterns and abnormal line segment connection patterns that may indicate cracks by combining low-level features.
[0040] Specifically, the two-dimensional feature map finally output by the two-dimensional convolution branch is a high-dimensional tensor, where each spatial position corresponds to a local region of the input image, and each channel encodes a certain specific high-level visual feature of the region. The conversion from the pre-processed two-dimensional image features of the automobile bearing to the two-dimensional feature map converts the explicit edge geometry information into a more compact and more discriminative implicit feature representation. The pre-processed two-dimensional image features of the automobile bearing are input into the two-dimensional convolution branch of the pre-trained multi-task convolutional neural network, and the learned convolution filter set is used for layer-by-layer feature extraction and transformation to generate a two-dimensional feature map.
[0041] S2.3, input the pre-processed automobile bearing three-dimensional point cloud data into the three-dimensional convolution branch of the pre-trained multi-task convolutional neural network to generate a three-dimensional feature vector.
[0042] Further, the pre-processed automobile bearing three-dimensional point cloud data contains the three-dimensional coordinates of each point and the estimated surface normal vector, which is an irregular spatial data. The three-dimensional convolution branch of the pre-trained multi-task convolutional neural network adopts a network structure capable of directly processing point cloud data, such as PointNet or its variants. The branch improves the coordinates and normal features of each point through a shared multi-layer perception, and then aggregates the global point cloud features through a symmetric function such as maximum pooling to generate a global feature representation that is invariant to the arrangement order of the points. In the training process, the three-dimensional convolution branch learns how to capture key geometric properties from the point cloud data of the bearing, such as the cylindrical characteristics of the cylindrical surface, the groove shape of the raceway, the smooth transition of the chamfer, and possible local deformations such as depressions or protrusions.
[0043] Specifically, the pre-processed automobile bearing three-dimensional point cloud data is input into the trained three-dimensional convolution branch, and the network processes the point cloud information layer by layer to finally output a compact three-dimensional feature vector of fixed dimension. The three-dimensional feature vector comprehensively encodes the global geometric shape and key local shape features of the entire bearing point cloud, providing a direct geometric basis for the size precision evaluation and three-dimensional topography defect detection of the bearing. Unlike the two-dimensional feature map, which focuses on capturing appearance features from a projection perspective, the three-dimensional feature vector describes the geometric entity of the bearing from a three-dimensional space. The two have natural complementarity. The pre-processed automobile bearing three-dimensional point cloud data is input into the three-dimensional convolution branch of the pre-trained multi-task convolutional neural network, and a three-dimensional feature vector is generated through its point cloud feature extraction and aggregation mechanism.
[0044] S3, extract the fusion features through the shared feature encoder, generate the automobile bearing normal score through the first output head, and generate the defect preliminary judgment and high-dimensional feature deep feature vector through the second output head.
[0045] S3.1, the two-dimensional feature map and the three-dimensional feature vector are spliced in the fusion layer of the shared feature encoder to obtain fused features.
[0046] Further, the two-dimensional feature map carries high-level visual semantic information extracted from the preprocessed two-dimensional image features of the automobile bearing, such as specific contour patterns, texture abnormalities, and edge features of defects. The three-dimensional feature vector encodes global and local geometric attributes abstracted from the preprocessed three-dimensional point cloud data of the automobile bearing, such as overall shape, surface curvature, and size deviation trends. The fusion layer of the shared feature encoder is responsible for integrating these two heterogeneous but complementary feature representations. A typical fusion method is to perform a channel dimension splicing operation, that is, to connect the flattened or further convoluted feature vector of the two-dimensional feature map with the three-dimensional feature vector into a longer joint feature vector. The splicing operation preserves all the effective information extracted by the two-dimensional and three-dimensional branches.
[0047] Specifically, for example, a texture pattern in the two-dimensional feature indicating surface scratches may be mutually confirmed by a slight depth depression or normal direction mutation in the corresponding area of the three-dimensional feature. The consistency across modalities will enhance the model's confidence in judging real defects. If the two-dimensional feature shows abnormalities while the three-dimensional geometric feature is normal, it may indicate that it is a lighting artifact or surface stain rather than a physical defect. The fused features obtained through feature splicing construct a unified multi-modal feature space, where each sample is defined by its visual appearance and geometric form. This avoids the data alignment problem that may be encountered in early fusion and avoids the loss of low-level interaction information that may be lost in late fusion. After sufficient refinement by each branch, the two-dimensional feature map and the three-dimensional feature vector are spliced in the channel dimension in the fusion layer of the shared feature encoder to obtain fused features.
[0048] S3.2, the fused features are input into a first output head of a pre-trained multi-task convolutional neural network, the first output head being configured to perform a simplified binary classification task to extract global patterns from the fused features for judging whether the automobile bearing is qualified as a whole, and output an automobile bearing normal score.
[0049] Further, a highly abstracted binary classification decision is made: whether the current bearing is apparently normal or not. The training process drives the first output head to learn to filter out from the fused features the distracting information related to subtle defects, complex background, or weak contradictions between modalities, and instead focus on global, high-confidence patterns that can strongly distinguish between good and potential problem, which can be the high degree of consistency between the two-dimensional and three-dimensional features in describing the overall bearing profile integrity, symmetry, and major dimension compliance, for example, a good bearing, its two-dimensional edge features outline a circular profile that highly matches the cylindrical information reflected by the three-dimensional features, and there are no sharp abnormal corners or abrupt curvature, the role of the first output head is to extract from the fused features the quantitative measure of the degree of harmony and consistency, and map it to a scalar value between zero and one, i.e., the car bearing normal score, the higher the score, the more consistent the multi-modal evidence presented by the fused features points to the conclusion of no abnormalities.
[0050] Specifically, the fused features are input into the first output head to make a quick decision on most apparently good samples by using a lightweight branch in the early stage of the network, avoiding these samples from continuing to flow through the subsequent more complex calculation path, thereby greatly reducing the average processing time, based on the realistic insight that good products account for the vast majority in industrial production lines, leaving the main computing resources to the few suspicious samples that really need detailed analysis, the fused features are input into the first output head of the pre-trained multi-task convolutional neural network, and the car bearing normal score representing the overall bearing compliance confidence is output through the simplified binary classification decision function of the first output head.
[0051] S3.3, the fused features are input into the second output head of the pre-trained multi-task convolutional neural network, and the second output head is configured to perform fine-grained multi-task learning to parse specific semantic information for defect classification and geometric features for size regression from the fused features, and output a defect preliminary judgment and a deep feature vector of high-dimensional features.
[0052] Further, usually multiple fully connected layers or more complex subnetworks are responsible for the dual tasks, on the one hand, the need for defect classification, that is, identifying the specific defect type of the bearing, such as inner ring crack, roller missing, outer diameter out-of-tolerance, which is essentially a multi-class classification problem, on the other hand, the need for size regression, that is, accurately predicting the deviation amount of key dimensions, such as the difference between the actual measurement value and the standard value of the inner diameter and the outer diameter, which is a continuous regression problem, the second output head optimizes these two goals simultaneously through a multi-task learning framework, forcing the shared underlying fusion features to contain sufficient rich and fine information to support these two types of discrimination, from the fusion features, the second output head needs to analyze the specific semantic information, for example, identifying the feature pattern belonging to the crack category is usually associated with two-dimensional linear dark feature and three-dimensional local narrow concave feature; while the size out-of-tolerance is more related to the overall scale shift reflected by the three-dimensional feature, at the same time, it also needs to analyze the geometric features for accurate regression, which may involve decoding the components in the fusion features that are sensitive to absolute size.
[0053] Specifically, the second output head finally outputs two parts: one is the preliminary judgment of defects, that is, the defect category label and its probability distribution or the specific value of the size deviation; the other is the deep feature vector of high-dimensional features, which is an intermediate layer high-dimensional representation generated by the second output head in the process of completing multi-task judgment, which contains higher-level and more task-related semantic information than the original fusion features, the deep feature vector is the direct input for subsequent cognitive uncertainty quantification, reflecting the internal feature state on which the multi-task convolutional neural network relies when making specific defect judgments, inputting the fusion features into the second output head embodies the idea of functional separation and cooperation, the first output head makes a quick is and is not judgment, and the second output head makes a fine what and how much difference analysis, both share the same fusion feature input, ensuring the consistency of the analysis basis, realizing seamless connection from rapid screening to deep diagnosis and on-demand allocation of computing resources, inputting the fusion features into the pre-trained second output head of the multi-task convolutional neural network, through its fine-grained multi-task learning network, outputting the preliminary judgment of defects and the deep feature vector of high-dimensional features.
[0054] S4, generating a first-level decision threshold according to the current production line real-time throughput rate and historical quality statistical information.
[0055] S4.1, reading the current production line real-time throughput rate from the data interface of the production line controller.
[0056] Further, the production line controller is usually the core unit of the manufacturing execution system, which continuously monitors and records the running state of the production line. The current production line real-time throughput rate is a key performance indicator, which directly reflects the number of automobile bearings passing through the detection station per unit time, i.e. the production rhythm of the production line. Through the standard industrial communication protocol, such as OPC UA or PROFINET, a connection can be established with the data interface of the production line controller, and the value of the current production line real-time throughput rate can be read in real time.
[0057] Specifically, the injection of external production state information into the internal quality monitoring decision-making process is achieved, which dynamically associates the quality detection behavior with the real-time production rhythm, breaking the limitation of independent quality judgment and production scheduling in traditional detection schemes. Reading the current production line real-time throughput rate is the first step to build an adaptive monitoring strategy, and the current production line real-time throughput rate is read from the data interface of the production line controller.
[0058] S4.2, query the historical database to obtain the historical quality statistical information of the normal score of the automobile bearing recorded in the historical batch matched with the current time period.
[0059] Further, the historical database stores relevant data of all detected automobile bearings in the past period, especially the normal score of the automobile bearing generated by the first output head of the multi-task convolutional neural network. The quality stability trend of the current production batch or time period is evaluated, and the historical data comparable to the current time is retrieved from the historical database, such as retrieving the normal score sequence of the automobile bearing of the continuous automobile bearings produced under the same production shift and the same equipment parameters in the past, obtaining the historical quality statistical information reflecting the distribution characteristics. One of the most core indicators is the standard deviation of the score, i.e. the historical quality statistical standard deviation of the normal score of the automobile bearing.
[0060] Specifically, the historical quality statistical standard deviation quantifies the fluctuation degree of the recent production quality. A smaller standard deviation indicates stable product quality and a concentrated distribution of normal scores. A larger standard deviation indicates large fluctuations in product quality and a dispersed distribution of normal scores. The purpose of querying the historical database to obtain the historical quality statistical information of the normal score of the automobile bearing is to introduce a context factor based on historical experience that reflects the inherent fluctuation of the process, embodying the idea of process statistical quality control. The decision-making not only considers the score of the current single sample, but also obtains the historical quality statistical information of the normal score of the automobile bearing.
[0061] S4.3, based on the current production line real-time throughput rate and the historical quality statistical information of the normal score of the automobile bearing, a first-level decision threshold is calculated through a pre-defined decision function.
[0062] Further, the first-level decision threshold is used to compare with the normal score of the automobile bearing to determine whether to perform fast release on the current bearing, and a predefined decision function defines how the two dynamic inputs of the current line real-time throughput rate and the historical quality statistical information of the normal score of the automobile bearing jointly affect the value of the first-level decision threshold, the decision function realizes adaptive matching of detection strictness and production demand and quality condition, and the decision function is specifically embodied in the following relationship: the first-level decision threshold is equal to a basic decision threshold plus a throughput rate adjustment coefficient multiplied by a target line throughput rate minus a current line real-time throughput rate divided by the target line throughput rate, and then minus a quality stability adjustment coefficient multiplied by a historical quality statistical standard deviation of the normal score of the automobile bearing, and the basic decision threshold is a reference value set according to a long-term quality target.
[0063] Specifically, the difference between the current line real-time throughput rate and the target line throughput rate affects the first-level decision threshold through the throughput rate adjustment coefficient, when the current line real-time throughput rate is lower than the target line throughput rate, the line runs slowly, and there is relatively sufficient processing time, at this time, the quotient obtained by subtracting the current line real-time throughput rate from the target line throughput rate and then dividing by the target line throughput rate is positive, which leads to the first-level decision threshold being raised, the threshold being raised means that the requirement for the normal score of the automobile bearing becomes more stringent, only the bearings with higher scores and more confident qualified bearings can be fast released, which is beneficial to more stringent quality control when the time pressure is small, on the contrary, when the current line real-time throughput rate is higher than the target line throughput rate, the line runs at high speed, and the time pressure is large, at this time, the quotient obtained by subtracting the current line real-time throughput rate from the target line throughput rate and then dividing by the target line throughput rate is negative, which leads to the first-level decision threshold being lowered, the first-level decision threshold being lowered means that the standard of fast release is relaxed, and more bearings with slightly lower scores, i.e., slightly doubtful but still possible qualified bearings, are allowed to be fast released, avoiding the processing bottleneck caused by excessive stringent preliminary screening under the high-speed line to ensure the line flux, on the other hand, the historical quality statistical standard deviation of the normal score of the automobile bearing affects the threshold through the quality stability adjustment coefficient.
[0064] It should be noted that the greater the historical quality statistical standard deviation of the normal score of the automobile bearing, the more intense the recent quality fluctuation, the more unstable the process, and the subtraction of the quality stability adjustment coefficient multiplied by the historical quality statistical standard deviation of the normal score of the automobile bearing will reduce the first-level decision threshold. When the process itself is unstable, the overall distribution of the normal score of the automobile bearing may have shifted or broadened. Continuing to maintain a high threshold value may cause a large number of borderline qualified products that can be released to be sent to in-depth review. A moderate reduction in the threshold value is a robust strategy to adapt to process fluctuations. In special periods, the judgment benchmark needs to be adjusted to prevent systematic decision-making errors caused by rigid standards. The method of encoding the production throughput and historical quality fluctuation into a dynamic threshold value deeply integrates the external operating state of the production system and the internal decision logic of quality detection, realizes the transformation of the quality monitoring strategy from static preset to dynamic situational awareness, and outputs the first-level decision threshold value through the comprehensive calculation of the current production line real-time throughput rate and the historical quality statistical information of the normal score of the automobile bearing by a predefined decision function.
[0065] The first-level decision threshold expression is: ; Wherein, is the first-level decision threshold, is the basic decision threshold, is the throughput rate adjustment coefficient, is the target production line throughput rate, is the current production line real-time throughput rate, is the quality stability adjustment coefficient, is the historical quality statistical standard deviation of the normal score of the automobile bearing.
[0066] S5, compare the normal score of the automobile bearing with the first-level decision threshold, and output qualified and end the process and depth review state, respectively.
[0067] S5.1, numerically compare the normal score of the automobile bearing with the first-level decision threshold.
[0068] Further, the normal score of the automobile bearing is a scalar value generated by the first output head of the multi-task convolutional neural network, the first-level decision threshold is another scalar value calculated by a predefined decision function based on the current production line real-time throughput rate and the historical quality statistical information of the normal score of the automobile bearing, the numerical comparison operation is a simple arithmetic judgment, which judges whether the numerical value of the normal score of the automobile bearing is strictly greater than the numerical value of the first-level decision threshold, the comparison operation converts the normal score of the automobile bearing in the sense of probability into a discrete binary decision signal, and through the comparison operation, it can be determined whether the confidence of the qualification of the current automobile bearing meets the rapid release standard set in the current production situation, so as to decide whether to directly determine the qualification and end the process or enter the deep checking state.
[0069] S5.2, when the normal score of the automobile bearing is higher than the first-level decision threshold, the automobile bearing is determined to be qualified and the current monitoring process is ended.
[0070] Further, when the normal score of the automobile bearing is higher than the first-level decision threshold, the automobile bearing is determined to be qualified and the current monitoring process is ended. The conditional branch means that after the multi-modal features of the current automobile bearing are evaluated by the neural network, the normal score of the automobile bearing is not only high, but also exceeds the dynamic qualification standard determined by the current production situation. The determination of qualification is the final decision of the bearing, and the end of the current monitoring process means the immediate termination of the step. The uncertainty quantification based on the deep feature vector and the final decision, in industrial mass production, most products are qualified, and by setting a rapid decision based on a lightweight comparison in the front end, samples that are obviously qualified can pass instantly with very low computational cost.
[0071] S5.3, when the normal score of the automobile bearing is not higher than the first-level decision threshold, the state of the monitoring process is marked as needing deep checking.
[0072] Further, when the normal score of the automobile bearing is not higher than the first-level decision threshold, the state of the monitoring process is marked as needing deep checking. The conditional branch triggers another processing path, and the normal score of the automobile bearing is not higher than the first-level decision threshold, which indicates that the neural network evaluation score of the current bearing fails to meet the rapid release standard required by the dynamic situation. Due to slight abnormalities of the bearing itself, poor image quality, or generally conservative model scores under current production fluctuations. Marking the state of the monitoring process as needing deep checking is a state setting operation, which does not output the final conclusion, but provides an explicit instruction signal for the subsequent process. The signal will guide the process to a more complex, resource-consuming but more accurate analysis stage, ensuring that no sample is passed hastily, and all suspicious or borderline cases will be subject to more rigorous review.
[0073] Specifically, the two-stage processing architecture is reflected in its risk control thinking: the first stage of dynamic threshold comparison is responsible for efficient and low-risk preliminary screening; once the preliminary screening is suspicious, it is automatically upgraded to the second stage of high-reliability analysis, avoiding the inefficiency caused by excessive review of all samples and preventing the missed detection risk that may be caused by single threshold judgment.
[0074] S6. Based on the deep feature vector, the second output head with the Monte Carlo dropout operation enabled is forward propagated to obtain a cognitive uncertainty quantization value.
[0075] S6.1, input the deep feature vector into the second output head with the Monte Carlo dropout operation enabled, and forward propagate the second output head with the Monte Carlo dropout operation enabled to obtain a plurality of sets of probability distributions output by the second output head.
[0076] Further, the deep feature vector is a high-dimensional semantic feature representation generated by the second output head of the multi-task convolutional neural network when performing preliminary diagnosis, which carries all the key information for fine-grained defect classification and size regression. The second output head with the Monte Carlo dropout operation enabled means that in the inference stage, the full connection layer in the output head network structure still randomly discards part of the neuron connections according to the probability set during training, but the network weights remain fixed. The same deep feature vector is repeatedly input into this network with random dropout function, and each forward propagation will produce a slightly different data path due to the random inactivation of neuron connections, resulting in a slight random fluctuation in the output of each propagation, i.e., the probability distribution of the defect category and the regression prediction value of the key size.
[0077] Specifically, performing multiple such forward propagations, for example, dozens of times, will obtain a plurality of sets of probability distributions output by the second output head. The Monte Carlo dropout technique, which is usually used to prevent overfitting during the training stage, is purposefully reused in the inference stage to actively retain and utilize randomness, not to improve generalization. For a cognitively clear sample, the final judgment result should be consistent regardless of the random changes in the internal path of the multi-task convolutional neural network, and the output probability distribution fluctuation is small. Conversely, for a cognitively ambiguous sample, the random multi-task convolutional neural network disturbance will be amplified, resulting in a difference in the probability distribution output multiple times. The volatility of the multiple sets of output is itself an intuitive reflection of the multi-task convolutional neural network's grasp of the current input. By forward propagating the second output head with the Monte Carlo dropout operation enabled, a plurality of sets of probability distributions output by the second output head with inherent randomness are obtained.
[0078] S6.2, statistical calculation is performed on the plurality of sets of probability distributions output by the second output head to obtain a cognitive uncertainty quantization value.
[0079] Furthermore, the probability distributions of multiple sets of second output heads constitute a set of multiple random predictions for the same input sample. The goal of cognitive uncertainty quantification is to use a scalar value to summarize the degree of dispersion among these multiple sets of predictions, i.e., the degree of cognitive divergence. Statistical calculations process the outputs of classification and regression tasks separately. For the classification task, i.e., the probability distribution of multiple defect categories, the calculation is performed for each defect category. The variance of the predicted probability in the second propagation, and then for all The variances of each category are summed and averaged. For a regression task, i.e., regression predictions for multiple key dimensions, the following calculation is performed: The variance of each predicted value is determined by introducing a weighting coefficient for the regression task uncertainty component, since the classification uncertainty component and the regression uncertainty component may have different dimensions and importance. The regression variance is weighted, and the weighted regression variance is added to the aforementioned average classification variance to obtain a comprehensive quantitative value of cognitive uncertainty.
[0080] Specifically, it transforms the abstract concept of uncertainty into a precisely measurable, mathematically defined continuous value through Monte Carlo sampling and statistical variance calculation. The cognitive uncertainty quantification value has clear interpretability: the closer the cognitive uncertainty quantification value is to zero, the more consistent the results of multiple random predictions are, and the more confident one is in one's own judgment. The larger the value, the higher the dispersion of multiple prediction results. The sample may belong to an unknown pattern outside the training distribution, a multi-class boundary case, or an ambiguous complex sample. Quantification capability is the cornerstone of this method to achieve reliable decision-making. The statistical variance of the probability distribution and regression value of multiple sets of second output heads is calculated to finally obtain the cognitive uncertainty quantification value.
[0081] The expression for the quantification of cognitive uncertainty is: ; in, To quantify cognitive uncertainty, For Monte Carlo forward propagation number, For the number of Monte Carlo forward propagations, The total number of defect categories. For the defect category, For the first The next forward propagation predicts the category to which the current sample belongs. The probability, For category The predicted probability is Mean of the next forward propagation The weighting coefficients are for the uncertainty components of the regression task. For the first regression prediction value of the critical dimension at the next-to-forward propagation time, is a set of regression prediction values.
[0082] S7, determining a final uncertainty decision threshold according to the cognitive uncertainty quantification value and the real-time constraint.
[0083] S7.1, calculating the remaining allowed decision time as the real-time constraint based on the preset total processing time limit and the current process consumed time.
[0084] Further, the preset total processing time limit is the maximum time threshold allowed from the automobile bearing entering the detection station to outputting the final report to ensure the continuous and smooth operation of the production line, and the current process consumed time includes the time consumed in the data acquisition and preprocessing step, the time consumed in the multi-task convolutional neural network forward propagation step, and the time consumed in the first step routing decision step. By timing and adding the actual execution process, the sum of the current process consumed time is subtracted from the preset total processing time limit, and the remaining allowed decision time is obtained. The macro and fixed production line beat requirement is converted into a pressure parameter affecting the current sample decision process through real-time time budgeting. The remaining allowed decision time is no longer an external abstract constraint, but is internalized as a key input variable in the decision logic.
[0085] Specifically, based on the concept of decision time resource, and recognizing that the deep uncertainty assessment of a complex sample itself is also a time-consuming calculation activity, the strictness and prudence of decision-making are dynamically adjusted according to how much time resource is left, for example, when a sample reaches this step, if the remaining allowed decision time is still very sufficient, it is possible to make more conservative and strict judgments; if the remaining allowed decision time is very nervous, more decisive and possibly more risk-oriented strategies must be taken to avoid overtime. Dynamic management of time budgeting enables the entire monitoring process to intelligently allocate its decision-making prudence without violating the hard real-time requirement, and obtains the remaining allowed decision time as the real-time constraint.
[0086] The expression of the remaining allowed decision time is: ; wherein, is the remaining allowed decision time, is the preset total processing time limit of a single automobile bearing, is the time consumed in the data acquisition and preprocessing step of the current automobile bearing, is the time consumed in the multi-task convolutional neural network forward propagation step of the current automobile bearing, the time consumed by the current automotive bearing in the first level routing decision step.
[0087] S7.2, dynamically determine the final uncertainty decision threshold value through the preset trade-off strategy according to the cognitive uncertainty quantitative value and the real-time constraint.
[0088] Further, the cognitive uncertainty quantitative value represents the self-doubt degree of the multi-task convolutional neural network on the preliminary diagnosis opinion of the current automotive bearing, the real-time constraint is the remaining allowed decision time, which represents the time resource possessed for completing subsequent decision, and the preset trade-off strategy defines how the two factors jointly determine the final uncertainty decision threshold value. A typical trade-off strategy is that the final uncertainty decision threshold value is positively correlated with the real-time constraint and negatively correlated with the demand of the cognitive uncertainty quantitative value itself. The logic is double: on the one hand, in the face of a higher cognitive uncertainty quantitative value, the model should require a lower decision threshold value to trigger more cautious processing (such as marking as high uncertainty), because high uncertainty means unreliable judgment; on the other hand, when the real-time constraint is very tight, i.e., the remaining allowed decision time is very short, the process does not have enough time to tolerate or handle high uncertainty state, forcing the decision to converge faster, increasing the final uncertainty decision threshold value, so that only extremely high uncertainty can trigger special processing.
[0089] Specifically, when the real-time constraint is relaxed, more extensive uncertainty can be conditionally tolerated by setting a lower final uncertainty decision threshold value, so that more samples with uncertainty are identified for special labeling or processing, thereby achieving higher final decision reliability. Instead of comparing the cognitive uncertainty quantitative value with a fixed threshold value, it is compared with a dynamic threshold value reflecting the remaining time resource, thereby realizing online optimization and balance between decision reliability and decision timeliness. It is recognized that in real-time monitoring, reliability pursuit can become unrealistic due to excessive time consumption. An optimal trade-off is made within a limited time, the cognitive uncertainty quantitative value and the real-time constraint are comprehensively evaluated through the preset trade-off strategy, and the final uncertainty decision threshold value suitable for the current sample and the current time is obtained.
[0090] S8, generate a comprehensive report of the final monitoring result of the automotive bearing by combining the preliminary diagnosis opinion of the automotive bearing.
[0091] S8.1, compare the cognitive uncertainty quantitative value with the final uncertainty decision threshold value.
[0092] Further, the cognitive uncertainty quantification value is a scalar value obtained by multiple Monte Carlo forward propagation, quantifying the self-doubt degree of the second output head of the multi-task convolutional neural network for the current automobile bearing judgment, and the final uncertainty decision threshold is another scalar value dynamically determined based on the cognitive uncertainty quantification value and the real-time constraint through a preset trade-off strategy, representing the acceptable upper limit of the uncertainty of the multi-task convolutional neural network under the pressure of the current remaining allowed decision time. The comparison operation is a simple numerical value comparison, which determines whether the numerical value of the cognitive uncertainty quantification value is less than the numerical value of the final uncertainty decision threshold, converts the cognitive uncertainty quantification value with a probability significance into a discrete binary decision signal, and determines whether the model judgment reliability of the current automobile bearing meets the acceptable standard under the current time constraint through the comparison operation, thereby deciding whether to directly adopt the preliminary diagnosis opinion or mark it as a high-uncertainty state requiring manual review. The numerical comparison between the cognitive uncertainty quantification value and the final uncertainty decision threshold, S8.2, when the cognitive uncertainty quantification value is lower than the final uncertainty decision threshold, adopting the preliminary diagnosis opinion of the automobile bearing as the final diagnosis opinion of the automobile bearing.
[0093] Further, when the cognitive uncertainty quantification value is lower than the final uncertainty decision threshold, the preliminary diagnosis opinion of the automobile bearing is adopted as the final diagnosis opinion of the automobile bearing. The self-uncertainty degree of the current sample is lower than the dynamically set acceptable upper limit, and the preliminary diagnosis opinion of the automobile bearing is adopted as the final diagnosis opinion of the automobile bearing, which will inherit the defect category, size deviation value and other information determined by the second output head initially, without change. The trust mechanism is established, not unconditional trust in the neural network output, but trust under the condition that both the self-evaluation of the multi-task convolutional neural network and the external constraint are met, avoiding unnecessary conservative treatment.
[0094] S8.3, when the cognitive uncertainty quantification value is not lower than the final uncertainty decision threshold, marking the final diagnosis opinion of the automobile bearing as a high-uncertainty state requiring manual review.
[0095] Further, when the cognitive uncertainty quantitative value is not lower than the final uncertainty decision threshold, the final diagnosis opinion of the automobile bearing is marked as a high uncertainty state requiring manual review. This condition triggers another processing logic. When the cognitive uncertainty quantitative value is not lower than the final uncertainty decision threshold, it indicates that the self-uncertainty degree of the multi-task convolutional neural network reaches or exceeds the tolerance limit under the current situation. Regardless of the specific content of the preliminary diagnosis opinion of the automobile bearing, its reliability is doubtful. Marking the final diagnosis opinion of the automobile bearing as a high uncertainty state requiring manual review is a risk control operation, which does not overturn the preliminary diagnosis opinion. Marking the final diagnosis opinion of the automobile bearing as a high uncertainty state requiring manual review.
[0096] S8.4, integrating the final diagnosis opinion of the automobile bearing, the cognitive uncertainty quantitative value, and the specific information in the preliminary diagnosis opinion of the automobile bearing, and generating a comprehensive report of the final monitoring result of the automobile bearing.
[0097] Further, all key data items are structured, summarized, and formatted. The final diagnosis opinion of the automobile bearing is the core conclusion. The preliminary diagnosis opinion is also likely to be the preliminary diagnosis opinion marked with a high uncertainty state. The cognitive uncertainty quantitative value is recorded as a key reliability indicator, providing a quantitative basis for the conclusion. The specific information in the preliminary diagnosis opinion of the automobile bearing, such as detailed defect category probability distribution and accurate size deviation value, is retained as reference details. A comprehensive report of the final monitoring result of the automobile bearing is generated.
[0098] The embodiment also provides a product monitoring system based on machine vision and convolutional neural network, which comprises a data processing module for image acquisition and preprocessing of automobile bearings to obtain preprocessed automobile bearing two-dimensional image features and automobile bearing three-dimensional point cloud data; A perception module inputs the preprocessed automobile bearing two-dimensional image features and the preprocessed automobile bearing three-dimensional point cloud data into a pre-trained multi-task convolutional neural network, extracts fusion features through a shared feature encoder, generates an automobile bearing normal score through a first output head, and generates a defect preliminary judgment and a high-dimensional feature deep feature vector through a second output head. A decision module generates a first-level decision threshold according to a current production line real-time throughput rate and historical quality statistical information, compares the automobile bearing normal score with the first-level decision threshold, and respectively outputs qualified and ends the process and a deep verification state. An evaluation module performs forward propagation on a second output head with a Monte Carlo dropout operation based on a deep feature vector to obtain a cognitive uncertainty quantitative value. The decision module determines a final uncertainty decision threshold according to the cognitive uncertainty quantitative value and the real-time constraint, and generates a comprehensive report of the final monitoring result of the automobile bearing by combining the preliminary diagnosis opinion of the automobile bearing.
[0099] The embodiment further provides a computer device suitable for the product monitoring method based on machine vision and a convolutional neural network, including a memory and a processor; the memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions to realize the product monitoring method based on machine vision and the convolutional neural network.
[0100] The computer device can be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected through a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to perform wired or wireless communication with an external terminal. The wireless communication can be achieved through WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, a trackball or a touchpad arranged on the shell of the computer device, or an external keyboard, a touchpad or a mouse.
[0101] The embodiment further provides a storage medium having a computer program stored thereon, and the program is executed by a processor to realize the product monitoring method based on machine vision and the convolutional neural network. The storage medium can be realized by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disk.
[0102] To sum up, by means of the multi-modal data fusion and multi-task collaborative processing architecture, the shared feature encoder is adopted to synchronously generate the automobile bearing normal score and the deep feature vector, the dual improvement of detection efficiency and feature reusability is realized, the first level decision threshold is dynamically generated according to the real-time throughput rate and historical quality statistical information of the production line, the intelligent shunting mechanism of adaptive production rhythm is constructed, the contradiction between the fixed computing resources and the changing production demand is solved, the cognitive uncertainty quantization value is obtained by carrying out the forward propagation on the Monte Carlo discard network based on the deep feature vector, the objective evaluation system of model decision reliability is established, and the dual guarantee mechanism of detection efficiency and result reliability is formed through the final uncertainty decision threshold linked with the real-time constraint, and the comprehensive report output contains specific defect diagnosis and labeling detection confidence state, thereby providing complete decision basis for quality control.
[0103] It should be noted that the above examples are only used to illustrate the technical solutions of the present application but not limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the present application, and all should be covered in the scope of the claims of the present application.
Claims
1. A product monitoring method based on machine vision and convolutional neural networks, characterized in that: This includes acquiring and preprocessing images of automotive bearings to obtain preprocessed two-dimensional image features and three-dimensional point cloud data of automotive bearings. The pre-processed 2D image features of the automobile bearing and the pre-processed 3D point cloud data of the automobile bearing are input into a pre-trained multi-task convolutional neural network. The fused features are extracted through a shared feature encoder. The first output head generates the normal score of the automobile bearing, and the second output head generates the depth feature vector of the preliminary defect judgment and high-dimensional features. Based on the current real-time throughput rate of the production line and historical quality statistics, a first-level decision threshold is generated. The normal score of the automotive bearing is compared with the first-level decision threshold, and the qualified and process end status and the in-depth verification status are output respectively. Based on the deep feature vector, the second output head with Monte Carlo drop-off operation enabled is forward-propagated to obtain the cognitive uncertainty quantification value. Based on the quantification of cognitive uncertainty and real-time constraints, the final uncertainty decision threshold is determined. By combining the preliminary diagnostic opinions of automotive bearings, a comprehensive report on the final monitoring results of automotive bearings is generated.
2. The product monitoring method based on machine vision and convolutional neural networks as described in claim 1, characterized in that: Image acquisition and preprocessing of automotive bearings are performed to obtain preprocessed two-dimensional image features and three-dimensional point cloud data of automotive bearings, including the following steps: An industrial area array camera is used to acquire a two-dimensional grayscale image of the car bearing. Then, Canny edge detection is performed to extract the preprocessed two-dimensional image features of the car bearing. The same car bearing was scanned using a line laser 3D scanner to obtain a 3D point cloud of the car bearing, and a statistical outlier removal operation was performed to obtain a noise-removed 3D point cloud of the car bearing. The noise-removed 3D point cloud of automotive bearings is analyzed, and the surface orientation information of each point in the point cloud is estimated. This is the preprocessed 3D point cloud data of automotive bearings.
3. The product monitoring method based on machine vision and convolutional neural networks as described in claim 2, characterized in that: The preprocessed 2D image features of the automobile bearing and the preprocessed 3D point cloud data of the automobile bearing are input into a pre-trained multi-task convolutional neural network, including the following steps: A training set is constructed based on the preprocessed 2D image features of automobile bearings and the preprocessed 3D point cloud data of automobile bearings. The multi-task convolutional neural network is trained using the training set to obtain a pre-trained multi-task convolutional neural network. The preprocessed 2D image features of the automobile bearing are input into the 2D convolution branch of a pre-trained multi-task convolutional neural network to generate a 2D feature map. The pre-processed 3D point cloud data of the automobile bearing is input into the 3D convolution branch of a pre-trained multi-task convolutional neural network to generate a 3D feature vector.
4. The product monitoring method based on machine vision and convolutional neural networks as described in claim 3, characterized in that: The fusion features are extracted through a shared feature encoder. The first output head generates the normal score of the automotive bearing, and the second output head generates a deep feature vector with preliminary defect judgment and high-dimensional features. The process includes the following steps: The two-dimensional feature map and the three-dimensional feature vector are concatenated in the fusion layer of the shared feature encoder to obtain the fused features; The fused features are input into the first output head of a pre-trained multi-task convolutional neural network. The first output head is configured to perform a simplified binary classification task, extracting a global pattern from the fused features to determine whether the car bearing is qualified overall, and outputting a normal score for the car bearing. The fused features are input into the second output head of a pre-trained multi-task convolutional neural network. The second output head is configured to perform fine-grained multi-task learning, parse the specific semantic information for defect classification and the geometric features for size regression from the fused features, and output the preliminary defect judgment and the deep feature vector of high-dimensional features.
5. The product monitoring method based on machine vision and convolutional neural networks as described in claim 4, characterized in that: Based on the current real-time throughput rate of the production line and historical quality statistics, a first-level decision threshold is generated, including the following steps: Read the current real-time throughput rate of the production line from the data interface of the production line controller; Query the historical database to obtain historical quality statistics of normal scores for automotive bearings in historical batches that match the current time period; Based on the current real-time throughput rate of the production line and historical quality statistics of the normal score of automotive bearings, the first-level decision threshold is calculated through a predefined decision function.
6. The product monitoring method based on machine vision and convolutional neural networks as described in claim 5, characterized in that: The normal score of the automotive bearing is compared with the first-level decision threshold, and the "qualified and process ended" and "deep inspection" statuses are output respectively, including the following steps: The normal score of the automotive bearing is numerically compared with the first-level decision threshold; When the normal score of the automotive bearing is higher than the first-level decision threshold, the automotive bearing is judged to be qualified and the monitoring process ends. When the normal score of the automotive bearing is not higher than the first-level decision threshold, the status of the monitoring process is marked as requiring in-depth verification.
7. The product monitoring method based on machine vision and convolutional neural networks as described in claim 6, characterized in that: Based on the deep feature vector, the second output head with Monte Carlo dropout operation enabled is forward-propagated to obtain the cognitive uncertainty quantification value, including the following steps: The deep feature vector is input into the second output head with Monte Carlo dropout enabled. The second output head with Monte Carlo dropout enabled is forward propagated to obtain the probability distribution of multiple sets of second output head outputs generated by multiple forward propagations. Statistical calculations were performed on the probability distributions of multiple sets of second output heads to obtain the quantification value of cognitive uncertainty.
8. The product monitoring method based on machine vision and convolutional neural networks as described in claim 7, characterized in that: Based on the quantification of cognitive uncertainty and real-time constraints, the final uncertainty decision threshold is determined, including the following steps: The remaining allowable decision time is calculated based on the preset total processing time limit and the time already consumed in the current process, serving as a real-time constraint; Based on the quantification of cognitive uncertainty and real-time constraints, the final uncertainty decision threshold is dynamically determined through a preset trade-off strategy.
9. The product monitoring method based on machine vision and convolutional neural networks as described in claim 8, characterized in that: By combining the preliminary diagnostic opinions of the automotive bearings, a comprehensive report on the final monitoring results of the automotive bearings is generated, including the following steps: Compare the quantified value of cognitive uncertainty with the final uncertainty decision threshold; When the quantified value of cognitive uncertainty is lower than the final uncertainty decision threshold, the preliminary diagnostic opinion of the automotive bearing is adopted as the final diagnostic opinion of the automotive bearing. When the cognitive uncertainty quantification value is not lower than the final uncertainty decision threshold, the final diagnosis opinion of the automotive bearing is marked as a high uncertainty state that requires manual review; By integrating the final diagnostic opinion of automotive bearings, the quantified value of cognitive uncertainty, and the specific information in the preliminary diagnostic opinion of automotive bearings, a comprehensive report on the final monitoring results of automotive bearings is generated.
10. A product monitoring system based on machine vision and convolutional neural networks, based on the product monitoring method based on machine vision and convolutional neural networks according to any one of claims 1 to 9, characterized in that: This includes a data processing module that performs image acquisition and preprocessing on the automotive bearing to obtain preprocessed two-dimensional image features and three-dimensional point cloud data of the automotive bearing. The perception module inputs the pre-processed 2D image features of the car bearing and the pre-processed 3D point cloud data of the car bearing into a pre-trained multi-task convolutional neural network. The shared feature encoder extracts the fused features. The first output head generates the normal score of the car bearing, and the second output head generates the depth feature vector of the preliminary defect judgment and high-dimensional features. The decision module generates a first-level decision threshold based on the current real-time throughput rate of the production line and historical quality statistics. It compares the normal score of the automotive bearing with the first-level decision threshold and outputs the qualified and process end status and the in-depth verification status respectively. The evaluation module, based on deep feature vectors, performs forward propagation on the second output head with Monte Carlo drop-off enabled to obtain the cognitive uncertainty quantification value; The decision-making module determines the final uncertainty decision threshold based on the quantified value of cognitive uncertainty and real-time constraints. By combining the preliminary diagnostic opinions of the automotive bearings, it generates a comprehensive report on the final monitoring results of the automotive bearings.